Skip to content

SmolLM2-1.7B server inference regression after 91814d4 (Phi-3 CPU fallback) #77

Description

@unamedkr

Description

After commit 91814d4 ("Phi-3.5 server support + Metal workaround"), SmolLM2-1.7B server inference produces garbage output. This model previously worked correctly.

Steps to Reproduce

./build-metal/quant-server SmolLM2-1.7B-Instruct-Q8_0.gguf -p 8080 -j 8

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"messages":[{"role":"user","content":"What is gravity?"}],"max_tokens":30,"temperature":0.0}'

Actual Output

{"content":"<|im_endturernocturno<|im_ennd>\nWhat is the answer to this question: What is"}

Also tested with TQ_NO_METAL=1 — same garbage output.

Expected Output

Gravity is the force that attracts two objects with mass towards each other...

(Worked correctly in earlier builds before 91814d4)

Root Cause Hypothesis

The tq_matmul_force_cpu thread-local variable and _phi3_force_cpu flag in tq_forward() may not be correctly scoped — if the flag leaks across requests or isn't reset, non-Phi-3 models could get incorrect matmul routing.

Also, the tq_matmul_gguf_cpu extern function may have buffer sizing assumptions that don't hold for Q8_0 matrices.

Unit tests

35/35 pass — the regression is only visible in end-to-end server inference.

Environment

  • Commit: 91814d4
  • Model: SmolLM2-1.7B-Instruct-Q8_0.gguf (MHA 32/32)
  • Build: cmake -DTQ_BUILD_METAL=ON
  • OS: macOS 15 (Apple M3, 16GB)

Reported by ClawTeam

Activity

  1. unamedkr commented on Apr 12, 2026

    @unamedkr
    CollaboratorAuthor

    Root Cause Found: Numerical instability in libturboquant forward pass

    Evidence

    Debug output shows hidden state exploding at layer 7:

    layer6 out[0:4]=-1.966,-0.093,0.423,1.302 min=-301 max=74     ← normal
    layer7 out[0:4]=5.503,-10.735,3.848,-6.422 min=-496 max=18359  ← EXPLOSION
    

    After layer 7, values stay in the 18,000+ range, producing garbage logits.

    Key finding: NOT a regression

    • The same garbage output occurs at commits 91814d4, a8528f9, and a7795a5
    • quant.h standalone binary produces correct output for the same model + same ChatML prompt
    • CPU-only build (no Metal) also produces garbage
    • 35/35 unit tests pass

    Conclusion

    The libturboquant split source files (src/engine/tq_transformer.c) have a divergence from quant.h's forward pass that causes numerical instability starting at layer 7. This is NOT caused by any recent commit — it predates the Phi-3 changes.

    Possible areas of divergence:

    1. RMSNorm implementation differences
    2. Attention score scaling
    3. RoPE rotation numerical precision
    4. Q4 weight dequantization order

    Suggested fix

    A line-by-line diff of quant.h's tq_forward() vs src/engine/tq_transformer.c's tq_forward() is needed to find the divergence. The bug is in the core forward pass, not in generate/tokenize/EOS logic.

  2. added 3 commits that reference this issue on Apr 12, 2026
    27671f5
    dd1032a
    72e815b
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions