Speculative decoding is supposed to be output-invariant. Rejected drafts are discarded, accepted tokens must match what the target model alone would have produced, and that guarantee is the entire point of the algorithm. This post reports a measured exception: on two models of the Qwen3.5-MoE lineage, running llama.cpp's native multi-token-prediction (MTP) path on the Strix Halo box, enabling MTP measurably reduced code correctness on hard tasks. For Ornith 1.5 35B A3B the cost was 17% overall and 25% on the hardest task. A paired control on Qwen3.6-35B-A3B replicated the direction at about a third of the magnitude.

The scope claim first, because it matters: this is two models, one lineage, one inference stack, small samples. I am not claiming speculative decoding is broken in general. I am claiming that on these two models it demonstrably is not free, that open llama.cpp bugs make the mechanism plausible, and that the shared lineage may be part of the story. If you run Qwen-lineage MTP under llama.cpp, this is a measure-before-trusting situation.

Where this came from

The Ornith 1.5 benchmark found that the model's native MTP head is a real throughput win on this hardware, +14% to +40% at every concurrency level, once the draft depth is tuned to n=1. That post picked Q8_0 + MTP n=1 as the production config on speed alone and explicitly deferred the quality question. This is the quality question.

The harness

Execution-based scoring only, no LLM judge. Every task ships a real pytest suite; the model's code passes real tests or it does not. Every suite was verified against a hand-written reference solution before any model touched it, which catches bugs in our own tests rather than blaming them on the model. Every config runs identical seeds, so differences are attributable to the config, not to which random generations got drawn. Scoring is partial-credit (fraction of tests passed), which gives materially more resolution than all-or-nothing at small sample sizes. Full raw data kept: every model response, reasoning trace, and pytest output.

The off-the-shelf suites (SWE-bench, BigCodeBench, LiveCodeBench, EvalPlus and friends) were surveyed and passed over: most assume cloud-scale compute or need adapters to point at a local llama-server endpoint, and the question here needed paired seeds against a live server anyway.

Phase one: five easy tasks, at ceiling, telling us nothing

The first suite was five single-function problems with canonical solutions (parse a duration string, fix a buggy LRU cache, aggregate transactions, merge intervals, a token-bucket rate limiter). Three seeds per config across BF16, Q8_0 with and without MTP, Q6_K, Q4_K_M.

Result: near-perfect everywhere. Three configs at 111/111, Q8_0 dropping a single seed on a single task, and MTP exactly quality-neutral: Q8_0 with and without it produced identical outcomes across all 15 seed-runs. I briefly reported that as closing the quality question. It closed nothing. A suite everyone aces cannot separate anything, and the one Q8_0 failure did not even correlate with precision, since the lowest-precision quant was flawless.

The ceiling effect is worth a paragraph of its own because it is the trap in this kind of work. Easy tasks produce clean, confident, useless numbers. If I had stopped there, the conclusion would have been "MTP is quality-neutral," backed by 75 seed-runs, and wrong in the direction that matters.

Phase two: three hard tasks, real signal

The second suite was designed for the model's capability frontier: a dependency-injection container (lifetimes, constructor injection, circular-dependency detection, reverse-order teardown), a two-phase-commit coordinator (the nasty case: a resource that already committed must still roll back when a later commit fails), and a recursive-descent expression parser (precedence, right-associative exponentiation, unary-minus disambiguation). Piloted on BF16 first to confirm real per-seed spread before spending GPU-hours on the grid. Five paired seeds per config, partial credit.

Ornith 1.5, mean fraction of tests passed:

ConfigOverallExpr parser
BF160.8120.620
Q8_0, no MTP0.9010.837
Q8_0, MTP n=10.7450.666
Q6_K0.8110.604
Q4_K_M0.8220.873

Two findings.

MTP has a real correctness cost on hard tasks. Q8_0 falls from 0.901 to 0.745 with MTP n=1, a 17% relative drop, concentrated on the hardest task: the DI container falls from 0.868 to 0.649, 25%. The easy-task neutrality was real but meaningless. A concrete flavor of the failures: under MTP, one run produced code referencing a _container variable that was never defined anywhere, an error class that appeared in no non-MTP run.

Quantization does not behave monotonically. Q8_0 without MTP scored highest overall, above BF16. Q4_K_M was worst on the DI container (0.592) and best on the parser (0.873); Q6_K mirrored it. No config was uniformly better across the three tasks, which matches the published large-scale quant studies: quantization magnifies a model's existing per-task unevenness rather than degrading cleanly with bit-width.

Why would MTP change outputs at all?

The invariance guarantee should make this result impossible, so I went looking for a mechanism. The evidence ranks like this.

llama.cpp's draft-mtp path has open, reproduced non-determinism bugs on this exact lineage. Two independent reports (#23302, #23335), both on Qwen3.6 MTP models under greedy, seeded, fully deterministic sampling, show the committed token stream changing with the draft setting. Under greedy decoding the invariance should be trivially exact (accept only if the draft equals the argmax), and it is not. The second report shows divergence at n_max=1, the tuned setting used here, not just at deeper drafts. Both issues are open with no public root cause: a reproduced symptom, not a diagnosed line of code, but sufficient on its own to explain wrong-but-plausible tokens landing in generated code.

The Qwen3.5 lineage looks unusually hard to speculate against. A recent paper on self-speculation in hybrid models measured draft/target divergence by architecture: parallel hybrids like Falcon-H1 sit at a total variation distance around 0.3, while the sequential hybrid design (their example is a Qwen3.5 model) measures 0.8, with perplexity 82x more sensitive to attention ablation. Attention is load-bearing and non-redundant in this family. The same paper found acceptance drops sharply on tasks needing long-range dependencies, which maps onto the easy/hard split here: the hard tasks are exactly the ones threading state across a whole file. Different model, same lineage; corroborating, not proof.

The general MTP literature predicts the shape. MTP heads are trained on near-term objectives and their proposal quality decays with distance from the last verified token; reported strengths concentrate on low-entropy, repetitive, structured continuations. Hard multi-constraint code is close to the worst case.

The control: same tasks, same seeds, Qwen3.6-35B-A3B

If the cost were purely architectural (lineage plus llama.cpp bugs), any model in the family should take a comparable hit. If it were purely Ornith's own MTP head being immature, a mature head should show little or none. Qwen3.6-35B-A3B is the same lineage, ships a separately published MTP variant, and is this box's production daily-driver: a natural control. Same three hard tasks, same five paired seeds, same harness.

ConfigOverallDI container2PCExpr parser
Q8_0, no MTP0.8320.6501.0000.847
Q8_0, MTP n=10.7750.7380.8860.701

The direction replicates: 6.8% relative cost overall, with the two-phase-commit task losing a previously perfect score. The magnitude is roughly a third of Ornith's, and the effect is not uniform: the DI container actually improved under MTP. Read at face value, that moves the weight of evidence away from "purely architectural" toward a mix: a small, family-wide cost, consistent with the open llama.cpp bugs firing at some rate, compounded by a larger model-specific cost from Ornith's own younger MTP head. Five seeds by three tasks per arm is a small sample and the mixed per-task signal is real, so treat the attribution as suggestive. What is not ambiguous is that on neither model was MTP free on hard tasks.

Update: cross-engine confirmation at temperature 0

After this post went up, the Trail Openers team running our H100 deployment reproduced the finding on a different engine, different hardware, and a stricter protocol: vLLM on an H100, running the same Ornith 1.5 35B A3B benchmarked on halo above, temperature 0, greedy, five coding prompts, three repeats per config, comparing output hashes.

Promptbaseline vs MTP k=1baseline vs MTP k=2
merge listsidenticalidentical
binary searchidenticalidentical
flattenidenticalidentical
LRU decoratordivergeddiverged (differently)
kv parserdivergeddiverged (differently)

Every config was bit-self-consistent across its three repeats, so this is not noise: the engine is deterministic, and MTP changes the output. That is a qualitatively stronger result than anything above. At temperature 0.6, my llama.cpp measurements always left a residual escape hatch of sampling subtleties. At temperature 0 there is none: correct speculative decoding is provably lossless there, because the target model verifies every draft token against its own argmax. Divergence at temperature 0 is a verification bug by definition.

Correction 2026-08-31: that last sentence is too strong. A later investigation found a benign mechanism that also produces temperature-0 divergence: the verification pass changes tensor shapes, which changes CUDA kernel selection and floating-point reduction order, flipping last bits in the logits and thereby the argmax on near-tied tokens. The losslessness proof assumes exact arithmetic; hardware does not provide it. Temperature-0 divergence therefore proves outputs change, not that they degrade. The scored quality drops in this post stand on their own measurements; the hash divergence alone is weaker evidence than this section implied.

Two details in their data sharpen the diagnosis. First, k=1 and k=2 diverge differently on the same prompts, which rules out "MTP just breaks ties another valid way": the corruption depends on how many tokens were speculated. Second, the split is not random. The three simple prompts survive; the two harder ones (a memoizing LRU decorator, a parser with error handling) diverge. Longer, higher-entropy generation means more rejected drafts, more rejections mean more rollbacks, and that gradient matches the hard-task concentration measured above, including the 25% drop on the DI container.

The mechanism now has a name. In a hybrid model, rolling back a rejected draft is asymmetric: the attention layers just drop KV entries, but Gated DeltaNet updates its recurrent state in place, and rewinding that state on rejection is where the implementations break. There is an open vLLM bug for exactly this: PR #51508, "GDN/KDA: silent recurrent-state corruption (and CUDA crash) for stale zero-accept spec rows." That reframes the llama.cpp issues cited above: not a llama.cpp bug and a separate vLLM bug, but the same architecture-level trap, incorrectly handled in both engines we tested. It also explains why the failure presents as silent quality degradation rather than a crash, and why hybrid linear-attention models specifically are the risk class.

Their throughput pass completes the picture from the economics side. On the H100 (their measurements): baseline 158.3 t/s single-stream and 720.1 t/s at concurrency 8, with 12.48x max concurrency headroom; MTP k=2 gains +40% single-stream (221.9 t/s) but both MTP configs are slower than baseline at concurrency 8, and the drafter's own KV cost cuts concurrency headroom by roughly 20% (12.48x to 9.93x). The crossover between concurrency 1 and 8 is an independent replication of the high-concurrency crossover measured on halo in the Ornith post, on entirely different hardware and engine. Draft acceptance decayed with depth (roughly 57-75% depending on k; their logged and counter-derived figures disagree on the exact denominator, so treat that as a range). Their conclusion for that deployment: MTP off. Even with the correctness bug fixed, at real agentic concurrency it is a single-stream latency lever, not a throughput win.

Update 2026-09-17: a later 15-pair validation campaign on Qwen3.8-Flash-Next's MTP head added a data point on the benign side of this ledger: 0 of 15 pairs byte-identical (divergence is universal), yet functional correctness preserved in all but one comparable pair, with the single divergence scoring better under MTP. Same conclusion from the other direction: whether divergence costs anything is a per-model question only an executed benchmark answers. Details in the MTP adoptions post.

What this does and does not claim

Does claim: on two models of this lineage, with the central result reproduced on the same model across two engines and two hardware platforms (llama.cpp draft-mtp on ROCm/gfx1151, vLLM on CUDA/H100), MTP changes outputs it provably should not change, at a correctness cost that concentrates on complex, multi-constraint code, with GDN recurrent-state rollback as the identified mechanism (open, unmerged fix in vLLM; no public root cause yet in llama.cpp). And: the easy-task eval that showed perfect neutrality was true and useless at the same time, which is a warning about eval design, not about MTP.

Does not claim: that this generalizes beyond the Qwen-lineage hybrid-GDN family, that the exact magnitudes would survive more seeds, or that speculative decoding with a separate draft model has the same problem (it was not tested here). The rollback mechanism is specific to architectures that carry recurrent state, which is consistent with the lineage being the story, and equally a reason not to extrapolate to pure-attention models without measuring.

One gap worth naming: everything published on MTP quality that I could find measures perplexity, acceptance rate, or QA benchmarks. I could not find a prior measurement of MTP's effect on execution-graded code correctness, the metric that actually matters for a coding agent. If you know of one, I would like to read it.

Takeaways

  • Verify spec-decode invariance on your own stack, against tests that execute. "Speculative decoding never changes outputs" is a property of the algorithm, not of every implementation of it. It held in neither engine we tested, including at temperature 0, where it is supposed to be a theorem.
  • Eval difficulty is a validity condition. The easy suite produced 75 seed-runs of confident neutrality that phase two overturned. If every config is near-ceiling, the eval is measuring the ceiling.
  • Speed and correctness need the same rigor. The throughput bench that picked this config was careful: seeds, paired arms, acceptance-rate diagnosis. It still almost shipped a config that costs a quarter of the hardest task's correctness, because quality was deferred. Bench both before wiring anything into production.

Reproducibility: AMD Ryzen AI MAX+ 395 (Radeon 8060S, gfx1151), 128 GB unified memory, llama.cpp draft-mtp (--spec-type draft-mtp --spec-draft-n-max 1). Ornith 1.5 35B A3B (bartowski GGUFs) on llama.cpp b10530; Qwen3.6-35B-A3B control on the production b10038 container. Custom execution-based harness: pytest-scored tasks self-verified against reference solutions, paired seeds, partial-credit scoring, temp 0.6 per the vendors' agentic-use recommendations. Throughput context in the Ornith post; base stack in the setup guide.