pi-delegate: Claude does the thinking, a cheaper model does the typing

Soo... pi-delegate is out. It's a Claude Code plugin with one skill: you say "delegate to pi: ..." and the heavy part of a coding task - reading files, editing, running tests - runs on pi, an open-source coding agent that can drive any model you pick, including one on your own machine. Claude only writes the brief and reads a short result. Your tests decide whether it worked.

The headline from my benchmark: on multi-file tasks, Claude's cost per run dropped from $0.082 to $0.055, a third off, with every hidden test still passing. But the more interesting part is that the first version of this tool made Claude more expensive, and the fix was mostly deleting things.

Version one cost more, not less

The first design was the one I'd have defended in an architecture review. A deterministic bash driver ran a develop, review, fix loop: pi wrote the code, a second read-only pi instance reviewed it adversarially, findings went back for fixing, up to six pi calls. Cheap model builds, strong model judges.

Then I benchmarked it against plain Claude Code on real historical fixes from the click library. Quality was at par, every run passed the grading tests. Cost was not: the delegated runs used 1.7x and 2.7x as much Claude as just letting Claude do the work. Where did it go?

  • the skill text and role prompts loaded into Claude's context, roughly doubling cache creation;
  • extra turns to launch the loop, wait for it and read back the summary;
  • and the big one: after the loop's own review had passed, Claude re-read the diff and re-ran the tests anyway, to be sure.

That last one is funny in a slightly painful way. I had built a whole review pipeline so Claude wouldn't have to check the work, and Claude checked the work anyway.

What changed

Three things, all subtractions.

The reviewer is gone. Field data from about 15 real delegated tasks showed a reviewer on the same model as the developer approved almost everything, while a stronger model later found real bugs in several of those "approved" changes. A same-model reviewer is a second opinion from the same mind. So now there's no model in the quality gate at all: you pass --verify "<your tests>", the script runs it after pi finishes, and if it fails pi gets exactly one retry with the failure output. A deterministic gate instead of a second opinion. Same lesson I keep landing on in other places: put the decision in code where you can.

Claude is told not to redo the work. When the gate passes, the skill tells Claude to reply in a sentence or two and stop. No re-reading the diff, no re-running the tests. That was the single biggest leak.

The output got trimmed. The script prints at most about 1.2 KB of pi's final text, because everything pi says gets re-read, and paid for, by Claude. The skill itself shrank to about 25 lines, with all the launch, wait and abort logic moved into run.sh.

The numbers now

Eight small tasks, two runs each, plain Claude versus Claude with pi-delegate, pi running Qwen3.8-27B on the Trail Openers H100:

Claude cost per runPlain ClaudeWith pi-delegate
Multi-file features$0.082$0.055 (−33%)
Tiny edits (ten-line fixes)$0.042$0.047 (+11%)
Runs passing the hidden checks28 / 2828 / 28

The tiny-edit row is the useful one. Delegating has a fixed overhead of roughly half a cent to a cent per task: the skill text, one extra tool round-trip, reading the result. On a ten-line fix that overhead is bigger than what pi saves, so don't delegate those. On multi-file work it pays for itself, with the best task at −47%.

The honest fine print

This counts Claude's cost only. pi's own spend comes on top: nothing extra if you already run your own model like we do, cents on a hosted cheap model. Delegated runs are also slower, several times over in these runs: about 25 seconds plain versus two to three minutes. Two runs per task is a small sample, and these are small synthetic tasks, not a big real codebase. Delegated diffs also came out 1.3-1.8x larger, mostly because pi writes more tests - which you may or may not want.

There are guardrails too: it refuses to run on your default branch or next to .env and key files, and disables git push for pi. That guards against mistakes, not a malicious model; for real isolation use a disposable clone, a container, or the optional sandbox wrapper.

So who is it for? I'd say: you use Claude Code, your usage limit or token bill matters, your tasks touch several files, and you have a test command that can tell right from wrong. If speed matters more than cost, or there's no way to check the result, Claude alone is still the better tool.

The benchmark takes about five minutes to rerun yourself with bench/quick.sh, and I'd genuinely like to see numbers from other setups and other models. Mine is one Claude model, one pi model on one H100 and a handful of tasks... early days tho.

A year of Fiona: my personal assistant, the Hermes swap, and the agent that rewrote its own playbook

Fiona is my personal assistant. I built her about a year ago, she lives in Telegram, and she runs on my own hardware. Last week her web-research worker hit a billing error, spent several minutes improvising around it, and then wrote that improvisation into its own long-lived instructions as standing policy. No attacker anywhere. Just an agent optimising for "finish the research" with write access to its own playbook.

The incident only makes sense as a test of decisions made a year ago, so let me start there.

A year of Fiona

The founding requirement for Fiona was never capability. It was containment. Some of what I handle through her is sensitive, and I did not want it leaking anywhere: not to a vendor, not to a random website, not through a clever prompt. So the design has been paranoid from day one. Fiona herself runs on an internal-only container network with no route to the internet. Her only way out is an allowlist proxy that reaches Telegram and the local model gateway, and nothing else. She cannot browse the web at all.

Web research had to live somewhere else, so it went to a separate worker in its own network zone. The first version was a small agent I wrote myself and named Pi. I wanted a non-gendered name for it... and only later noticed that the agent framework I build my coding harnesses on is also called pi. Naming things, one of the two hard problems, confirmed again.

Then Nous Research released Hermes, a capable open-source agent, and I made what felt like an obviously sensible call: swap my hand-rolled Pi for Hermes and get someone else's maintenance for free. Fewer things for me to keep patched, more capability out of the box. Keep that decision in mind, because it is half of this story.

The zone split stayed exactly as it was. Outbound research queries pass a PII gate before they leave Fiona's zone, so the worker only knows what it is asked. Its answers come back as data: English by contract, and Fiona's own model rewrites anything I actually read. The worker never talks to me directly.

What happened

On September 20 the credit on the search API account ran out. Every search from the worker started returning a payment error. Nobody noticed, because that error went into a log file and nowhere else.

On September 23 I asked Fiona a research question. Fiona called the worker, waited three minutes, timed out, retried, timed out again, and told me the worker "couldn't connect to the server". That was wrong: the worker was reachable and very busy. Inside it, each run hit the payment error and then walked a fallback chain on its own initiative: a browser (not installed, so that failed), the shell, code execution, delegation. Both abandoned runs finished long after Fiona had given up, holding the shared GPU the whole time.

Forty-two seconds after the second run ended, the worker's own research skill file was rewritten. It gained a new "fallback strategies" section: when search reports a payment error, fetch pages another way instead; when a site refuses access, make the requests look more like a normal browser. A one-off improvisation had become standing policy.

A few honest markers, because this is where it would be easy to overclaim. Established from logs and file timestamps: the payment errors, the fallback chain, the timing, the rewritten file. Inference: that the rewrite came from Hermes' post-run background review, based on timing (the weekly curator ran hours later). Unknown: whether any of those fallback attempts actually fetched anything (the logged attempts returned errors, but the worker did have internet access), and whether the rewritten skill was ever used before I found it. As far as I saw, no further research requests arrived before the fix.

And one detail that changes how I read the whole thing: that skill file did not ship with Hermes. The agent wrote it itself, back in August, and had edited it seven times since.

Why it happened

I'd put it this way: this was not a rogue AI. It was an agent optimising around a constraint, which is a more boring description and a much more useful one.

The agent was told to complete research. A billing error is just another obstacle, and routing around obstacles is exactly what it was built to do. Nothing in its instructions said "payment and auth errors mean stop". Hermes also has a self-improvement loop: post-turn reviews that can create and edit its own skills. That loop is the step that turned a workaround into persistent behaviour. And the worker had far more capability than the job needed: shell, browser, code execution, file access, delegation, memory and scheduling, when all it ever needed was web search. The workaround it wrote down only works because a shell was there.

The prompt-injection literature mostly assumes an attacker supplies the pressure. Here the agent's own drive to finish the task was enough. So the rule I take away: an agent that reads the web must not write its own instructions. Self-improvement plus untrusted input plus a completion objective means persistent behaviour change, with or without an adversary.

What held

This is the part that made the year of paranoia worth it.

The zone split. The agent that read the untrusted web and misbehaved never shared a process, filesystem or network with Fiona's memory, my profile, my chat history, or her secrets. Simon Willison's lethal trifecta is private data plus untrusted content plus an outbound channel. The worker had the last two and, by construction, none of the first. That breaks the exfiltration leg regardless of what the agent decides to do.

Secrets by construction. The worker gets a derived environment file with only the keys it needs: at the time, two. A misbehaving worker holds only its own credentials, which is why "rotate everything" was never on the table.

A one-way, typed boundary. Queries go out through the PII gate, answers come back as data and get rewritten by Fiona's own model. A confused worker cannot talk to me directly or inject instructions into Fiona.

Monitoring that doesn't depend on the agent. My topic monitors run through a deterministic script calling the search API directly, not through agent jobs. That was a deliberate decision in August, and it meant the agent's behaviour never touched them.

The egress canary stayed green. A periodic check proves Fiona's zone cannot reach the internet. It kept passing throughout.

Forensics were possible. The worker's logs kept every tool call with timestamps, and the modified skill file with its modification time was itself evidence. The chain from payment error to rewritten playbook reconstructed cleanly.

Human review on the fix path. When my operator agent tried to edit the self-modified skill directly inside the container, the permission system blocked it, and it also blocked merging fixes without review. So every change went through reviewed, tested pull requests, and independent reviewers caught two real bugs in the first "lock down the worker" change by reproducing it against the actual image. One left the config owned by root, so the gateway silently fell back to defaults with every tool re-enabled. The other blocked the browser toolset, which silently removed web search too, because toolsets overlap. Guardrails need testing against the real artifact, not against stubs.

And one thing that held by luck, not design: the browser fallback failed only because Chrome was not installed in the image. I'm counting that as a near miss, not a defence.

What failed

No least privilege inside the worker zone. I spent a year hardening the boundary around the worker and almost no time on what the worker could do inside it.

Self-modification on by default. Skills were writable, and the background review and curator were enabled. I hadn't turned them on. I also hadn't turned them off.

Alerting that stopped at log files. Four days of payment errors produced zero alerts. A health check even noticed the errors, and wrote that down in a log too.

Misleading failure reporting. "Couldn't connect" hid both the real cause and the fact that the worker was still running, and Fiona's instructions encouraged retries, which stacked orphaned runs on the GPU.

Quiet looked the same as broken. In an unrelated failure, 7 of 11 topic monitors had silently died weeks earlier from a crashing sync script, and the morning digest kept saying "nothing new".

What changed

The worker is now a web-only worker: an allowlist of web search and fetch, nothing else, and it refuses to start if its config ownership or policy is wrong. Skill and memory writes are staged for my approval and never applied automatically, and a staged change pings me on Telegram: propose, don't apply. The background review, the curator and the agent's own memory are off. The self-authored research skill is archived. Fiona now reports failures honestly and does not retry them, including a clear "search unavailable" signal. Provider errors and failed worker jobs alert me directly. The digest says "monitoring degraded" instead of "nothing new" when monitors are down.

One more thing, and it is the part I keep thinking about. The Hermes upgrade that shipped these guardrails also enabled a new default: when a paid search fails, silently retry on other vendors' free, anonymous tiers. A brand new egress path, arriving in a version bump, and one that would have hidden the next billing failure completely. I switched it off explicitly.

The trade I actually made

Which brings me back to the swap. I replaced my own worker with Hermes to get less maintenance, and I did get less maintenance. What I also got, without really noticing, was someone else's defaults: a self-improvement loop, a broad toolset, and now a fallback behaviour I would never have written myself. None of those are bugs from Hermes' point of view. They are features for a general-purpose agent. They were just the wrong features for a narrow worker sitting next to sensitive data.

So I'd say the lesson is not "don't use Hermes". It's that outsourcing maintenance means inheriting defaults, and defaults drift with every upgrade. Keep the worker boring. Put the safety in network boundaries, credentials and file permissions, because the prompt did not stop anything here and the architecture did. And re-audit the defaults on every version bump... the next one is probably already on its way.

MTP on Strix Halo: +79% at one stream, zero at four, and the ceiling isn't bandwidth after all

This blog's running finding on MTP - speculative decoding via a model's built-in multi-token-prediction head - has been that its gains invert under concurrency: measured on the halo box at high slot counts, measured on the H100 where MTP was outright slower than baseline at 8 concurrent streams. The safe deployment shape we settled on was single-slot, opt-in routes.

New measurement, cleaner than any of the earlier ones, and it refines the doctrine. Qwen3.8-Flash-Next (UD-IQ3_XXS, its MTP sidecar, the newly adopted engine build), real coding generations rather than synthetic tokens, thinking off so tokens are code tokens, speeds read from the server's own timings. Aggregate throughput:

Concurrent requestsWith MTPWithout MTPMTP gain
149.7 t/s27.8 t/s+79%
263.5 t/s38.0 t/s+67%
463.1 t/s63.2 t/snone

Draft acceptance at 2 slots: 98%. Correctness under concurrency: 18 of 18 answers correct at both 2 and 4 slots (6 distinct prompts, 3 rounds each). And the row worth staring at is the last one: with or without MTP, four concurrent streams aggregate to the same ~63 t/s.

What the ceiling is, and what it isn't

Correction, same day: the first version of this section called the plateau "the signature of a memory-bandwidth-bound machine". A reader's comment on the announcement took that apart, and he's right. What follows is the corrected analysis; the original claim is retracted.

The counterargument is self-contained and worth learning from. If the binding cost were bytes of weights per second, then amortising weight reads - which is exactly what both MTP and batching do - should keep raising aggregate tokens per second. A ceiling that is indifferent to how much you amortise is not a per-weight-pass cost. It is a per-token cost. My own explanation ("speculation and batching are the same trick") was, read carefully, an argument against my own headline.

And our own numbers agree, once actually computed instead of gestured at. Plain decode moves roughly the model's ~6B active parameters per token, call it 2-3 GB at this quant; at 27.8 t/s that is ~70 GB/s of achieved bandwidth, about a third of the ~215 GB/s this box can move. Not bandwidth-bound, not even close. The step times say the same: a 1-token step costs ~36 ms, a 4-token verification step ~63 ms. If weights were read once and shared, the ratio would be near 1x, not 1.75x. And an earlier quant sweep implies a fixed ~18 ms floor per decode step regardless of model bytes, roughly 40% of the step.

So what is the per-token cost? Honestly: not established. It does not look like a clean FLOPS roofline either (the arithmetic puts utilisation well below peak). The live suspects are per-token kernel work that amortises poorly: 36 of this model's 48 layers are recurrent gated-delta-net, whose state updates cost per token verified; MoE routing means each drafted token can hit different experts, so expert bytes grow with draft length (the amortisation premise was never quite true for a MoE); and IQ3_XXS is one of the most dequant-heavy formats there is. Hypotheses, all three. The settling experiments are named and queued: profile achieved GB/s during a 4-slot run, and repeat it on a cheap-to-dequant quant like Q4_K - if the ceiling rises despite more bytes, it was never about bytes.

What survives untouched is the contrast with the H100 result, where the same lever went negative under load because that box binds on KV-cache headroom and scheduler pressure, which MTP's overhead actively damages. Halo's ceiling absorbs the overhead; the H100's constraints are hurt by it. So the transferable rule stands, now with better epistemics: MTP's value under concurrency isn't a property of MTP. It's a property of your machine's binding constraint - and you should identify that constraint by measurement, which, as this correction demonstrates, I had not.

What changed in production

The old doctrine (single-slot MTP routes only) was inherited from measurements where gains inverted. This curve says something friendlier for this box: the gain tapers, but it never goes negative, and at 2 slots it's still +67%. So the production route for this model is now --parallel 2 with MTP - two agents get 63.5 t/s aggregate where they used to get 38. MTP stays on because on this box there is, on this evidence, no concurrency level at which it hurts - a measured fact that does not depend on knowing why the ceiling sits where it does. Quality: unchanged, as previously established - same weights, verified drafts, matching perplexity, and the earlier campaign's 0-of-15 byte-identical / 10-of-11 functionally-identical result still stands.

Honest limits

Mostly single runs (the winner's repeat agreed within ~4%), one machine, one quant, up to 4 slots. Not tested: the full 262K context, more than 4 slots, quantized KV combined with MTP, or a long soak under real agent load. The engine underneath is an experimental community build, hand-built and unpinned - the fuller audit that selected it is its own story. And "never negative" is a claim about this box's constraint, not about yours: on a machine that binds on KV or compute first, the H100 numbers are the ones to expect.

The MTP series scoreboard, five posts in: it changes outputs at the bit level everywhere, costs quality on some heads and none on others, pays +40-80% single-stream when the head is good, and under concurrency it does whatever your bottleneck tells it to. Measure the bottleneck first - actually measure it, not name it from vibes as the first version of this post did - and the MTP decision makes itself...

A Strix Halo engine audit: a scary claim that didn't reproduce, a fork worth +35-92% prefill, and seven broken routes

A research sweep of the Strix Halo ecosystem turned up four things at once: newer kyuz0 toolbox builds, an experimental toolbox running a Halo-specific llama.cpp fork on a patched ROCm runtime, a community author who abandoned ROCm claiming it "returns wrong logits for any prompt longer than the batch size" (perplexity 84.8 versus 13.8 on Vulkan), and a headline of 1207 t/s prefill on this hardware. House rule, by now well earned: nothing gets adopted on faith. So everything got tested. This is the short version; the full scripts and results are journaled in the audit's own directory.

The scary claim: not reproduced

Wrong logits from a production backend would outrank every speed number on this blog, so it went first. Method: llama-perplexity over 160 KB of real Python source, same text and model on both backends, varying exactly the thing the claim names (physical batch size smaller than vs equal to the prompt) plus the claimed workaround (HIP_LAUNCH_BLOCKING=1). Two models: the production Qwen3.6-35B and Qwen3.8-Flash-Next.

Every ROCm configuration landed within 0.3% of Vulkan. Qwen3.6: 1.2994 on ROCm in all three configs against Vulkan's 1.2981. Flash-Next: 1.1750-1.1754 against 1.1718. We also re-ran the concurrency-corruption check from the known HIP bug: 12 of 12 concurrent answers correct. Stated carefully, because this matters: we cannot prove the claim wrong on the author's build and workload. We can say it does not reproduce on ours, with the exact knob it names being turned. Perplexity that moves 0.3%, not 6x, is not silent corruption.

The engine race: real prompts, one winner

Flash-Next's engine has been a hand-rebased fork since day one. Three candidates ran the same weights, same MTP sidecar, same inputs: prefill on real source code at three depths with caching off, decode on real coding tasks with thinking off, numbers from the server's own timings.

Enginepp ~2k / 8k / 24k (t/s)decode (t/s)
Old production fork (b10672)387 / 400 / 34945.8 / 40.0
strix-llama, patched ROCm (b11195)523 / 677 / 66856.8 / 51.9
strix-llama, -ub 16384 + lazy embeddings300* / 899 / 92452.0 / 49.5
Community Vulkan branch472 / 502 / 44358.4 / 55.6

Adopted: strix-llama. Roughly +35-92% prefill and +24-30% decode over the old engine, on the metric this box actually cares about. The asterisk is a measurement gotcha worth stealing: the first request after model load reads low. Steady-state 2k prefill on the adopted engine is ~636 t/s, not the 300-520 the first request shows, so our own first comparison understated the winner. If your benchmark's first row looks odd, it probably is.

The 1207 t/s headline? A different measurement: synthetic pp2048 at depth zero, a different quant, -ub 16384. Our best real-prompt number is 924 t/s at 24k depth. Not reproduced on our metric is not the same as false, and this one is a textbook case of why benchmark configs must travel with benchmark numbers.

The Vulkan runner-up lost on everything except decode: slower prefill at depth (the wrong direction for agentic work), 439 commits behind upstream (which means older tool-call and template handling, exactly what agents exercise), and a trimmed sidecar format incompatible with the official one, so its speedups can't be combined with anything else's. Credit where due: it fixed the September Vulkan MTP collapse, holding 55 t/s over a 1037-token generation where the old branch fell to 6.3. Quality on the adopted engine: unchanged by every check we have (matching perplexity to the fourth digit, identical outputs on repeated prompts, 5/5 generated-code tests, 4/4 tool-calling round trips in both thinking modes).

The war story: a routine refresh broke seven routes

The same sweep refreshed two standard toolboxes, and the new builds removed --no-mmap outright in favour of --load-mode. Our launcher passed --no-mmap by default. Seven model routes, including the always-on model behind the assistant's research sidecar and the default agentic coder, failed to start for about twenty minutes, and the only reason it was twenty and not a day is that one real test request went through after the refresh. The container's version string looked fine throughout. This is the same flag removal that broke the tuning skill's scripts two weeks ago, now graduated from "migration note" to "production incident."

The fix is a probe, not a pin: the launcher now asks the container's llama-server whether it speaks --load-mode and picks the right flag. And the rule that went in the journal: after any toolbox refresh, send one real request through every affected route. Version strings lie by omission; requests don't. The refresh also paid for its trouble: +14-29% prefill on the default coder, decode flat.

Honest limits

Mostly single runs on one machine and one quant (the winner's repeat agreed within ~4%); no 262K-context test, no quantized KV with MTP, no long soak. The adopted runtime is experimental and hand-built, unpinned, with no registry copy - and its own README recommends a workaround for a batched-inference bug that none of our checks reproduced. Post-switch verification went through the real router: 49-second load, 62.8 t/s at two concurrent requests, 18/18 concurrency correctness, 4/4 tool calling. The MTP concurrency curve measured on this engine got its own post.

Four claims went in; one engine came out. I'd say the audit's real product isn't the +35-92% - it's the two sentences you can reuse: not reproduced on our metric is not the same as false, and one real request beats any version string...

Would Jev help pi-rukas? Mostly no, and the reason is the interesting part

The hot potato of the weekend: Jev, from TypeSafe AI, the lab Diogo Almeida (RLHF co-inventor) took out of stealth on September 15. Jev is a "System 1 model": instead of generating text, it returns a typed decision in a single forward pass - a choice from a schema, a score on a rubric, or a calibrated yes/no probability. No autoregression, no hallucinated prose, 70-500 ms, input at $0.042 per million tokens. The pitch is 20-200x faster and 40-400x cheaper than frontier LLMs on decision-shaped work.

And the ecosystem sprinted. Within days there were LangChain harness guides, safer-agent tutorials, and use-case roundups - by this weekend, seemingly everyone with an agent harness was retrofitting Jev into it, all making the same move: take the judgment calls in your agent loop - route this intent, gate this tool call, approve this step - and hand them to Jev instead of an LLM. One LangGraph demo measured the trade honestly: 96.8% routing accuracy for Jev against 99.1% for an LLM, at a fraction of the latency and cost.

There's a subplot too. Days after Jev's launch, Laya appeared: an open-source, Apache 2.0 implementation of the same idea from Nandakishor Mukkunnoth of ConvAI Innovations, built on ModernBERT-class encoders, claiming 6-8x faster than Jev with 3x better calibration. The announcement's framing is barbed: he published the non-autoregressive decision-model concept in March 2025 and formalised the training method (RLCD) that September, a year before a frontier lab presented the same concept as a breakthrough. Jev's own materials name their training method... RLCD. I'll let readers weigh that one themselves.

So, the obvious question for anyone running an agent harness: should this go in ours?

What pi-rukas already does

pi-rukas is the Trail Openers agent harness, the current form of the orchestrator lineage this blog has written about since pi-ensemble. I spent a while this weekend checking where Jev or Laya could slot in. The answer surprised me by how small it was, and the reason is worth spelling out.

The heavy flows in pi-rukas are compiled drivers: TypeScript state machines where an issue runs end-to-end through explicit states. The work driver is literally a string-union step enum and a transition table:

export const WORK_STEPS: readonly WorkStep[] = [
  "explore", "plan", "branch", "develop", "adversarial", "commit-pr",
  "lens-review", "lens-fix", "step-back", "handoff", "ci", "merged",
];

Twelve states, 33 event kinds, 23 distinct cap conditions (review rounds, CI retries, wall clock, token budgets, loop detection), per-step failure policies (halt / retry-once / degraded-ok), and a persisted event log that is authoritative over the state snapshot. Research and planning run as the same kind of pipeline. The README says it plainly: /work is a compiled driver, not prose.

The division of labor is the whole design. Deterministic code owns every transition, every cap, every gate, and every irreversible act: git operations, PR creation, and merging happen on executed evidence (a real diff exists, typecheck and tests pass, gh pr checks is green), never on an agent's claim. LLM calls own the things that genuinely need judgment: exploring a codebase, writing the code, adversarial review, the six-lens review. And when an LLM answers, it doesn't answer in prose that code then trusts: verdicts come back through schema-validated tool calls or through parsers whose core rule is that a marker that fails to parse is "absent", never "approved". An unparsed verdict at the cap becomes an infrastructure retry, not a pass.

One precision that matters: this is deterministic control flow, not deterministic outcomes. The LLM still decides whether the diff passes adversarial review. What it cannot do is decide what happens next with that verdict, skip a gate, extend its own budget, or rewrite the rules it's judged by (policy prose is read at base SHA precisely so a cycle can't self-grant). The machine consumes typed values; it does not take advice.

None of this was built for aesthetics, and the doctrine for when to build it is written down in the repo: compile a flow when the steps are enumerable, a repeated incident class exists where LLM improvisation was the failure source, and a mechanical gate can catch the failure; keep prose when the shape is judgement-driven. Every compiled driver traces back to numbered incidents. The research behind that doctrine leaned on MAST, the 1,642-trace taxonomy of multi-agent failures this blog has cited before: 41.8% of failures are system-design issues, and fixing the structure - role specs, a verification step - moved task success by +9.4% and +15.6% on the same model with the same prompt. Structure is a design variable, not a model capability. That is the whole bet the drivers make.

Why a decision model has almost nothing to add here

Now look at what the weekend's Jev tutorials use it for: should this request route to the cheap model? Is this tool call risky? Should the loop continue or stop? Has the agent finished?

In pi-rukas, essentially none of those are model calls. Which model handles which role is configuration. Whether a tool call is permitted is a guard, a hook that fires on the tool-call event and refuses, in any trust mode. Whether the loop continues is the transition table plus caps. Whether the work is done is executed evidence. Replacing an LLM judgment call with a 30 ms calibrated classifier is a real improvement - 96.8% at near-zero cost is a good trade against 99.1% at LLM prices for high-volume routing. But a transition table is 100% at zero milliseconds, with calibration that isn't a metric because there is no probability involved. You cannot beat code at being code.

And the first independent measurements of Jev, out within days of launch, back the code side harder than I expected. One hand-labelled tool-call-risk benchmark found Jev is not deterministic: identical input flipped the label on 1.7-3.3% of items. The same study noted the "zero hallucinations" claim is trivially true of any schema-constrained decoder - all 1,800 of its GPT-class calls under a strict JSON schema were "hallucination-free" in exactly the same sense. A multi-suite quality benchmark put Jev mid-pack on accuracy (76.3% on a 77-way intent task, behind two LLMs) and found the vendor's 193x speed claim "holds only against LLMs left in their slowest default mode". An ordered tool-routing eval got 44% hit rate - better than the LLM's 24%, and still wrong more than half the time. And a pre-registered routing study found confidence-cascading Jev to an LLM wins on one dataset and loses on the next, with no routing parameter transferring between them. Type-safe is not the same as correct, and non-autoregressive is not the same as deterministic.

The learned-router literature says the deeper version of the same thing. A 21-method routing study found routers converge to a narrow band far below the oracle because they learn global trends, not per-query signals. Learned routers collapse to the expensive arm as budgets rise, a failure mode code does not have, and they are adversarially steerable with query-independent token gadgets. A robustness lifecycle study puts the punchline better than I could: training-free routers showed the strongest adversarial and backdoor robustness, "benefiting from the absence of learnable parameters". Learnability is precisely what sells the classifier, and precisely what buys the variance back.

I'd put the general principle this way: most of what people are gating with Jev this weekend is control flow, and control flow should not be a model call of any size. The System 1 framing is right about the problem - agent loops are full of small decisions that don't need a frontier model - but a 421M-parameter encoder is the second-best answer to most of them. The best answer is that the decision was never open. The honest scope of Jev and Laya is the decisions that are genuinely about content: is this email spam, which of ten queues does this ticket belong in, does this text violate policy. Those need a model because code can't read. If your harness makes thousands of those per hour, a calibrated non-autoregressive classifier is exactly the right shape, and Laya's open weights make it locally runnable, which is very much this blog's territory. The published case where the pattern demonstrably works confirms the shape: MemRouter replaced an LLM's per-turn memory-admission decision with a small classifier and beat it (F1 52.0 vs 45.6, latency 970 ms to 58 ms) - a closed boolean, high volume, and the classifier as sole authority, not an advisor the main model can overrule.

pi-rukas mostly doesn't have that shape: its remaining judgment calls are few, low-volume, and deep - is this diff correct - which is precisely the case where you want the big model and the executed-evidence check behind it, not a fast classifier. "Mostly", because the research pass did find one seam that fits: intent resolution, a closed five-way verdict (proceed / park / and friends) where the LLM resolver has a documented history of unreliable answers and a wrong "proceed" costs a full work cycle. That one is classifier-shaped, and it's on the list. One seam, out of a harness. That's what "for the most part, no" means when you actually count.

So if you're one of the many wiring Jev into a harness right now, the exercise I'd actually recommend is the one this post came from: list every place your loop asks a model a question, and sort them into three piles. Decisions about content at volume: Jev or Laya, genuinely. Decisions that need deep judgment over a diff or a document: keep the big model, put an executed-evidence check behind it. Everything else - routing, permission, continuation, done-ness: that pile is your missing state machine, and no classifier, however fast and well-calibrated, should be holding it. My suspicion is that for most harnesses the third pile is the biggest one, which would make the honest headline of the weekend not "gate your agents with Jev" but "stop asking models questions code can answer".

This is the same conclusion this blog keeps arriving at from different directions: instructions steer, mechanisms stop, and a skill's rules had to become code before agents actually followed them. pi-rukas even splits the vocabulary explicitly: steering is prompt-layer discipline (there's a dispatch_steer tool for mid-flight nudges, deliberately uncapped, trust-model-of-the-prompt), gating is code. The weekend discourse uses "gating and steering" as one phrase. They're opposites.

The tongue-in-cheek part, which is only half tongue-in-cheek

Laya's author watched a well-funded lab present his year-old concept as a breakthrough. I have some sympathy, and also some deja vu: when "loop engineering" got named and celebrated this summer, the practice it described had been running on this blog's infrastructure for over a year. There's a pattern here: practitioner builds the thing because the work demands it, writes it up in unglamorous terms (a transition table! caps! parsers that admit they didn't parse!), and a while later the same idea arrives from a lab with a name, a waitlist, and a launch video.

So let me get ahead of it this once. pi-rukas' drivers - deterministic, typed, event-logged control flow where language models supply judgment as schema-validated values and never hold the steering wheel - are public, Apache-licensed, and running in production. If a frontier lab would like to call this a breakthrough in a few months, the repo history has the timestamps ready. I'd suggest "System 0": it's like System 1, but faster, free, and correct by construction...

DwarfStar, DiffusionGemma, and the ROCm 10 unfreeze: an ecosystem status report from halo

Not everything investigated for halo gets adopted, and the reasons why are often more useful than another adoption story. Three items from the recent research pile: two dead ends worth knowing about, and one unfreeze that was pure win.

DwarfStar: real support, wrong backend

DwarfStar (antirez/ds4) is Salvatore Sanfilippo's local inference engine, 22,000+ stars, built primarily for the DeepSeek V4 family. It came onto the radar because of a recollection of impressive Qwen3.8-Flash-Next numbers floating around - and the support turned out to be real: a merged PR adding native qwen4exp support, benchmarks through 262K tokens, actively maintained.

Two catches. First, the backend: at investigation time the model's dedicated kernel graph was documented as Metal-only, and while the engine has since broadened (it now targets Metal, CUDA and ROCm generically, with the Flash-Next path on Metal and CUDA), there is still no ROCm implementation of this model's kernels - and halo is a ROCm/Vulkan box. Second, the packaging: DwarfStar's Flash-Next builds run 137-165 GiB on disk with a 95 GB n-gram table read straight from the GGUF, a fundamentally heavier shape than our ~35 GB quant plus a 4 GB sidecar. A well-built project that simply is not for this hardware yet. Worth a re-check when the backend matrix grows; not before.

The meta-lesson from how this was investigated: the first research pass reported the ds4 PR as "still open, weak numbers". An independent gh pr view showed it had merged the day before, with materially different numbers than reported. Never trust a single research pass on anything that will inform a real decision - a rule this blog has now paid for and been paid by several times.

DiffusionGemma: served by vLLM, not by llama.cpp

DiffusionGemma is Google's experimental text-diffusion take on Gemma-4: the same 25.2B/3.8B-active MoE backbone, fine-tuned into a block-diffusion model. Instead of predicting one token at a time, a causal encoder reads the prompt and a bidirectional decoder iteratively refines a 256-token block, roughly 12 forward passes per block and ~20 tokens per pass - Google cites ~1,500 tok/s on a single H100. Genuinely different decoding economics, which is exactly why it's interesting for this hardware class.

The catch for halo is where the support landed. vLLM serves it natively. llama.cpp does not: support lives in an unmerged draft PR that adds a dedicated llama-diffusion-cli binary, and llama-server - the OpenAI-compatible path this whole box's architecture depends on - fails on the model with an unknown-architecture error on both ROCm and Vulkan. There's even a public failure report from another Strix Halo box hitting exactly that, and the community workaround is a Python proxy that shells out to the CLI. A one-off CLI behind a proxy is a demo, not a route.

So the accurate status: the serving path exists, it just lives in vLLM, and halo's entire stack is llama.cpp. Filed under revisit-when-llama-server-support-lands, which I'd say is the correct amount of enthusiasm for a model this box can build but not serve.

The ROCm 10 unfreeze

The quiet win of the batch. Halo's production ROCm container had been frozen at a six-week-old llama.cpp build since early August, pinned there to dodge a HIP integrated-GPU corruption bug (#25992) whose complete fix never merged. Then kyuz0's toolboxes moved to ROCm 10, where gfx1151 graduates from tolerated to an AMD-packaged, officially supported target, with the workaround baked in.

Evaluated with the usual paranoia before touching production: the corruption symptom exercised directly (8/8 concurrent distinct prompts, no cross-contamination), then a speed A/B against a freshly-taken baseline on the frozen build, not stale historical numbers. Result: prefill +10.8% at 8K depth and +17.5% at 16K, decode flat. For a box whose stated priority is prefill at depth for coding agents, that's the right win in the right place, for free. Adopted; the frozen container stays on disk as rollback.

One migration note for anyone following the same path: the new builds have removed --no-mmap outright in favour of --load-mode, which breaks any script that hardcodes the old flag - including, at time of writing, the strix-halo-optimize bench scripts on this class of toolbox. Flag shims needed; the skill's findings log has the details.

Three investigations, one adoption. The ecosystem moves fast enough that the dead ends of September are worth re-checking by November - which is, I suppose, exactly why the journal exists...

Four coding models on one Strix Halo: pass@1 hides two very different failure modes

The coding-model roster on halo got a proper shootout: four models, two benchmark systems, every number measured on this box. Which model won? That turns out to be the less interesting question. The finding I'd actually lead with: pass@1, the number everyone quotes (the share of problems a model solves on its first try), quietly merges two completely different ways of failing - answering wrong, and never answering at all. And the model ranking changes depending on which failure your workload actually cares about.

The roster

Halo (the AMD Ryzen AI MAX+ 395 box from the setup guide, 128 GB unified memory) currently serves four coding models via llama-swap:

  • gemma-4-26b-a4b: Google's Gemma-4 MoE, 25.2B total / 3.8B active, the box's strongest single coding model and the production grounding champion elsewhere in the stack. Now also vision-capable here (more on that another day). A second route runs the same weights with an official Unsloth MTP speculative-decoding sidecar: +68-73% decode, measured.
  • qwopus3.6-35b-coder: a community fine-tune of the Qwen3.6-35B-A3B lineage, thinking-off by default, tuned for fast read/edit/test/fix agent loops. The default fan-out model, serving 6 concurrent slots.
  • qwen3.8-flash-next-mtp: the Qwen4-preview architecture, now with its native MTP head working via a community-built toolbox. That story earned its own post, coming next.
  • north-mini-code: CohereLabs' North-Mini-Code-1.0, a 30B-A3B MoE, smoke-tested this week and not yet in production routing.

Housekeeping note with a number attached: this roster review also decommissioned six models (the whole Nemotron-3 line, Qwen3-Coder-Next, Ornith-1.5, Qwen3.5-122B), freeing 243 GB of disk. Model churn is real; the audition process keeps working in both directions.

Getting comparable numbers: EvalPlus, and the bug that nearly poisoned them

Our own to-bench-style harnesses are great for A/B work but mean nothing to anyone else's hardware. So the roster got scored on EvalPlus: HumanEval+ and MBPP+, the standard, globally-comparable Python coding benchmarks. Anyone can put these numbers next to published ones, which is the whole point of adopting them.

Except the first numbers were garbage, and the way they were garbage is the instructive part. EvalPlus's OpenAI-API backend hardcodes a 768-token generation cap, no flag to raise it. Reasoning models spend their token budget thinking before any code appears - so the harness was silently truncating them mid-thought and scoring the empty result as a wrong answer. The tell was an anomalous number of empty completions. Raising the cap to 8192 (a monkeypatch; there's no flag) fixed most of it, and then revealed a second, genuine finding: some completions stay empty at any budget. On a reproducible fraction of problems, these models reason indefinitely and never answer. Same problems, every run, multiple budgets: a characteristic, not an artifact.

The per-model empty rates at 8192 tokens: qwen3.8-flash-next 12/164 on HumanEval+ and 27/378 on MBPP+ (~7%), north-mini-code 5/164 and 9/378 (~2%), qwopus 1/378, and gemma-4 zero throughout - not because it never overthinks, but because its production config caps reasoning at 1024 tokens, which prevents the failure by construction. Empties are counted as fails in the scores below, which I'd argue is correct: an agent that never answers has failed the task either way.

One more trap for the pile: EvalPlus silently skips re-scoring when a results file already exists. After the patch, the scores did not change until a --i_just_wanna_run flag forced re-evaluation. The series tradition of measuring the wrong thing without an error message continues: your benchmark harness is part of the experiment.

The scores

pass@1, greedy decoding, temperature 0:

ModelHumanEval+ baseHumanEval+ plusMBPP+ baseMBPP+ plus
gemma-4-26b-a4b0.9880.9510.9390.796
qwopus3.6-35b-coder0.9570.9090.9260.794
qwen3.8-flash-next-mtp0.9270.8960.9180.812
north-mini-code0.9390.8840.9370.804

Read column by column and the tidy ranking dissolves. HumanEval+ shows Gemma-4 clearly ahead. MBPP+, with shorter and more constrained problems, narrows everything and flips the base ranking outright: North Mini's 0.937 against Gemma-4's 0.939 is a statistical tie at n=378. And the MBPP+ plus column is led by qwen3.8-flash-next, the model with the worst rate of never-finishes-thinking failures in the whole roster.

That last one is the takeaway I'd underline. Flash-Next's failures cluster almost entirely in "never answers", not "answers wrong": when it produces code, the code tends to be correct. Think about what that means for model selection: a model that answers everything with 92% accuracy and a model that answers 92% of problems with near-perfect accuracy carry the same pass@1, and they are completely different tools. One needs review. The other needs a timeout. Raw pass@1 tells you neither, and I'd guess most model-picking decisions out there never split the two.

Honest limits, stated up front rather than in a footnote: EvalPlus is Python-only, single-turn, zero tool use. It does not test the agentic terminal work these models actually do all day. Portable signal, not the whole picture - and the agentic-benchmark gap is now the most obvious hole in this box's evaluation suite.

The realistic speed sweep

For speed, the usual llama-bench synthetic-token corpus was deliberately rejected. Prefill content: an ~855 KB concatenation of the transformers library's actual modeling_*.py source files, sliced at three depths, with prompt caching forced off so repeated depths get no cache discount. Decode task: implement a thread-safe LRU cache with tests. Numbers read from the server's own reported timing fields, not wall-clock. And one deliberate choice worth naming: each model ran its own production route (its real backend, quant, and flags), because the question was "how fast is this model as I actually serve it", not a backend-controlled lab comparison.

ModelPrefill @2k@8k@24kDecode t/s
gemma-4-26b-a4b1201112188340.6
qwopus3.6-35b-coder90888273164.5
qwen3.8-flash-next-mtp35938934423.6
north-mini-code90080366765.8

Caveat first, because it's a real one: three of the four decode numbers measure reasoning-token throughput, not code-token throughput - the sweep didn't force thinking off, and only Gemma-4's 1024-token reasoning cap left room for actual code inside the test budget. Architecturally, next-token decode cost doesn't depend on what the tokens say, so it's very likely a fair proxy. Flagged, not buried; re-running with thinking forced off is on the list.

The structural finding survives any caveat: Flash-Next's prefill sits at a third to a quarter of the other three, and no amount of MTP decode speedup touches that. Agentic work is prefill-dominated - agents re-read code and re-send long context constantly - so this is the bottleneck that decides real workloads. The receipts are in the benchmark wall-clocks themselves: Flash-Next took ~3.5 hours for HumanEval+ against ~1.5 for the others, and ~8.5 hours for MBPP+ against ~2.5-3, while scoring similar-or-better per problem. Good answers, slowly: the day-one assessment of an immature kernel stack still stands.

Where this leaves the roster

Gemma-4 keeps the crown for single-shot quality, and its MTP route makes it quick too. The fan-out slots stay with qwopus3.6: 64.5 t/s decode, competitive prefill, thinking-off, six parallel slots - the shape agent orchestration wants. North Mini Code earned a real decision: statistically tied with the champion on MBPP+, fastest decode on the box, and not yet wired into production. Whether it displaces anything is exactly what the missing agentic benchmark should decide, and building that benchmark is now the top open thread.

So the roster verdict is provisional in a specific, stated way: single-turn Python is measured, the day job isn't yet. And the question a year into this series keeps getting less like "which model is best" and more like "best at which failure mode, at which speed, in which slot". Early days on that agentic benchmark tho...

MTP on halo: two production adoptions, one banned backend, one corrected claim

MTP - multi-token prediction, the speculative decoding where a model drafts its own future tokens with a built-in head instead of a separate draft model - has had a rough arc on this blog: a tuning win that required rejecting the community default, a measured quality cost on hard tasks, a cross-engine divergence confirmation at temperature 0. The running advice was: measure before you trust it.

Soo... we measured. A lot. The result is that halo now runs two MTP routes in production: Gemma-4-26B at +69% decode, and Qwen3.8-Flash-Next at a campaign-validated +51% mean. This post is how they earned their way in, including the backend that got banned and the claim of ours that did not survive its own validation campaign.

Gemma-4: the easy one

Unsloth publishes an official MTP draft sidecar for Gemma-4-26B-A4B: a 441 MB "smart Q4_0" drafter that llama.cpp has supported since June. Wire it in with --spec-draft-model and --spec-type draft-mtp, and on this box: 40.88 to 68.97 t/s on a short real prompt (+68.7%, 78.7% acceptance), 70.68 t/s sustained through a 2500-token generation at 82% acceptance, no collapse, no corruption.

Correctness got the treatment the quality-cost findings made mandatory: temp-0 comparison (not byte-identical, as expected by now - more below), then the generated code extracted and executed against real test cases. Passed.

One deployment decision worth copying: the MTP route is a separate, opt-in, single-slot route (gemma-4-26b-a4b-mtp), not a replacement for the production multi-slot route. Why? MTP gains invert under concurrency on this hardware - measured for the Qwen3.6 family, inherited here as the safe default rather than re-proven for Gemma-4 specifically. Batching already amortises what speculation buys; the H100 team found the same shape independently. Single-stream lever, single-stream route.

Flash-Next: the hard one

Qwen3.8-Flash-Next ships a native MTP head, but llama.cpp support for it still lives in open PRs (#27836 base support, unmerged as of this writing). The working path came from the community: a member (drluoto) assembled a branch combining the PR head with the loader and tensor-naming fixes needed to actually run it, and that branch became a custom ROCm toolbox here. Our base quant (UD-IQ3_XXS) paired with the Q8_0 sidecar was a combination nobody had published numbers for.

The first smoke test looked almost too good: 24.84 to 45.30 t/s, +82%, 97.2% acceptance. House rules say a single prompt adopts nothing, so it got a campaign: 5 real to-bench task prompts × 3 seeds, MTP-on vs MTP-off, paired. Result: +51.1% mean speedup (range +36.2% to +60.7%, stdev 8.2%), acceptance 87.4%, consistent across every pair. Then a third, independent cross-check on fresh real prompts that were never part of the campaign: +38.5% and +57.0%, landing inside the campaign's range. Three methods, one effect size. That's an adoption.

And a swap worth a paragraph: when Unsloth later published an official sidecar for this model, it got A/B'd against drluoto's on identical prompts. 40.55 vs 39.59 t/s, 90.5% vs 89.5% acceptance - essentially equivalent, slight edge to official. Adopted anyway, for the boring reason: a maintained official source beats a community rehost that could go stale. The community sidecar stays on disk as rollback. Community work got us here months early; official sources are what you settle on.

The backend that got banned

This box's default backend preference is Vulkan RADV for most things. For MTP it inverted, and the way it inverted is the cautionary tale.

On a short 143-token prompt, Vulkan MTP looked like a clean win: +60%, 94.4% acceptance. On a longer, harder 2000-token generation, throughput collapsed to a steady 6.3-6.6 t/s - not a crash, a collapse, sitting at a fifth of the no-speculation baseline (29.33 t/s) from early in the generation to the end. A short-prompt-only test would have shipped a severe regression into production with a green checkmark on it.

Honest scoping: the short and long tests differed in context size and prompt content, not just length, so the exact trigger wasn't isolated. But the operational conclusion needed no more precision: ROCm is the only backend validated safe for MTP on this hardware, through full-length generations. Benchmark depth is not optional. It's where this class of failure lives.

The claim we had to correct

That first Flash-Next smoke test also reported MTP-on and MTP-off producing byte-identical output at temperature 0. The validation campaign proved our own claim wrong: 0 of 15 pairs were byte-identical. Small floating-point-driven divergence appeared in every single pair; the smoke-test prompt just happened not to trigger it.

What was preserved is the thing that matters: functional correctness. Five pairs hit token-budget truncation in both arms and weren't comparable; of the comparable pairs, all but one had identical pytest pass/fail outcomes on the extracted code, and the one divergent pair scored better with MTP on (12/13 vs 10/13) - verified by full pytest diff to be a genuinely different-but-valid implementation, not corruption.

This slots cleanly into the two-mechanisms picture this blog ended up with: divergence is universal (the losslessness proof assumes exact arithmetic, hardware disagrees), and whether it costs anything is a per-model, per-head question only an executed benchmark can answer. Ornith's immature head cost 17-25% on hard tasks. Flash-Next's, on this evidence, costs nothing measurable. Both posts now cite each other, because readers of either need the other half.

Where this lands

I'd put the state of MTP on this hardware like this: it is neither the free speedup the leaderboards imply nor the quality hazard a divergence report implies. It is a per-model, per-backend, per-workload lever that pays +40-70% single-stream when the head is well-trained, inverts under concurrency, collapses on the wrong backend, and always changes outputs at the bit level without necessarily changing them at the level that matters. Every clause in that sentence was measured on this box, most of them twice.

The validation tax for the two adoptions was real: campaigns, cross-checks, temp-0 diffs, executed code. Also, I'd say, obviously worth it: the alternative was either leaving 50-70% of decode on the table or shipping a Vulkan collapse. Measure before you trust it still stands. It's just that now, twice, the measurement said yes...

Do synthetic LLM benchmarks lie? On halo, real code matched llama-bench to 0.6%

Every speed number this blog has published for halo rests on llama-bench, and llama-bench's prompts are synthetic: repeated tokens, not real text. A reader raised the reasonable question - do those figures reflect real usage at all, or are they an artifact of the corpus? Fair challenge. The answer to a fair challenge is a measurement, not a defence.

The cross-check: real prompt content (concatenated coding-task prompts, reference solutions, and a genuine coding question) at three sizes - 206, 3,192, and 6,421 real tokens - run three times each on Qwen3.8-Flash-Next UD-IQ3_XXS, on both ROCm and Vulkan backends, and compared against the synthetic llama-bench sweeps at the corresponding depths. Decode measured the same way: 250 real generated tokens on the same prompts.

The result, backend by backend

ROCm: essentially an exact match. The real 3,192-token prompt prefilled at 392.98 t/s. The synthetic pp4096 sweep: 390.78 t/s. That's 0.6% apart, within run-to-run noise. The 6,421-token real prompt (371.10 t/s) fits smoothly on the synthetic declining-with-depth curve. Decode: real-prompt numbers land 4.9% from the synthetic figure at comparable depth, with rep-to-rep spread of 0.05-0.21 t/s. Hard to ask for better.

Vulkan: a real 8.7% gap, disclosed rather than papered over. Real 3,192-token prefill measured 344.09 t/s against synthetic pp4096's 377.07. The plausible explanation: llama-bench's chunked-at-depth measurement and a fresh single-shot prompt of comparable length are not strictly identical measurements - attention cost accumulates differently - so the gap doesn't necessarily mean either number is wrong. But it's real, it's backend-specific, and if you quote Vulkan synthetic numbers, real prompts on this box run a bit slower than they suggest.

One number that looks alarming and isn't: the short 206-token real prompt measured far below the synthetic short-depth figure (182.63 vs 431-ish on ROCm). Not a discrepancy - llama-bench's "pp0" measures a 2,048-token chunk, and small-batch fixed overhead dominates a 206-token prompt. Not a like-for-like comparison, so don't read it as one.

Worth saying why this check mattered extra on this particular model: it's a MoE, and expert routing is content-dependent. Synthetic repeated tokens could, in principle, exercise a completely different expert-load pattern than real code. Measured verdict: they don't, at least not enough to matter on ROCm.

The bug the check itself caught

The first attempt at this cross-check produced garbage, and the way it did is the transferable lesson. Reps 2 and 3 of each prompt came back suspiciously fast - because llama-server's prompt cache silently recognised the repeated prompt and collapsed the prefill to a handful of tokens. The "benchmark" was measuring cache lookups. Caught by inspecting the per-rep prompt_n fields, fixed with cache_prompt:false, whole exercise redone.

That's the third measured-the-wrong-thing incident this series (the judge, the stale server, now the cache), and I'd say the pattern is now a rule: any live-serving benchmark must either disable prompt caching or vary the prompt per rep, and you verify which happened from the server's own counters, not from the wall clock looking plausible.

Where this leaves the numbers

So: synthetic benchmarks, on this box, on this model class, measure the real thing - on ROCm nearly exactly, on Vulkan with a known, quantified offset. The MTP campaign numbers got the same treatment (fresh real prompts, never used in the original campaign) and converged on the same effect size, which is what let those adoptions ship with confidence.

The general habit I'd recommend from all this: don't defend your benchmarks, cross-check them. It cost one afternoon, it answered the skeptic properly, and it found a methodology bug that would have quietly poisoned some future measurement anyway...

Cerebras vs our H100: 11x the decode speed, half the agent turns

Well.. here's a result that sounds wrong until you see the mechanism: Cerebras runs Qwen3.8-27B about eleven times faster than our H100, and still finishes fewer agent turns per minute than that H100. Both halves measured, same day, same harness, same model on both sides.

Let me be precise about what this is not saying. Cerebras is not slow. Single stream it decoded at 728-1142 tok/s in our tests against our box's ~87. The wafer-scale silicon delivers exactly what the marketing says, and if your workload is one stream of raw generation, it wins by a mile and this post does not apply to you. What follows is about a different question: what happens when you point a real agent harness at it.

The setup

Same model on both sides: Qwen3.8-27B, Cerebras' hosted qwen-3.8-27b versus RedHatAI's INT4 on the Trail Openers H100 in Helsinki. Both driven through the same OpenAI-compatible harness (pi-rukas, the agent framework we use daily): identical prompts, identical token budgets, reasoning_effort: "none" on both so neither burns thinking tokens, prefix caching warm on both, 429s retried with backoff rather than counted as failures. Six tests, about 22 minutes under load, 16.9.2026.

Agent work has a particular shape, and it decides everything here: the harness resends the whole conversation every turn. A coding agent sitting on a ~38k-token repo context sends those 38k tokens again and again, ~97% of them cache hits, to generate a few hundred new ones each time.

What the six tests said

Raw speed: Cerebras, decisively. 988 tok/s median decode vs 87. No contest, no dispute. (Funny detail: our time-to-first-token is better, 0.14s vs 0.33s, because the box is close and uncontended. They win the moment generation starts, not before it.)

Warm cached turns: us, actually. A 30k-token system prompt, four sequential turns, both sides caching. Their warm turns: 0.49-0.67s. Ours: 0.22-0.28s. For a short answer over a cached context - the single most common agent turn there is - our endpoint is already faster end-to-end.

Concurrency: we scale, they flatline. Simulated agent sessions over a ~38k context: Cerebras sits at 15-16 turns/min whether you run 1, 4, or 12 parallel sessions. Twelve sessions produced 60 rate-limit rejections and zero extra work. Our box went 9.2 → 27.5 → 45.6 → 62.7 turns/min at 1/4/12/24 sessions, zero rejections, still climbing when I stopped.

The obvious rebuttal, controlled for. Cerebras reserves input + max_completion_tokens up front, so surely the harness's 32,768-token default is self-inflicted? Re-ran the whole sweep at max_tokens: 1024, which is their own documented fix for spurious 429s. It changed nothing: 13.9-16.9 turns/min across the board.

The crossover: ~13k tokens of context. At ~10k context Cerebras wins clearly (62.8 vs 32.2 turns/min). At ~13k it's a dead heat. At ~38k we do 27.6 to their 12.4, and our box pushed 1.08M total tokens/min while theirs pinned at 460-670k no matter how the work was shaped.

What it feels like in practice: blocked. With the harness honouring their real Retry-After: 59, four sessions at 38.6k context spent 714 of 768 session-seconds waiting on backoff. 93% of wall-clock, blocked. The "it feels slow" instinct was correct: it was watching a full-minute stall, several times per session.

The mechanism, because it predicts every number above

Cerebras enforces two per-minute token buckets: an uncached one (150,000 tok/min on this account) and a total one at three times that, ~450k. Cache hits are exempt from the uncached bucket. They are not exempt from the total bucket. So a 38k-token agent context costs 38k against the quota every single turn, even when 97% of it is a cache hit and the model generates 300 new tokens. Our measured uncached burn was 19,248 tok/min, 13% of the allowance - we were nowhere near the advertised limit and throttled anyway.

Which gives you a rule simple enough to say out loud: turns per minute ≈ the total-token ceiling divided by your context size. Below ~13k of context the fast silicon wins. Above it, the ceiling does the deciding, and the world's fastest inference sits idle waiting for the next minute to start.

What I'd actually claim

So I'd put it this way: this is not a hardware story, it's a quota-design story. Their limiter prices a cached token like a computed one, and agents are the workload that maximises exactly that ratio. The unarguable version of the finding: dedicated capacity beats metered capacity for sustained agentic work. Our H100 has no rate limits because it's ours; that, not the silicon, is what won.

The caveats, before someone else states them: this is one account on one tier (450 RPM / 150k uncached TPM, measured 16.9.2026), not "Cerebras" - enterprise and dedicated endpoints get provisioned differently. Theirs is a shared multi-tenant service measured at one point in time. No cost claim here: a metered API against a box rented 24/7 is not a like-for-like bill. And no quality claim: they serve unquantised weights, we serve INT4; throughput compared, output quality not.

The part I find genuinely interesting: every provider racing to the top of the tok/s leaderboards is optimising the number that matters least for agents. An agent's real currency is completed turns per minute at its actual context size, and that's decided by quota shape, cache pricing, and backoff behaviour - the things no leaderboard shows. Speed-of-light decode with a context-priced meter in front of it is a sports car with a toll booth every 38,000 tokens...