pi-delegate: Claude does the thinking, a cheaper model does the typing
Soo... pi-delegate is out. It's a Claude Code plugin with one skill: you say "delegate to pi: ..." and the heavy part of a coding task - reading files, editing, running tests - runs on pi, an open-source coding agent that can drive any model you pick, including one on your own machine. Claude only writes the brief and reads a short result. Your tests decide whether it worked.
The headline from my benchmark: on multi-file tasks, Claude's cost per run dropped from $0.082 to $0.055, a third off, with every hidden test still passing. But the more interesting part is that the first version of this tool made Claude more expensive, and the fix was mostly deleting things.
Version one cost more, not less
The first design was the one I'd have defended in an architecture review. A deterministic bash driver ran a develop, review, fix loop: pi wrote the code, a second read-only pi instance reviewed it adversarially, findings went back for fixing, up to six pi calls. Cheap model builds, strong model judges.
Then I benchmarked it against plain Claude Code on real historical fixes from the click library. Quality was at par, every run passed the grading tests. Cost was not: the delegated runs used 1.7x and 2.7x as much Claude as just letting Claude do the work. Where did it go?
- the skill text and role prompts loaded into Claude's context, roughly doubling cache creation;
- extra turns to launch the loop, wait for it and read back the summary;
- and the big one: after the loop's own review had passed, Claude re-read the diff and re-ran the tests anyway, to be sure.
That last one is funny in a slightly painful way. I had built a whole review pipeline so Claude wouldn't have to check the work, and Claude checked the work anyway.
What changed
Three things, all subtractions.
The reviewer is gone. Field data from about 15 real delegated tasks showed a reviewer on the same model as the developer approved almost everything, while a stronger model later found real bugs in several of those "approved" changes. A same-model reviewer is a second opinion from the same mind. So now there's no model in the quality gate at all: you pass --verify "<your tests>", the script runs it after pi finishes, and if it fails pi gets exactly one retry with the failure output. A deterministic gate instead of a second opinion. Same lesson I keep landing on in other places: put the decision in code where you can.
Claude is told not to redo the work. When the gate passes, the skill tells Claude to reply in a sentence or two and stop. No re-reading the diff, no re-running the tests. That was the single biggest leak.
The output got trimmed. The script prints at most about 1.2 KB of pi's final text, because everything pi says gets re-read, and paid for, by Claude. The skill itself shrank to about 25 lines, with all the launch, wait and abort logic moved into run.sh.
The numbers now
Eight small tasks, two runs each, plain Claude versus Claude with pi-delegate, pi running Qwen3.8-27B on the Trail Openers H100:
| Claude cost per run | Plain Claude | With pi-delegate |
|---|---|---|
| Multi-file features | $0.082 | $0.055 (−33%) |
| Tiny edits (ten-line fixes) | $0.042 | $0.047 (+11%) |
| Runs passing the hidden checks | 28 / 28 | 28 / 28 |
The tiny-edit row is the useful one. Delegating has a fixed overhead of roughly half a cent to a cent per task: the skill text, one extra tool round-trip, reading the result. On a ten-line fix that overhead is bigger than what pi saves, so don't delegate those. On multi-file work it pays for itself, with the best task at −47%.
The honest fine print
This counts Claude's cost only. pi's own spend comes on top: nothing extra if you already run your own model like we do, cents on a hosted cheap model. Delegated runs are also slower, several times over in these runs: about 25 seconds plain versus two to three minutes. Two runs per task is a small sample, and these are small synthetic tasks, not a big real codebase. Delegated diffs also came out 1.3-1.8x larger, mostly because pi writes more tests - which you may or may not want.
There are guardrails too: it refuses to run on your default branch or next to .env and key files, and disables git push for pi. That guards against mistakes, not a malicious model; for real isolation use a disposable clone, a container, or the optional sandbox wrapper.
So who is it for? I'd say: you use Claude Code, your usage limit or token bill matters, your tasks touch several files, and you have a test command that can tell right from wrong. If speed matters more than cost, or there's no way to check the result, Claude alone is still the better tool.
The benchmark takes about five minutes to rerun yourself with bench/quick.sh, and I'd genuinely like to see numbers from other setups and other models. Mine is one Claude model, one pi model on one H100 and a handful of tasks... early days tho.