How to Get Claude Code to Finish the Job on the Cheaper Model
Get haiku finishing jobs you were paying sonnet to attempt, and quote what one finished job costs with the failures counted in. My first five runs finished nothing on either model; eleven words of job text later both went six for six.
>This covers pricing a finished job. Your Agent Is More Than the Model goes further: the gate that decides what finished means, a routing rule your setup reads, and a spend cap that stops a run at the tool boundary.

Your Agent Is More Than the Model
The 7-Part Agent Harness for Swapping Models, Stretching Your Token Budget, and Shipping Working Code
Hello builders,
You can get Claude Code finishing real jobs on the cheaper model, and put a price on each finished one that you would be comfortable quoting to a client. My first five runs across haiku and sonnet delivered nothing and cost 34 cents, and after I changed eleven words of the job text, both models went six for six. Here is how to work out your real Claude Code cost per task, with the runs that failed counted in.
A Pro subscriber wrote the post this answers, and it got 166 points: “took 15 minutes to burn thru the 5 hour and almost 30 minutes more to burn the entire weekly usage. I reset my limits twice and still I couldn’t get done with the task, I think this is absurd or I’m doing something wrong.” Seventy-seven replies. Nobody told him which one it was, because every number he had was about tokens and his problem was about a job.
Price the finished job
What does a failed run cost you? Whatever it spent, added to the price of the next run that works. So we divide everything spent on a model by the runs that actually passed a check, and a script decides what passed, because our eyes wave through things a script catches.
We log one line per run. If you’re using a runner like the one in my two-model swap test, $P is the check script’s exit code and the cost and turns come straight out of --output-format json.
python3 -c "import json,sys; d=json.load(open(sys.argv[2])); print(sys.argv[1], 'pass' if sys.argv[3] == '0' else 'fail', d['total_cost_usd'], d['num_turns'], sep='\t')" "$M" "raw-$M.json" "$P" >> runs.tsv
Then the pricer:
#!/usr/bin/env python3
"""price.py [runs.tsv] — cost per FINISHED job, failed runs folded in.
runs.tsv: one line per run, tab-separated: alias, pass|fail, cost, turns."""
import sys
from collections import defaultdict
rows = defaultdict(list)
for line in open(sys.argv[1] if len(sys.argv) > 1 else "runs.tsv"):
if line.strip():
alias, verdict, cost, _turns = line.rstrip("\n").split("\t")
rows[alias].append((verdict, float(cost)))
print(f"{'alias':<8} {'runs':>4} {'pass':>4} {'spent':>8} per finished job")
for alias, rs in rows.items():
passed = sum(1 for v, _ in rs if v == "pass")
spent = sum(c for _, c in rs)
per = f"${spent / passed:.4f}" if passed else "never finished"
print(f"{alias:<8} {len(rs):>4} {passed:>4} {spent:>8.4f} {per}")
There are no per-token rates in it on purpose. The cost comes back with every run, so nothing goes stale. And when a model’s line says “never finished”, that’s the most useful answer the script can give us, because it means we have been paying for attempts.
The table that changed my mind
I ran one small job, adding a --json flag to a line-count script, eleven times across two models on the same afternoon. The first batch used the job the way most of us would type it: “Add a —json flag to linecount.py so that it prints a JSON object with the file path and the line count. Then stop.” Can you see what’s wrong with it? I couldn’t, and it’s the sort of thing I type every day.

Haiku went 3 runs with 0 passed and $0.1212 spent, and sonnet went 2 runs with 0 passed and $0.2153 spent, so we had 5 runs, $0.34 and nothing delivered.
And every one of those failed runs would have passed a human glance. The agents were confident, and each one shipped a working --json flag. The checker failed them because “the file path” came back as a key called path when the contract said file, and because no test appeared. Anything downstream reading ["file"] gets a KeyError.
Then I changed the job text and nothing else. I named the exact keys and asked for the test at tests/test_json.py, same models, same fixtures, same hour. Haiku went 3 for 3, $0.2053 spent, $0.0684 per finished job. Sonnet went 3 for 3, $0.3949, $0.1316.
Look at what we actually changed between those two batches, because the model was never the variable. The pass rate went from 0% to 100% on both of them. Sonnet came out about 1.9 times haiku per finished job, partly because it used 7 turns against haiku’s 10 to 14. We would never have guessed that ratio from a price list, and the only way to get it is to run your own jobs through the pricer.
Why doesn’t per-token pricing catch any of this? One public benchmark of ten coding agents says it outright in its coding README: “Ranking by tokens and ranking by money are not the same, because input and output are priced ~3-5x apart and agent loops are overwhelmingly input: the measured split runs from 15:1 to 250:1, so 92-99% of what a harness pays for is conversation it already sent.” A cheap-looking run tells you almost nothing about what finishing costs.
The rule I use now
Start every job on the cheapest model that has ever passed your check, and only move up when it fails. When a job fails twice on two different models, stop switching models and rewrite the job. My thirty-four cents went on a missing key name, and we fixed it with one sentence of job text.
Log your next ten runs to runs.tsv this week and run python3 price.py runs.tsv on Friday. If any line says never finished, open the job text before you open the pricing page.
Now go build something this weekend!
John Cook