How to Know If Sonnet Is Worth 3x Haiku on Your Code
Pick the Claude model you pay for on a real job from your own repo, and have an answer the same afternoon a new model ships. Both models passed my test job and one cost three times as much, and the same model re-run moved 15% with nothing changed.
>This covers the swap test on one job. Your Agent Is More Than the Model goes further: an adapter that points the same job at Codex or a local model, and an eval set of ten of your own jobs you run the day a model ships.

Your Agent Is More Than the Model
The 7-Part Agent Harness for Swapping Models, Stretching Your Token Budget, and Shipping Working Code
Hello builders,
You can know which Claude model is worth paying for on your own code by the end of today, and answer the same question the afternoon the next model ships. I ran one real job on haiku and sonnet, both passed, and sonnet cost three times as much. Here is how to compare models in Claude Code with one job, one check script and one flag that changes.
The best question I found on this got 55 upvotes and zero answers. It sat under a 620-point thread about a new model’s prompting guide: “If we optimize our repos for astra, what happens when you use sol, fable or opus in that repo?” Forty-seven comments, nobody ran it. We are going to run it on your repo.
Pick a job a script can judge
The Claude Code docs give you the experiment without meaning to. From the model configuration page:
The —model flag and ANTHROPIC_MODEL environment variable apply only to the session you launch with them. To run different models in different terminals at the same time, launch each one with its own —model flag rather than switching with /model.
Two terminals, two flags, one job. The aliases on that page are the ones you’d use: haiku for “simple tasks”, sonnet “for daily coding tasks”, opus “for complex reasoning tasks”.
The job has to be something a program can decide. “Make the code better” can’t be scored. Mine is tiny on purpose: add a --json flag to a small linecount.py, print exactly the keys file and lines, and add a test at tests/test_json.py. Put the starting files in fixtures/ and the prompt in job.md, and we never touch either again.
Then write the checker before either model runs, so you cannot grade in favour of the answer you expected:
#!/usr/bin/env bash
# check.sh <dir> — exit 0 only if the job is done. 1 = wrong JSON, 2 = tests.
set -uo pipefail
cd "${1:?usage: check.sh <dir>}" || exit 3
python3 - <<'PY' || exit 1
import json, subprocess, sys
r = subprocess.run([sys.executable, "linecount.py", "--json", "sample.txt"],
capture_output=True, text=True)
d = json.loads(r.stdout)
assert set(d) == {"file", "lines"} and d["lines"] == 3, d
PY
[ -f tests/test_json.py ] || exit 2
python3 -m unittest discover -s tests >/dev/null 2>&1 || exit 2
Run ./check.sh fixtures first. It has to fail, because the flag doesn’t exist yet. Mine exits 1. A checker that passes before any work happens isn’t checking anything.
The runner holds everything still
Every run gets a fresh copy of the fixtures, the same prompt, the same timeout and the same checker. Only --model moves:
#!/usr/bin/env bash
# run.sh <model> — run the fixed job once on one model, print the numbers.
set -uo pipefail
M="${1:?usage: run.sh <model>}"
rm -rf work && cp -R fixtures work
S=$(date +%s)
(cd work && timeout 600 claude -p --model "$M" --output-format json \
--permission-mode acceptEdits "$(cat ../job.md)" < /dev/null) > "raw-$M.json"
W=$(( $(date +%s) - S ))
./check.sh work >/dev/null 2>&1; P=$?
python3 - "$M" "$W" "$P" "raw-$M.json" <<'PY'
import json, sys
m, w, p, f = sys.argv[1:]
d = json.load(open(f)); u = d["usage"]
tin = (u["input_tokens"] + u["cache_read_input_tokens"]
+ u["cache_creation_input_tokens"])
print(f"{m} check={p} in={tin} out={u['output_tokens']} "
f"turns={d['num_turns']} {w}s ${d['total_cost_usd']:.4f}")
PY
The < /dev/null matters: without it print mode waits on standard input and the script hangs. timeout ships with GNU coreutils, so on a Mac that’s brew install coreutils. And we add all three input fields together, because on my runs the plain input_tokens read about 90 while the real total was over a quarter of a million.
What mine printed

./run.sh haiku then ./run.sh sonnet. Haiku passed on 254,862 input tokens, 2,754 output, 12 turns, 34 seconds, $0.0501. Sonnet passed on 290,456 in, 1,982 out, 9 turns, 31 seconds, $0.1510. Same result, a third of the price, three seconds slower. I couldn’t tell you which diff came from which model without the sheet.
Then I did the thing nobody does. I ran haiku again, same job, nothing changed. It came back at 309,085 in, 2,989 out, 14 turns, 43 seconds, $0.0576. That is 21% more tokens and 15% more money, and nothing changed but the roll of the dice.
So anything inside about 15% is not a result. My 3x gap is real. A 1.2x gap would be noise, and I’d have picked a winner off it.
Is that noise band stable? No. I ran the same job on Claude Code 2.1.282 and haiku passed again at 423,198 input tokens and $0.0938. Same job, same fixtures. Two runs tell you which way to lean, not what’s true, which is why I would run each model three times before switching anything I pay for.
Keep the numbers
Append every run to a file, ./run.sh haiku >> results.md, not to a scrollback you will close. The day a new model ships you add one line to that file, and you have got an opinion before lunch while everyone else is still reading launch posts.
Pick one job you gave Claude Code this week that a script can judge. Write check.sh, watch it fail, then run it on two models twice each.
Now go build something this weekend!
John Cook