How to Resume a Long Claude Code Job Without Starting Over
Hand Claude Code or Codex a long job and let it crash, because a fresh session picks up at the next unfinished step on a 122-byte brief. One Codex run spent 7.13 million tokens asking its workers the same question 47 times, because nothing wrote the answer down.
>This covers the state file that lets a job resume. Your Agent Is More Than the Model goes further: a token budget with an eviction rule, a claim checker, and a stop that fires at the tool boundary, all in the same swap/ directory.

Your Agent Is More Than the Model
The 7-Part Agent Harness for Swapping Models, Stretching Your Token Budget, and Shipping Working Code
Hello builders,
You can hand an agent a job with forty steps, let the session die at step twenty-two, and have a fresh one start at step twenty-three without being told anything. When I killed a four-step job halfway and started a brand-new session, it read a 122-byte brief and went straight to the next step, on half the turns. Here’s how to resume a long Claude Code job with a state file that lives outside the conversation.
47 checks, nothing new
The best evidence for this came from somebody who pulled his own run’s ledger and went through it line by line. His five-hour Codex allowance went from 53% to 100% in about 33 minutes. The orchestrator was waking every thirty seconds to ask two workers whether they were done.

All 47 checks came back with nothing. They cost 7,130,181 parent input tokens, 68% of the session, and 23 minutes 30 seconds of the 33-minute run went to timeout waits. One correction almost every retelling gets wrong: 99.8% of those were cached reads. Say quota, not burned. What he actually lost was a five-hour allowance, in about half an hour.
The part that got me is further down his post. The parent interrupted its first worker twice, then went and looked, and found that worker had already written 2 files, 127 insertions. Real work, done, and the coordinator had no record of it.
What do we actually hand the model at step 40? Only what’s in the window. Every turn re-sends the whole conversation, so step 40 costs what steps 1 through 40 weigh, and anything that fell out of the window never happened.
Anthropic landed on the same answer
The Claude team’s own write-up, Effective harnesses for long-running agents, has its first session set up three things:
an init.sh script, a claude-progress.txt file that keeps a log of what agents have done, and an initial git commit that shows what files were added.
And they say plainly, “compaction isn’t sufficient.” One detail is worth stealing outright: they keep the feature list in JSON because “the model is less likely to inappropriately change or overwrite JSON files compared to Markdown files.”
The store is two commands
We keep the plan in swap/store/plan.json, { "steps": ["a", "b", "c", "d"] }, and the history in an append-only steps.jsonl next to it. The model only ever sees brief, so that’s the one output we keep small.
Here’s the whole store:
#!/usr/bin/env python3
"""swap/store/store.py: run state on disk instead of in the conversation.
record <step> <status> [note] append one finished step
brief fixed-size digest; the only thing the model reads
"""
import json, pathlib, sys, time
HERE = pathlib.Path(__file__).resolve().parent
LEDGER, PLAN = HERE / "steps.jsonl", HERE / "plan.json"
def steps():
if not LEDGER.exists():
return []
return [json.loads(l) for l in LEDGER.read_text().splitlines() if l.strip()]
cmd, args = (sys.argv[1] if len(sys.argv) > 1 else ""), sys.argv[2:]
if cmd == "record" and len(args) >= 2:
note = " ".join(args[2:])[:200] # a note is not a log
with LEDGER.open("a") as f:
f.write(json.dumps({"ts": int(time.time()), "step": args[0],
"status": args[1], "note": note}) + "\n")
print(f"recorded {args[0]}={args[1]}")
elif cmd == "brief":
plan = json.loads(PLAN.read_text())["steps"]
seen = {s["step"]: s["status"] for s in steps()}
done = [p for p in plan if seen.get(p) == "done"] # walk the plan, not the ledger
todo = [p for p in plan if p not in done]
print("PLAN:", ", ".join(plan))
print("DONE:", ", ".join(done) or "(nothing yet)")
print("NEXT:", todo[0] if todo else "(all steps complete - stop)")
for s in steps()[-1:]:
print("LAST:", s["step"], s["status"], s["note"][:80])
else:
sys.exit("usage: store.py record <step> <status> [note] | brief")
That comment on done is the bug I shipped first. I built the DONE line from the ledger, so the digest grew every time the agent did anything, which is the conversation again with extra steps. Walk the plan and it can’t grow.
We prove it before we spend a token. Append 200 junk records and measure both files. When I ran it the ledger went 73, 1,306, 6,786, 27,336 bytes, 375x, and the brief read 58 bytes at the first record and sat at 122 from record 10 on. Yours will differ by a few bytes and it should be just as flat.
Then we put the protocol in the job itself, because the agent can’t read our minds.
Here’s mine:
You are resuming a job that may already be partly done. Do NOT assume you are
starting fresh.
First, always, run: python3 swap/store/store.py brief
Then loop: do the ONE step its NEXT line names, record it, and run brief again.
Keep going until NEXT says all steps are complete, then stop and report.
After finishing step X, run:
python3 swap/store/store.py record X done "<one line about what you produced>"
Write it as a single loop. My first draft said “do one step and stop” and also “keep going”, and the agent obeyed the stop.
Kill it halfway and watch
So what happens when a session dies halfway? A cold run of four steps took 14 turns and 347,383 input tokens. Then I deleted the session, left a.txt and b.txt on disk with their two ledger rows, and started a new one with the same prompt. It ran brief, read NEXT: c, and finished in 8 turns on 192,466. It didn’t reread a.txt and it didn’t ask me what happened. I re-ran that resume on Claude Code 2.1.282 before writing this and got 8 turns again.
The store is just a command the agent shells out to, so the same script runs unchanged under codex exec. What goes in it? Anything expensive to discover again: a fact that cost a tool call, a decision and its reason, a pointer to a big file. Never the file itself.
Our test for whether it works is simple. Take the longest job you’ve restarted from zero this month. Write its steps into plan.json, kill it at the halfway mark on purpose, and start it again.
Now go build something this weekend!
John Cook