Free playbooks in your inbox

How Long Does It Take to Switch From Claude Code to Codex?

Try Codex, a cheaper model or a local one whenever you like, because you already know what the move costs you in hours, layer by layer. My whole setup came to 3 hours to move, and the expensive row was the run state I had never saved.

From the youcanbuildthings catalog ▸ Build-tested

Hello builders,

You can know by tonight whether trying Codex, or any newer model, costs you an afternoon or a rebuild, and make that call from a number you can defend. When I priced my own setup, moving everything that existed came to 3 hours, and the only expensive row was the run state I had never saved. Here is how to price a switch from Claude Code to Codex one layer at a time, with a command or a checkable fact behind every hour.

The strongest question on this got 252 upvotes and 250 comments: “Which agent harness do you use and why? … has anyone switched from these?” Nobody produced a table. They couldn’t, because the answer depends on your machine and theirs.

And their picks don’t transfer to you anyway. One public benchmark ran ten coding agents on two models, and its README reports that swapping the agent moved pass@1 from 23.2% to 52.4% on one model, “and the ranking does not transfer between them (Spearman -0.05).” So when we read a recommendation in a thread, we are reading about somebody else’s model and somebody else’s repo.

Five inputs, each one checkable

We price six layers: observation, context, control, action, state and verification. Every hour we write down has to come from one of these five inputs, or it’s a mood.

1. The licence. Open the LICENSE file, not the README badge. If it is unclear, your exit starts with a lawyer, so add hours.

2. Does the model swap by config? Codex has a provider table built in. The config reference lists base_url, wire_api (“responses is the only supported value”) and env_key, the “Environment variable supplying the provider API key”:

model_provider = "my-gateway"

[model_providers.my-gateway]
name = "My Gateway"
base_url = "https://gateway.example.internal/v1"
wire_api = "responses"
env_key  = "MY_GATEWAY_KEY"

The gateway URL is a placeholder for yours. The docs add that “Built-in provider IDs (openai, ollama, and lmstudio) are reserved and cannot be overridden.” We can test the block without touching our real config. Put it in config.toml inside an empty directory and point CODEX_HOME at that directory for one command: CODEX_HOME=/tmp/codex-test codex --strict-config exec --skip-git-repo-check "hi" < /dev/null. Leave off the < /dev/null and it sits waiting on standard input. On Codex 0.152.0 ours stopped on Missing environment variable: MY_GATEWAY_KEY, which means the table parsed. Setting wire_api = "chat" fails at load, and so does a block named openai. If your agent has a table like this, the action row is one to three hours. If it does not, you are writing a translating proxy and it is days.

3. Does your stop transfer? Good news, and the only structural portability win I found. A PreToolUse hook that exits 2 blocks the call in Claude Code, and the Codex hooks page says the same thing in its own words: “You can also use exit code 2 and write the blocking reason to stderr.” Moving mine was rewriting six lines of JSON as TOML.

4. Does your event set transfer? Bad news, the exact opposite. Here is the free trick for finding what Claude Code accepts. Put a fake event in .claude/settings.json:

{ "hooks": { "BogusEventName": [] } }

Then run claude doctor. On 2.1.282 it rejects the fake name and prints every valid event, 33 of them, with no model call and no API key. The Codex hooks page lists 12. If you built anything on PreModelSwitch, TaskCompleted or FileChanged, there’s nothing on the other side, so put four hours against each one you actually use.

5. Where your state lives. If run history exists only in the conversation, there’s nothing to port. That reads like a zero and it is the most expensive row, because you redo every long job from the start. Price it as your longest job times the number of times you’d redo it.

What mine came to

Switching cost from Claude Code to Codex, in hours per layer: observation 0, context 0, control 1, action 2, verification 0, for 3 hours to move what exists; state 6 hours per long job, excluded from the 3-hour total. The dread was real and the number was not

Here is how we price a real one. This is a Claude Code setup I use most days, one small Python repo, one hook and about 90 lines of CLAUDE.md, priced against Codex:

  • Observation, 0 hours: both have file and search tools (config swap)
  • Context, 0 hours: I wasn’t managing it anywhere, so nothing to move (where state lives)
  • Control, 1 hour: rewrite one PreToolUse hook from JSON into TOML (stop transfers)
  • Action, 2 hours: re-point the model, and the provider table exists (config swap)
  • Verification, 0 hours: nothing to port, because nothing was there (where state lives)
  • State, 6 hours per long job: nothing to port, so every long job restarts from zero

3 hours to move what exists, with the state row excluded from that total. On the long job I actually care about it is nine, and the six repeats.

That surprised me, because I’d spent a year assuming the number was huge, and the dread was real and the number was not. Everything I thought was welded on was either portable by design, like the stop, or never built at all, like state and verification.

When should that number make you stay? If the total comes back under about four hours, there is nothing to defend and we can go back to work. If you are on a frontier model and your jobs are landing, the setup around it changes your results least. And if your current agent already talks to your tracker and three internal services, that’s worth more than any pass rate.

Price your own six rows tonight, write the input next to every hour, and name the agent you’d actually move to. Then the next time someone tells you to switch, you answer with a total.

Now go build something this weekend!

John Cook

Why trust this? Every youcanbuildthings guide is pulled from a build-tested book: code that ran in production before it was written down.