How to use Claude Code and Codex together
A practical workflow for putting Claude Code and Codex on the same development objective: when a second model helps, what each agent should be told, how worktrees keep them apart, and how one model reviews another's change.
Updated · Winskel
The short answer
Yes, you can use Claude Code and Codex on the same project. Both are command-line coding agents that read and change files in a local Git repository, run commands, and report what they did. Nothing stops you from running both against one repository.
Running them is easy. Coordinating them is the work: deciding which agent does which part, telling each one what it needs to know, keeping them out of each other's files, and knowing whether the combined result is correct. The rest of this guide covers that coordination, using one bug fix from start to finish.
- Keep tightly coupled work with one agent. Split only when parts are independent or need a second opinion.
- Hand off facts, not reasoning: objective, findings, relevant files, constraints, changes, and verification results.
- Give every agent its own Git worktree and branch.
- Use a different model for review than the one that wrote the change, in a fresh session.
- Run the checks yourself, or record exactly which commands ran and what they returned.
Why use more than one coding model
- Independent review. A model reviewing its own change tends to repeat its own assumptions. A different model, starting from the diff rather than the author's conversation, is more likely to catch what the author missed.
- Parallel work. Two parts of a task that touch different files can be built at the same time by two agents.
- Fit. Some work needs deep reasoning and some is routine. Using a lighter model for routine work and a stronger one for hard problems spends your plan limits where they matter.
- Availability. Claude Code and Codex run under separate Anthropic and OpenAI plans. When one is unavailable or out of quota, the other can continue.
When one model is enough
Every handoff costs something. The receiving agent has to rebuild an understanding of the code that the previous agent already had, and anything left out of the handoff is lost. For most small and medium changes, one agent that investigates, fixes, tests, and repairs its own failures is faster and more reliable than a relay.
- A small, well-understood change
- Diagnose, fix, and test steps on the same files
- Work where the second model would only repeat checks that already pass
- A review of low-risk changes that tests already cover
Example: an authentication bug
Users report being signed out at random. The pattern: it happens when the app is open in two tabs. The suspicion is the token refresh path. This is a good candidate for two models because authentication is high risk: a fix that looks right can still open a security hole, so an independent review is worth its cost.
The manual workflow below uses four handoffs so each boundary is visible. Claude Code investigates, Codex implements the fix, Claude Code reviews the change, Codex repairs what the review finds, and verification runs before the work is called done. The repository, file names, commits, and test counts are illustrative.
Set up separate working copies
Create a branch and a worktree for the fix, so the implementing agent never works in your main checkout. A worktree is a second working directory attached to the same repository.
git switch main && git pull
git worktree add ../acme-refresh-fix -b fix/refresh-race mainStep 1. Claude Code investigates
Start Claude Code in the main checkout and ask it to find the cause without changing code. Its output is not a fix; it is a finding the next agent can act on. Ask it to end with a handoff note in a fixed shape.
Objective: Users are signed out when two tabs refresh the session at the same time.
Findings: refreshSession() in src/auth/refresh.ts reads the stored refresh token, rotates it, then writes the new one. Two concurrent refreshes both read the same token. The second rotation presents a token that was already used, reuse detection fires, and the whole session is revoked.
Relevant files: src/auth/refresh.ts, src/auth/session-store.ts, test/auth/refresh.test.ts
Constraints: Keep refresh token rotation and reuse detection. No database schema change. Do not change the public API of session-store.ts.
Verification so far: Reproduced with two concurrent refreshSession() calls in a scratch test. No code changed.Step 2. Codex implements the fix
Start Codex inside ../acme-refresh-fix and give it handoff 1 as the task. Ask it to add a regression test for the concurrent case, run the auth tests, commit on fix/refresh-race, and report what it changed and what it ran.
Objective: (unchanged)
Constraints: (unchanged)
Changes: Commit 4e1c2a9 on fix/refresh-race. src/auth/refresh.ts serializes refreshes per session with an in-process lock keyed by session id. New test: test/auth/refresh.test.ts "concurrent refresh keeps the session".
Verification: npm test -- auth: 41 passed, 0 failed. npm run typecheck: passed.Step 3. Claude Code reviews the change
Open a new Claude Code session so the reviewer does not carry the investigation conversation. Give it the objective, the constraints, and the diff, and ask for specific problems with file and line. The reviewer reads; it does not edit.
git diff main...fix/refresh-raceFinding 1 (correctness, blocking): src/auth/refresh.ts:38. The lock lives in process memory. Production runs several app instances behind a load balancer, so two tabs that reach different instances still race and still trigger reuse detection.
Finding 2 (minor): the new test only exercises a single process, so it cannot catch finding 1.
Suggested direction: make the rotation itself atomic in the session store, for example a compare-and-set on the current token, instead of locking around it.Step 4. Codex repairs
Return to Codex in the same worktree with handoff 3. It already has the branch and its own change in front of it, so it does not need the investigation again. It replaces the lock with an atomic swap of the stored token and changes the test to simulate two instances sharing one store.
Step 5. Verify before calling it done
Run the checks yourself in the fix worktree, or keep the exact commands and results the agent ran. An agent writing that tests pass is a claim, not evidence.
cd ../acme-refresh-fix
npm test -- auth
npm run typecheck
npm run lintThen read the final diff yourself, merge fix/refresh-race the way your team normally does, and remove the worktree with git worktree remove ../acme-refresh-fix.
What context crosses each boundary
Each agent starts with none of the previous agent's memory. A handoff note replaces that memory with the parts that matter, in a form the next agent can check against the code.
- Objective: the outcome, in one or two sentences, unchanged from start to finish.
- Findings: what was established, with the evidence that established it.
- Relevant files: where to look first, so the next agent does not re-survey the repository.
- Constraints: what must not change, carried forward at every step.
- Changes: the branch, the commit, and a short description of what changed. The code itself travels through Git, not through the note.
- Verification results: the commands that ran and what they returned, including anything that was not run.
Objective:
Findings:
Relevant files:
Constraints:
Changes: (branch, commit, what changed)
Verification: (command: result, and what was not run)
Open problems:When work should run in parallel
Run two agents at the same time only when their work does not depend on each other and does not touch the same files. In the example, a second agent could update the sign-in help text while Codex fixes the refresh path: different files, no shared interface.
- Parallel: disjoint files, a stable interface between the parts, and tests each agent can run on its own.
- Sequential: one part needs the other's result, both touch the same files, or the interface is still being decided.
- Dependent work starts from the commit it depends on, not from main, so it sees the change it builds on.
When one model should keep ownership
The example splits investigation from implementation to show the boundary. In practice, the agent that found the race is usually the best one to fix it: it has already read the code and reproduced the failure. Keeping diagnose, fix, and test with one owner avoids a handoff that loses detail.
The repair in step 4 follows the same rule. Codex wrote the change, so Codex repairs it, in the same worktree, with the review findings as new input. What should not stay with the author is the review.
How one model reviews another
- Use a different model family from the author, in a new session.
- Give it the objective, the constraints, and the diff. Do not give it the author's conversation.
- Ask for findings with a file, a line, a severity, and a reason. Ask it to say plainly when it finds nothing blocking.
- Keep the reviewer read-only. Send its findings back to the author to fix.
- Save review for changes where a defect is expensive: authentication, payments, data migrations, concurrency, security boundaries.
Branches and worktrees keep agents apart
Two agents in one working directory will edit the same files, pick up each other's half-finished changes, and run tests against a mix of both. git worktree gives each agent its own directory and branch while sharing one repository, so commits made in one are visible to the others.
git worktree add ../acme-refresh-fix -b fix/refresh-race main
git worktree add ../acme-help-text -b docs/sign-in-help main
git worktree list
git worktree remove ../acme-help-textA worktree separates Git state. It is not a security sandbox: an agent running under your user account can still read other files that account can read.
What the manual workflow costs
The example worked, and every step in it was coordination you did by hand.
- Deciding to split the work, and where
- Choosing a model for each part
- Creating and cleaning up worktrees and branches
- Writing a handoff note at every boundary
- Starting a fresh session for the reviewer
- Tracking which part waits for which
- Collecting the changes and checks from several sessions into one picture
Where Winskel fits
Winskel is a multi-model coding orchestrator for macOS that coordinates coding work across Claude Code and Codex. It does the coordination listed above. It does not replace either tool: it runs the Claude Code and Codex installs already on your Mac, under your accounts.
What Winskel does today with an objective like the example:
- Plans the objective with Claude Code and decides whether the work stays with one model or splits. Diagnose, fix, and test usually stay together.
- Chooses a model for each part, only among the models your accounts can run. See supported models and model routing.
- Can add an independent review by the other model family for high-risk work, or when you ask for one in the objective.
- Gives each run its own Git worktree and branch. Dependent work starts from the commits it depends on. See projects and worktrees.
- Writes the handoff for you: the goal, constraints, what the earlier step reported, its changes and commit, and any open problems. Agents are told not to include private reasoning.
- Can add a fix step when a review finds a real problem, and can retry a failed step once on a stronger model.
- Shows every part, model, branch, diff, and observed check in one window. A check counts only when Winskel saw the command and its result. See review and verification.
Holding dependent work until the work it relies on has passed its checks is still being qualified in the Alpha. Today a dependent step starts when the step it depends on finishes, and Review shows which checks were observed.
Winskel compared with separate sessions
- Separate sessions: you split the work, pick models, create worktrees, and write every handoff. Winskel: you write one objective.
- Separate sessions: each agent starts cold. Winskel: each step receives the earlier steps' summaries, changes, and open problems.
- Separate sessions: you track what waits for what. Winskel: dependent work waits for the step it needs.
- Separate sessions: results are spread across terminals. Winskel: one review with the changes and checks.
- In both cases the code, the branches, and the final merge stay yours.
Controlling which models Winskel uses
Winskel decides by default. Today you can:
- Name a model in the objective, such as Opus 5.5 or GPT-5.6 Sol. Winskel uses it, and stops and asks if it cannot run that model or cannot tell which part it is for. It never swaps in another.
- Choose the default model Winskel falls back to, during setup.
- Ask for an independent review in the objective.
In development, not available yet: limiting Winskel to a pool of models, excluding models, preferring models by kind of work, defining your own workflow such as the four steps in this guide, and per-repository rules. See choosing models.
Next steps
- Download Winskel for Mac. The Alpha needs an invitation to sign in; see install Winskel.
- Connect Claude Code and connect Codex.
- Read how Winskel breaks down work and handles Needs You decisions.
- Read what runs on your Mac and what Winskel stores in local execution, Security, and Privacy.
Common questions
Can Claude Code and Codex work on the same repository at the same time?
Yes, as long as they do not share a working directory. Give each agent its own Git worktree on its own branch. Two agents editing one checkout will overwrite each other's files and confuse each other's test runs.
Should I pass one agent's reasoning to the other?
No. Pass the objective, findings, relevant files, constraints, the changes made, and verification results. Private reasoning is long, unverified, and anchors the next agent on the first one's assumptions, which defeats an independent review.
Do I need both Claude Code and Codex?
No. One model is often enough. A second model is useful for an independent review of risky changes, for parallel work on parts that do not touch the same files, or when one provider's limits or availability get in the way.
Does Winskel replace Claude Code or Codex?
No. Winskel is an independent macOS app that coordinates the Claude Code and Codex installs and accounts already on your Mac. Claude Code is required because Winskel plans each objective with it; Codex is optional. Billing and usage limits stay with each provider.