03 · Evidence
See what
actually happened.
Winskel records who ran the work, why they were chosen, what changed, and which checks actually ran.
How do I know what AI coding agents actually changed?
Every objective ends in one review. It lists the changed files by mission, model and branch, shows the diffs, and marks each check Passed, Failed or Not verified. A check only counts when Winskel observed it run.
- Objective
- Fix the flaky webhook test
- Mode
- Autopilot
- Model
Claude CodeSonnet 5.5- Source
- Strong default
- Why
- One owner keeps the context
- Checks
- Tests · typecheck · observed
- Review
- Approved
The record
What Winskel records for every route
- Model
- The model selected and the model that actually ran, on Claude Code or Codex.
- Why
- The reason in one sentence, shown as Why this model in the workspace.
- Route source
- What decided it, for example a model you named, an agent system you named, your preferences, your Custom workflow, Autopilot's default, or an escalation after a failed step.
- Mode
- Autopilot, Preferences or Custom, and which models were candidates or excluded.
- Branch and worktree
- The branch and commit for each step, in its own Git worktree.
- Changed files
- Every changed file with its status and line counts, grouped by mission.
- Diff
- The diff for each file, read from your Mac.
Checks
A check only counts when Winskel observed it run
An agent saying the tests pass is not a check.
Winskel counts a check from two sources: the checks it runs after a mission from your project's package.json scripts (test, typecheck, lint and build), or a check command the agent ran whose exit status its CLI reported.
Everything else stays Not verified. Winskel does not run your project's checks on work Codex writes inside its sandbox, so those steps show Not verified unless Codex reported a check it ran.
The objective as a whole is never labeled verified. Review the code and run the checks your project needs before merging or deploying.

Decisions
Needs You, review and history
- Needs You
- Questions, permission requests and decisions that need a person wait in one place. System failures show as Failed, not as Needs You.
- Approve
- Records that you accept the result. It merges, pushes and deletes nothing.
- Request changes
- Sends your instruction as a follow-up objective in the same chat.
- Routing history
- Counts of routes by kind of work and model over the last 90 days, in Settings.
- Feedback on a routeIn development
- Rating a routing choice so it shapes future routing.
- Learning from outcomesIn development
- Routing that adjusts from the results of past routes.
