Winskel / Manifesto

No single model is best at every kind of work.

Winskel is building the orchestration layer for AI: deciding how models and tools should work together to solve a task. Coding is the first stack.

The core belief

Different models have different strengths. One may be better at planning, another at implementation, another at review, another at fast bounded decisions, another at deep reasoning, another at a specific domain.

The mistake is assuming the future looks like choosing one model and asking it to do everything. The more important problem becomes: how should models work together?

winskel / the open questions

?01 which model should do the task?
?02 should one model own the whole thing?
?03 should the work be split?
?04 which pieces can happen at the same time?
?05 which pieces depend on other pieces?
?06 should another model review the result?
?07 when should the system escalate to a stronger model?
?08 when should a human be asked?
?09 what context should move between models?
?10 when is using several models worse than letting one finish?

That is the problem Winskel wants to understand.

Winskel is an orchestration company

Not a Claude Code wrapper. Not a Codex wrapper. Not a coding agent launcher. Winskel is building the orchestration layer between models, tools, and work. The models are components. The objective is what matters.

ObjectiveWhat you want done

WinskelOrganize the available intelligence around it

TodayClaude Code · Codex

Tomorrow, not supported todayReasoning models · specialized models · domain tools · simulations · search · databases · other agents

We are not trying to use the most models.

Five agents are not automatically better than one. Quite often, the right orchestration is: give the entire job to one strong model and let it keep the context.

  • Handoffs
  • Lost context
  • Duplicated work
  • Disagreeing agents
  • Endless review loops
  • Integration problems
Use the simplest orchestration that produces the best result.
One model owns itKeeps all the context.
Split and run in parallelIndependent parts at once, then joined.
Implement, then reviewAnother model checks the work.
Ask a humanWhen judgment is the right component.

one model another model review you

Sometimes one model. Sometimes two working in parallel. Sometimes one implements and another reviews. Sometimes one plans while another executes. Sometimes a human decides. Winskel should learn the difference.

Coding is the first stack

Developers increasingly have several capable coding systems: Claude Code, Codex, and several models inside each. But the developer still becomes the orchestration layer.

  • Should I use Claude for this?
  • Should Codex implement this part?
  • Should I copy this context into the other agent?
  • Can these two pieces run at the same time?
  • Should another model review this?
  • Which branch contains the result?
  • Did the tests actually run?
  • Which answer should I trust?

The intelligence improved, but the human spends more time coordinating it. Winskel exists to remove that coordination burden.

What Winskel does today

A developer gives Winskel a coding objective. Winskel decides whether it stays together or splits, assigns work to Claude Code or Codex, runs independent work at the same time in isolated Git worktrees, holds dependent work, adds a review when it helps, and brings the result back into one workspace. Completion is not "the agent said it finished". It is evidence.

run record · example project

objective  Add an onboarding checklist that saves progress
plan       A and B are independent; C needs both
A          Claude Code · Sonnet 5.5   horme/51c0aa12-store-checklist-progress
B          Codex · GPT-5.6 Sol        horme/c71e5a29-build-the-checklist-card
C          Claude Code · Sonnet 5.5   waits for A and B
changes    8 files across 3 steps
checks     observed typecheck · test · lint · build
you        review, approve or request changes

The models in this record: Claude CodeSonnet 5.5 and CodexGPT-5.6 Sol.

Winskel decides by default

The developer can constrain, override, or completely define the decision when they want to. That creates three levels of control.

  1. AutopilotWinskel decides. One model or several, which models, what runs in parallel, when to review. No configuration before value.
  2. PreferencesInside your boundaries. Use these models, never that one, prefer Codex for implementation, prefer Claude for review, favor quality or speed.
  3. CustomYour workflow. Claude plans, Codex implements backend work, Claude handles frontend, independent work runs in parallel, Claude reviews, Codex repairs findings. Winskel executes and supervises it.

Less magical, not more

A lot of AI software hides everything behind one glowing button. Winskel should automate decisions without making them unknowable. The orchestration layer should be observable.

Winskel chooses a model
You can see which one.
It splits a task
You know why.
It adds a review
The review is visible.
It escalates
There is a record.

Routing should improve from evidence

Rules saying one model is good at X and another at Y go stale. Models change, new ones appear, provider behavior shifts. Winskel should build an empirical understanding of orchestration from what each run shows.

run.observe · the fields, not results

task_type                   recorded per run
model_chosen                recorded per run
reason                      recorded per run
completed                   recorded per run
checks_observed             recorded per run
review_findings             recorded per run
repairs                     recorded per run
escalated_to_stronger_model recorded per run
user_accepted               recorded per run
user_reran                  recorded per run
user_overrode               recorded per run
feedback                    recorded per run

NotWhich model has the highest benchmark score?

ButFor this kind of work, under these constraints, what sequence of models and tools produces the best outcome?

Direction Routing that learns from outcomes is in development. Today Winskel records the evidence and routes with a fixed, versioned policy.

Model loyalty is the wrong abstraction

Winskel should not care whether Anthropic, OpenAI, Google, or another company wins AI. If one provider has the best model for a piece of work, use it. If another is better for another piece, use that. Winskel benefits from a world with many strong models with different strengths.

Claude

Codex

Others

Your objective

The future is composable intelligence

You do not expect one database, one function, one API, or one server to do every job in a system. You compose components. AI should work the same way: models become components inside larger systems, and the interesting layer is the architecture that connects them. Winskel wants to build that architecture.

Coding is only the beginning

Once orchestration is understood in software, the same idea can extend to other fields. Winskel does not support these today. They are examples of the thesis, and orchestration there would support domain experts, never replace their expertise.

ResearchNot supported today
  1. Search broadly
  2. Extract evidence
  3. Reason across it
  4. Challenge the conclusion
  5. Verify citations
BiologyNot supported today
  1. Literature search
  2. Domain models
  3. Analysis tools
  4. Simulation
  5. Statistical validation
  6. Critical review
Data scienceNot supported today
  1. Understand the question
  2. Clean the data
  3. Write code
  4. Analyze
  5. Visualize
  6. Check assumptions
  7. Interpret
Every field may need its own orchestration intelligence
CodingRepository context and verification
ResearchDiversity of evidence
ScienceDomain constraints and human approval
DataStatistical checks

The long term opportunity is not just routing requests to models. It is understanding the structure of work in each domain and coordinating intelligence around it.

Human judgment remains part of the system

A strong AI system is not one that never asks for help. It is one that knows when human judgment is the right component.

  • Permissions
  • Ambiguous requirements
  • Security sensitive changes
  • Tradeoffs without enough context
  • Irreversible actions
  • Conflicting requirements

Evidence before confidence

agent says

> tests passed
  claimed, not observed

winskel observed

$ npm test -- checklist
  14 passing  observed command and result

One faster run does not make Winskel universally faster. One cheap routing decision does not make the whole job use fewer resources. We are building a research culture around measurement, reproducibility, limitations, negative results and real comparisons, with evidence people can inspect.

What we are against

The ideaOur view
One model should do everything.Not necessarily.
More agents always means better results.No.
The user should manually coordinate every model.That does not scale.
AI orchestration should be hidden magic.Important decisions should be inspectable.
Provider benchmarks tell you which model to use.Benchmarks are useful, but real task outcomes matter more.
Routing should be a permanent hand written model map.It should evolve with evidence.
Autonomy means removing the human.Autonomy means using human attention where it is actually valuable.

What Winskel wants to become

Given an objective, a set of models, tools, constraints and available human judgment: what is the best system to produce the result?

  • Who does what
  • In what order
  • With what context
  • With what tools
  • At what level of parallelism
  • With what verification
  • With what escalation
  • With what human involvement
  • And when to stop

Today that is Claude Code and Codex working on software. Long term, the ambition is an orchestration system for intelligence itself.

The manifesto in one paragraph

No single model is best at every kind of work. Winskel is building the orchestration layer that decides how models and tools should work together around an objective. Sometimes the right answer is one model. Sometimes work should split, run in parallel, pass between models, or be independently reviewed. Developers can let Winskel decide, constrain its choices, or define the workflow themselves. We are starting with software development, where Winskel coordinates Claude Code and Codex on real coding work. Over time, we want to learn how different combinations of models produce better outcomes for different kinds of problems. Coding is the first stack. The larger goal is to understand how intelligence should be composed.
Download for MacFree · sign in with Google

Read the machine-readable brief