Research program · Coding first

Winskel Lab

Research on how models work together.

We are building a way to compare models on specific tasks, both alone and in different workflows. We want to find when a combination improves the result, when it reduces cost, and when one model is enough.

experiment.specExample study definition

Task: fix a regression in a repository
Candidates: one model; two passes; cross model review
Compare: quality, total cost, completion time
Keep fixed: starting code, tools, acceptance checks
Validate: separate tasks after workflow selection
Publish: method, failures, limitations, measured results
An example, not a running experiment

01 / Workflows

The workflow is part of the experiment.

A model name does not describe the whole system. We also need to decide who plans, who does the work, what context moves between steps, and how the result is checked. Lab is where we want to test those choices.

workflow.explorerExample workflows · conceptual

Numbers show the order of handoffs, not time. Select a handoff to see what moves.

Same for every workflowTask

The same requirements, starting code and tools for every workflow.

Single model baseline

One model does the whole task and keeps all of its context.

Model Adoes the whole task

Cross model review

One model implements and another reviews. A repair step runs only when the review finds a problem.

Model Aimplements
Model Breviews
Model Arepairswhen needed

Parallel work

Two independent parts run at the same time, then are integrated. Only for parts that do not depend on each other.

Model Apart 1
Model Bpart 2
Integration
Outside the workflowIndependent evaluation

Tests, rubrics and, where needed, human review. Hidden checks stay here, outside every workflow.

Cross model reviewHandoff 2 · Model A, implements to Model B, reviews

Patch

Moves
The changed files as a diff, sent with the task requirements.
Does not move
The implementing model's private working context. The reviewer sees the output, not how it was produced.
Conceptual, not a recorded run. Three example workflows for the same task. Each starts from the task's requirements and ends at an independent evaluation, outside the workflow being tested. In the single model baseline, Model A does the whole task. In cross model review, Model A implements, Model B reviews, and Model A repairs when the review finds a problem. In parallel work, two independent parts go to Model A and Model B and are then integrated. Model A and Model B stand for any two models; they are not a ranking. Who implements and who reviews is part of the experiment, so both orders are candidates.

Combining models here means coordinating separately accessed models, tools and files around one task. It does not mean merging models into a new one, and a handoff passes specific output, never another model’s private context.

02 / What we measure

What we measure

Each dimension is reported on its own.

Quality
Does the result meet the task's requirements? Judged with the checks or review appropriate to that task.
Cost
What resources did the whole workflow use, including coordination and repairs?
Time
How long did it take to reach the evaluated result, rather than just the first answer?
Human effort
How often did someone need to intervene, correct a handoff, or resolve a blocker?
Reliability
Did the approach work across repeated tasks, and how did it fail?

These are what we evaluate, not claims that anything has improved. We do not fold them into one score.

03 / Method

How we compare

We compare whole configurations, not provider names: the model version, its role, tools, context, instructions, settings, retries, verification and the workflow around it. A coding agent such as Claude Code or Codex is more than its model, so agent comparisons are kept apart from comparisons of models called directly.

  1. A concrete task

    Requirements, permitted tools, starting state and acceptance criteria are written down before any workflow runs.

  2. Strong baselines

    A well configured single model is always a candidate, with a less expensive one where it matters. A claimed benefit from a second model is checked against a second pass by the same model.

  3. Equal conditions

    Every candidate gets the same starting code, tools, context budget, timeouts and retries, and every run records its exact model versions and settings. Model API comparisons and coding agent comparisons are kept apart, because an agent is more than its model.

  4. Independent evaluation

    Tests and explicit rubrics, with qualified human review where needed. Hidden checks and reference answers stay out of every candidate's context.

  5. Separate confirmation

    Workflows are explored on development tasks. The one we select is then tested on tasks that played no part in choosing it, with the selection rule fixed in advance.

  6. Everything counts

    Total cost includes planning, review, coordination and repairs. Failures, timeouts and invalid outputs stay in the denominator. An unknown value is reported as unknown, never as zero.

quality.costMethod · no data

Research objective · not measured results

Quality versus total workflow cost: research directionsConceptual diagram with no data points. A quality target line and a fixed budget line, with two arrows: along the target towards lower cost, and along the budget towards higher quality.Quality target, set before comparingFixed budget1 · Maintain quality, reduce cost2 · Improve qualitywithin a fixed budgetTotal workflow costTask outcome qualityQuality versus total workflow cost: research directionsConceptual diagram with no data points. A quality target line and a fixed budget line, with two arrows: along the target towards lower cost, and along the budget towards higher quality.Quality target, set firstFixed budget1 · Maintain quality,reduce cost2 · Improve qualitywithin a fixed budgetTotal workflow costTask outcome quality
Conceptual diagram with no data. Across is the total cost of a workflow, including planning, review, coordination and repairs. Up is how well the outcome meets the task's requirements. Direction 1 holds a quality target, set before the comparison, and looks for lower total cost. Direction 2 holds the budget fixed and looks for a better outcome. A measured version will plot one point per evaluated workflow configuration, with a table of the plotted data.

Comparable quality needs a rule. The quality target is set before costs are compared. Two similar averages do not show that two workflows are equivalent, and “best” only holds for the tasks, settings and date it was measured on.

What counts as a useful finding

A combination can help. It can also add delay, cost, or new mistakes. A single model winning, a review making an answer worse, or a result remaining uncertain all tell us something useful. We want the evidence to decide.

04 / Experiments and findings

What we are investigating

  1. Does a reviewer catch meaningful mistakes, and does it matter which model reviews?
  2. Does a handoff between models lose context that the next step needs?
  3. Does splitting a task help, or does it only add coordination work?

experiment.matrixProposed study

Experiment design · not a result table. Each cell is a comparison we plan to run.
ConfigurationBug fixesFeature changesCode review
Model A aloneProposed comparisonProposed comparisonProposed comparison
Model B aloneProposed comparisonProposed comparisonProposed comparison
Model A, then a second Model A passProposed comparisonProposed comparisonProposed comparison
Model A implements, Model B reviewsProposed comparisonProposed comparisonProposed comparison
Model B implements, Model A reviewsProposed comparisonProposed comparisonProposed comparison
Independent parts in parallel, then integrationProposed comparisonProposed comparisonProposed comparison
Rows are workflow configurations, columns are kinds of coding task. Model A and Model B stand for any two models, so every comparison is planned in both orders and against each model alone. No results are published yet.

We will publish measured findings with the method and limitations.

05 / Coding first

Coding first

Coding is where the Lab starts. Research and data analysis are directions we want to study later. They are not products.

Current focus

Coding

Implementations, bug fixes, independent work, review and verification. Coding gives the Lab concrete changes, tests and dependencies to study.

See the coding product

Future direction

Research

Workflows for searching, extracting evidence, comparing sources, checking citations and challenging conclusions. A source check or agreement between models is not proof of a discovery.

Future direction

Data analysis

Workflows for preparing data, writing analysis code, running statistical checks, interpreting results and reproducing them. Quality criteria come first, before any combination is compared.

The Lab does not decide what the Winskel product does today. Product routing follows a fixed, versioned policy. A promising Lab result becomes a candidate for product testing, and it changes the product only after integration, security and regression review.

Winskel Lab is part of Winskel. The product you can use today coordinates Claude Code and Codex on coding work.

Read the Manifesto See the product