Current focus
Coding
Implementations, bug fixes, independent work, review and verification. Coding gives the Lab concrete changes, tests and dependencies to study.
See the coding productResearch program · Coding first
Research on how models work together.
We are building a way to compare models on specific tasks, both alone and in different workflows. We want to find when a combination improves the result, when it reduces cost, and when one model is enough.
Task: fix a regression in a repository
Candidates: one model; two passes; cross model review
Compare: quality, total cost, completion time
Keep fixed: starting code, tools, acceptance checks
Validate: separate tasks after workflow selection
Publish: method, failures, limitations, measured results01 / Workflows
A model name does not describe the whole system. We also need to decide who plans, who does the work, what context moves between steps, and how the result is checked. Lab is where we want to test those choices.
The same requirements, starting code and tools for every workflow.
One model does the whole task and keeps all of its context.
One model implements and another reviews. A repair step runs only when the review finds a problem.
Two independent parts run at the same time, then are integrated. Only for parts that do not depend on each other.
Tests, rubrics and, where needed, human review. Hidden checks stay here, outside every workflow.
Cross model reviewHandoff 2 · Model A, implements to Model B, reviews
Patch
Combining models here means coordinating separately accessed models, tools and files around one task. It does not mean merging models into a new one, and a handoff passes specific output, never another model’s private context.
02 / What we measure
Each dimension is reported on its own.
These are what we evaluate, not claims that anything has improved. We do not fold them into one score.
03 / Method
We compare whole configurations, not provider names: the model version, its role, tools, context, instructions, settings, retries, verification and the workflow around it. A coding agent such as
Claude Code or
Codex is more than its model, so agent comparisons are kept apart from comparisons of models called directly.
Requirements, permitted tools, starting state and acceptance criteria are written down before any workflow runs.
A well configured single model is always a candidate, with a less expensive one where it matters. A claimed benefit from a second model is checked against a second pass by the same model.
Every candidate gets the same starting code, tools, context budget, timeouts and retries, and every run records its exact model versions and settings. Model API comparisons and coding agent comparisons are kept apart, because an agent is more than its model.
Tests and explicit rubrics, with qualified human review where needed. Hidden checks and reference answers stay out of every candidate's context.
Workflows are explored on development tasks. The one we select is then tested on tasks that played no part in choosing it, with the selection rule fixed in advance.
Total cost includes planning, review, coordination and repairs. Failures, timeouts and invalid outputs stay in the denominator. An unknown value is reported as unknown, never as zero.
Research objective · not measured results
Comparable quality needs a rule. The quality target is set before costs are compared. Two similar averages do not show that two workflows are equivalent, and “best” only holds for the tasks, settings and date it was measured on.
A combination can help. It can also add delay, cost, or new mistakes. A single model winning, a review making an answer worse, or a result remaining uncertain all tell us something useful. We want the evidence to decide.
04 / Experiments and findings
| Configuration | Bug fixes | Feature changes | Code review |
|---|---|---|---|
| Model A alone | Proposed comparison | Proposed comparison | Proposed comparison |
| Model B alone | Proposed comparison | Proposed comparison | Proposed comparison |
| Model A, then a second Model A pass | Proposed comparison | Proposed comparison | Proposed comparison |
| Model A implements, Model B reviews | Proposed comparison | Proposed comparison | Proposed comparison |
| Model B implements, Model A reviews | Proposed comparison | Proposed comparison | Proposed comparison |
| Independent parts in parallel, then integration | Proposed comparison | Proposed comparison | Proposed comparison |
We will publish measured findings with the method and limitations.
05 / Coding first
Coding is where the Lab starts. Research and data analysis are directions we want to study later. They are not products.
Current focus
Implementations, bug fixes, independent work, review and verification. Coding gives the Lab concrete changes, tests and dependencies to study.
See the coding productFuture direction
Workflows for searching, extracting evidence, comparing sources, checking citations and challenging conclusions. A source check or agreement between models is not proof of a discovery.
Future direction
Workflows for preparing data, writing analysis code, running statistical checks, interpreting results and reproducing them. Quality criteria come first, before any combination is compared.
The Lab does not decide what the Winskel product does today. Product routing follows a fixed, versioned policy. A promising Lab result becomes a candidate for product testing, and it changes the product only after integration, security and regression review.
Winskel Lab is part of Winskel. The product you can use today coordinates
Claude Code and
Codex on coding work.