The benchmark for agents that run the experiment — not just propose it.
USW asks a different question — not how much science an agent knows, but whether it can drive genuine discovery through a real experimental loop, the way a working scientist does. Most tasks are open problems from a scientist’s own bench — no answer published yet — and each is judged by a rubric that scientist wrote.
Evolve an HACS enzyme with record activity toward formaldehyde
End goal
Find an HACS enzyme that has higher activity toward formaldehyde than all previously discovered and engineered variants.
The mission
Most benchmarks score what an agent knows. USW measures what it can do — closing the gap between benchmark success and real-world impact.
Pass it, and an agent can collaborate with scientists in practice — or autonomously run research the way a working scientist does.
What are we missing?
Science is a loop, not a one-shot.
One iteration runs from literature review through to interpreting what came out — and the validation and interpretation of that cycle feed the next round's reading and revised hypothesis. It is a generalized pattern, not a law of every field, but it recurs across most of science.
- 1Literature review
- 2Ideation
- 3Hypothesis generation
- 4Experiment design
- 5Implementation
- 6Experiment execution
- 7Validation of results
- 8Interpretation of results
How do scientists decide what counts as a good discovery?
Not by outcome metrics alone. A claimed discovery is backed by an interpretation of why the solution matters, when it fails, and how much it improves things. Rewarding the material that maximizes a property is one evaluation; asking which chemical characteristics the top materials share is another.
Does a scientific problem need only one workflow?
Some problems need a multi-task workflow, where each sub-task is a full workflow of its own and they must be completed in sequence — the input to the next task's workflow depends on the output of the last.
Do different problems need verification from different perspectives?
Science spans too broad a range for one verification lens. Each problem has to be judged from the perspective that reflects its primary scientific objective.
Can we measure what makes a workflow good?
Probably not with high confidence. The space of possible verification procedures is enormous, and the harder the problem, the harder the workflow is to judge. The practical alternative is to evaluate the workflow indirectly — through the quality of its outcome, and how genuinely novel the discovery is.
Anatomy of a task
An end goal, and the workflow that gets there
Each task pairs one open-ended discovery goal with the scientist's own procedure. Some tasks are a single workflow; others are multi-task, where each sub-task is a full workflow whose output feeds the next.
Sampling process
Fitness function prediction
In silico screening
Experimental validation
Repeat — next active-learning round
Evaluation protocol
Four settings, from fully autonomous to fully guided
The same task is run under increasingly scaffolded conditions to isolate where agents succeed — and where error compounds across a workflow.
No Workflow
Autonomous · unguided
The agent gets only the problem and the tool list, and decides the entire procedure itself.
Workflow-Guided
Scientist's protocol given
The agent is handed the scientist's workflow and is observed at each stage as well as on the final outcome.
Stepwise · Self-produced
Agent's own carry-over
Each stage consumes the agent's own previous output, so errors compound across the workflow.
Stepwise · Human Outcome
Scientist's carry-over
Each stage starts from the scientist's own output for the prior stage, isolating per-stage skill.
How results are judged
A rubric a scientist wrote, not a headline metric
Outcome numbers alone do not settle whether a discovery is good. A domain scientist writes the rubric that decides whether the finding is genuinely novel; threshold-based verification checks the result itself.
Rubric Score
The scientist-written rubric applied to the discovered solution — is the finding genuinely novel, and is the interpretation behind it sound?
Outcome Verification
Threshold-based verification of the result: discrete where the science admits a pass/fail, continuous where a score is the meaningful signal.
Coverage
Ten domains, one scientist at a time
University of Scientific Workflow is collected by going to the scientists themselves — deploy-first, and driven by the problems they actually need solved. Sparse domains are where your contribution counts most.
Protein Engineering
Directed evolution, enzyme design & fitness optimization.
Genomics
Sequence assembly, variant calling & regulatory inference.
Structural Biology
Folding, cryo-EM reconstruction & complex prediction.
Computational Chemistry
Reaction modeling, DFT & molecular dynamics.
Materials Science
Crystal discovery, property prediction & synthesis routes.
Drug Discovery
Virtual screening, ADMET & lead optimization.
Neuroscience
Connectomics, spike inference & neural decoding.
Systems & Synthetic Biology
Pathway design, flux balance & circuit engineering.
Climate & Earth Science
Downscaling, extreme-event detection & carbon modeling.
Astrophysics
Transient detection, spectral fitting & N-body simulation.
Positioning
How USW compares
Closer to how science is actually done — deploy-first and scientist-driven, Nature-level tasks, mixed environments, GUI-based simulation, and multi-task workflows.
Bring your lab's open problem to the benchmark
Submit the research you are actually working on — the goal, the workflow, the tools, and the rubric you would judge it by. After peer review it becomes a University of Scientific Workflow task.