The dream of autonomous wet-lab discovery hits a hard wall with the release of LabBench, a new benchmark from Gamow Labs that tests whether AI agents can make the critical decision of what experiment to run next. The results are stark: while frontier models like GPT-6 Astra and Claude Opus 5.5 can interpret existing data with surprising competence, they fail catastrophically at the core task of experimental design. The benchmark, published on September 30, 2026, utilizes 20 held-out tasks derived from real-world drug discovery and genomics records, creating a scenario where the agent must commit to a next step without knowing the lab's actual choice.

The Illusion of Competence

LabBench constructs a 'snapshot in time' for each task, providing the agent with the evidence available at a specific decision point while withholding the lab's interpretation and chosen action. The grading is rigorous, using 20–22 binary criteria derived from the true answer. The performance of top-tier models is mediocre at best. GPT-6 Astra achieved a mean score of 40.8%, passing just two out of every five criteria. Claude Opus 5.5 trailed closely at 40.5%, but with a significant efficiency penalty: Astra completed tasks in a median of 5 minutes, while Opus required 35 minutes. Grok 4.7, Muse Spark 1.3, and Gemini 3.8 Flash lagged further behind, with Gemini 3.8 Flash scoring only 18.4%.

Analysis vs. Agency

The breakdown of failures reveals a fundamental architectural limitation in current LLMs. The agents excel at passive interpretation—identifying what a measurement is or explaining an analysis—with pass rates hovering around 47%. However, when forced to 'choose, commit, or rank' a next step, the pass rate plummets to 21%. Notably, no agent successfully passed any of the 13 criteria specifically asking 'which experiment should come first.' This suggests that while models can retrieve and synthesize biological knowledge, they lack the causal reasoning required to prioritize actions in a resource-constrained environment.

The Latent Knowledge Bottleneck

Gamow Labs tested whether these failures stemmed from a lack of knowledge or an inability to access it. In a revealing experiment, researchers appended a single sentence to prompts for five core decisions that all agents failed. This hint did not provide the answer but merely redirected attention to evidence the model already possessed. The result? GPT-6 Astra’s score jumped from 20% to 75% on one task, and from 23% to 64% on another. The models clearly know enough biology to work alongside human experts, but they cannot reliably recall or apply that knowledge to make autonomous strategic decisions without significant prompting scaffolding.

Key Takeaways

  • Frontier models fail to independently select the next experimental step, with 0% success on 'which experiment first' criteria.
  • GPT-6 Astra and Claude Opus 5.5 perform similarly in accuracy (~40%) but differ vastly in speed (5 min vs 35 min).
  • Agents are strong at interpreting existing data (47% pass rate) but weak at designing new actions (21% pass rate).
  • 'Hinting' at relevant evidence already in the context can dramatically improve decision-making scores, indicating a retrieval or reasoning bottleneck rather than a knowledge gap.

The Bottom Line

We are still far from AI agents that can run their own wet labs; current models are brilliant librarians but terrible principal investigators. Until agents can autonomously weigh cost against potential insight, the human scientist remains the indispensable driver of experimental design.