SciBench-Bio v0.4: a public benchmark for autonomous wet-lab planning
We open-sourced the eval suite we run against every model release. 312 protocols across 11 sub-fields. Why we built it and what we learned.
Engineering write-ups, benchmark reports, and the occasional opinion piece — from the team building Sciento.
We open-sourced the eval suite we run against every model release. 312 protocols across 11 sub-fields. Why we built it and what we learned.
The honest version: frontier LLMs are bad at catching their own mistakes. Here is the model we trained to do it for them, and the 11pp accuracy lift it produced.
From hypothesis to submission. The workflows they ran, the bottlenecks Sciento removed, the bottlenecks it didn't.
We do not believe AI will replace scientists. We believe it will replace the parts of being a scientist that nobody got a PhD to do.
A walkthrough of our per-document ACL enforcement at query time. Every retrieval is checked. Every result is auditable.
We measured 14 agent systems on the same 200 tasks. The gap between demo accuracy and production accuracy is 38 points. Here is why.