Benchmark · 2026-05-12

SciBench-Bio v0.4: a public benchmark for autonomous wet-lab planning

We open-sourced the eval suite we run against every model release. 312 protocols across 11 sub-fields. Why we built it and what we learned.

14 min
Engineering · 2026-04-30

Why we wrote our own verifier model

The honest version: frontier LLMs are bad at catching their own mistakes. Here is the model we trained to do it for them, and the 11pp accuracy lift it produced.

9 min
Customer story · 2026-04-18

How Helix Bio shipped a Nature Methods paper in 11 days

From hypothesis to submission. The workflows they ran, the bottlenecks Sciento removed, the bottlenecks it didn't.

7 min
Opinion · 2026-04-02

Against "AI for science" maximalism

We do not believe AI will replace scientists. We believe it will replace the parts of being a scientist that nobody got a PhD to do.

5 min
Engineering · 2026-03-22

Permission-aware RAG, or: how we stopped leaking patents

A walkthrough of our per-document ACL enforcement at query time. Every retrieval is checked. Every result is auditable.

11 min
Research · 2026-03-08

The reliability gap: why agent demos lie

We measured 14 agent systems on the same 200 tasks. The gap between demo accuracy and production accuracy is 38 points. Here is why.

12 min