Nicholas Larus-Stone is Benchling’s Head of AI and coauthor of a new benchmarking study putting large language models (LLMs) like Claude and ChatGPT to work redesigning wet lab protocols.
The models did just fine, but didn’t blow the Benchling team away. “They can certainly be helpful at certain biological tasks, but they’ve by no means solved some of the kind of biology problems that people really care about,” Larus-Stone said.
That said, he believes that pharmaceutical companies should be spending way more on AI usage than they currently do. That’s just one of several spicier takes we got into in our conversation.
Listen for more of Larus-Stone’s thoughts on why benchmarking LLMs is important, how open-source models can be competitive against proprietary ones, how he has fostered cross-pollination between software engineers and scientists by founding Bits in Bio, and whether he likes the term “TechBio.”
You can also listen on Apple Podcasts and Spotify
Show Notes
Benchling’s Bench-Bench protocol study white paper
Benchling blog post about the benchmarking study
Bits in Bio Website
Bits in Bio 2025 Annual Report
Other AI benchmarking studies
From Jure Lescovec’s lab and the Biomni team:
Qu et al. “BiomniBench: Process-level Evaluation of LLM Agents for Real-world Biomedical Research” bioRxiv. May 12, 2026.
Guo et al. “PromptBio-Bench: Benchmarking LLM-based Bioinformatics Agents for End-to-End Data Analysis.” BioRxiv. June 1, 2026.
Other podcast episodes referenced











