AI Models Outperform Virologists in the Lab | ChatGPT app | ChatGPT OpenAI | OpenAI free | Turtles AI
A recent international study finds that AI models like OpenAI’s o3 and Gemini 2.5 Pro handle wet lab procedures better than PhD virologists, highlighting research benefits, potential misuse risks, biosecurity, and future adjustments needed.
Key Points:
- AI models like OpenAI’s o3 (43.8%) and Gemini 2.5 Pro (37.6%) outperform PhD virologists (22.1%) in wet lab troubleshooting tests.
- The test was designed by the Center for AI Safety, MIT Media Lab, UFABC, and SecureBio, with questions not found in academic literature.
- Experts like Dan Hendrycks and Tom Inglesby call for pre-release audits, differentiated access, and mandatory unfiltered versioning.
- “Unlearning” techniques like CUT aim to weed out dangerous knowledge from AI without compromising performance.
The joint study by the Center for AI Safety, MIT Media Lab, Brazil’s UFABC University, and SecureBio put language models like OpenAI’s o3 and Google Gemini 2.5 Pro through a wet-lab troubleshooting test, built in collaboration with expert virologists to recreate complex viral culture scenarios and protocols not described in academic literature. While PhD-level researchers averaged 22.1 percent correct, o3 scored 43.8 percent and Gemini 2.5 Pro 37.6 percent, significant improvements over previous versions of the models and clear progress for both Anthropic’s Claude 3.5 Sonnet and a preview of GPT-4.5, reflecting rapid accumulation of practical expertise. The findings have prompted companies like xAI to implement virus safety risk management frameworks and OpenAI to conduct red-teaming campaigns to block 98.7% of potentially harmful conversations, while Anthropic updated its performance documentation without specific mitigation details and Google has refrained from commenting. Experts like Dan Hendrycks of the Center for AI Safety and Tom Inglesby of the Johns Hopkins Center for Health Security are calling for differential access rules, pre-release audits, and mandatory regulatory standards to evaluate unfiltered versions, arguing that third-party enforcement can balance scientific utility with abuse prevention. In parallel, there are emerging “unlearning” strategies like CUT, developed by Scale AI and the Center for AI Safety to remove dangerous knowledge without affecting general functionality, and academic analyses that distinguish risks between large language models and biological design tools, suggesting multi-level interventions like independent assessments, controlled access, and universal screening of genetic materials for effective governance.
In a context where biosecurity defenses and regulatory policies are struggling to keep pace, the maturation of these technologies requires constant attention.


