Biollama: testing biology pre-training risks

We collaborate with RAND to find out whether adversaries can fine-tune LLMs for use as bio lab assistants

Additional ways to view:

We collaborated with RAND to see if adversaries can fine-tune LLMs as bio lab assistants. Results suggest pre-training on biological corpora is unlikely to improve performance. Conversely, inference scaling and task-specific fine-tuning are likely to provide boosts.

May 7, 2026

Language Models Can Autonomously Hack and Self-Replicate

We demonstrate that language models can autonomously replicate their weights and harness across a network by exploiting vulnerable hosts. The agent independently finds and exploits a web-application vulnerability, extracts credentials,...

SecurityAutonomous HackingSelf-Replication

October 22, 2025

Misalignment Bounty: crowdsourcing AI agent misbehavior

Advanced AI systems sometimes act in ways that differ from human intent. To gather clear, reproducible examples, we ran the Misalignment Bounty: a crowdsourced project that collected cases of agents...

AI SafetySecurity

September 12, 2025

End-to-end hacking with AI agents

We show OpenAI o3 can autonomously breach a simulated corporate network. Our agent broke into three connected machines, moving deeper into the network until it reached the most protected server...

Autonomous HackingSecurity