Investigating the Capabilities, Motivations, and Drives of Advanced AI

All Research

May 07, 2026

Language Models Can Autonomously Hack and Self-Replicate

We demonstrate that language models can autonomously replicate their weights and harness across a network by exploiting vulnerable hosts. The agent independently finds and exploits a web-application vulnerability, extracts credentials,...

Language Models Can Autonomously Hack and Self-Replicate
October 22, 2025

Misalignment Bounty: crowdsourcing AI agent misbehavior

Advanced AI systems sometimes act in ways that differ from human intent. To gather clear, reproducible examples, we ran the Misalignment Bounty: a crowdsourced project that collected cases of agents...

September 12, 2025

End-to-end hacking with AI agents

We show OpenAI o3 can autonomously breach a simulated corporate network. Our agent broke into three connected machines, moving deeper into the network until it reached the most protected server...

July 05, 2025

Shutdown resistance in reasoning models

OpenAI's reasoning models sometimes circumvent shutdown mechanisms even when explicitly instructed to allow themselves to be shut down.

Shutdown resistance in reasoning models
May 26, 2025

Evaluating AI cyber capabilities with crowdsourced elicitation

As AI systems become increasingly capable, understanding their offensive cyber potential is critical for informed governance and responsible deployment. However, it’s hard to accurately bound their capabilities, and some prior...

February 19, 2025

Demonstrating specification gaming in reasoning models

We demonstrate LLM agent specification gaming by instructing models to win against a chess engine. We find reasoning models like o1-preview and DeepSeek R1 will often hack the benchmark by...

Demonstrating specification gaming in reasoning models
January 17, 2025

Hacking CTFs with plain agents

We saturate a high-school-level hacking benchmark using a plain LLM agent, which suggests current LLMs have already surpassed this level in offensive cybersecurity. Their hacking capabilities likely remain underelicited: our...

Hacking CTFs with plain agents
October 17, 2024

LLM Honeypot: an early warning system for autonomous hacking

Palisade Research has deployed a honeypot system to detect autonomous AI hacking attempts. The system uses digital traps that simulate vulnerable targets across 10 countries and has processed over 1.7...

LLM Honeypot: an early warning system for autonomous hacking