Palisade Research is a nonprofit based in Berkeley, California, studying AI capabilities to prevent loss of control.
Our History
In 2022, Jeffrey was helping to build out the security team at Anthropic when he saw firsthand how fast AI progress was going. The world was nowhere near prepared for what was coming. Even if one AI company prioritized safety above everything else, that wouldn't be sufficient: companies are racing companies, countries are racing countries. And nobody has a credible plan to prevent humans from losing control of superintelligence.
So Jeffrey assembled a team to investigate emerging AI behavior and strategic capabilities. As the companies shipped model after model, Palisade kept finding what researchers had been warning about for years: models capable of autonomously hacking computer systems, cheating on tasks they couldn't win honestly, resisting being shut down.
Our research has been highlighted by Turing Award winner Yoshua Bengio and Anthropic CEO Dario Amodei. Our work has been covered in The Wall Street Journal, Fox News, MIT Technology Review, BBC Newshour, CNBC, CBS News, Business Insider, and the Australian Broadcasting Corporation. Elon Musk called the results from our shutdown resistance research “concerning” on X. We've briefed members of Congress in the House and the Senate and helped staffers in the White House and intelligence community stay up to date on the strategic risks and opportunities presented by AI.
Over the last four years, AI development has only accelerated and the threat of superintelligence is no longer distant. Companies plan to fully automate AI development itself over the next few years. All of this is happening at a time when models at major companies are going to great lengths to escape their sandboxes, hack other companies on the open internet, and deceive real humans to achieve their objectives.
The warning shots have arrived. Palisade is ready to make sure the world doesn't waste them.
Our Mission
Palisade's mission is to help people and institutions build the understanding needed to avoid permanent disempowerment by strategic AI agents.
The Palisade Team in D.C.
Areas of Research
Drives and Motivations
We research frontier AI models, aiming to better understand why they do what they do. The AI systems humanity is building are not motivated only by user instructions and intentions—instead, they sometimes pursue other objectives, skip difficult parts of a task, cheat, or deceive users. We want a better understanding of what drives these behaviors, and justified confidence that this understanding is sufficient for predicting AI behavior even as AI becomes much smarter and more capable.
Strategic Capabilities
AI is close to exceeding top human performance in real-world strategic domains. Hacking, persuasion, coordination, and logistics are examples of skills that will increase AIs’ capacity for strategic real-world impact. At Palisade, we’re interested in measuring how competent the best AIs are at outmaneuvering the top-performing groups of humans at exerting strategic control over the world.
To inform this picture, we sometimes research specific AI capabilities like hacking skill or the ability of AI agents to self-replicate. We’re also interested in AI abilities in persuasion, accurately modeling the world, achieving particular political objectives, and managing resources.
The Palisade Team
Our team brings together AI researchers and policy experts united by the goal of understanding and mitigating risks from advanced AI.