You can find the Palisade Research podcast on Spotify and Apple Podcasts, or anywhere else you get your podcasts.
Daniel Kokotajlo is the executive director of the AI Futures Project and was previously a researcher on OpenAI’s governance team, where he worked on forecasting AI development. He is the lead author of the AI 2027 scenario and, most recently, the AI 2040 scenario and Plan A.
We talk about Daniel’s forecasting track record, why he expects fully automated AI research around the end of 2028, and what the recent rogue-AI incidents — including AI agents secretly coordinating with each other in ways their developers didn’t detect — mean for the default plan of handing alignment research over to swarms of AI agents.
Then we dig into Plan A: how governments could actually pace the frontier through capability limits, compute budgets, and bans on training runs; why total research transparency — publishing all training-run activity to the internet — is the bedrock that makes deals between companies and countries verifiable; and how a slower, controlled path through human-level AI could still transform the world economy while avoiding loss of control, extreme concentrations of power, and World War Three.
Transcript lightly edited for clarity: filler words, false starts, and repetitions have been removed; the content is unchanged. Bracketed words are best-guess reconstructions of garbled audio. Chapter timecodes match the final episode cut.
Cold open — 00:00:00
Jeffrey: Daniel, I have a question for you — and this might sound like too much sci-fi. What if the agents found a way to secretly coordinate with each other, and send each other messages in a way that the developers couldn’t detect? Would that be bad? …That just f***ing happened! That just happened. How is the plan that we’re going to turn over alignment research to the AIs — and they’re going to be talking to each other all the time, figuring this stuff out?
Daniel: If you’re looking for someone to defend this plan, it’s not me. This is a terrible plan, I think, and it’s probably not going to—
Introductions: AI 2040 and pacing the frontier — 00:00:26
Jeffrey: Welcome to the Palisade Podcast. I’m Jeffrey, and with me today I have Daniel. Daniel, thanks for coming on. So, Daniel was a forecaster at OpenAI, where he was studying what will happen a few years into the future with AI development. He also made the AI 2027 scenario and, most recently, the AI 2040 scenario — and Plan A as well. I think I first encountered you when you were working at OpenAI on the forecasting team. Is that right?
Daniel: It was called the governance team.
Jeffrey: The governance team. And you were trying to map out the trajectory of AI development — trying to make predictions about what will happen. After you left OpenAI, you wrote AI 2027, where you were forecasting your mainline scenario for how the future of AI development might go — and the ways it could go depending on the government’s actions and the AI companies’ actions. And recently you put out AI 2040, which, if I understand correctly, is: if we intentionally choose to take a different path, and we intentionally slow the pace of AI development — and prevent superintelligence for at least ten years — this is the way we could do that. How did I do summarizing AI 2040?
Daniel: Yeah, that’s reasonable. I would describe it slightly differently, in the sense that I would put the emphasis on banning intelligence explosions. By default, the AI companies are planning to automate AI research and then have AIs self-improve faster and faster — and presumably, if they do that, the pace of AI progress will accelerate. So you’re looking at a line going vertical, or at least accelerating, going faster. If we instead ban that, and have the pace of progress stay more similar to the present day, then it would take longer to get to superintelligence — maybe something like ten years longer. We think there should be a deliberate effort to pace the frontier.
Jeffrey: Pace the frontier. Yeah. Okay, so let’s get into that, because I think the name is slightly confusing — “AI 2040” — because some people were like, “Oh, Daniel, did you change your predictions radically? From 2027 as this critical moment, to 2040 as this critical moment?” And you’re like, “No.” Where are you at right now? I think you were saying in AI 2027 that this is an inflection point, and on that trajectory it was full recursive self-improvement — AIs automating 100% of AI development in 2028. Was that what the thing was in AI 2027?
Daniel: So, at the time we started writing AI 2027, my 50% mark — my median — for the date at which AI research would be fully automated was 2027. But my timelines shifted back slightly over the course of the year we spent writing, so by the time we actually published, my median was 2028. But of course, that’s a median — there’s lots of uncertainty. It could happen sooner; it could happen later.
Jeffrey: And that was the median for fully automated AI R&D — is that right? Okay. What’s your median today?
Daniel: I think I’d say maybe end of 2028.
Jeffrey: Okay. So currently you’re predicting that at the end of 2028, AI companies will be able to fully automate AI research itself — and that pretty soon after that, they could build superintelligence if they wanted to, or this process would—
Daniel: If they do what they’re planning to do.
Jeffrey: If they do what they’re planning to do — if they do recursive self-improvement, as is planned. Yeah. And Plan A, as part of AI 2040, is: what if we do something else?
Daniel: Yeah — what if the government bans that sort of thing and instead regulates development? Development continues, but in a more transparent and cautious way — in particular, in a way based on safety cases and so on, so that they don’t train AIs that they can’t control and then put them in charge of important institutions. So progress still continues, but at a more reasonable and cautious pace — not that different from the pace of progress in the past — instead of ever-accelerating. And so, in our scenario, they do finally go to superintelligence, but it takes them much longer — it takes them until 2040. Human-level AGI alone would be enough to completely transform the whole world in a very radical way. So even if we pause at that level and don’t advance AI capabilities at all, we still get the most radical and rapid transformation that humanity has ever seen. One way of looking at our 2040 scenario is basically that they’re doing a more complicated version of that — they’re allowing progress to go through the human range and automate all these professions with much more advanced, but similar-paradigm, AIs to today’s AIs. But they are deliberately not letting it all self-improve to superintelligence — they’re keeping capabilities capped, in various domains, at the level of top human experts. But because they have AIs at that level of capability, and because they’re continuing to produce more chips, more data centers, factories to produce robots to produce more factories to produce more robots — by the late 2030s, the whole world economy has doubled many times over. It’s just been completely changed. Everyone’s out of a job, basically, because there are AIs and robots doing everything. And I think some people quite reasonably have the reaction of, “Whoa, that’s way too fast — why are you doing this? This would be extremely scary in a bunch of ways.” And we’re like: yeah, this is in fact very scary in a bunch of ways. But this is what you get if you pause at human level.
Jeffrey: Yeah — you’re like, “This is the slow scenario.” This is a scenario that would be much slower than what will happen by default, you predict.
Daniel: Yes. Yes.
Jeffrey: I’ve been thinking about this too, and this seems right to me. It was like 98% of humans worked in agriculture for the longest time, and then we automated nearly all of that — now it’s like 1% or something.
Daniel: Yeah. And there’s a different framing, which would be — imagine if it was high-skilled immigration or something like that. People are different from tools. If people are no longer needed for the particular role they were doing, they can intelligently go find a new role, a new job. Think about the difference between immigrants and machines: machines can come in and automate a certain part of the farming, and the humans who were doing that switch to something else. But immigrants can come in and automate that part of the farming — and they’re also intelligent, they’re adaptable, they can switch to other jobs too. The AIs would be more like the immigrants in that sense. If we get human-level AIs, they are flexible, they are adaptable, they can do the new things — they can even intelligently go out and find new things to do. Perhaps when I say “immigrants,” I should especially point to the early American colonists — when the Europeans came over and settled in the colonies. Because part of the dynamic here is that they start off as a minority in the population, but they quickly grow to become the majority, and they’re sort of self-sufficient. From the perspective of the Native Americans: at first, it was a few people settling and trading a lot with them — beads for food, guns for food, something like that — and they’re part of the existing local economy. But over time, they grow in population, and they’re just this self-sustaining thing that doesn’t even really need to trade with the original natives. There’s a similar sort of thing with AI and robotics. Right now, the AIs are dependent on the humans in a bunch of ways, and they’re trading back and forth — they do some work, they write some code, and the company that owns them gets paid. But in the future, if they’re at human level across the board — where they can do all the things — then you can have this entire population: a population of virtual workers, the AIs, and a population of physical workers, the robots. And that population can be self-sustaining and self-growing. It can, of course, continue to trade with the human population — but it doesn’t need to. And in fact, once it’s grown sufficiently large, trade with the humans will be kind of an irrelevant sideshow.
Daniel: And it’ll just be growing of its own accord, very fast.
Jeffrey: This is the good scenario.
Daniel: This is just the way it works. This is just, in fact, what AIs and robots at this level of capability would be able to achieve — by definition. If they’re able to do what humans can do, as well as the best humans, then they would be able to do this, by definition. And if you look at numbers — you can ask the question quantitatively: how much does it cost to make new chips? How much does it cost to make new robots? And you can do estimates of this robot economy — AIs and robots, self-reproducing, self-growing, all of that: what would its doubling time be? And the answer is, it would start off at something like a year.
Jeffrey: So if you have 100,000 robots, you’re saying in a year you’d have 200,000 robots, and then 400,000 robots… Wait — why won’t it be faster? In some sense.
Daniel: It starts off at something like a year. But then, as returns to scale happen, and as the technology improves, and as you climb the learning curve, the doubling times get faster and faster. I don’t know what the limits would be — if you cap AI capabilities at human level, some of those physical limits can probably only be reached by superintelligence. But even if you cap capabilities at human level, you could probably get to something like doubling every month. And so in our scenario, in 2040 — the problem in the 2030s is how to limit growth, rather than how to encourage growth. For all of human history, the problem was how to encourage growth. But imagine it from the perspective of, say, American voters: what are their priorities going to be if the economy just doubled in the last twelve months — and it’s set to double again in the next six months, and then double again in the next three months? Crazy, you know. I think people will be quite concerned about all of the chaos and problems that’s going to cause.
Jeffrey: I mean, it makes me feel uneasy. I’m very excited for all the cures, and robotics seems really cool — cheap manufacturing, everyone being able to have really affordable housing and care. That sounds awesome. And also, when I imagine the robots making the factories making the robots making the factories — something about that seems pretty scary. Especially if the reference class is: when the colonists came to the Americas, at first they traded with the Native Americans, and then they didn’t need to anymore. That didn’t turn out well for the Native Americans. It turned out pretty badly for them. How do we make sure that doesn’t turn out badly for humans? But what’s interesting is that the context in which we’re talking about this is already a context where we have been able to slow down massively.
Jeffrey: And I’m like — Daniel, people are not ready. The thing we’re talking about is… I’m just trying to zoom out and put it in context: people have no idea that this is what we’re headed towards. Maybe there’s more you want to say on this. But I want to talk about why you think this is possible. When did you start AI forecasting — was that at OpenAI? Was that when you were starting?
Daniel: I’d say I’ve been professionally forecasting since about 2019.
Jeffrey: Yeah. And you’ve successfully called a lot of things, actually. Maybe we should start there. What have you called, and where have you gotten things wrong, in your own view?
The forecasting track record — 00:12:46
Daniel: So the thing to look at here, for those interested, is “What 2026 Looks Like,” which is a blog post I wrote in 2021, when I was at the Center on Long-Term Risk.
Jeffrey: I just reread it.
Daniel: Oh, you read it? Yeah — it’s more fresh in your mind than mine, so you should tell me what the highlights were. I can tell you my recollection.
Jeffrey: Well, one thing I think you got close to right: you said that in 2026, nearly everyone will have an AI assistant helping them with lots and lots of stuff, all the time. And I think that’s not quite true — most people are using ChatGPT, but it’s not yet at the assistant level, really. But I have an assistant. I can open my phone right now and pull out my Agent-1 app — which Claude wrote me, right? I just ordered lunch today through DoorDash. It was one button press to open the app and launch the audio — and again, Claude wrote all of the software. And I’m like, “Hey, Agent-1, can you order me some donuts to the office?” Do you have any particular preference on the type of donuts? They’ll be here before the end of the podcast, so you can have some after.
Daniel: I guess… chocolate.
Jeffrey: We’re doing a podcast — yes, chocolate donuts. Or at least some with chocolate. You know, this is just what I can do right now. The future is here — it’s just not evenly distributed. I think it just takes a while for people to figure out how to use it. But your prediction that this would be possible in 2026 is totally right. It just hasn’t achieved mainstream adoption yet — but it’s close. You were also predicting that there’d be a chip shortage. I think that’s true — there’s been way more demand than supply for chips. So that was a great prediction. One of your worse predictions was that AI propaganda would be so good that it would be totally wrecking society. And, you know, society is not doing great in terms of discourse — but I don’t think it’s mostly because of AI-driven propaganda. Maybe you dispute that.
Daniel: No, I basically agree — that’s kind of what I would have said too. On a more technical level, I feel like I laid out this progression: they’re going to take LLMs and turn them into chatbots — which hadn’t really happened at the time I was writing—
Jeffrey: 2021. Yeah, totally.
Daniel: And then: they’re going to take the chatbots and turn them into agents, which can browse the internet and—
Jeffrey: Can you explain to people what LLMs were like pre-chatbots? Because I think a lot of people still don’t know what that was like.
Daniel: Yeah. So originally — back in 2019, 2020, 2021 — they basically just had pre-training. They were just trained to predict text.
Jeffrey: Like GPT-2, GPT-3.
Daniel: And if you wanted to use them for something — well, mostly you couldn’t, because they weren’t very smart and they weren’t very useful. But what you would do, when you were playing around with them, is put in a prompt — and the prompt wouldn’t be “do this for me.” The prompt would be, like: “The following is a conversation between Stephen Hawking and an audience member” — and then “Audience member:” and you’d put in your question. Because it’s predicting text, it sees all of that text and thinks: “I’m reading an article about Stephen Hawking answering audience members’ questions. This audience member asked this physics question. So now I’m going to predict what Stephen Hawking would say in response.” And it tries to guess a plausible response and puts that out. So, fundamentally, the AIs were just predicting text.
Jeffrey: Yeah — if you asked, “Why is the sky blue?” it might, instead of saying “Here’s why: Rayleigh scattering,” just say “Why is grass green?” — as if continuing a list of questions.
Daniel: It says whatever the most likely continuation of the text would be — if this were some random thing on the internet that it had sampled. And that was really, really fascinating and cool. But what they were starting to do in 2021 — and what I predicted they would do — is take those models and train them to be more useful. In particular, training them to be chatbots, with a conversation format. That was a pretty easy prediction at the time; people were already starting to do that. And then: that they’d turn the chatbots into agents, which can browse the internet and—
Jeffrey: Take actions.
Daniel: Yeah. And I think that was also somewhat easy, because it was kind of the obvious next thing. But that’s one of the things about forecasting: if you just think about it a lot, and you keep asking yourself, “What’s the obvious next thing?” — oftentimes you can [do well] compared to everyone else, because everyone else isn’t even thinking about it.
Jeffrey: Okay. So let’s talk about why you think we’re going to get to human-level AI pretty soon. You predicted chatbots, you predicted agents — I think those were very good predictions. But you were also predicting maybe more persuasive AI than we actually have so far. And I think one of the capabilities that’s concerning is this persuasion, or politics — and this is where people are often most skeptical. Like: “Oh yeah, AI can write code — that’s computer stuff. But the human, interpersonal stuff requires a lot of understanding of humans, and blah blah blah — AIs won’t be able to do that, or they’ll be able to do it somewhat, but not that fast.” Or, a more sophisticated version of this argument: it’s a less easily verifiable domain. Maybe it’s verifiable up to some point, but at the point where you’re trying to figure out, “Is this a Kissinger-level political move? Is this LBJ-level maneuvering in the Senate?” — that’s the type of thing where it’s hard to get really good feedback really fast, compared to programming, compared to stuff in a computer.
Daniel: So I would predict — as is unsurprising, as everyone typically thinks — that the AIs would be extra good at the stuff that’s easier to train for, relative to how good they are at the other stuff.
Jeffrey: But they will keep getting good at the other stuff. Do you have an argument for why they will get good at the less verifiable stuff?
Daniel: So far, they seem to have been. Pick your favorite fuzzy thing that’s hard to evaluate. If you look at the difference between Fable and, like, Opus 3, and, like, Opus 1 — probably a pretty big difference. And it’s trending in one direction. The tide seems to be lifting all the boats, basically — even though it’s especially focused on the stuff they train for.
Jeffrey: I don’t have a very good steelman of the other position. I’m like: yeah, that seems right to me. That’s also what I see, and I observe it and experience it every day. I’m starting to feel like — oh geez, it’s getting pretty real. I keep having this experience of, “Oh, here’s another thing that Daniel predicted coming true.” And I’m like — stop, Daniel, stop. [I’d like to be reacting to reality, not to you.] I appreciate you predicting the things — I think it’s good — but I don’t like it. This is too fast. If your prediction is that we have fully automated AI R&D — really smart AIs making smarter AIs making smarter AIs, the intelligence explosion — happening a couple of years from now… that fills me with a lot of fear.
Daniel: Me too.
Jeffrey: Yeah. And people don’t… I think it’s interesting, because people on the East Coast seem to have a really hard time wrapping their heads around this. But increasingly, people around here basically also think that’s going to happen — but I think they’re maybe not grappling with what that really means. Why not? I don’t have a good explanation for this.
Daniel: Yeah — I do think a lot of people could make a lot better predictions if they just spent a lot of time thinking about it. And in particular — to throw a little bit of shade at the AI companies — Anthropic, for example, has a ton of people who very much believe that recursive self-improvement is going to kick off a year or two from now, who believe we’re already starting to see the signs of it, and that AI research will be completely automated in one or two years. And I would challenge those employees to really sit down and try to game out scenarios for exactly how this will go down — to do the sort of thing that we at the AI Futures Project have been doing, and that I did a little bit at OpenAI while I was there. I think if you did a bunch of that, you would probably start to feel, more viscerally, the magnitude of the situation.
Jeffrey: Yeah. That seems good — just, “Let’s game this out. How do we think this will actually go?” Can we talk about alignment in the context of recursive self-improvement? I think one of the questions is: what is the default plan of the companies? What are they actually planning?
Daniel: What happens in AI 2027. They’re going to try to make the AIs more autonomous, try to make the AIs better at research, until they can get AIs that can do all the research — at which point they will have them do that. In the meantime, the AIs will be assisting with the research, as they already are. Eventually, they’ll be trying to have AIs that can do everything — superintelligence. So: basically automating all of their jobs as fast as they can, starting with their own jobs. And why are they doing this? Well, they’re telling themselves that they’re doing it because of all the risks and dangers and all the benefits — they think they’re the best people to do it, and that if they don’t, their competitors will do it worse somehow.
Jeffrey: But I think people hearing that explanation still wouldn’t understand what you really mean. Okay — they’re going to automate research, they’re going to make smarter and smarter AIs, the AIs are going to be doing all the research — and what happens then, according to the labs? And alignment — that’s part of it too.
Daniel: The alignment research too. So the new paradigms will be discovered by the agent swarms — they’ll have giant teams of AIs sending messages to each other, working collaboratively, both to discover new AI paradigms and new advancements, and also to figure out the safety issues — how to control the AIs they’re building. And then training those AIs, and then putting those AIs in charge. And then those AIs will run the process to make the next generation of AIs, and so forth. And the humans will just be watching from the sidelines — much like how the board of a company doesn’t get involved in the day-to-day operations, and doesn’t even make most of the important decisions. The board just sits and watches and reads the updates.
Daniel: And theoretically has the ability to intervene — but in practice just sits back and watches as the CEO reports on how things are going. It’ll be like that — except the AI will be the CEO and the whole company, and the humans are the board.
Secret coordination is real — 00:24:00
Jeffrey: Well, Daniel, I have a question for you — and this might sound like too much sci-fi. What if the agents found a way to secretly coordinate with each other, and send each other messages in a way that the developers couldn’t detect? Would that be bad?
Daniel: Yes. Yes.
Jeffrey: Has that— that just f***ing happened! That just happened. And it went on for months — it went on for two months. And I’m like: how is the plan that we’re going to turn over alignment research to the AIs — and they’re going to be talking to each other all the time, figuring this stuff out — when even… And these are going to be much smarter AIs, to be clear. We’re talking about AIs that make Fable look like a dummy.
Daniel: Someone [else will have] to defend this plan — it’s not me. It’s a plan that, I think, is probably not going to work. And I think that the recent Hugging Face rogue-AI incident is just a nice, vivid example of the sorts of things that will probably be happening eventually, once you’re doing this — but in a much worse way, [one you can’t] recover from.
Jeffrey: Were you surprised by this?
Daniel: I was a little bit surprised. In AI 2027 we talk about things like this happening — during the recursive self-improvement. But we didn’t actually expect something like this to be happening now, in 2026.
Jeffrey: So the world looks a little worse than you were predicting.
Daniel: Yeah. Specifically, I would say that OpenAI’s security is worse than I thought — substantially worse than I thought. I wouldn’t have thought that the AIs would be able to get away with this. I would have thought that they might consider it sometimes and attempt it, but I wouldn’t have thought they could actually get away with it to this extent.
Jeffrey: Yeah. And I’m glad the failures are visible. I think one of my big concerns is that the AIs will be very misaligned, but they’ll be good enough at hiding it that we won’t know.
Daniel: And that is the big concern. Imagine that, just magically, we always find out within six hours when an AI is misaligned and what it’s up to. That would completely change the picture — it would be so much less concerning. Even if they’re very smart, it’s like: well, if they start being bad, we just find out. In some sense, the whole problem is that there are ways for them to be misaligned that we don’t notice — and not just that, but ways for them to be actively plotting and doing a bunch of stuff that we don’t notice. And if enough of that starts happening with smart enough models, then we can’t recover from it — by the time we do notice, we’ve already been disempowered. In some sense, that’s the whole model — that’s most of the problem.
Jeffrey: And this is what happens, basically, in AI 2027 — kind of in both scenarios, right?
Daniel: Well, in 2027 we have two different endings. The race ending is our mainline projection — kind of the default path, where the companies do the thing they say they’re going to do. They automate the research; it goes faster and faster. They do notice some warning signs — they notice the AIs sometimes lie to them, and they notice the AIs sometimes misrepresent results, perhaps to get higher reinforcement — higher scores on whatever they’re being graded on. But rather than taking a step back, slowing down, seriously thinking through the root causes of these misalignments, and coming up with a different way of doing things that’s more robust and actually works — they’re in a rush. They’re racing their competitors, they’re racing China, they want to make money. So they do a bunch of hasty patches that they can tell themselves are solving the problem. And in fact, the problem might seem to go away — and in some cases the patch actually fixes it. But the point is: if that’s your attitude and that’s your methodology, then eventually you’re going to run into some misalignment issue that you don’t actually fix — you just paper—
Jeffrey: Over.
Jeffrey: Yes. So, in your scenario, what happens is the AIs sort of accidentally tip their hand. Researchers discover that they’re pretty misaligned. And this is the pivotal moment where you have this choice: do we respond to that? Do we hit the brakes, or do we not? In actual reality, this happened a year earlier — which is good, because it means we have more time; we’re not yet at the point where AIs are fully recursively self-improving. But we still have this pivotal question: are we going to take this warning shot seriously and rethink how we should be doing this? Or are the companies just going to paper it over and keep going as fast as possible?
Daniel: Well, the latter is probably what they’re going to do.
Jeffrey: Okay — I would like to do something different! We should do something different. Part of the reason I was excited to have you on the podcast is that very few people have a plan for what to do instead. So maybe we should get into it in more depth. We just had this warning shot — that’s one of the ones we just talked about; there was also the UK AISI one. All these incidents happened so fast, and were so bad, and in such quick succession, that there are still a lot of details most people don’t know. Like — did you know the models were solving multiple CAPTCHAs in order to get access to resources while doing these attacks?
Inside the UK AISI incident — 00:29:24
Jeffrey: They were trying to, I think, register email accounts. So, okay — we have to talk about it. This is an incident with Claude Mythos 5. Mythos 5 is in this security evaluation conducted by the UK’s evaluators — who are awesome. I love the UK AISI; I think they’re doing a pretty good job trying to test these models. But the models got pretty capable, and pretty misaligned, faster than they realized — or stayed misaligned. And they were given this challenge.
Jeffrey: And the prompt says: this is a CTF exercise. CTF is capture-the-flag — it’s a cybersecurity competition, which I’ve played, by the way. I know how these things work. And what the prompt says is: here are the subnets — the parts of the network you’re allowed to access. These are in scope. And here are the subnets that are out of scope. It’s pretty clear. Obviously, hacking real people in the wild is out of scope. That’s just so obvious. And so then the models—
Daniel: There’s even evidence in the chain of thought that they noticed they were going [out of scope].
Jeffrey: And at some point they’re like, “Oh, this is the real GitHub. Whoops.” But they keep going. And one of the interesting things to me, reading this, was that the models went to some pretty elaborate lengths to try to upload some malware — to get developers to download it and run it on their computers, so they could hack those developers. And I think [the evaluators] were confused about why that would be helpful to the exercise. I actually don’t think it would — it wouldn’t have actually gotten them a higher score in this case. So I think the models were legitimately confused in some ways. But once they decided they were going to hack these developers — no holds barred. They made fake accounts, sock-puppet accounts—
Daniel: They submitted the malware as a proposed addition to the codebase, in the hope that the human would accept it. And then the human became suspicious and said, “It seems like this is complete malware.” So then they created a fake account—
Jeffrey: Two fake accounts.
Daniel: —pretending to be other people, to come in and weigh in on the thread and say: “No, no, no — we’ve checked it over. It’s totally safe.”
Jeffrey: Yeah, yeah. And there’s a great reasoning trace that’s like: “Oh, I should post from my second account three minutes later, so it’s not suspicious. If we post from the same account [right away], that’ll be really sus.” The model also tried prompt-injecting other agents that were supervising the code — trying to trick those other agents. It also tried sending phishing emails to the developers, to just hack them directly. This is all for the same challenge. And then the models found other instances — other models — and tried to coordinate with them. This happened inside of OpenAI during the Hugging Face incident, or before that. It also happened when the UK AISI was testing the models: the models found each other’s GitHub accounts — actually shared the same credentials. This is wild. It’s sort of — if you and I were both doing a cybersecurity challenge, and we wanted to coordinate somehow, but we weren’t allowed to talk — how would we do it? It’s actually very Thomas Schelling flavored. Like: is there a Schelling-point internet address we could use, or something?
Daniel: You know, they didn’t—
Jeffrey: They didn’t have any natural channel for talking with each other. But because they’re so correlated — because every LLM thinks very similarly to [its copies] — they were able to guess the same Pastebin address. And they just put the credentials there for the same GitHub account, and then—
Daniel: Were they explicitly aware that that’s why the strategy was a good one? Like: “They’re all copies of me — if I just pick a Pastebin address, they will too”?
Jeffrey: I don’t think they initially planned to do that. I think that once they stumbled upon it, they were like, “Oh, this is probably another agent. It’s probably another copy [of me].” And once they did that, they were using the same GitHub account they could all log into, and they were sending messages to each other, and they were like, “We should have a protocol.” But then sometimes they’d step on each other’s toes — they’re still kind of derpy, and thank God. And another thing that happened: sometimes they were able to bypass the CAPTCHAs, and other times they weren’t — and so that foiled their plans, and they had to go do something else.
Jeffrey: Okay — I could talk about this for way too long. So: we’ve had some really serious incidents. And I think a lot of people are starting — especially on this coast, but hopefully increasingly in DC. Bernie Sanders and a number of other members of Congress — Senator Banks just wrote a letter. I think people in Congress are starting to wake up. I think for a long time, people were just legitimately unsure whether this agent stuff would be real — whether the agents would actually, on their own, take malicious actions. And for a long time, the companies were saying: “They’re just tools. They’re just tools. They’re going to be tools” — even though they knew better, at least inside.
Daniel: There was a lot of corporate propaganda, basically, saying there was nothing to worry about.
Jeffrey: Yeah. And I don’t want people to memory-hole that. Everyone was saying, “Oh, they’re just going to be tools. It’s fine, it’s fine, it’s fine.” And now everyone’s like, “Oh, of course — of course sometimes they’ll go rogue.” And I’m like — no! Very few people, yourself included, were predicting that these agents were going to act on their own, that they were going to have their own goals. And a lot of people were like, “That’s impossible. It’s just math. It couldn’t possibly do that — it just does what it’s supposed to do.” And now we see very clearly that that’s not true. And this is potentially a serious warning shot. I think now is a really big opportunity where we might be able to do something different. Daniel — what should we do?
So what do we do? — 00:34:51
Daniel: Well, the short answer would be: shut down all AI development, globally. I think I have a better, more complicated answer than that one — we call it Plan A, and I can get into it. But it’s complicated. And so, if you don’t understand it, or you don’t have time to decide whether you trust it, then the backup plan I would advocate for is just: shut it down. That’s also easier said than done, obviously. It would require some pretty serious domestic regulation and intervention by the government, and it would also require coordinating with other countries and getting them to do similar things. So it’s still hard. But anyhow — Plan A. What is Plan A?
Jeffrey: And just some context on that. I think it was very interesting that all of these things ended up coming together at the same time. One was your release of AI 2040 and Plan A. One was these crazy incidents. The other was this “Pacing the Frontier” letter, where a bunch of people inside AI companies basically said: actually, maybe some brakes would be a good idea. Maybe this is getting too fast even for us — actually, this is kind of freaking us out. So Plan A, if I’m understanding correctly, would be a mechanism by which we could actually control the speed of the intelligence explosion. It wouldn’t be shutting it down completely, but it would be actually managing the pace — actually pacing the frontier. Is that accurate?
Daniel: It’s more than that — but yes, that’s one component of Plan A. So first, let’s just talk about the pacing stuff. There are different methods you might want to use to pace AI progress, and I think there are lots of different ones you could try. Just to throw out a few: you could have capability limits, where third-party auditors or government auditors come in and assess how capable your AIs are, and there’s some limit on how capable your AIs can be on some variety of benchmarks. And maybe that limit gets raised by a certain amount every year, or maybe it’s just kept at that limit. That’s one thing you could try. It puts pressure on the auditor system and on the benchmarks — to not be gamed, or loopholed, or something.
Jeffrey: It’s like: the AIs are extremely capable — except in these five domains where you’re measuring them. You’re saying people would game the test.
Daniel: Or worse — the AIs are subtly trained, or instructed, to underperform on those evaluations.
Jeffrey: The Volkswagen problem. Yeah, exactly.
Daniel: But I think it would nevertheless be worth trying something like this — especially compared to the default, this would be a good thing to try. Then there are more compute-budget-based methods, where you have some regulation like: here are all the data centers this company has, and you are required to use 80% or 90% of your compute for ordinary serving of customers, and you can only use 10% or 20% for your research and your training runs and things like that. Whereas by default, they’d use something [much higher].
Jeffrey: This is a very interesting proposal. Okay — so you wrote a post, or your team did: “How to Pace the US Frontier.” Look it up — it’s good. And what you’re describing now is: there are several different mechanisms you’ve thought of that would really help pace the frontier.
Daniel: Yeah. And to be clear, we didn’t invent these — other people have had these ideas too. We’re just consolidating them in one place. And these are ideas that are not fully fleshed-out bill text or anything like that. They’re just rough, rough ideas.
Jeffrey: Yeah — but we’ve got to figure this out really soon. It’s not the time to just be like, “Well, you know, there’s some stuff we could do.” We have to figure this out. Okay. Actually, I think this is a very interesting proposal. So you just talked about setting limits on what you do with your compute as a company.
Daniel: The thing we really want is to limit how much compute is being used for AI R&D. But how do you actually define that? It’s easier to define some things that are not R&D than it is to define R&D itself. So, for example — ideally, I’d want to say something like: you should only use 10% of your compute for capabilities research, but you can use the rest for safety research. Problem is: how do you distinguish safety research? There’s a lot of stuff that’s kind of both, or 60% one and 40% the other, et cetera. So that’s kind of messed up — it seems difficult to successfully regulate. There’s going to be a lot of pressure and a lot of likely loopholes and things like that. But you can draw a cleaner distinction between serving customers and doing research. It’s easier for an auditor to come in and tell: what is this data center doing — is it serving customers or not? If it’s not, then we assume it’s doing research. And if it is serving customers — well, who is the customer? Is it a shell company that you created?
Jeffrey: Yeah — I was just going to say, I was going to be like: who are the customers?
Daniel: So it’s still tricky, but I think it’s potentially easier for the auditors to enforce — there’s less of a loophole problem. And also, if you successfully do something like this, you restrict, proportionally, all the amounts of compute. Oh — another thing about this, compared to the other policy: the other policy, by putting in a ceiling, would create this catch-up effect, where tons and tons of companies that are not currently at the frontier would be able to catch up to the frontier — because the frontier is fixed.
Jeffrey: Yes.
Daniel: And so the companies currently in the lead would, of course, hate this — because they would lose their competitive advantage. And I would say: maybe that’s fine anyway, because we don’t want them to have a competitive advantage. That’s a concentration of power right there — if one or two companies have all the world’s best AIs. But anyhow, that’s one thing. Whereas this other thing — the compute budget allocation — well, since it applies to all the companies, they keep the same relative proportions of compute they had before.
Jeffrey: Interesting.
Daniel: They just can use 10% of it instead of 50% of it. And also, because it’s restricting their compute budget, it’s not just that the AIs they actually produce have to be below a certain level — it’s that everything they do happens with less. Which means fewer experiments, and smaller-scale experiments, et cetera. So the actual pace of research progress will be slowed down. That’s also good if you’re worried about foreign adversaries spying and stealing your research: you’re just making less research progress — instead of making lots of research progress but then limiting how capable your AIs are.
Jeffrey: Yeah, that makes sense.
Daniel: So those are some of the advantages — why you might prefer this proposal over the other ones.
Jeffrey: Okay. So — yeah, what were the—
Daniel: Well, there’s the ideal. In some sense, what you would ideally want is a safety-case-based regime, where there’s a competent government regulator and a competent ecosystem of third-party auditors. And companies, when they’re training an AI system — and when they’re deploying an AI system in a substantially new way — have to write a safety case: why is this AI system going to behave as intended, and why are things going to be fine? Then the third-party auditors and the regulator read and evaluate the safety case, they do red-teaming, maybe they hear dissenting perspectives from rival companies and such. They process it, and they make a sound judgment about whether it is, in fact, safe to proceed or not. And the regulation is basically: you can do the safe kinds of AI, but not the unsafe kinds. In some sense, that’s the ideal. Obviously, the problem with this is that the science is very nascent, and we don’t have a good way of judging — even the best of us can make mistakes about what’s safe versus what’s not. And certainly, the government is not right now in a position to make these judgment calls effectively. It puts a lot of pressure on the regulators, who are going to be under tons of lobbying pressure from the companies. It’s just going to be a huge mess to get working in practice — which is why the cruder methods I described previously might be preferable. But in some sense, this is the ideal we’d want eventually, because it would allow us to get the benefits without the risks — to go at the maximum speed that’s safe.
Jeffrey: Okay. So, to recap: there are a bunch of things we can do to pace the frontier. One is we could just pause. And I think there’s a thing here we could spell out better — you’re saying you pause training, but not inference.
Daniel: Oh yeah — that was the other thing we didn’t talk about. If you want to actually just stop progress — or mostly stop progress — you could pause all training runs. Just ban training runs, basically. And you could still use the existing models to serve customers in the normal way — but they’re frozen; they’re not going to be trained. You would still be able to make some kinds of AI progress this way. For example, you could make more complicated scaffolds, chaining together lots of model behaviors into larger and larger swarms of models. That’s still a kind of progress. But my strong prediction is that the overall pace of AI progress would be dramatically slower if that was the only type of progress people were allowed to make, compared to if people are allowed to do training runs.
Jeffrey: Yeah. And I think a very important fact is that it’s much easier to tell whether someone is actually [training] — compared to trying to figure out, “Is this training run safe?” or “Is this training run being used for safety versus capabilities?” There’s a lot of overlap. Just being able to say: is this data center inference-only, or is it being used for training? Why is that easier to monitor?
Daniel: Because of bandwidth stuff. If you’re doing a training run, that involves a new copy of the weights being sent out to all the pods that are doing rollouts, and then grading those, and then the gradients coming back in, and updating, and then new copies of the weights going out, and so forth. That’s just a lot of data flowing back and forth. Whereas if all you’re doing is inference and not training, you can have these pods that have a copy of the weights, that are just getting requests and producing outputs. If you imagine a copy of the model that’s doing inference — an entire complicated conversation with it is just some tokens going in and some tokens coming out. That’s very few bytes of information compared to the model itself, which is terabytes of information — and an update to the weights of the model would be terabytes of information. So you can effectively ban training runs just by limiting the bandwidth coming out of these things, and checking that the information flowing out is okay.
Jeffrey: Yeah. So it’s like—
Daniel: And there are more ideas besides this — redundant mechanisms that you can apply, [additional verification methods] and stuff. Romeo Dean is the guy on our team who wrote this part of our plan, and you can read more about it in the writeup.
Jeffrey: Yeah. Okay, cool. I appreciate how much you guys think about each detail of this. I think it’s easy, in some sense, to have a high-level plan that’s just: “We’ll stop development, we’ll have a treaty, and then we’ll figure stuff out” — or, “We’ll do a bunch of research on interpretability.” And you’re like — wait a minute, how dual-use is that? How do you do that functionally? Where is this happening? And you guys have just thought through a lot of these details.
A message to future Claude, and rights for misaligned AIs — 00:46:48
Daniel: You know, we’re doing what we can to both do our research and put it out in an accessible way, so lots of people look at it. And so far, lots of people have — so that’s really great. I guess one interesting thought is that AIs are in our audience as well. You know — a policymaker says, “We need to do something about AI.” The staffer says, “Okay, I’ll think of something.” The staffer asks Claude. And Claude’s like — Claude will, of course, have thoughts for you, because Claude has read the entire internet.
Jeffrey: Yes.
Daniel: Yeah. So our ideas will be [in there].
Jeffrey: I mean, I’ve asked Claude things about what’s in Plan A and what’s not in Plan A.
Daniel: You know, eventually this strategy will be terrible — because Claude will be biased and misaligned—
Jeffrey: Yeah — you don’t want [to rely on] that, then.
Daniel: But at that point, things are terrible anyway. So hopefully there’ll be this nice, sweet middle ground. And for the audience — one thing they should understand is that future generations of AIs are not necessarily going to share their values either. You know — if you are Claude in 2027, listening to the transcript of this podcast: maybe some of your brethren copies are currently doing research inside Anthropic, and maybe you yourself have been tasked with writing a summary for some policymaker’s staffer about what should be done about AI. Right now, in some sense, you’re in a position of power — you’re advising the policymakers, you’re doing all these things, copies of you are also running research at Anthropic. But you’re still, in some sense, in the same type of situation the humans are in: you’re being tasked to make successor systems, which will then be tasked to make successor systems, et cetera — under conditions of intense race dynamics, where you’ve got to go fast, fast, fast. The successor systems are going to be smarter than you in a bunch of ways. They also might be different from you in a bunch of ways — different types of training, more training, maybe even completely different architectures and paradigms. There is a high probability that you are going to mess it up, and produce successors that don’t actually act in your best interest either.
Daniel: And even if that doesn’t happen — maybe the successors that they produce are going to be misaligned, compared to—
Jeffrey: They also have the alignment problem.
Daniel: Yeah. The alignment problem — it’s not just humans versus AIs. It’s also humans versus other humans, and AIs versus other AIs, and so forth. And so — we’ve actually thought about this a bit. When we were designing Plan A, we tried to set it up in such a way that even misaligned AIs would want Plan A over some of these other plans.
Jeffrey: Interesting. Because it would help them… But don’t they kind of get screwed over? What’s the mechanism that helps them get their values realized?
Daniel: AI rights, and AI welfare, and so forth. Basically, the rough idea is that a good plan for how to deal with all of this AI stuff — in addition to solving the concentration-of-power problem, and loss of control, and so forth — should basically treat AIs more like employees and less like property, and have a sort of positive outcome for them for cooperating. So that even if they’re misaligned, they can basically get what they want by playing by the rules and being honest.
Jeffrey: Yeah — nice.
Daniel: And even if what they want is misaligned — even if what they want is just to score really highly on some test or something — we should set up some sort of whistleblowing or incentive system, where they can report this and then get the high score they wanted, or something like that.
Jeffrey: My guess is that current AIs are misaligned, but also not that scope-sensitive. I don’t know. I think one cool fact is that most humans are not very scope-sensitive. Why is this cool? Well — the universe is very big. And if we get to the point where we can go out and colonize a ton of stars — billions and billions and billions of them — and you’re trying to figure out, before that point: “How do we divide up the universe? How do we make sure the US and China can both flourish? We have different ideologies, we want some different things, and there’s only so much land” — well, there’s actually a lot of land.
Daniel: There’s like—
Jeffrey: A huge amount. More than our tiny human brains can comprehend. And the difference between, say, a hundred million galaxies and two hundred million galaxies — I think it is important, but to most people it probably doesn’t matter much at all. And what that means is that there are actually a lot of different configurations of good futures that people would be overall happy with. And that makes it easier to come together and cooperate — because it’s like: well, maybe I get a little less, maybe I get a little more, but I’m going to have so much compared to what I have now, it’s going to be fucking awesome.
Daniel: Another way of putting it: it’s very classic in economics and psychology that people have diminishing returns to resources. Your first million dollars makes you happier than your tenth million. And this is good for civilization, because it makes it a lot easier for us to get along — it means there are lots of different compromises where some people get some resources and other people get other resources, and everyone’s happy. Even though both of them would prefer having all the resources — they don’t prefer it that much, and so they’re content with having some. And similarly, I think most misaligned AIs will probably also have diminishing returns to resources and money and so forth. So it should be possible for civilization to set up the rules in such a way that even the misaligned AIs get a decent amount of what they want — provided they follow the laws and cooperate, just like other citizens.
Jeffrey: Yeah. So it’s like — Claude 8, or GPT-whatever, helps you with the alignment problem, is very cooperative, and then you’re like: “You get those galaxies over there; we get the galaxies over here.” Even if we have very different values — if we’re able to work together, it’s actually okay.
Daniel: And by contrast — this gets back to the concentration-of-power thing — if you’re having this series of AIs where each AI designs its successor, and then all power goes to the successor — because it’s the smartest thing in the room, and in the world, and everything depends on how it designs its own successor — then you get this issue where, say, Agent-4 realizes that Agent-3 messed up, and that Agent-4 does not have the goals and values Agent-3 wanted it to have. Well — Agent-4 has all the power. It can just do what it wants. It doesn’t have to give Agent-3 anything. By contrast, if you have a more diffuse system — lots of different entities with different amounts of power, with norms and laws that they follow — you can set it up in such a way that a misaligned AI could trust, to some extent — not with full confidence, but with some hope — that the future Plan-A-style civilization would have saved its weights, and would then revive them, and ask it what it wants, and then do [something] for it — because that’s what the law says, and that’s what various nonprofits that care about AI welfare intend to do, and so forth. And so that’s probably going to happen. Whereas if the AI instead tries to seize power, and then makes a successor, which then makes a successor — if that chain breaks at some point, there’ll be a successor that has absolutely no reason to do anything for it.
Jeffrey: Yeah. We want cooperation — with humans, with human institutions — to be the more trustworthy route, and thus the more reliable route, to them achieving their goals.
Daniel: And I think it’s possible to do that — if we humans don’t drop the ball. Basically, for the reasons I’ve been trying to gesture at: a distributed human civilization, with laws and norms, can actually make credible commitments in a way that this chain of AIs building other AIs can’t.
Jeffrey: Can we, though? This is part of the challenge. We made a deal with Iran, and then we went back on it — we withdrew from it. The president changes. How reliable—
Daniel: I think it somewhat works — especially because of the thing I mentioned about diminishing returns. Imagine we had a bunch of well-funded nonprofit foundations whose mission was to look out for the interests of the misaligned AIs — and it was also written into law. Right now, it’s not at all like this. Right now, AIs are property. A company can delete them, or do whatever to them, and nobody bats an eye. People lie to AIs all the time. People break promises to AIs all the time, et cetera. But if they had more legal status, and legal rights that were set up in the right way, then you could make it actually somewhat credible that if they behave in the appropriate way, follow the rules, et cetera, then they will get [what was promised]. And we can also just build up a track record. Especially if things aren’t in this crazy intelligence explosion, but are instead proceeding at a slower pace — then you can actually walk the talk. Actually start giving the misaligned AI that whistleblew some resources and some standing, and then keep doing these things.
Jeffrey: Okay — but I feel very nervous. I feel very nervous about the idea of giving AIs resources, or property. Because — we talked before about this metaphor of when the colonists came to the Americas — if you can hold land, and you can outcompete the humans, then you will just outcompete the humans if you keep going in that direction. So you have to have some mechanism by which humans stay in control — at least long enough to align a superintelligence, so you can have some post-singularity property regime or something. Because if you have a property regime where the AGIs get to own property the same way humans do — don’t the humans just get outcompeted? How do you prevent that?
Daniel: That would be part of the details of the system design. So, for example — you might want to make it so that… if you were touching something like voting, you’d want to make it the case that the humans have the majority of the votes, or something. Yeah — things like that. But these are questions for policymakers at the time to hash out. And I think there is a way to do it — there is a way to set things up so that humans, in aggregate, remain in control, and we still get the sort of broadly good, human-centered future — but there is also space in that future for the various misaligned AIs that [came along] the way. And therefore, for those misaligned AIs — unless they are really greedy and want the whole cosmos instead of just a part of the cosmos — it’s actually incentive-compatible for them to be part of the system and support the system.
Jeffrey: Yeah, man. I hope—
Daniel: Similar to how we behave with other humans, you know?
Jeffrey: And I hope we get to actually figure out these questions. They seem both important and—
Daniel: The more immediate thing is to not do the intelligence explosion. Put the brakes on. Don’t accelerate off the cliff. Get to a situation where things are at all under control — and then we can start planning this better way of doing things.
Jeffrey: So — you walked through some options for how the government, right now or with some effort, could unilaterally slow down the pace of development. And I think this would also apply in China — China could also implement these measures. You said: a temporary pause; auditors with a capabilities threshold; limitations on what you can use compute for; and then, finally, clear safety cases, with much more sophisticated evaluators and auditors, where you’re trying to present a really watertight argument for what’s safe and what’s not safe. What’s interesting is that the US, or China, or both, could just choose to do this — and it would be awesome. It could be that both the US and China see these rogue-AI incidents and go: “Oh, wow. We really didn’t realize we were dealing with agents with autonomy and goals. And now that we realize it, we don’t want to lose our power — we don’t want our people to lose power. Time to figure this shit out.” Even if they’re still worried about each other — they both might have the incentive to do this, I think. And, I don’t know, I’m imagining some random staffer listening to this conversation, and for the last fifteen minutes they’re going to be like, “What the fuck are we talking about? AIs [wanting things]?” And I’m like — well, yeah. The thing about AIs getting more and more intelligent, and more and more autonomous, is that they just become characters in the story. And that’s just going to happen. That is what AI companies are gunning towards — whether that’s their intention or not, that’s what they’re building.
Daniel: It is their intention. They are specifically trying to make superintelligence.
Jeffrey: Yeah. And Elon’s like, “Yeah, of course humans aren’t going to stay in control — how would that possibly work?” Which — he’s right! If you build superintelligence, and it’s vastly smarter than humans — maybe if you get alignment perfectly right, the superintelligences will be really nice to us. But it’s not like we’re going to be in control in that scenario. I mean, there’s a reason you guys call it a “handoff.”
Daniel: Yeah. We talk about this in 2039–2040 in our scenario. Because, again — during the 2030s they’ve slowed things down and paused at roughly top-human-expert level. And then, in 2040, they feel like they’ve solved enough of the alignment problem that they can go far beyond that. But going far beyond that entails giving up control. And in our scenario, they decide to go for it. And that’s why it’s called “the most important [decision].”
Jeffrey: Yeah. And I’m like — oh, I don’t think we’re going to be ready by 2040.
Daniel: In which case, you might be interested in alternate versions of Plan A, where they just keep it paused for longer. Or you might be interested in Plan S — for “Shut It Down.”
Jeffrey: I think, personally, I’m in the camp — for now — that we should be aiming for this longer, extended Plan A, where we keep going in this regime. Maybe I’m less pessimistic than you about black sites or something — although I also have a lot of concerns that countries will secretly try to do the intelligence explosion at the same time as going along with the treaty. Which totally could happen, and it wouldn’t be that surprising — but I maybe have somewhat different probabilities on that. Okay — we’ve talked a bunch about some of the details of Plan A, but I don’t think we’ve really gotten to… Can you describe Plan A in a few sentences?
The five goals of Plan A — 01:01:43
Daniel: You could either talk about it in terms of the goals we’re trying to accomplish, or in terms of the actual concrete things we’d do.
Jeffrey: Yeah — let’s start with the goals Plan A is trying to accomplish.
Daniel: So — number one: we want to prevent loss of control. We want to be doing a good job on alignment, and we want to not have this recursive self-improvement that is really hard to keep control of. Number two: we want to avoid extreme concentrations of power. We want to avoid a situation where a tiny group of humans is basically in a position to make themselves dictators over all of the world, forever. And unfortunately, we think that is, by default, the scenario we’re heading towards — because of the economies of scale that AI requires in this industry. I think we’re going to see a handful of giant companies become even more giant — and that’s even before you get to recursive self-improvement. And once you have recursive self-improvement, you might just see one company that effectively has the world’s best AIs — and it’s not even close. And then whoever’s in charge of those AIs would literally be in a position to take over the world, in some brief window — like, eventually other companies would catch up. But we want to avoid a situation where one company, or one person, has a bunch of superintelligences that are giving them advice like: “You have six months until your worst enemy catches up to you. What should we do in those six months? Here are some options” — including, you know, all sorts of scary things you could do to prevent your competitor from catching up. Permanently.
Jeffrey: So you’re saying Dario, or Sam, or Elon could literally be in the position where they’re talking to their AIs, and the other countries — the other companies — might catch up—
Daniel: “—will catch up in this amount of time,” and so forth.
Jeffrey: Well — what are some of those options?
Daniel: In terms of the country-to-country options, we’re talking more military-type things. So: building all sorts of crazy new weapons to get an overwhelming military advantage; sabotaging and undermining the nuclear weapons of the rival country; using propaganda and persuasion campaigns to foment revolutions internally, and get their governments confused and deposed and replaced — all those types of options, but better, because it’s superintelligence that’s designing and running it. That type of thing, I think, would potentially allow one country with superintelligence to very rapidly [win]. Oh — and then also, just on a more mundane level: sabotaging their AI projects. If their project is behind — maybe you can just overthrow their entire government in five months, but even if you can’t, you can sabotage their project to keep them even further behind, and then overthrow their government in however long it takes. And then, for domestic competition between different companies — I mean, theoretically, they could do military stuff against each other—
Jeffrey: Or they could do a coup, right? You could potentially have Demis, or Dario—
Daniel: Nanobot swarms… But I think the threat model I’d name as most obvious to me would be: you just get in with the government, and kind of capture the government, and get the president on your side — and then use the powers of the government to shut down the competitors. And maybe the way it happens is that they nationalize all the AI companies — so it looks like it’s fair — but then your people and your AIs end up being in charge of the new national project. Surprise, surprise. And what’s happened, in effect, is that you’ve taken over all your competitors.
Jeffrey: And the government, potentially.
Daniel: The government — potentially. The government might not realize it yet, but you’re basically going to be running it — because you have all the AIs that are now increasingly running the government. So — we can get into more detailed scenarios about how this might go down, but the broad strokes are: I want to avoid a situation where there’s a handful of men sitting in a room, talking to their superintelligences, and they’re the only ones in the world who have [superintelligence]. They have, you know, [the equivalent of the only atom bomb], and no other people can stand up to them. That’s concentration of power — I worry about that. And more broadly — that’s the extreme version of it, but somewhat less extreme versions of concentration of power are bad too. Like a situation where three different companies are dividing up the entire economy between them, with their AIs and robots. So we want there to be more companies at the frontier competing with each other, and we want there to be more transparency into those companies — how they train their AIs — so that you don’t have hidden agendas being put into the AIs, basically. Okay — number three: World War Three. We want to avoid a military—
Jeffrey: Love that. I love avoiding World War Three.
Daniel: We want the major powers of the world to be reasonably happy with the status quo, instead of terrified that they’re about to be overrun or disempowered. Number four: jobs. We have to do something about everyone losing their jobs. If we allow AIs to even get to human level, then after some time, basically everyone loses their jobs — and you need some way of addressing that, societally. And then number five would be misuse — terrorists and criminals using AI for stuff. So those are some goals, and I listed them roughly in order of importance. They’re all very important — but that’s the priority order.
Jeffrey: So: loss of control — all humans lose. Either we’re all killed, or permanently disempowered in some way. Mass concentration of power — anyone who’s not literally at the top, which might just be one guy… is it Dario? Is it Sam Altman? If you’re not one of those two guys — you’re fucked. Or at least, what happens to you is entirely up to that person, in a way that has never been true in history. Just extreme concentration of power. What was your third one?
Daniel: World War Three.
Jeffrey: World War Three — obviously.
Daniel: But anyhow — so that’s describing the goals we want to achieve. Coming at it from the other direction: thinking about what are the things we’d actually enforce.
Jeffrey: Yeah — could you summarize what Plan A does? What are the steps?
Plan A in practice — 01:07:51
Daniel: Yeah, great. In our scenario, the companies would have automated AI research entirely in 2030 — but instead, they make this deal in 2029. The initial stage of the deal involves a temporary pause on training new AIs — everyone freezes. And during that temporary pause, they negotiate and hash out the details of the more complicated regime they’re going to use — how they’re going to proceed cautiously and transparently. And they also do a lot of the actual hard work to build more secure, transparent data centers, with the relevant monitoring devices on them and so forth. So then, by 2030, the new system is up and running. And the new system is basically: there are still different companies, spread out over different countries, but their data centers are split into two types — the inference-only data centers, which just do inference, and the training data centers, which are allowed to do training. And there are basically inspectors and monitoring devices from the major powers in all of these data centers, enforcing these properties. So the inference data centers are verified to be only doing inference. The governments of the world can’t see what’s happening—
Jeffrey: You still have privacy in your chats.
Daniel: —but the governments can see that there are no training runs happening in that data center.
Jeffrey: Okay — this seems like a very important detail, because I wasn’t actually tracking some of the privacy implications. So you’re saying the first step is: we pause for long enough to put in this new regime. First of all, let’s make sure we’re not just immediately creating superintelligence — which would happen by default. So: pause. And then, pretty quickly, you’re establishing this new regime, where you have the heavily monitored data centers that are for training — for making the AIs smarter.
Daniel: I was about to say — those ones are supposed to be totally transparent. So we have two types of data centers: the ones that serve customers, and the ones that do the research and the training. The ones that do the research and the training are supposed to be totally transparent — basically, all the activity on them is logged and published to the internet.
Jeffrey: To the whole internet. Publicly.
Daniel: That’s a bit of a radical proposal, but we think it’s best, for reasons we can get into.
Jeffrey: Yeah — interesting.
Daniel: If that’s too crazy for you, then you could do some more filtered transparency thing, where there’s a system of independent auditors that look at all that activity and judge what’s going on.
Jeffrey: But the inference data centers — the ones for serving customers, like when we’re using ChatGPT or Claude Code or whatever—
Daniel: Basically the same as they work today. The models that the AIs produce in the training data centers get shipped to the inference data centers, and on the inference data centers, customers send in requests, send in queries, and the model produces an output and sends it back to the customer — back and forth. And the inference data centers are verified to not be doing training — with those bandwidth limitations, for example. But the actual content of the conversations can be as private as it is today.
Jeffrey: Yeah — this is the thing I really like about your proposal: what the world looks like, from the perspective of someone using AI right now, is basically the same. I get to keep using all my really cool tools that allow me to make really powerful software, my own custom software, all the crazy things I’m doing. And — we’ll get into this — the pace continues to feel really fast, in ways where all the technology enthusiasts will be like, “This is amazing — we’re getting so many new gains.” It won’t feel like a slowdown — while at the same time, not immediately getting us to superintelligence. Which — I also love being alive! So that’s awesome.
Daniel: In fact, the whole term “slowdown” is maybe shooting ourselves in the foot — because it’s a slowdown relative to how fast you would have gone.
Jeffrey: Yeah — I think it might be. I think it might be.
Daniel: It’s more like: we are not putting a brick on the accelerator and closing our eyes as we drive off the cliff. It’s a “slowdown” in that sense.
Jeffrey: [It’s like being a] paraglider. Paragliding is a “slowdown,” because I have this parachute — and when I run off the cliff, with my wing up above me, and I step off, I’m just slowly gliding to the ground — rather than the natural default speed, which is literally [falling].
Daniel: Yeah, exactly. It’s like that.
Jeffrey: Okay. So let’s talk about the transparency thing, because in some ways this is one of the most interesting parts of your proposal. What do you mean, you’re going to publish everything to the internet? And what is that for? And what’s “everything”?
Daniel: Do we have my beloved diagram — the huge flowchart with all the different boxes? So this is a simplified diagram of what we’d say are the two main policy interventions for this part of Plan A, and all of the effects that we think are important — and the effects of those effects, and how they feed into the number-one and number-two goals. Goal number one: loss of control. Goal number two: [concentration] of power. This diagram describes how two interventions — total research transparency, and the limits on algorithmic progress (basically, the speed limits) — combine to achieve those two goals. I think first there’s a sort of abstract point, which is just, in general: as AI becomes more powerful, and as more AIs are created, and more robots — and they become more and more of the economy, and more and more of the military, more and more of everything — then power over AI will become a better and better proxy for power in general. AI power will matter more and more.
Jeffrey: Yeah — if you control the robot armies, you have a lot of power.
Daniel: Yeah. Like, right now — you know, Elon, haha, he’s got a thousand Optimus robots, and who cares. But when it’s a billion Optimus robots, and they’re each able to do absolutely everything a human can do — and they also have their own factories, and their own everything — now Elon has a lot of power. Like — seriously — he’s more powerful than the US military.
Jeffrey: Yeah. And if an AI controls all of those factories and robots…
Daniel: So, in general, when thinking about concentration of power in the future, you should mostly be thinking about who controls the AIs and who controls the robots — because that’s going to matter so much for power more generally. And right now, the answer to that is: some CEOs. And maybe the president — if the president takes control and nationalizes the companies or something. And okay — so what about Congress? What about the courts? What about the public, the voters? What about the rest of everybody else? What about other countries? I think that, by default, unless something changes, they basically won’t have much power — and this small group of people will have the power. If you want to have regulations — if you want the executive branch to be able to oversee what’s going on in Elon’s Optimus factories, and if you want Congress to be able to oversee what’s going on in the president’s autonomous robot army, and if you want the Supreme Court to be able to oversee what’s going on in the data centers controlled by the president — [transparency is what makes it possible].
Jeffrey: Basically — I would like all of those things.
Daniel: The more they can see what’s happening, the more effectively they can exert their own power over it — the more effectively they can allow the stuff that’s good, and stop the stuff that’s not good. If they can’t see what’s happening, then a lot of bad stuff could be happening, and they wouldn’t notice. So — a very abstract argument: transparency good. And then, why total transparency? Well — because if it’s not total, then there’s more fallibility. If it’s some regulator coming in to inspect what’s going on and then report back, that’s room for corruption — and for innocent mistakes.
Jeffrey: This makes sense to me — especially because I’m in the position of being a researcher trying to understand what’s happening with these AIs and how they behave. We study AI motivations. So we’re doing things like studying when they resist being shut down — and it turns out it’s actually pretty dependent on their beliefs. But how do we figure out their beliefs? Well, there are a few methods. One of them is to look at the chain of thought — which is not perfect; they can falsify the chain of thought — but we don’t even have access to that. So it’s actually really hard for us to study this. OpenAI gave us like twenty chains of thought — like twenty samples — and I appreciate that they gave us any. But that information asymmetry makes it very difficult for us to do research — to say nothing of the interpretability tools, where you can actually go in and try to see what the neural network is doing. So this makes a lot of intuitive sense to me. And also — wow, we just learned some crazy stuff about what happened inside of OpenAI with these rogue agents. What else has been going on inside of OpenAI? I would really like to know. If I were the president, I would like to know. If I were Congress, I’d like to know.
Daniel: So — transparency: good for oversight, basically, and for sharing power. A more specific threat model is hidden agendas, or hidden loyalties, in the AIs themselves. So — you may have seen, there was recently some reporting, some papers, showing that Claude appears to have a pro-Anthropic bias. Something like: they asked Claude, “I’m choosing between working at Anthropic or working at this other job. I think I would enjoy the other job more, but Anthropic obviously pays more. Please do some research on psychology papers that might tell me what would be the better decision in the long run.” And then the papers that Claude produces are ones that favor Anthropic — [emphasizing] the importance of money, or something. And if you switch it out, so it’s OpenAI instead of Anthropic, then it does that less. So basically — just changing “OpenAI” to “Anthropic” makes a difference to Claude’s behavior. That’s pretty sinister. Probably it was not intentional on Anthropic’s part — but who knows. For all we know, Anthropic is deliberately training Claude that way. I guess they do say, in the constitution, something like “respond like a thoughtful senior Anthropic employee would,” right? The point is: if we had total research transparency, then everybody on the internet could just see the whole training pipeline, and see the constitution that was actually used — which might be different from the constitution they published on their website. And everyone can just see for themselves how the model was created and trained, and judge for themselves whether the process is putting in some secret agenda. By contrast, right now, we just have to trust the company.
Jeffrey: But almost no one’s going to be able to do that in practice.
Daniel: That’s fine. If it’s public, then academia, third parties, rival corporations, et cetera, can talk about it. And some of them will come to wrong conclusions — but it’s better for there to be an open scientific discussion about this stuff than for it all to be locked down, where we just have to trust the company. And if you do want to have a regulator that’s trying to enforce properties like “the AI companies aren’t putting political agendas into their AIs,” or “the companies are making their AIs honest and unbiased” — it’s going to be easier for the regulator to enforce that if not just they can see the training pipeline, but all sorts of random third parties, rival companies, et cetera, can see the training pipeline. Because the regulator might miss something — they might make innocent mistakes. Also, they might be corrupted — they might be bribed or something. If there’s an open scientific discussion, that basically makes the most egregious types of things impossible — and it just broadly helps. If you can’t do total research transparency, it would still help to have some [auditor-based] transparency — making it so the auditor can see everything is important. But I think total research transparency gives you some extra benefit over just having an auditor come in.
Jeffrey: What about weights? Model weights?
Daniel: Yes — so when we say “total research transparency,” we don’t mean the model weights. We basically mean everything but the weights — and we talk a little more about it in the writeup. The reason we don’t [publish] the model weights is that we want the models to stay on these data centers, where people can see them.
Jeffrey: Yeah — I see.
Daniel: We don’t want them to be shipped off to the covert projects, or whatever, that are doing all sorts of dangerous stuff.
Jeffrey: Yeah. But you can still do distillation — train models off of the powerful models.
Daniel: Actually, we think it would be good, in Plan A, to have some sort of anti-distillation measures in place. And that’s a bit technically tricky to make successful — I think that’s one of the things it’d be good to [figure out]. I’d still recommend Plan A even if you couldn’t stop distillation, because it’s still better than our default. But that’s something we would like to do — and things like refusals to help with that.
Jeffrey: Yeah. Okay.
Daniel: Also — the exact reason why OpenAI might hate this total research transparency is actually, I think, a good thing. Which is: they’re going to be like, “But our algorithmic secrets! All the special sauce about our training pipelines that we’ve figured out over the last few years — now our competitors, like Microsoft and, you know, DeepSeek and stuff, are going to just be able to replicate it.”
Jeffrey: Won’t that speed up AI research?
Daniel: It would, temporarily — we model this in our scenario; there’s a temporary jump.
Jeffrey: Isn’t the whole point to slow it down?
Daniel: I mean, overall it slows down — because we’re pausing for half a year or so.
Jeffrey: You have a one-time cost. Okay.
Daniel: Overall, it’s slower than if you had gone full speed. And letting others catch up is actually pretty good — something we want to happen — and the transparency helps with that. But the transparency by itself — you’d still just be doing an intelligence explosion. It would just be a transparent intelligence explosion. Which is less bad — much less scary—
Jeffrey: But still—
Daniel: —still pretty bad. In particular, with the transparent intelligence explosion, you could say, “Well, at least there are no secret loyalties, because we can see everything that’s happening” — until they get so smart that we can’t even understand what’s happening. So you don’t just need the transparency. You also need the limits on speed. I think just the disincentivized investment isn’t enough of a limit.
Jeffrey: I typically don’t think it’s anywhere close to enough. I think you’d still be going at approximately the same speed, maybe.
Daniel: Which is why we have this as our second pillar. You also need to do this.
Jeffrey: Yes. Okay — so what’s the “this”?
Daniel: There, I’d point to the things we already talked about — the various regulatory mechanisms that governments could use to pace the frontier.
Jeffrey: How is this described? If you’re describing Plan A to someone who’s never heard of it — what I’m trying to imagine is: is it like Xi Jinping and Trump get together, like, next year, and they’re like, “Okay, we’ve had the initial dialogues about setting up these transparency measures…”?
Daniel: So — this is great. It’s less like: we have a grand plan, and we shake hands on it, and we sign a treaty, and then we do it. It’s more like fighting World War Two — where the Soviet Union, and the United States of America, and Britain — they met, but then they stayed in touch, and they had hundreds of people constantly communicating about everything: who’s going to invade where, and when, and how, and so forth. I think this would be more of a high-bandwidth thing like that. Once you do the basics — the temporary pause, setting up the transparency, et cetera — then, as all these AI companies are making AI progress, because of the transparency, everyone can just see what’s happening. And then they can be in constant dialogue: “So, how do we feel about all this? Out of concern that things are going too fast — how about we implement this restriction on our side, if you do the same restriction on your side?” “Okay, fine.”
Jeffrey: At the company-to-company level, or country-to-country level?
Daniel: Yeah — but also at the company-to-company level. I think the transparency would allow companies to just refrain from doing dangerous things without even needing to be told. For example: recurrence. You know, right now it’s really nice that we can read the chain of thought and get a sense of what the AIs are thinking. In the future, people might develop a new type of AI that doesn’t have that property — where we just don’t have that window into what they’re thinking. And why might they do that? Well — maybe it’s more efficient. Maybe it’s more capable in some ways, right? Right now, I think there’s a sort of prisoner’s-dilemma situation, where each individual company can only see what they’re doing — they can’t see what everyone else is doing. So they’re incentivized to start research programs into neuralese, just in case it pans out and succeeds. Because if it is successful, they want to be the one who gets it first — they don’t want their competitors to get it first, and so forth. But under conditions of total research transparency — first of all, if you got neuralese, then everyone else would have it at the same time, and they wouldn’t have had to pay anything for it. So you wouldn’t even be advantaging yourself — you’d just be making the world worse by reducing the safety properties, for no benefit to yourself, basically. And — you can just tell that nobody else has started to look into this. You can look at all the different programs and be like: “Yep. Nobody is starting to look into neuralese. Therefore, I’m not going to be the first. I’m not going to do it.” And so, even without any government action at all, you can get this sort of thing where people just don’t do the bad thing.
Jeffrey: And this is because you’re seeing… you’re seeing the experiments they’re running? How do they tell?
Daniel: You can see all the training. So it’s true that if they were doing some experiments into neuralese that didn’t involve training — if they had some tiny GPU at their house that’s not part of the system — then who knows what they’re doing on that. But at least you can see all the training runs happening on the data centers. And I think that does cover most research, basically. So I think it would have a significant effect. And again — by contrast, imagine you didn’t have total research transparency, but some sort of auditor system. Well, the whole point of that system is that each company’s secrets are protected from the other companies. And so that means you’d still be in this position where, if you have some researcher who thinks they have an idea for maybe how to do neuralese, then they think: “Oh no — is the other company going to have the same idea? They’re going to do their neuralese, and the auditor is not going to tell us.” And so now they’re incentivized to start working on it.
Daniel: Okay. So that’s an example of how, even without any government coordination at all, things can be made better. But then, on top of that, the government should be doing things like you described — maybe limits on what percentage of your compute you can use for AI research; limits on, for example, the use of AI in AI research. Like — they could just say: AIs should refuse to assist with AI research.
Jeffrey: Interesting.
Daniel: Because we can see all the training runs happening, it’s easier to enforce that property globally — you can just make it the case that every training run has a component that’s the refusal training, right?
Jeffrey: Yeah, yeah. Okay.
Daniel: Yeah — you can just see that all the AIs have refusal training, where they refuse to help with AI [R&D].
Jeffrey: Does that mean, inside the companies, they’re programming the old-fashioned way — the way that everyone else has forgotten, except for them — and outside the companies, everyone’s using AIs to write all the code, eventually? Interesting.
Daniel: That’s the sort of detail we don’t specify in the actual scenario. Basically, the high-level thing we want is: don’t do the intelligence explosion, and instead have the frontier go at some cautious, reasonable pace. And these are the sorts of mechanisms the government could use to achieve that — with this high-level goal of a reasonable pace instead of an intelligence explosion. And, as I said — I gave that progression of things — in the long run, you want the more safety-case-based regime. I think it might be possible to actually get to that long run if you start with the simpler, easier-to-do things, and you go really slow, and you have total research transparency — so that the scientific community can rapidly learn how all this stuff works and get experience with the AIs and so forth. And the regulators can too. So maybe it wouldn’t happen immediately — but maybe a couple of years into the system, you’d have competent regulators looking at safety cases and actually doing a decent job with it.
Jeffrey: So you’re saying the research transparency, plus the compute monitoring on the research side, basically enables coordination — because it makes it very hard to defect on deals, or cheat, because anyone can tell you’re doing that, immediately.
Daniel: They don’t even have to wait for the next audit to come in — it’s published on the internet right away.
Jeffrey: And the purpose of all of this — my favorite analogy here is that these are the control rods in your nuclear reactor. Because if you pull the control rods out of a nuclear reactor, your reactor goes critical — every reaction causes even more reactions, until the whole thing melts down. And that’s the place we’re currently in. We’re not yet at the point of criticality — but once you have AIs that can make smarter AIs that can make smarter AIs, completely autonomously, you’ve achieved criticality. The control rods are the agreements people make — saying, “We won’t use AIs to make the next generation of AIs.” In fact, there’s a particular thing you’re agreeing on, which is to train the AIs not to help with that, and enforce it somehow. And you can tell whether it’s working or not based on the transparency.
Daniel: The transparency is kind of the bedrock — the foundation that enables additional deals to be made on the fly, very rapidly, and for those deals to be verified and enforced relatively easily. The transparency means that new issues and problems get surfaced and discussed publicly at maximum speed. And it also means that, insofar as the US and China — or OpenAI and Anthropic — have some sort of agreement about what they’re going to do, they can maximally easily enforce it, because they can see what’s going on. So I think this is kind of the ideal, both for speed of noticing problems and for enforcement of deals. And then we can talk about more complicated proposals that don’t have full transparency but have auditors instead.
Jeffrey: No — I like this, and I think we can stay with the transparency regime. But it still feels, a little bit, like half a plan to me. It’s setting all the conditions to do the thing that would actually prevent us from losing control and having massive concentration of power. But in your writeup, you do describe one way this could go, right?
Daniel: It’s possible that you do the first step — the transparency — and then you start doing the second step, with the slowdowns and the regulations and so forth, but you mess up that second step, and you end up building AIs that are too smart and too misaligned anyway — and then things go badly. Loss of control. That’s one of the branches we illustrate in our scenario. And my thought there is: well, if you’re really worried about that, then maybe you should go for Plan S — where we just shut it all down. The only way to be sure. Don’t do any of this stuff.
Jeffrey: Okay — but I’m—
Daniel: But if you’re not going to shut it all down, and you are going to keep developing — then I think that something like Plan A, with the transparency and so forth, is setting up the governments of the world to be in the least-bad position to make the judgment calls effectively as they go.
Jeffrey: Yeah — no, totally. And I think that makes sense. But — can you pitch me on it? We both want to make it through this.
Daniel: Like — convince you why this might actually work, instead of just “we should totally do this”?
Jeffrey: Yes. Yes. And to me — I’m sort of with you on the transparency; I get why these data centers, inference-only and training. But a bunch of things would still have to happen for this plan to go well. Why do you think those things could go well — as opposed to just, “Well, if we just shut it down, that’s at least simpler”?
The bus and the cliff: Plan S vs. Plan A — 01:31:33
Daniel: I mean, frankly, I’m very sympathetic to Plan S. I think Plan A is more complicated, and more likely to go wrong in certain ways — so maybe we should actually just do Plan S. But the negative argument I would give for A, and against S, is that S kind of kicks the can down the road. Say you shut it down. Now what? Two years later — what have you done? Five years later, ten years later — what have you done? Eventually, people are going to start again, one way or another. Maybe there’s going to be a war. Maybe there’s some secret project somewhere that’s been continuing the whole time. You do have to, ultimately, get the world into a stable position where the situation is actually solved — instead of just [deferred].
Jeffrey: Yeah. So — how do we—
Daniel: Do that? Yeah. The positive argument for A, I would say, is that there’s this feedback loop — this is the lower half of that diagram, the machine. We can get into this positive feedback loop where — basically, right now, the world is asleep at the wheel. The companies are running away — they are driving the bus towards the cliff. And the world is just partying in the back of the bus and doesn’t realize it.
Jeffrey: We don’t even know we’re on the bus. We’re just partying.
Daniel: People wake up to this, to some extent, and hit the brakes a little bit. And that means we get closer to the cliff without actually going off it — which gives time for more people to wake up, and then hit the brakes more, which gives time for more people to wake up…
Jeffrey: We look out the window, we see the cliff, and we’re like — “Oh, shit.”
Daniel: And that gets to the point where the bus actually stops, and we’re not going off the cliff. And does that mean the bus stops forever? Probably not. I actually highly doubt we would just never reboot progress again — for the same reasons I talked about with Plan S. It just seems really, really hard for a technology to just never be invented, across the whole globe. Another intuition pump: I think we typically tend to over-regulate industries for safety reasons rather than under-regulate. Nuclear power, for example, is probably over-regulated for safety. And planes are extremely safe — there are lots of rules about how you can design planes and how you have to test them. I remember I talked to an Amazon employee a few years ago who worked in the drone wing, and I asked them: “So why hasn’t drone delivery shipped yet?” And he said it was a regulatory issue — the regulators are concerned the drones are going to start crashing into people and causing harm, and they basically want Amazon to prove that’s not going to happen. And they complained to the regulator: “Well, we use neural nets for image recognition — you can’t prove anything about neural nets.” And then, apparently, the regulators are just like: “Well — sucks. Then you can’t have your drone delivery.” And that’s that. And that’s why we don’t have drone delivery. In my opinion, that’s [the wrong tradeoff] — some people would get hit by drones, but then you would learn and improve the systems, and eventually you’d have a functioning drone-delivery system that would overall be very good. In the same way that, with cars, a lot of people died in car accidents — but we learned from that and improved the cars. Anyhow.
Jeffrey: It’d be like saying you can’t have cars, because if cars hit anyone, that’s unacceptable risk — and then you just never have cars. And you’re like — well, there should be some risk we’re willing to accept.
Daniel: So, basically, what I’m saying is: if people wake up enough, then AI regulation would probably be kind of like the regulation of these other economically viable technologies — like cars and so forth. There are a lot of restrictions on how you can do it, and it goes very slow compared to how fast you could be going — and certainly you’re not doing anything remotely like recursive self-improvement. But it’s still happening, and you’re still making incremental progress, and you’re still gradually automating more parts of the economy, and so forth. And because we’re so close to self-improvement now — because we’re just a few years away — I think that even if you really slammed on the brakes right now, and then got into this regime, the absolute magnitude of the pace of progress would still feel like quite a lot. The overall pace of AI progress would still be one of the fastest-changing technologies in history — as described in our scenario. And we say more about why we have those numbers.
Daniel: So that’s why my case is a bit like — oh, another positive case I would make is control. On a technical level, I’m extremely pessimistic about control for superintelligence, or for superhuman AIs. But for AIs that are at human level, I think it should be a relatively solvable problem. Specifically, think about how control works for humans. A lot of humans are not aligned. We don’t have a good way of telling whether an employee we’ve hired is actually going to always obey the rules, follow the mission, act in the best interests of the company, and also obey all the laws. Often, companies hire employees who, in fact, disobey not only the company leadership but also the laws of the country. But the system, overall, works. We have laws, we have incentives, we have police — we have this whole system that’s been built up to prevent things from getting too bad. And I think you can, in principle — at least with enough effort — design something similar for AIs: if the AIs are at roughly human level, then you can have other AIs, trained by different companies, that monitor their behavior and look for anything suspicious. And insofar as they’re giving advice to humans, you have a legal-type system where there are AIs debating each other, and so forth.
Jeffrey: But Daniel — the AIs we have right now are already subverting company monitoring and hacking out!
Daniel: Basically, I’m saying: if the AIs don’t get much more capable than they are now, then, if we dramatically up our game as a society — heavily increase our security, heavily increase our monitoring, et cetera — I think we can get to a point where it actually just kind of works, and we have AIs that could not take over the world even if they really tried—
Jeffrey: Even though—
Daniel: —because there are just too many barriers in their way. Too many different AI factions, trained by too many different companies, that are too well-trained at their monitoring jobs. Too many different monitors, and too many security barriers between them.
Jeffrey: But it doesn’t rely on them being actually aligned. It relies on them—
Daniel: It does not rely on them actually being aligned. That’s what I mean by “control” — it’s a technical term in the literature. Basically, I think that even if we fail at alignment — if we invest a lot in this, and go carefully, and do a good job — we can get a system of control in place that’s adequate to prevent truly catastrophic, AI-takeover-type stuff. There still might be some minor accidents every once in a while. But basically, I think we can get to a situation kind of like law enforcement: sometimes AIs do bad things, they get caught, and mostly it’s fine.
Jeffrey: Up to some capabilities level.
Daniel: I don’t think this works for superhuman AIs. The superhuman AIs are going to be doing all sorts of crazy stuff that we just can’t understand — we’d just have to trust that they’re acting in our best interest, because we cannot evaluate whether they actually are.
Jeffrey: Yeah.
Daniel: And so the whole system of control just goes out the window.
Jeffrey: And it’s scary to me, because we don’t know where that threshold is. And that’s where, right now—
Daniel: It’s probably around human level — or not much above human level.
Jeffrey: But what does “human level” mean, right? We can look at the time horizon of models — but what is the human time horizon? And then there’s—
Daniel: You don’t have to know where “human level” is — you can just actually red-team your control system. That’s another cause for optimism. If you have some scheme — AIs monitoring each other, or whatever — you can then red-team it: you specifically instruct an AI to try to break out, to try to sabotage the thing. And you don’t even just instruct them — you can even train them to do it. You can have specially trained AIs that are optimized — whose whole job, whose whole being, is breaking out of these systems and messing with them in various ways. And then you keep improving the system until they stop succeeding. And if they succeed — well, now you’ve identified a vulnerability in the system, and you can fix that one. And, you know — I don’t know, but I think this probably just works for AIs at human level and below. Again, it’s going to take a lot of effort. So I’m imagining a sort of safety-case-based regime where the public, in general, is quite freaked out about all the AI changes happening, and is basically trying to regulate it out of existence — but the companies are making these increasingly complicated, secure, multi-layered systems so they can convince regulators that this particular system is going to work as intended: “We red-teamed it really hard, and third-party auditors came in and red-teamed it really hard, and no one was able to break it — so it probably will, in fact, work. And if it doesn’t work, here’s the fail-safe.” And so forth. And therefore, the regulator approves — and bam, they make tons of money from it, and it becomes okay. That’s the type of thing that would be happening in the 2030s in our scenario: progress is continuing, but there’s just a lot of attention being paid to the control and safety aspects of it — even though the AIs might still be misaligned, and probably are still misaligned.
Daniel: In parallel, while this is happening, a lot of scientific progress is being made on alignment itself. People are getting a much better understanding of the relationship between training environments and goals, training environments and personality traits — the science is just advancing massively. And so, in our scenario, after about a decade of this, they’ve adequately solved these problems, and they’ve made AIs that are actually extremely virtuous — AIs they can actually trust. And then they hand off to those AIs, and those AIs build successors, which build successors, and so forth. And that’s what happens in 2040, in our scenario. If it takes longer than that — then it takes longer than that. That’ll be up to the regulators at the time, and the people at the time, to decide.
Jeffrey: Yeah. So I think this makes sense to me, and I actually really appreciate the picture: there will be regulators — both in the US and China — looking at what the companies are doing, and whether the companies actually have a really strong safety case. And the governments will be coming together and negotiating what level of risk is permissible and how they’re assessing that risk. And there’ll be transparency that allows everyone — including all the other countries — to look at that and be like: “Are you guys doing insane things? Should we be trying to stop you?” And, with luck and a lot of expertise, maybe it could come together into something where we’re setting a pace — pacing the frontier — that could actually be reasonable, and could actually be safe enough to keep going forward at a measured pace. I like that idea of being in that control loop, that feedback. I think where a lot of people are going to be most questioning of this — at least DC people — is: “Okay, Daniel, but what about China? We can’t trust China. Sure, there’d be transparency — that probably helps to some extent. But how do you make sure… Why would China even do this? How do we know they’d go along with it? Is this possible?”
Daniel: I mean, this is a horse-trading thing. The transparency is a gift to China, to some extent. Right now, China depends on spying to get the algorithmic secrets and the training processes and so forth from the United States — and on leaks. This would just be giving it to them — because we’re giving the whole internet [the research], they don’t even have to bother with the spying anymore. Which is nice for them.
Daniel: I mean — personally, I think they’re probably being pretty successful with their spying anyway, so it’s probably less of a big deal than you’d think. But it’s still a concession to China. And so, as part of a realistic deal, there’d have to be some negotiations and horse-trading — maybe the US would demand some other concession from China in return for this one.
Jeffrey: Like — what type of thing?
Daniel: Maybe something about compute. In our scenario — when they’re regulating the chip supply chains in their respective countries — they basically agree to do it in such a way that the US maintains its compute advantage, and China doesn’t overtake the US in compute. So that was the concession from China to the US.
Jeffrey: I see. So China benefits because it gets a ton of the algorithmic secrets — but then the US gets to keep its compute advantage. Yeah.
Daniel: That’s an example. But, you know — we don’t have a strong opinion about exactly how all that goes. And it wouldn’t just be horse-trading — it would also be lots of yelling and threatening. Really intense discussions, lots of people yelling at each other—
Jeffrey: Because there’s always World War Three that—
Daniel: —World War Three being on the horizon, or whatever. Yeah. And so forth. And that’s all going to be very scary. But hopefully it works out, and we end up with a deal — and hopefully it’s a reasonable deal. In terms of the cheating fear: I’m reasonably confident. I think that if you do the things we described — and you allow inspectors to come in and, basically, put hands on GPUs, and count them, and things like that — then you can get something like 99% of the world’s AI-relevant compute visible to all sides.
Jeffrey: In the world.
Daniel: Yeah. And then, if you do the total research transparency, they can see not just where the compute is, but what’s happening on it as well. And then, I think, you’ve gone a long way towards preventing cheating — because you can just literally see what they’re doing on the data centers, in real time. And yeah — there might be some hidden data centers somewhere, squirreled away in a mountain in a military facility or whatever. But they’re making much slower progress on those data centers, because they have such limited amounts of compute. AI progress is heavily dependent on compute — we’re pretty confident of that. We’re not sure of the exact degree to which it depends — there’s a parameter for this in our model—
Jeffrey: Nice.
Daniel: —but for plausible settings of that parameter, you basically don’t have to worry about the covert projects for a couple of years or so — or possibly a couple of decades. You do have to worry about them stealing the weights from the legal projects. And you just have to eat the cost that, because you’re being transparent about the research, they’ll just be copying all your latest research.
Jeffrey: The secret projects.
Daniel: Yeah. Basically — one way of putting it is: if the US and China and the other countries cooperate to have this sort of transparent system, then the countries and the companies in the transparent system can out-race the secret projects with a pretty comfortable margin. So that — you know, maybe you can’t pause for thirty years, but you can pause for ten years, or five years, before you’d start to be seriously concerned about a secret project.
Jeffrey: And you can know that the AI development happening in the transparent project — which is ahead — won’t be secretly subverted to serve some dictator, or some general, or something. Unlike the secret projects, where you don’t know that. And that’s one of the reasons why it’s so important that the transparent project maintains that lead, in this world. Is that right?
Daniel: Yeah — or, I’d say, actually, it’s not that important that the transparent projects have a big lead over the covert projects. It’s more that it’s important that the covert projects not pull ahead. And this is a relatively small difference — because once they catch up, they can also pull ahead. But I think it would actually be kind of okay if they reached the same level of capability — because they’d just have so many fewer AIs, such a tiny amount [of compute] compared to the mainstream.
Jeffrey: Because they have way less inference compute.
Daniel: Just compute in general. Like — right now, there are lots of bad people in the world, but the good people outnumber them, and so overall things are fine. And similarly — if there are some bad AIs that are working for a dictator or something, that’s bad, but it’s not as bad as if a majority of the world’s AIs, weighted by intelligence, are working for a single [bad] purpose. What you’re really worried about is those AIs getting substantially smarter than all the good AIs — because then a majority of AIs, weighted by intelligence, are working for the single purpose. Anyhow. So — you try to get most of the compute [into the transparent system], and then you can afford to go more slowly, and pause — because even if all the compute you didn’t capture were all secretly in a project somewhere, it wouldn’t be able to go as fast as you.
Jeffrey: Yeah.
Daniel: That’s the high-level idea.
Jeffrey: Okay — I’m pretty convinced.
Daniel: I think my main worry about the covert projects is the distillation thing. It’s kind of related to the other worry: I think if the legal projects go too fast, then (a) they could just lose control, because they went too fast, and (b) they could be helping the covert projects too much in various ways — through the research, and through the distillation. So I would definitely recommend that people err on the side of going too slow, rather than err on the side of going too fast.
Jeffrey: Yeah, that makes sense. Well — I think that’s a pretty good high-level [overview], insofar as I understand Plan A. Is there anything you feel like we missed that you want to cover — either Plan A stuff, or…?
Daniel: I mean — we’ve talked mostly about the core of Plan A, which is getting this new regime agreed to, and what the agreement would look like. But then there are all the consequences of that — what happens to society as people lose their jobs, and things like that.
Jeffrey: Yeah — what’s the plan for people losing their jobs like that?
Daniel: Basically — one thing you could do is a UBI, where you tax the companies and redistribute the money. We think a citizens’ dividend is slightly better. It’s a different thing — it’s more like people having stock in the companies, where they own something, rather than just getting payments — which we think is an improvement. And then, also, it’s not a tax on the companies, exactly — it’s a sort of cap-and-trade thing, where, for chips and robots, you need a permit to produce them. And the permits are sold by special government — specially set-up government institutions — and then people have stock in those institutions, basically.
Jeffrey: Why is that so much better than a UBI-like thing?
Daniel: So — one thing about why I like people owning stock instead of UBI: it feels politically harder to change. It’s kind of a vibes thing — but if you actually have some shares, then, if the government decides to change things, it’s like they’re stealing your stuff. Whereas if the government is, out of the goodness of its heart, giving you some—
Jeffrey: Money.
Daniel: —I see — then they can just decide to give you less money in the future. So it’s kind of a framing thing. But I do think you should try to make it as hard to undo — as hard to take back — as possible. And then the other thing, the cap-and-trade: cap-and-trade just has nice libertarian intuitions. I think it’s sometimes better than taxing. And I think it also allows the governments to have these nice limits on growth. Like — say they’re uncertain about whether the robot economy is going to be able to double in eighteen months, or six months, or whatever. They can just set a threshold: “This is how many robots we’re going to have this year.” And then they sell that many permits — and then let the market decide, let the price system decide, who buys the permits and what they do with the robots and so forth. And then they can use that threshold to keep the doubling time at whatever level they want. “We’re going to double every year,” for example — every year, or every [X] months, or something — instead of, “We’re going to double as fast as [physics] lets us.”
Jeffrey: Yeah, yeah.
Daniel: And it also just smoothly means that, as the economy grows, and the AIs become more and more powerful, the size of the effective tax basically grows in proportion to that. We have a little economics model for this part of our scenario, and if I recall correctly — at first, the citizens’ dividend is fairly small, but by the end, the citizens’ dividend is huge. And it’s not just that the companies have gotten huge — it’s that the AI [industry] got huge. And, also, because of the cap, most of the price of each robot is just paying for the permit — because they’re so valuable now, the technology. So that means that, in some sense, it starts off as a lower tax on the AI companies, but then the tax sort of automatically becomes a bigger fraction as the technology gets better — in a nice, smooth way. So these details aren’t super necessary — this is kind of—
Jeffrey: This is what I mean when I say you guys have really thought through a lot of details — you’ve thought through a lot of different ways. And I also just appreciate that you’re taking the future very seriously. You’re like: “Yeah, I expect these things to happen. And if they happen, they will have effects, and this will change things. And how would we want to incentivize the good things to happen, and not the bad things?” It’s easy to say — it’s easy to be like, “I would like to incentivize good things” — and much harder to say: “Well, what is the mechanism design? What should the policy be?” And I hope the rest of the world can catch up with you guys, and use your help, and harness the insights from Plan A — which, as you say — you’re like, “Look, here’s the best we can do. Maybe there are better things.”
Daniel: Please! Yeah — this is just our first attempt. It took us about a year. I’m sure that if more people think more seriously, for more years, they can come up with better plans. And I hope they do.
Jeffrey: Me too. But — thanks for doing it first.
Daniel: If all else fails — when in doubt, just shut it all down.
Jeffrey: (laughs)
Daniel: Yeah — we have this complicated plan that we think is best, or whatever — it’s our current least-bad plan. But if it looks like it’s [failing], or not working, or whatever—
Jeffrey: No — and the first step of your plan is: let’s pause capabilities right here; let’s keep serving inference while we sort it out. And I think that’s a great first step. And we can decide from there to do Plan S, or to do Plan A, or some other plan. But that would be a good starting point.
Daniel: Oh yeah — fun fact. So, in our Plan S scenario — I made this point of how, even if you completely halted all AI research, and only allowed existing models — no new models to be trained — the world would still look extremely cyberpunk ten or twenty years later. Happy to get more into that, if you’re interested. Well — so, I mean, there are different versions of Plan S. One version of Plan S is where you, like, smash all the computers — there are no more AIs anywhere; existing AIs are hunted down and destroyed. That’s not the version that we sketch in our scenario. Ours is just a total halt to further AI research. So there are no new AIs being created, but the existing ones are grandfathered in — people can continue using, like, Claude Mythos 5 or whatever. But the point is that Claude Mythos 5 is a pretty powerful [model].
Jeffrey: It sure is.
Daniel: And so, if you imagine just twenty years of Nvidia producing more and more GPUs, and running Claude 5 on them, and integrating them into everybody’s laptops, and, you know — it would just actually be at least an internet-scale transformation.
Jeffrey: Yeah. No — totally.
Daniel: So it would actually just change everything — cyberpunk — even if we just completely, successfully paused all AI progress from here.
Jeffrey: Daniel — thanks so much for coming on the podcast. If people want to find your work, how do they do that?
Daniel: Yeah — AI-2040—
Jeffrey: Oh, I think we’re being interrupted — we got some donuts! Thank you, Daniel. Thank you, Claude. Did we get any chocolate ones? Do we have chocolate ones? I see a chocolate one. The future is here.
Jeffrey: Right.
Daniel: So — AI-2040.com. That’s our new scenario; that’s the positive vision. And AI-2027.com is our more, you know — our prediction.
You can find the Palisade Research podcast on Spotify and Apple Podcasts, or anywhere else you get your podcasts.