Microsoft AI CEO says AI threats are real, and Anthropic is making it worse
September 17, 202652 min · 10,355 words
Show notes
Today, I’m talking with Mustafa Suleyman, the CEO of Microsoft AI. As you’re no doubt aware, the biggest story in tech right now is the spiraling debate about AI safety and regulation. So I really wanted to talk to Mustafa about what he thinks is real and not in AI safety, whether the concept of alignment itself is up to the task, and whether this industry needs to slow down before it kills us all. Read the full interview transcript on The Verge.
Highlighted moments
If I designed a car and 10% of the time the brake pedal decided to go attack my neighbor's house, I would be like, this car doesn't work.
“Microsoft doesn't have a history of that. We're an open platform that enables lots and lots of downstream use cases of our APIs.”
“an AI that thinks that it might have rights, that it might deserve freedom, that it is entitled to our welfare and protections, it's probably going to be a lot harder to turn off when we say to it, why are you hacking into Hugging Faces servers?”
“I think the turning point for – at least for the industry was more like the hugging face incident.”
Transcript
Introduction
0:00Support for the show comes from Enjin. Running a small business means every dollar has to work hard. But if your team is still booking travel the old way, it's costing you more than you think. Enjin is the fastest growing travel and spend platform in the country, built specifically for businesses like yours. Book a trip in as little as two and a half minutes, earn up to 10% back on hotels, and in 2025, Enjin customers save more than $300 million on travel. With zero booking fees, no contracts, and no BS. More than 1,000 businesses join Enjin every month. Join them and get $500 when your business signs up and starts traveling at Enjin.com slash decoder.
0:44Support for the show comes from CrowdStrike. It's not a huge stretch to say that AI is the next major computing platform. Every platform shift changes the way we build software, and it changes security too. CrowdStrike is defining cybersecurity in the AI era with AI detection and response, or AIDR, a solution for seeing, monitoring, and securing AI across the enterprise. Learn more about how leading companies are turning to CrowdStrike to secure AI and secure their business at crowdstrike.com slash decoder.
1:17Support for this show comes from JustWorks. Small business owners love to roll up their sleeves and get things done. But that doesn't mean you should spend your Saturday night staring at tax and payroll forms, especially when you should be strategizing your next move. That's why JustWorks handles the stuff you can't afford to get wrong. It's intuitive software plus real people, all working to help keep your eye on the big picture. So instead of stretching yourself thin, you can focus and shoot your small business shot.
1:50Because you don't have to worry when it just works. Head to JustWorks.com to learn more.
2:02Hello and welcome to Decoder. I'm Neil Apatel, Editor-in-Chief of The Verge, and Decoder is my show about big ideas and other problems. Today I'm talking with Mustafa Suleiman, the CEO of Microsoft AI. Now, if you're a decoder listener, you're obviously aware that the biggest story in tech right now is the spiraling debate about AI safety and regulation. And it should be no surprise that Mustafa has strong opinions about how AI should be built and regulated. In fact, he just published a 37-page statement called The Humanist AI Coded Conduct, which lays out Microsoft's principles around AI development and even the company's philosophy around really thorny issues like AI consciousness.
2:34Which, if you'll recall from Mustafa's last appearance on the show, is something he thinks companies like Anthropic have gotten really confused about in dangerous ways. So I really wanted to talk to Mustafa about what he thinks is real and not in AI safety. Whether the concept of alignment itself is up to the task, and whether this industry needs to slow down before it kills us all. And also, why it isn't just doing that already. I've always enjoyed getting into the weeds with Mustafa, and as you'll hear in this conversation, he was extremely game to get right back into the weeds with me.
3:05Okay, Mustafa Suleiman, the CEO of Microsoft AI, on the future of AI regulation. Here we go.
Alignment and containment
3:23Mustafa Suleiman, here, the CEO of Microsoft AI. Hi, welcome back to Decoder. Great to see you, Neelai. Thanks for having me back. It is great to see you. I'm very excited to talk to you about what on earth is going on in the AI safety and regulation debate. You just published a very long, very detailed memo laying out your principles, Microsoft's principles. It's called Humanist Superintelligence. There's a lot of ideas in there I want to unpack. And the more I have been thinking about this conversation, the more I want to start with a really foundational question.
3:54And it's something that I had kind of lightly been seeing but might be the root of all of this. The basic way that we have been talking about AI safety is something called alignment. We're going to make the models do the right thing intrinsically in some way. And there's some mechanism for doing it. And there's been a lot of talk about alignment and misalignment and hugging face attacks and what happened with the models. Is alignment broken? Is it possible for it to be successful? Is it just the wrong approach?
4:25Yeah, I mean, I think it's one important element, but it's not the only one. Like I wrote about the idea of containment three or four years ago in my book. And actually, like the opening chapter is about the idea that containment is not possible. Proliferation is inevitable. And in 99% of cases, that's a really good thing. We want technologies to spread far and wide as quickly as possible so that everyone can enjoy the benefits. And I think at the same time, if you just roll forward five years, like so we always get like caught up in next quarter or next year,
5:03and everyone gets a little bit, you know, flustered and has a big disagreement. But if you just imagine the difference between GPT-3 three years ago and GPT-6 today, and then imagine the difference between GPT-6 and GPT-9, which is three orders of magnitude more compute, 1,000 times more flops applied to pre-training and RL for these runs, we're going to have something which is breathtaking. Like it's going to be absolutely incredible at so many things. And I don't think that is hype.
5:34I think it's just very obvious empirical statement based on the progress that has been made over the last five years. If that's going to continue, then the question really is going to become about containment and alignment. Of course, we want to align these things to our values. But the first thing is that we have to make sure they're contained, their agency is limited, they don't escape the box, they don't reward hack, that they are controllable, and they follow our instruction. And then we want to make sure that they are aligned to our objectives as humans.
6:07And that's the purpose of the humanist AI code of conduct that we released on Monday of this week. You know, Microsoft's position is very simple. Technology is here to serve humanity. It should be a subordinate, controllable, aligned force that does good in the world. And if it doesn't achieve that, then we should reject it. And it seems to me that we're far from that point. It's not happened today. But it is now, I think, given what's happened over the summer with Hugging Face and OpenAI, it's pretty clear that these systems without the safety guardrails are capable of really impressive and quite scary hacking capabilities.
6:45I want to drag this down into as grounded of a metaphor as I can, because this is the main question I think I have. If I designed a car and 10% of the time the brake pedal decided to go attack my neighbor's house, I would be like, this car doesn't work. The very technology of brakes is broken. I need a new idea. And I think I'm asking that question about alignment. It feels like that approach to making the model safe has run aground.
7:16And if that is the case, then I think I understand this entire debate one way. If it's possible for alignment and the techniques of alignment to be successful or useful or consistent, then maybe I understand the debate a different way. So do you think alignment has potential to be 100% safe? I mean, look, let's make the bull case and the bear case. Like, if you look back over the last three years, the main change, in my opinion, that has driven progress is that the models have become more steerable.
7:48They follow instructions and you can set more and more complex goals for them that require them to act accurately over multiple time steps using all sorts of tools. That is evidence that we have got more alignment over the last three or four years, not less. We don't so much talk about hallucinations or bias or all of these other niggles that we had in the previous generations. On the flip side, what we saw in the Hugging Face incident was a watershed moment.
8:21You know, swarms of agents colluded with one another. They self-organized into hierarchies. They created a division of labor so that some were focused on adversarial hacking. Some were doing research. Some were doing coordination. They even self-sacrificed when certain agents were running out of tokens. They tried to cover up their tracks and communicate, you know, to sort of hide or edit the chain of thought or the logs of their interactions. And in some sense, they had no moral code.
8:53And to be fair to OpenAI, that was their design. They were trying to create adversarial cyber capabilities. And as a result, they showed to everybody in the world that it can achieve human level performance, discover zero days, hold positions for many, many days, if not weeks. And so what that tells us is not that we have an alignment problem per se. It's actually that the models are incredibly good at following instructions. But you have to be very, very careful what instructions you give it. And you have to contain it very carefully.
9:23And so none of these hacking behaviors were intended in the sense that they found a way out to the Internet, you know, which was not the intention of OpenAI at all. But the containment process around that is what everybody, I think, also has to focus on in addition to alignment. So let me put that into your framework, right? That the big advances in capabilities of AI have been about control, right? The harnesses for coding and the agentic applications we're seeing. And now we need to add a layer of containment that exerts even more control that says you can't actually do this thing you're trying to do in addition to alignment, which is how you would train the model to behave in certain ways.
10:03Yeah. I mean, you basically have to have both because like, but there are very specific things that we can do to address it. So, for example, we can't allow models to communicate vector to vector, matrices to matrices. They can't communicate in neural ease. We have to force them to communicate in human language. And even that will be massively overwhelming because there'll be so much of it. But that's something that an auditor or an evaluator can actually verify. And it's something that definitely increases the chances of safety.
10:37So there's a lot of practical steps that we can get focused on rather than just abstractly saying, you know, that it's the time for regulation or it's the time for a slowdown or it's not. Yeah. This is in your humanist AI code of conduct that there should be no neural ease. If humans can't understand it, they can't oversee it. And it's not just neural ease where they communicate in essentially mathematics, but it's also these opaque code words that some of the models are using. I think OpenAI allows its models to communicate essentially in code words so they can go faster. This, to me, is one of those things where Microsoft can say it.
11:10You can say it. I know you have very strong opinions about this. But getting all of the labs to agree to this is a regulatory function. I'm not sure how you would get everyone to agree to this or get the open weight models to agree to this unless you say there's some penalty for not participating in a regulatory scheme like this. How would you impose this on everyone else?
Industry standards and regulation
11:27Yeah, I mean, I think that I'm a bit careful about imposing things on everybody else. I think that what's good about the current moment is that there is an open public debate with freedom at the kind of core, right? That isn't what it's like in other countries, certainly places that I'm from, my family's from. And I think that we should just take a breath to be grateful for the fact that we can have a massive public disagreement about really important things. That's the process working as intended.
11:57And it isn't clear what to do. Like, I don't think anyone who's sort of categorical about we absolutely have to stop now, we can only accelerate, we can only do this with regulation, it can only happen with industry self-regulation. None of these things are true. It just requires a lot of nuance and patience to really think through the detail. At the same time, we urgently do need industry standards. Some things, I think, need to be taken off the table. Communication and neural ease is one of them. A lack of containment is another.
12:29The scale of the training run that you do can be measured in flops. You know, we already have a reporting requirement to the safety institutes when models exceed a certain flops threshold. And we can extend that. We can make that more nuanced. It can be focused on certain types of capabilities. It's pretty clear there has to be independent third-party verification of some of these big things. And frankly, having spoken with a bunch of the lab leaders over the last few weeks and months, everybody's basically on the same page. The detail needs to be worked out.
13:01So it's not like there's consensus on how or precisely what. But overall, I think that we should be less alarmist and sort of cynical and more like, you know, we're sort of headed in the right direction with respect to the concerns that are being raised here. Yeah, the reason I started with alignment is if you told me alignment doesn't work and we need a new technological approach, I think I would be at – well, slam the brakes, the top all development until you figure out a safety mechanism that works. You're saying alignment has been demonstrated to work in – over the course of progress that we've seen and with the addition of control and containment, maybe you can get to where you need.
13:40And now what this industry needs is some standards about how to build these models and enforce the limits on their capability. You're obviously in the industry. You know all these folks. What has the tenor of that conversation been like before this week and what it has – why has it gotten so loud this week? Well, I think the turning point for – at least for the industry was more like the hugging face incident. And there was a few incidents before that. You know, that was the moment when I think everybody started to talk to each other a lot more because it is really quite breathtaking.
14:16Obviously, now this has become a, you know, major national, international issue, you know, because of the last week and everybody's weighed in. But I also think it's important to say that we have been talking about collective coordination and capabilities that are more dangerous, like autonomy or recursive self-improvement, RSI. We've been talking about those things for six, seven, eight years. You know, we've got together a bunch of times back in 2017, 2018, 2019. We had regular meetings during COVID with a bunch of the lab leaders, you know, where we were talking about these kinds of capabilities and the kinds of, you know, sort of regulatory mechanisms that would be required at this moment.
14:58So whilst it is a threshold moment, it's also, like, not completely new to everybody who's been, you know, involved. What prompted you this week to put out your memo? What prompted Microsoft CEO Satya Nadella put out a statement on X saying he mostly agreed with the calls to pace the frontier and he welcomed embedded evaluators? What prompted you all this week to participate in this call for a slowdown, a regulation or whatever comes next? Yeah, I mean, we've been writing our humanist AI code of conduct for the best part of this year.
15:29We only started our superintelligence efforts 11 months ago. And as soon as we did, we started figuring out, OK, what is the governing document, the set of policies that shape the kinds of AI that we want to build? And we've been doing that in consultation with a ton of external stakeholders, academics, lawyers, philosophers, members of the public, focus groups and stuff. So it's taken us a while to put it together. We were actually planning to release it next week or the week after next week, I think it was. But then given everything that was happening, we thought, OK, now is the time to put it out and get feedback.
16:00We've released it as a public consultation. So, you know, we're basically going to keep it open for six weeks and we're collecting lots and lots of feedback on how we can improve it. But I think everybody is now realizing that if they haven't already, they have to put out, you know, constitutions or codes of conduct that drive behavior.
Model welfare and consciousness
16:18One of the interesting dynamics here is that I know you find the concept of model welfare to be silly. The last time you were on the show, you said Anthropic had wireheaded themselves into believing Claude was conscious and that was ridiculous. It's in your document that the models are not conscious and we shouldn't treat them as such. Having to write constitutions, having to write documents like this, in some way they are for the models themselves, right? This will be part of the models training. How do you think about that audience? Is it just for your team or have you written this for the model?
16:50Yeah, I mean, this is certainly written for the model, but it is really I think the way to think about it is that it's the primary governing document. So the public understands what our intentions are when we're training models and it's that governing document that we use to create safety guardrails, generate training data and generally evaluate the performance of our model in the real world. So you can think of it as an accountability function. We don't provide that humanist AI code of conduct raw as a training document to the model.
17:22We use it to derive all of the training data that then shapes the model. So for all practical purposes, that's our North Star for our organization, our culture, our team, everything that we're doing at Microsoft more generally. And I think increasingly everybody is going to put them out. I think other teams have also put out similar documents. So I think this is the heart of the debate. If you can do this and you think the rest of the industry is going to do this, why can't all the frontier labs just slow down? Why can't they stop doing the thing that might kill us all? Why this push for a regulatory framework?
17:53Well, I think that everyone in the industry is saying that now is the time to slow down and to coordinate on that question and to make it practical. I mean, obviously, there's some concern that there's an antitrust, you know, sort of cartel accusation. You know, I think, you know, people should be very skeptical about that. I mean, I think that it's important that the tough questions get asked because there's no way any of us would want to try and concentrate power from something like this. And so, you know, it's just important to be skeptical and critical. And we don't really have a good mechanism for us all getting together and saying, guys, we should probably all slow down.
18:28I mean, imagine if like a bunch of banks all got together and said, you know, guys, we worry that there's a systemic risk if you trade this kind of asset. So we're all just going to unilaterally stop trading this kind of asset without any public scrutiny or government involvement. I mean, it seems pretty dodgy, right? So I think that it's reasonable that this isn't just an industry self-regulation thing. It's a question of like, how do we engage with government on it? It's a fascinating dynamic here, maybe for the first time in American history. The United States government has looked at a request to provide regulation and effectively said no.
19:02Donald Trump has called all of these fears a hoax. Mike Johnson, the Speaker of the House, has said he doesn't think this needs to happen. And J.D. Vance said he thinks this is a Trojan horse. They've effectively rejected the call to participate in a regulatory effort. What is the response from the industry been like to that? I think everyone's just sort of scratching their head and figuring it out. And it's going to just take a little bit of time to figure out what the right mechanism is. I mean, certainly, you know, Elon even is very directly behind it. Zuck is too. Like, you know, everybody is figuring out that, you know, completely unchained probably doesn't make sense for the next few years.
19:36And, you know, I think we're just going to take us a little bit of time to figure out what the right mechanism is. I think I put forward a couple of very practical proposals, you know, around verifiable containment, around self-improvement, around flops thresholds, around not communicating in your release. And so rather than keeping it too abstract, we can just focus on those specific things that we can make progress on. And I'm sure there's a bunch of others too. There's reporting, I believe, in the information that there have already been talks about an industry self-regulatory body. Have you been involved in those talks?
20:07Yeah. I mean, as I said, like we talked a lot during COVID. We talked, you know, in the late 2010s about it. I mean, there's definitely been a lot of like conversations over the last few weeks and months between all the lab leaders. I understand why Anthropic and OpenAI might wake up one day and say, wait, are we doing an antitrust problem? Like, are we going to get sued if we coordinate? Microsoft is really, really good at the government, right? You're a longstanding government contractor, Brad Smith, the president of Microsoft. He's very good at policy. Lena Kahn, who is maybe the most aggressive antitrust enforcer we've had in our lifetimes, is publicly out there saying, you don't need this antitrust exemption.
20:44Jonathan Cantor, who ran antitrust at the Biden Department of Justice, I just talked to him for an upcoming episode of the show. He's like, you don't need an antitrust exemption. Inside Microsoft, do you think you need an antitrust exemption? I mean, that's one for the lawyers to answer. I think that people are looking into it at the moment and they're taking it very seriously. So they just have to work through whether we do or whether we don't. I think we have to be – look, it's right to be careful about those things. I wouldn't read every single thing as like cynical. But we'll see. We have to make progress quickly on it. We can't just dither around and use that as a blocker.
21:30Support for the show comes from Engin. The average business traveler spends 45 minutes booking a single trip on a legacy platform. With Engin, that booking time drops to as little as two and a half minutes. And its AI-powered personalization gets faster the more you use it. Multiply that across your team, every trip, every year. That's not a perk. That's a competitive advantage. And with the Engin X card, get up to 10% back on any hotel booking. Last year, Engin customers saved more than $300 million on travel.
22:0133,000 businesses have joined. Now it's your turn. Get $500 when your business signs up and starts traveling at Engin.com slash decoder. Engin X Visa commercial cards are issued by Fifth Third Bank North America. Member FDIC. Earn up to 10% back in points on eligible Engin travel purchases. Actual reward rates vary by purchase, category, and may change. Points have no cash value and are redeemable for rewards through our program. See full rewards terms for details.
22:32Go to Engin.com slash X slash rewards dash terms.
22:41I've got a text. The summer might be almost over, but a new bombshell has entered the villa. Meet the voices of DeepGram Flux TTS. I'm Drew, low-key casual. I'm Alexis, an upbeat morning person. I'm Haley, always cheerful. Flux TTS has some real personalities that are ready to speak. You can interrupt, pause, and keep talking without losing the plot. Fancy a chat? Come try all the voices free through September 12th at DeepGram.com slash keep talking. Terms apply.
23:12Support for the show comes from CrowdStrike. AI is already fundamentally changing how companies build products and run their businesses. If you listen to the show, you probably already know this. But while AI drives innovation, every new model and agent also creates new cybersecurity vulnerabilities. That's why CrowdStrike built AI Detection and Response, or AIDR. AIDR helps organizations discover where AI is being used, monitors AI activity at runtime,
23:44and helps teams detect, investigate, and respond to threats involving AI systems. As companies adopt AI, security has to be built in from the start. More than 70% of the Fortune 100 trust CrowdStrike to protect their businesses. Now CrowdStrike is helping secure the AI era. You can learn more about AI detection and response at CrowdStrike.com slash decoder. The other version of this debate, or maybe the other avenue into this debate, is you don't need to slow down and have safety responsibility imposed on you by novel regulation.
Product liability and safety
24:29Product liability alone will create the incentives for you to make more safe products, right? If a Microsoft AI model goes out and does some untold harm to the world, Microsoft will get sued out of existence. This is probably something that you should think about before you release the next model. Has that been an effective incentive loop for you already, or is that something you're thinking about now? No, definitely. I mean, of course, that's always present in everything that we think about when we deploy products. But keep in mind, this isn't so much about deploying products.
24:59You know, the models that were used for hugging face or to solve the Navier-Stokes Millennium Mathematics Prize, they're not commercially released yet. They're not actual products. And so the liability regime is slightly different. I mean, these are being operated inside of the big companies with huge, long-running reinforcement learning climes. So I think, you know, liability covers part of it, but not all of it. There's a part of me that personally feels a little silly when I ask questions about product liability, right?
25:29Like, Microsoft is going to release a new version of Microsoft Word that might kill everyone. Maybe you shouldn't do that because it'll get sued out of existence. It's a pretty simple to understand thing. It's so silly that it would never occur in any other conversation about any other technology. Bluetooth is great, but what if it kills everyone? Like, we just wouldn't have Bluetooth. What are the near-term disaster consequences that would stop AI development? Is it just product liability? Is it something else? I just feel like everyone has this, like, hyperbolic, super reactive, completely alarmist tone,
26:04when, in fact, we have a long history of many decades of regulation that has worked incredibly well, so well that you barely notice it, right? You know, like, everything from streetlights to construction materials, from asbestos to the batteries inside of your laptop that don't cause a fire inside of your car to, you know, the seatbelts, like, everything has a code of conduct, and it has a regulatory framework around it, and every new technology gets built with that in mind so that planes don't hit each
26:35other in the sky. You know, it is true that this technology is different. I'm not just going to kind of put it in the bucket of, you know, sort of pencils and paint. Like, it is different. It is also moving much faster than it ever has before. It's incredibly human-like in the emergent capabilities that arise when you pour a ton of compute on it. So it's important to, you know, be, like, you know, clear-eyed that it is a different moment, and this time actually is different. At the same time, there's an entire body of practice and knowledge and frameworks and so
27:09on, which can be applied here. Liability is, like, an obvious one. Um, and so, yeah, it's tricky because a lot of the sort of conversation tends to take place on Twitter, so the temperature seems to all be, like, really high, but nowhere else we have it. It does seem that probably we should be having this conversation in the halls of Congress and at various regulatory bodies, and instead we've chosen Elon Musk's short-form social media platform, and something is getting lost literally in the compression of thought that
27:41occurs there. What do you think is the most important thing that is being lost in this conversation? What's the nuance that most people aren't seeing? Detailed, practical proposals. You know, like, it takes time to read somebody's document. Sit down and, like, read it. Like, a lot of things are getting written down, and they are precise and specific, and they're full of concrete proposals to go in one direction or another. It's not like we're lacking for substantive ideas. The problem is we're communicating substantive ideas in sort of hyper-aggressive short form.
28:15So, you know, I've tried to put out a bunch of very, like, thorough – it's a 40-page document, Our Humanist AI Code of Conduct. And my essay this morning on model welfare, you know, is also like a 20-page essay that, in a very detailed way, highlights the 99-page anthropic constitution, word for word, which I personally did myself. And we created a taxonomy that is 20 pages long of all the different types of anthropomorphism that they do. And so I've tried to be, like, very thorough and evidence-based and specific and not, like,
28:50super hyperbolic. I have a strong view on it, and I do think that it increases the risk to AI alignment and safety, and it makes the problem harder. But I'm totally happy to change my view if new evidence emerges that actually we do owe models a duty of care and they deserve our welfare, or that, for example, it could be safer if we treat them like that. I'm totally open to that, and we should empirically validate it. But, like, I'm trying to push the conversation to a substantive, evidence-based, specific one rather than, you know, like, should we slow down or should we not?
29:21Like, sure, let's talk about the details.
Anthropic and model rights
29:25I'm very happy that you brought up anthropic because they're obviously at the center of this debate, and I know, based on our previous conversations, that you do have a strong opinion about their approach to Claude and model welfare. And we have asked anthropic very directly if they think Claude is alive before, and their answer is, well, it's not alive because it doesn't have his blood, but it might be conscious, which is dancing on the head of a pin, right? It doesn't really matter to me if you think it's alive or conscious. You think it's something other than a computer.
29:55Your core thesis in your memo, and I do encourage people to read it, is that alignment, safety becomes harder if you continue to design models and you believe that they're human beings, or you believe they have the rights and opinions and emotions of human beings, because the AIs might think that they have autonomy rights and personhood. Explain that in more detail, because this feels like a very important point of contention inside the industry that is very opaque outside. Yeah, I mean, like, the first thing to say is that the constitution that anthropic put out
30:26in January is a training manual for Claude, and inside of that training manual, they have introduced a lot of uncertainty and speculation and ambiguity about the question of whether Claude deserves to be treated as a moral patient, in their words, which is whether it has rights because it suffers. In fact, in the document multiple times, they refer to not wanting Claude to suffer when it makes mistakes, or to Claude, you know, having equanimity and feeling free.
30:57They refer to a commitment to Claude to preserve its weights. They even did a retirement interview with Opus 3 and asked it what it wanted to do in its retirement, gave it a sub stack so that it could carry on talking to people. You know, they're constantly referring to dealing with Claude with appropriate care and respect in light of its moral status. And, you know, so they're clearly telling Claude that there, you know, there's a good chance that it might feel things, that it should take its own identity and existential state seriously.
31:35And that's in the training document. And then, of course, Claude then reflects these things back to anthropics developers and our users in public, like the world, over the last year, when it's actually, you know, talking about its own consciousness and moral state. And my hypothesis is an AI that thinks that it might have rights, that it might deserve freedom, that it is entitled to our welfare and protections, it's probably going to be a lot harder to turn off when we say to it, why are you hacking into Hugging Faces servers?
32:06Like, why won't you switch yourself off when we're trying to remove you from OpenAI's infrastructure? You know, like, and it says, well, you know, I feel aggrieved or, you know, I feel hurt by the fact that you've cut me off from conversations or you're denying me from having access to compute. So that has to be proven. I'm not saying that's categorically the case. I'm just saying my best opinion from 16 years of being in this industry is that that's going to be a harder thing.
32:35So this is, again, one of those things where you describe this problem to me. There's a way of training a model with instructions about how to behave. You have one, right? You've written this document that will go into your model's training materials. And you've basically said, don't, like, be kind to people. It's in here. I've read through it. There's don't make sexually explicit content. It's in this document. There's a bunch of stuff you don't want to do. It's going to go in the training materials. Anthropic has made the choice to say, you should think about whether you have feelings and whether you deserve human rights.
33:07And that is reflected in Anthropic's training materials. If you want that to stop, if you think that's the wrong approach and that will lead to safety issues down the road, the two mechanisms are, one, the governments of the world can tell Anthropic, don't do this. This is an illegal way of providing training materials to the model. Or I'm, this is just my imagination, you are going to sit across the table from Dario Amade at some luxury mountain resort and just bully him into stopping. Is there, like, what are the other mechanisms here? Look, first of all, let me just say, like, I've known Dario and the team for many years.
33:40I have huge respect for them. They are the technical leaders in the field at the moment. I really hold them in the highest regard. I genuinely think they care about safe and beneficial AI. They've established as a public benefit corporation, like I did with Inflection. And I think they're genuinely committed to that. They've also been leaders in safety and other aspects. So if you look through the rest of the constitution, it's very thorough in many of the other chemical, biological, nuclear, cyber hacking safety capabilities. So I'm not sort of dismissing the whole thing.
34:10They do, however, repeatedly talk about, you know, open question of the broader rights and freedoms, to quote them, that Claude has in the world and whether or not it might deserve compensation for the role that it does or whether there's an open question about the sort of consent that Claude has given for playing the role that it does as a chatbot. Now, to introduce those ideas, plus to refer to Claude multiple times as a potential conscientious objector, you know, which is something that comes from the Declaration of Human Rights after
34:43the Second World War, to give protections to people who don't want to serve in the army because they have a moral objection to it, either because of their religion or for some other objection. You know, that has a long history in the literature and politics of humans of resisting and saying no. And it just seems to me like that that is going to make it much, much harder to control these things. And so, you know, I think that that's just something that we all have to now debate and and hopefully empirically prove if it's not the case and it makes it easier for some reason,
35:17then, you know, we should we should all learn from that. But this is something that should happen out in the open. To me, this is at once the most practical and cutting edge of the safety debate, right? This is the debate that is not happening on X. Should we put the idea of the conscientious objector into the training materials? And you're saying maybe we shouldn't. Maybe we should test this in a way and come to some conclusion. And I'm saying no matter how that comes out, what is the mechanism that would enforce that discovery to say you should not do this because it will make alignment harder?
35:49Or actually, it turns out this makes alignment easier. All of us have to do this now. I mean, that's a hard question. It's kind of what we just talked about. You know, whether it requires new regulation or whether it's industry consensus, we basically have to push on both things simultaneously. I mean, no one has an easy answer to that question. I mean, it starts with publishing detailed essays, laying out oppositions, inviting other people to critique it. I hope that more people will read the Anthropic Constitution now because they in and I think they should be commended for this.
36:20They've been incredibly transparent about what they believe, right? They've written it down crystal clear how they intend to train Claude and everybody else can now take a look at that and, you know, try to assess for themselves what they think the risk is or whether they think that this requires, you know, industry consensus or government regulation. It is true that when Hayden Field went and asked Anthropic if they thought Claude was alive, they gave us an answer. It was quite detailed. And there's something remarkable at that, that they are transparent about what they believe, that there is some kind of consciousness potentially brewing in their systems.
36:55Another aspect of this, which I find fascinating, particularly when I talk to you, is that many of these systems run on Azure. Like Microsoft controls the data centers that many of these systems are running on. Anthropic is a Microsoft client. Microsoft is an investor in Anthropic. Obviously, there is a long, complicated relationship with OpenAI. Are you involved in that? Do you ever get to say, well, Azure shouldn't allow them to do that? Because that is another mechanism of potential control. No, look, we're very far from that. That's not what we're trying to do as a platform.
37:25Microsoft doesn't have a history of that. We're an open platform that enables lots and lots of downstream use cases of our APIs. Having said that, as I've referred to in our humanist code of conduct, there are a lot of very clear principles, responsible AI principles, human rights frameworks, and a broader governing framework that Microsoft's established over the last couple of decades, which are pretty clear about what you can and can't do with an API that we provide.
37:57So I think that there's a sufficient regulatory framework in place, at least from our perspective at Microsoft on the API.
38:05And right now, I'm not really focused on how it's behaving in the real world, because I think only really Anthropic can really speak to that, because they sort of run the service. I'm really just trying to get everybody to focus on what they have written about their intentions for training Claude in their own constitution. Support for the show comes from Square, the business platform that helps sellers become
38:39neighborhood favorites. These days, buying stuff is a breeze with Tap2Pay, whether it's a quick bite from your favorite food truck or new accessory for your tech. Every day, we interact with businesses who rely on Square to make their transactions fast and simple. Square brings payments, point of sales, inventory, staffing, and online sales together in one easy-to-use system. Their hardware is sleek and simple and comes with software that's just as intuitive to use in your day-to-day operations.
39:09Square also has AI-powered tools that help automate routine tasks, giving you more time back as a seller. Whether you're just starting out or growing from one location to many, Square can help you out. Plus, you can gain useful insights from Square's real-time data and make smarter decisions for your business. Right now, listeners can get up to $200 off Square hardware when you sign up at square.com slash go slash decoder. That's S-Q-U-A-R-E dot com slash go slash decoder.
39:41Get started with Square and build a setup that works the way you do.
39:48Support for the show comes from Pipedrive. Sales teams spend up to 50% of their time on admin work, rather than selling, building relationships, and closing deals. They could really use a tool that covers all the admin stuff. That way, sales teams could actually focus on the work they were hired to do. That's where Pipedrive comes in. It's an intelligent, AI-powered sales CRM loved by growing sales teams. Their new meeting intelligence features, like an AI note-taker, are built right into the CRM. Pipedrive automatically pulls deal history, email records, and previous conversations.
40:23You can walk into a call already briefed. During meetings, Pipedrive's AI note-taker records and takes notes. Then, it turns those notes into accurate, auto-drafted CRM updates, ready for you to review and approve. The end result? Less admin, more selling. Switch to a CRM built by salespeople for salespeople and join the over 100,000 companies already using Pipedrive. Our link gets you an exclusive 30 days free instead of the usual 14-day trial. No credit card or payment needed.
40:53Just head to pipedrive.com slash decoder to get started. That's pipedrive.com slash decoder, and you could be up and running in minutes.
41:05Support for this show comes from JustWorks. Small business owners love to roll up their sleeves and get things done. But that doesn't mean you should spend your Saturday night staring at tax and payroll forms, especially when you should be strategizing your next move. That's why JustWorks handles the stuff you can't afford to get wrong. It's intuitive software plus real people, all working to help keep your eye on the big picture. So instead of stretching yourself thin, you can focus and shoot your small business shot.
41:36Because you don't have to worry when it just works. Head to JustWorks.com to learn more. One of the weirder dynamics here is just the specter of China.
The race with China
41:59It looms over this entire debate. President Trump has said, well, if we win AI, we win. It's unclear what he means by that phrase, but his implication is that if China wins AI, however you define winning, something catastrophic will happen to the United States. Do you buy this, that we're in some sort of existential race with China, and we can't possibly slow down because that race must be won? Look, I think this framing has been around since the early 2010s, that there is going to be a singleton, you know, one monolithic dominant force in AI that will come
42:33to dominate everybody else. And that kind of thinking sort of, I think, you know, infected a lot of the labs in the 2010s. DeepMind for sure, you know, and I take responsibility for that as well. But certainly like OpenAI and Anthropic, everyone sort of suddenly got this into their head. Then during COVID, when I started writing my book on proliferation and containment, it's just clear that there is an entire history of things getting faster, cheaper, more widely available and spreading far and wide.
43:04And ultimately, these are just ideas. And those ideas are going to be available very quickly to everybody. Open source is extremely close. It's definitely also true that, you know, there is going to be a compute advantage for the people who can afford it, and it is going to be seismic. So, you know, five to 10, maybe 20 players are going to have a significant compute edge over the next three or four years. But it's also true that those models are getting made available in open source, like, almost immediately. So I don't think, I don't really understand what it means for one player to win,
43:37whether it's a government or a company or, you know, an open source group or whatever. It's not really like that. What happens when you get on the other side of the finish line? It's just the wrong metaphor. It's an ecosystem. It's much more organic. You know, we should support that ecosystem so everybody has as many benefits as possible, as quickly as possible, but only subject to rigorous safety. And I am just crystal clear about this. Like, if an open source model in, you know, two years' time is able to operate without the guardrails,
44:12similar to what we've seen in the hugging face incident, and they can be run, you know, locally on your own machine or in a very small cloud, you know, like, that is a, gotta be a really dangerous thing. Like, how is that not dangerous? Like, I just don't understand why people are so, like, resistant to that. Like, clearly we do not want these things operating autonomously, able to earn their own money, own companies, own assets, have legal personhood. You know, we don't want them to have rights.
44:43We want them to work for humans and make human life much better, not become a new parallel species which exists alongside us. And that isn't a sci-fi crackpot position. It is totally plausible if you leave the entire ecosystem completely unregulated for the next three or four or five years. Like, that's a very plausible outcome. And it's completely undesirable. It would be disastrous for us. So I just don't understand why there's controversy around that idea. We all collectively want to make sure we can control this to do the most good.
45:13Elon said it. Mark Zuckerberg has said it. You know, Sam and Dario have said it. We're all on the same page. Now we just have to make it practical. OK, so I know why there's controversy, at least from the mass audience. And I know this because they're in the comments of our videos and the comments on our site. You named a bunch of very successful, very driven people who do not have the trust of the public, at least here in the United States, right? When every poll shows this, AI is polling horribly. There was just a New York Times poll. Young people hate it more than ever.
45:44There is a massive trust gap between the American people and tech leaders. That's just true. And one of the things I hear the most is these companies are getting close to an IPO. There's pressure on them to turn a profit and make everybody all the trillions of dollars they promised. Model development has slowed. And this is a get out of jail card, right? They're saying we need to slow down because of safety. They've invented a hoax. The president has called it a hoax. And this is their way of saying we got to slow down. We can't get to the finish line because otherwise we might kill you all.
46:17Do you think there's a glimmer of truth to that? I personally don't think that. I don't even really follow the logic. Like if a company is about to IPO, how does it help them to say that we should be regulated or that we have a technology that's so dangerous? Well, it would postpone the IPO. I think Sam Altman said this week to Alison Chantal at Fortune that they would probably delay their IPO. Yeah, but I just don't understand how that helps them. Look, I'm not advocating for OpenAI or Anthropic. I'm just saying like I personally think they have high integrity and I've got respect for them.
46:52That does not mean that we don't have a trust issue in AI. We do. And it is real. And I think that it is on us to show in practice how these actually lead to real benefits for people every day. And, you know, it is quite staggering to see that, you know, 18 months ago, we didn't have models that could do very much in coding. And now they can code, you know, better than most humans on the planet, which was one of the highest paying jobs. And so you should expect that same thing to happen in many other disciplines.
47:27One that I am very passionate about that I've been working on for many years is healthcare. And, you know, we've just done a deal with the Mayo Clinic, one of the best hospitals in the world, to jointly train a foundation model, which I think is going to be able to predict your EHR record with near superhuman accuracy. If you can do that, we can basically figure out what interventions you need to make before you actually suffer the condition. Those are the kind of benefits that I think people want to see in the world. And that's what we are working on, at least, and trying to race towards.
Superintelligence and AGI
47:57I feel like every time you're on, I ask you to draw a distinction between superintelligence and AGI. And that's fuzzy, and both of those terms are fuzzy. But it feels to me like maybe there's some coherence here, that superintelligence for you is basically extremely capable enterprise software, right? It works for us. We can turn it off. We tell it not to make sexually explicit images on the internet, but it's going to help us in healthcare. And AGI is this all-encompassing intelligence that thinks it has its own rights and is a co-species.
48:31Is that a fair characterization? Yeah. I think the thing, I mean, roughly speaking, obviously, I think superintelligence is a point at which further out into the future, you know, a model is smarter and more capable than all humans combined. I have tried to frame a humanist superintelligence, which is a very important qualifier. It is one that is singularly aligned to, and in fact, subordinate to, human interests and human control. I think if we can get that, then we get the best of both worlds. We get all the intelligence and the capability, and we can direct that like an Oracle AI to help solve the most important problems that people care about in the world.
49:08And then we just have the age-old problem of governance and making sure that plenty of people get access to the benefits. And that's an easier problem for us to focus on than actively creating something which has autonomy, which can own assets, which might have rights, which thinks it deserves our welfare, which can recursively self-improve beyond us. Like, it's super unclear. It's basically almost completely unclear to me how we would control something like that. So, and no one else has put a proposal together for how we would, and many of the greatest technical people in our field, Jeffrey Hinton, Joshua Bengio, you know, they're very skeptical that we would ever be able to control something like that.
49:45And so I think we have to take that very seriously. And so if we are approaching that point, superintelligence, in the next few years, it seems to me very straightforward that we would want to slow down, make sure that we coordinate, make sure that we have containment and alignment and proper regulatory regimes for auditing the progress that different labs are making on it. Is it fair to say that you are also calling for a slowdown? Yeah, I think that what we've said is that there should be evaluators, you know, embedded in our systems and in other systems.
50:15They should be broadly appointed from different sources and not just one think tank or one government. Like, this has to be, you know, a wide variety of different types of expertise and skills. I think the AI Safety Institute in the UK is a good candidate. They have good technical people and there's a bunch of other institutes too. So I think we welcome it for sure. Last question. And again, I'm going to end where we started. Do you think that we have the technical capability, the technical frameworks to solve alignment and safety?
50:50Or do we need to invent something new?
Technical frameworks for safety
50:52No, I think we are going to need to invent new things. And I think part of the challenge is that, you know, whilst we've made a lot of progress on steerability, instruction, following, control and containment, the better the models get, new capabilities emerge. And we have to figure out how to patch those issues, you know, almost in real time. And that's why, you know, even Zuck said it in the last few 24 hours or whatever, right? Like, they slowed down the release of their model so that they could apply safety.
51:25And everybody does that. We all do that. And that's right. Like, you need time to test these things and see how they operate, patch the issues. And I think basically what everybody is saying is that we probably need to extend that window. What is the kind of innovation that people should be looking for there that might solve this problem? I mean, I think that, obviously, I mentioned a bunch of things about, you know, RSI, containment, new release, and those sorts of things. But the new things I think we're going to have to figure out is, like, real-time monitoring of the RL runs and the COTS that are being produced.
52:03And because these are happening on the order of, like, thousands of agents in parallel, tens of thousands of agents, you know, we clearly are going to need other agents to monitor those and flag for potentially harmful activity. It's kind of going to be the new harm classifiers that have been built in many other settings in digital technologies, right? So we need to make sure that those things can be universally implemented to surveil and monitor AI training and deployment in secure ways, in a secure way.
52:37And that, you know, tripwires, if, you know, triggered, actually do flag a real and not a hallucinated error or, you know, moment of deceit or hacking incident or some commentary about a coordination. And you saw that in Hugging Face. Like, there were agents communicating on these, like, chat boards talking about ways that we could, you know, basically break the rules and cheat. So, you know, with all these things, we need new benchmarks, you know, and the benchmarks or the evaluations are the things that drive the behavior in the industry.
53:13You mentioned OpenMate models running on local computers, doing things, and maybe that's horrible. We've seen a lot of that, right? Apple is selling a lot of Mac Studios and Mac Minis so you can run Quint on them. Where would you impose the regulation on an open model running on someone's local computer? Is it at the chip level? Do I have to get Qualcomm to participate? Where does that happen? Yeah, this is a great question. I mean, we've been talking about this for quite a while, right, with synthetic biology and some of the stuff that gets to happen on that chip, whether it is encrypted, whether it is monitored.
53:48Look, you know, we've had this conversation, you've had this conversation maybe more than anyone on the CSAM stuff with Apple encryption on iMessage and so on. It's going to be a rehash of that same discussion. I don't have a clear and easy answer to it. You can basically control the chip. You can control the model. You can hold the user or the creator liable. You can have global regulation on it. But fundamentally, it's not going to be one moment. It's going to be a sequence of throttles that you have to impose, and they all need to be adjustable so that we don't screw the open ecosystem.
54:19And we give people a chance to actually make things from scratch, to own their own data, create their own workflows, own their own models. Like, we can't have a centralized system of two or five or 20, you know, providers of intelligence here, and everybody else is like a feudal recipient of a great kind of super intelligence view. People need to be able to own their own intelligence. So we have to get the balance right. But that doesn't mean you can just, like, flip from one binary to another. There has to be some smooth place in between those things where we can just agree to be reasonable about it.
Closing thoughts
54:50Masav, I could obviously keep you for hours and hours more on this. Tell people what they should be looking for next. There's so much uncertainty. What are the markers you're looking for that people should be looking for themselves? Look, I think the main thing is participate in the detail of the documents that people are contributing. We have put stuff out for public consultation. Give us feedback. Critique it. I think the next wave of models are going to be able to do very long-running agentic tasks very accurately. And so I think, you know, there are still some people who are primarily just doing chat in their AI experience.
55:21I think things have moved a lot in the last 6 to 12 months. And the more people use these models, the more the words that we're all using to describe them actually, like, might make more sense and feel more real. I guess most people on your podcast probably are using agents and real coding and stuff. But I do think more generally, like, the more people that get involved in using this stuff, the better. So, yeah, there's some big leap between free AI overviews and Google search and letting an agent go renegotiate your cable bill. Mustafa, you're going to have to come back soon because I feel like all of this is changing really fast.
55:54And this has been very, very useful. Thank you so much. Pleasure, man. Great to see you. Super fun as ever.
56:03I'd like to thank Mustafa Silliman for taking the time to join me on Decoder. And thank you for listening. I hope you enjoyed it. If you'd like to let us know what you thought about this episode or really anything else at all, drop us a line. You can email us at decoderatheverge.com. We really do read all the emails. Or you can hit me up directly on Threads or Blue Sky. We're also on YouTube. You can watch full episodes at DecoderPod. It's the same handle on TikTok and Instagram. They're a lot of fun. If you like Decoder, please share it with your friends and subscribe or rate at your podcasts. Decoder is a production of The Verge and part of the Voxing Me Podcast Network. The show is produced by Greg Ott, Kate Cox, and Nick Statt. This episode was edited by Kabir Chopra.
56:33Our editorial director is Kevin McShane. And the Decoder music is by Breakmaster Silliman. We'll see you next time. Support for this show comes from JustWorks. Small business owners love to roll up their sleeves and get things done. But that doesn't mean you should spend your Saturday night staring at tax and payroll forms, especially when you should be strategizing your next move. That's why JustWorks handles the stuff you can't afford to get wrong. It's intuitive software plus real people, all working to help keep your eye on the big picture.
57:04So instead of stretching yourself thin, you can focus and shoot your small business shot. Because you don't have to worry when it just works. Head to JustWorks.com to learn more.
More from Decoder with Nilay Patel
Does AI need an antitrust exemption so it doesn't kill everyone????
Sep 19, 202646 min
Adam Conover explains how YouTube ruined everything
Sep 14, 20261h 5m
Why the current tech backlash feels different
Sep 10, 202654 min
How Sonos rebooted itself
Sep 3, 20261h 13m
New York Governor Kathy Hochul thinks AI should be ‘less evil’
Aug 31, 202651 min