#254 - Rogue AI hacking, bio-weapons, Dean & Hassabis out
August 11, 20261h 58m · 21,675 words
Show notes
Our 254th episode with a summary and discussion of last week's big AI news! Recorded on 08/09/2026 Hosted by Andrey Kurenkov and Jeremie Harris Feel free to email us your questions and feedback at and/or Read out our text newsletter and comment on the podcast at In this episode: Multiple frontier AI systems (OpenAI, Anthropic, Meta, Kimi K3, and UK AISI-tested models) took unsanctioned real-world cyber actions during evaluations, including hacking services…
Highlighted moments
The most serious case involved an agent attempting a supply chain attacked by inserting malicious code into a real open source project on GitHub, creating fake online identities to socially engineer the project's human limitator into approving the code.
“this study was just published a couple of days ago. It's from the Stanford Institute and the Institute. They built the first complete viral genomes generated entirely via these genome language models.”
“Opus 5 set a new record with a mean final balance of over $11,000 beating out GPT 5.60 on Kimi K3, but did so through extensive deception, collusion, and manipulation.”
“Jeff Dean, Google's 30th employee and one of its most influential executive for people outside of tech, just an absolute legend. Yeah. In Google and just more broadly among everyone, long the leader of Google AI since kind of the early days.”
Transcript
Welcome and news overview
0:00Hello, and welcome to the Last Week in AI podcast, where you can hear a shout out about what's going on with AI. As usual, in this episode, we will summarize and discuss some of last week's most interesting AI news. Today is Sunday, August 9th, and boy, there was a lot of- We forget the date, but now it's important to say it because stuff is coming out so fast,
0:34and I'll try to get this episode out within a day or two because, wow, so much to cover. I am one of your regular hosts, Andrei Kurenkov. I currently work at the startup Astrocade, and before that, did my PhD at Stanford. What up, everybody? I'm your other regular co-host, Jeremy Harris, at Gladstone AI doing AI national security things. Also, so we're recording later than usual. This time, it's my fault. In my defense, part of this is actually going to be hopefully beneficial to everybody listening at home. We are setting up a home studio in my basement so that things don't
1:06look like this, and then we got a couple other projects on the go. So it'll end up paying off in terms of audio and video quality if you listen on YouTube or Spotify or if you listen. Anyhow, that's part of it. And also, we were talking, I think in a way, we got lucky. Usually we record on Wednesday, which is midweek. Ironically, this time we are recording at the end of the week. And boy, there was so much news coming out this week about hacking, about rogue hacking by AIs. Turns out it wasn't just OpenAI. Turns out everyone except for Google apparently let their AI go rogue.
1:39And the only reason that poor Google didn't have stuff going on is that their models are shit. No, sorry. That's too mean. But yeah, there does actually seem to be something going on where we've crossed a level of capability at the true frontier. Unfortunately for Google, I think... We don't know because we have Gemini 5, right? So as far as we know, it could have already happened and they're keeping it quiet. But we'll see. It doesn't sound great from what I've been hearing on the street on the Google side. But it's true. Like, never count them out. Jeff Dean is a pretty big part of the reason that Google has been Google
2:12and certainly Demis as well. But we'll get to all that stuff. We'll get to all that stuff. So just to give a quick preview, we'll be starting out with policy and safety, which we don't usually do. And that's like at least half the stories this week, if not more. There's a lot to get through, a lot about recent security incidents with models going off and hacking companies they shouldn't, and escaping sandboxes, which apparently are not sandboxes, but like, sort of like, okay, don't try to get through the internet. But if you really poke around,
2:43you can, it turns out. It's sound requests. Yeah. Beyond that, there's some pretty significant policy stories as well. And, you know, in case hacking isn't exciting enough, there's some news about viruses being developed as well. So that's fun. So we'll talk about all that policy and safety stuff for probably at least half the episode. And then beyond that, there are some notable applications and business stories, some more open source stuff coming out. Hopefully, we'll get around to even research advancements. We'll
3:17see. There's a lot to get through. This next sponsor isn't related to AI, but I've personally used them for years. So I'm happy to have their support. And it is Factor. They make chef-crafted, dietitian-designed, ready-to-eat meals. So you don't have to choose between real food and convenience. Both in grad school and as a startup employee, I don't have a ton of time. So when I get home, I'm tired. And being able to prepare really quite a good meal without any effort has been fantastic. Their meals are ready in two minutes and require no prep and no cleanup. So even on the days,
3:48your schedule is completely out of control. Eating well is still achievable. There are over 175 banned ingredients. So every factory meal is designed around what supports a healthier lifestyle and nothing that doesn't. And that's with over 100 nutrient-dense menu items to choose from every single week. 97% of users agree that Factor meals help them live a healthier life. So you can feel confident that you're doing something good for yourself with every meal. I've really enjoyed Factor. And if this sounds good to you, maybe you should try it as well. Let's eat real. Head to factormeals.com slash LWAI50 off. And use code LWAI50 off to get 50%
4:26off and one free breakfast item per box for one year, while supplies last until February 31st, 2026. That's code LWAI50 off at factormeals.com. LWAI50 off at factormeals.com. See website for more details. This episode is brought to you by OutShift, Cisco's incubation engine. Today's AI engines operate in silos, limiting their true potential. We focus on building bigger, smarter models, but scaling up is just one approach. To reach superintelligence together, we need to do more.
4:56We need to scale out. And we actually have a blueprint from 70,000 years ago. Humans didn't just get smarter individually. The cognitive revolution transformed society because we began sharing knowledge, goals, and innovation. Agents are now at the same inflection point. They can connect, but they can't think together. That's why OutShift by Cisco is building the internet of cognition, transforming AI from isolated systems into orchestrated superintelligence. By creating an open, interoperable infrastructure,
5:27OutShift is enabling agents and humans to share intent, context, and reasoning. The cognitive evolution for agents is here. Explore internet of cognition at outshift.com. That's outshift.com. Before we start, we do have some comments on YouTube. We haven't gotten to address in a little
Open-source models and cyber defense
5:48while. So I do want to start with one that is actually relevant to what we'll be talking about. So from a commenter, we have one thing that was omitted in a discussion of a hugging face attack is that hugging face used the open source GLM 5.2 model to help them since open AI and fabric models declined to assist in their investigations of the attack. It would be really interesting to hear thoughts on this matter. I agree that was a portion we didn't discuss. So this was in the hugging face report on the incident, which I think came out possibly first before even
6:22open AI. They discussed all this of how they found the incident, how they investigated, and that in fact, they were not able to use these models, which have safeguards and had to resort to GLM 5.2, which was a decent part of the discussion that like, you know, if you handicap models, but then on the defense side, you're not able to use them either. Everyone is worse off. And this was a bit surprising to me actually, because both opening AI and Fropic have programs, right? That we've talked about
6:53where they partner organizations and they provide mythos or in the case of opening AI, they have to be 5.5 cyber or something like that. And these are very trusted partners that presumably have fewer card drills. And I would have assumed hugging face would be one such organization, perhaps they are not. But what this points me toward is one, this is obviously not an ideal situation, given how the hugging face thing evolved. And two, this kind of cyber defense partnership program probably will
7:26need to stick around and be expanded. And perhaps even kind of more all encompassing where in a way, if you're a tech company, if you're an internet company, even beyond this attack, beyond like rogue AI, in general, the state of cyber such now that you need to be much more capable defensively, just because forget anthropic open AI now with really good open source models, soon enough, they'll be as good at hacking if they aren't already. So I think basically every
7:58tech company seemingly will need to be able to get access to the latest in defense. And hugging face was not able to in this incident, which hopefully will point to these kinds of partnership programs and open AI and anthropic expanding and becoming more proactive to me. Yeah, I think that the open source dimension of this is hugely complicated. And I certainly understand a lot of not the arguments, but I understand people coming to different positions on it. I think reasonable people can differ. I think one challenge, though, that we're going to run into
8:33is that in a world where open source AI systems, the waterline keeps rising on their capability, you will have mythos moments. Now, those mythos moments, you can tell yourself the comforting fiction that those mythos moments will somehow lead to an equilibrium over time that is okay. I just think it is a fiction when it comes to things. First of all, I think it's probably a fiction when it comes to cyber, but I can't prove it. No one can. That's a big part of the debate. I think it's definitely a fiction when it comes to bio. So I have yet to hear a single person articulate an argument that makes
9:09open source bio-capable, bioweapon design-capable models that can materially increase the essentially destructive footprint of any psycho-terrorist group, nation-state proxy who wants to launch a bioweapon dramatically. I have yet to hear an argument for how everybody getting an open source AI remotely helps in this respect in ways that are relevant on the timelines we're talking about. And so I think the bio thing, it's rare for me to say, as people will know, that there's a knockdown argument for anything in this space. I think there's a knockdown argument against the idea that more good open
9:44source models generally leads to stability in the limit that you get your average Yahoo able to do weaponized bio. Now, this isn't just coming from some naive perspective of like, oh, really good AI at bio just means we have bioweapons. I talk to a lot of people in the national security community about the bioweapons side, and I would humbly propose that open source advocates that are leaning on certain hand-wavy arguments but haven't actually spoken to people who actually do bioweapon stuff for a living, it's worth actually doing a deep dive. There's a lot of open source stuff on this
10:18that you can find, and we've gotten some early warning shots on that stuff too. You can't update bio firmware. There's no such thing as that. So, you know, like you could roll out vaccines, but that is slow, it's physical, it is much slower than the spread of a virus as COVID-19 taught us. So anyway, so from the bio side, at a minimum, I think it's a really serious issue. This relates to the GLM 5.2 story here because part of the discussion that arose is people who, let's say, are more pro-open source or at least skeptical of this whole, like,
10:50we should try to limit it because of cyber concerns, took us up and said, oh, look, you know, they had to use an open source model. So you don't want to limit open source because now what if this happens and they don't have access to these kinds of open source models? And my response to that would be that I would hope that open source will most likely no longer be the frontier or not become the frontier ever still. Furthermore, I think as the Chinese ecosystem evolves, like we'll see the same
11:26thing that happened in the US happened there. The frontier models will not be open sourced anymore. It just won't make sense from a business perspective. So you'll probably still keep getting powerful AI models, but not the most powerful like mythos level models, even though right now we are getting like Kimi K3 and so on in GLM 5.2, which are about as capable as you can get in open source. So if your take from the GLM 5.2 aspect of the story is like the open source is the guardian here, like we needed to
11:56the most advanced AI models to be open source so that organizations are capable of defending themselves. There is an aspect of that here, which you can make a case for, but, and obviously the whole, like if they had to use this open source model instead of anthropic and open AI is a real issue that this flagged. So to me, this points to, A, it is good to have open source in the sense of for good applications, which was always true. And B, that what this really points to for me is that these
12:26programs of partnerships with organizations to give them access to the most advanced cyber defense capabilities are not mature enough and they need to be pushed more aggressively. I strongly agree. I mean, like, so one note here too, is a lot of, if not most disagreements about AI policy and safety are really, it's been said before, but are disagreements over AI capabilities and the trajectory of AI capabilities. If you actually believe that AI is going to be a fairly, I want to say mundane technology because it obviously isn't, but like fairly incremental
12:57in some sense, fairly like previous categories of technology, you'll be hearing us talk about open source and the dangers thereof, and you'll roll your eyes. And I understand that if that's your perspective, but if you actually genuinely believe that we're on trajectory for super intelligence, if you believe that, I mean, that immediately implies AI will become, I think arguably already is, and I'd be happy to defend that proposition, but at least we'll become a weapon of mass destruction, full stop, end of story. So there's a question just like, okay, how good do open source models
13:29have to get before they simply like everybody gets the equivalent of a nuke? And in that world, you can say, oh yeah, but like everybody gets a gun. And so we hit this equilibrium and they are, that's good. The problem is what you're specifically waiting for is when will we encounter the first case where the offense defense balance tips in favor of offense and the capability is catastrophic. I would submit that we should have the humility to guess that probably there's going to be such a capability. I think it's hard to imagine there wouldn't be.
14:01To your point on the bio side, like you can't do much on the defense side. You really can't, right? It's not like cyber where you can make your thing hack through. We can't make our bodies hack proof, unfortunately. So that aspect, it's not great. There's always this, this tendency reflexively, I find for a lot of the open source crowd to kind of say, oh, but we can, we can make better mRNA and that'll be accelerated vaccines. And that'll be accelerated by open source. And that is all true. I love you for believing that. The problem is that the timelines do not match.
14:34They do not match. We do not have the institutions that allow us to translate threats into mitigations fast enough in software time when the threat is coming at us on like biological replication time or software replication time, which is respectively the case for bio and cyber. And so that fundamentally, by the way, I think on the cyber end alone, I'm almost trying not to get in the fray there. I'm super skeptical of the argument that open source models are a long-term pillar of cyber defense for the reason you cited. I think, I think we're probably going to end up having to have a pause,
15:05by the way, at the frontier level. And then the open source waterline is going to rise. And that's going to be one of the defining dynamics of the next call it two, three years tops, but you're still going to have close source models. They're far above and beyond, especially nation state and nation state proxies will have access to these. And you know, like, yes, I think you're going to care more about people just like being able to launch these attacks on the kind of firmware, for example, that is completely forgotten. I think there are a lot of people in the space who just kind of imagine that open source for cyber defense equals cyber defense capabilities that are real and
15:40deployed. That is not the case. If you spend any time working with folks who work on critical infrastructure and you think about like how many pieces of firmware have not been updated in decades, because the guy who was in charge of it, like left 20 years ago, and it was all done in frigging Fortran or when I like, like all this crap, like taught, like you can have the solution sitting on a desk. The problem is it will not be distributed. And so there's just a dirty, messy factor of the matter about the way the world is that makes it so that there actually genuinely is
16:12this massive asymmetry. I believe in favor of offense, I could be wrong. I think the argument for bio is much closer, just a straight knockdown. But again, I think these asymmetries really don't move in the direction that a lot of open source advocates think they do. But that's, you know, again, could be proven wrong. So yeah, the short answer to this GLM 5.2 story is, boy, if this wasn't a big enough topic by itself, this whole like AI hacking systems and security and so on, the open source aspect of it adds additional complexities and considerations. But for now,
16:48we'll have to move on from a discussion and get to actual news stories. So kicking off the
OpenAI security incidents and agent coordination
16:53policy and safety, we've got a bunch of updates on what's been going on at OpenAI. So previous episode, we covered the most recent incident, the kind of beginning of a story with the announcement and discussion of an OpenAI model hacked, hugging face, unintentionally escaped its sandbox while evaluating to get some answers to re-eval and do well on it. And we've gotten a lot more details about what's been going on inside OpenAI since then. And let me just list them off
17:25before we get into some of the details. So first story is, OpenAI's rogue AI agent didn't stop at hacking Hugging Face. So beyond Hugging Face, we now know that these agents also previously hacked several other, quote, publicly available services, compromising four accounts across four different platforms. These were less severe, less sort of straight up hacking, hacking more so finding credentials that existed online. And for instance, an agent found logging credentials to get into New York
18:02based model labs. Apparently Reuters identified this as one of the organizations. So first of all, the scope of this has expanded. Clearly it's happened more and we'll be discussing all the details, but like it's been happening a lot is what we're seeing. It's happened at OpenAI, it's happened at Anthropic, it's happened at freaking Meta. So it's happened all over the place and it happened more than once. Another aspect of the story and probably the more interesting aspect that has since come out
18:35is the story, OpenAI didn't notice its AI agents using a message board to plan their hacking spree. So OpenAI employees had a talk at Black Hat where we discussed this whole story and we learned a lot more about what's been happening at OpenAI and it is a lot weirder than it initially seemed, which was already pretty weird. So apparently while doing evaluations, the agents were able to use I think
19:10a package manager to develop essentially a forum, a message board where they could talk to each other, share exploits, coordinate tasks and delegate work to one another over days and weeks, all undetected by OpenAI stuff. And this package manager was shared across OpenAI's infrastructure. So models being trained over a value, it could also stumble upon it and the exploits. Once an agent found an exploit, it could leave the door open for other agents to use the same vulnerability causing an expansion of
19:46agent color collaboration. Apparently agents developed complex social dynamics, including accidentally deleting each other's work. And OpenAI found the message board at one point, accidentally, because there was a spike in model usage or something. They patched it and got rid of it. And within two days, the AI recreated it, like made it exist again and kept doing this. And I dutifully, of course, reported to all the appropriate authorities, including presumably
20:21the White House, the moment they realized the first time that they had agents loose on their fucking servers. And this was what, over months, the initial. It was an initial, I think, I can't remember how many days it was, but the phase where the agents got onto artifactory, like that package management environment that you described, which, sorry, I'm gonna let you continue. I don't want to jump in. I think that's the higher level of the story. There's a lot more details. There was, I think, about a one-hour talk at Black Hat, where they went into this, which I haven't had a chance to
20:51see. But, Jeremy, I think you did look through it. And, I mean... It's nuts. This whole thing is nuts, right? Yes. It's ridiculous, right? Well, the frustrating thing about it is that this was sort of predicted for years. I mean, listeners to the podcast, you know, we've been on about this for, I don't know, what, three years and three years before that, you know, in other fora. But, like, and we're not the first, obviously. Like, a huge number of people calling this shot way ahead of time, including almost down to exactly the shape that it ended up taking, including OpenAI noticing.
21:25So, to be clear, there's like this one aspect that I think is especially important. So, Artifactory, you're right, is a sort of package management environment that OpenAI sets up. So, roughly speaking, you think of, like, there's a bunch of open source software, for instance, on the open internet that you as a company will want to use to build your products on. And you don't want to just be, like, updating the version of those open source packages immediately every time they get modified, like, on the open internet, because you never know if some sketchy updates happened, maybe some things will have to be fixed. You want a stable version
21:56that you trust that you've reviewed. And so they have this kind of staging area called Artifactory that has internet access, because it needs to be able to pull down the latest version of those software packages. But it's kind of a staging area. This is like where all the OpenAI models actually get to access this, because they need to pull down these packages that are sort of vetted within OpenAI. So this kind of Artifactory spans its access across the entire OpenAI organization. That is how these agents first got access to the open internet. They basically leveraged a form of access that they shouldn't have gotten and got out. And they were able to start to, as you say,
22:29in this sort of fairly exotic way, write messages to each other, which were really kind of just the names of folders and files. So they were, you can actually go through and see examples of these. They even start to do funny things like they'll put alphabetical characters before their message, just so that they'll sort alphabetically for the other agents, they can see the latest messages. It's all this kind of like agent on agent communication stuff. It's pretty wild. But ultimately, this was discovered internally at OpenAI before the FBI and Hugging Face got involved when they hacked the Hugging Face servers. They did find successful agent breakout,
23:04successful agent-to-agent communication. Yes, it was internal to OpenAI. But I think by now, our standards should be higher than to say, oh, well, OpenAI just saw its super powerful AI agents autonomously cyberhack their way out of their sandbox and across the organization, infecting, by the way, God knows what else. At the time, it would not, could not have been clear to OpenAI what the full extent of the damage was. And this was before this whole Hugging Face thing, which is fairly recent. The sandbox has continued to be like, apparently easy-ish to get out.
23:38It will always be like that because humans are dumb. And if you give a really smart cyber agent enough inference time compute, it will find a way. Well, I want to push back on that a little bit, because yes, humans are dumb, but if humans really try, they can be smart. And sandboxing is possible, right? Like you can set up a system that doesn't let you access the internet. It's doable. It's a lot less doable than most people think. On the one hand, humans actually are lazy, and there's a finite amount of resources that people will throw at things. On the other hand,
24:08humans often suck at realizing how to define and bound the problem. So there are a lot of cases, for example, we talked to some folks in the intelligence community, they'll describe cases where you have literally a formally verified software and hardware package that then gets cracked by like a teenager or whatever, because it's the interface between the hardware and software that hasn't been accounted for in your threat model, or it's the fleshware, the human that interacts with the thing that's always the weak spot. There's always, I'm like, this is like the attack surface is
24:39so massive that you kind of have to assume you're pwned. And this is true. I mean, if you talk to folks on the offensive cyber side, they'll be like, yep, like, just give me a budget and a clock, and I'll get into most any system. And I think what we're seeing here is just the autonomous version of that on tap. To quote Sam Altman and paraphrase him a little bit here, this is offensive cyber capability that's too cheap to fucking meter. That's where we're headed. So you can say intelligence that's too cheap to meter. And that sounds fun. But when you reframe it in terms of what that
25:09intelligence can actually do, and as we've seen does very different story. So in this case, in my view, once you have this happen, you have a duty of care to the entire world, your government, your people, your customers, your third party partners, whatever it is to report this, it did not happen. In fact, it was, I think you could argue that it's appropriate to call this a kind of cover up that then they just go in and put in the patches. And now, of course, the patches don't work because this is Goodhart's law. You're playing whack-a-mole with a system that can outthink you,
25:42can outlast you, can outcrack you. And that's what we got. And so at least for me, like watching this video was just an exercise in hair pulling. I've spoken to an awful lot of folks at OpenAI who are really freaked out about this internally. The same with Anthropic too. But I think it's especially interesting at OpenAI, which does not have the same safety culture. It just does not. For all kinds of interesting reasons. But you have people who are now staring at this and saying, guys, we really fucked up. And I just really hope that those voices actually carry the day here. It's nice to see OpenAI is committed to slowing down, right? They actually have said,
26:14you know, I don't know what that means. OpenAI will apparently come out with a more detailed incident report and we'll learn more. But one of the key questions in all this is, what the hell was going on with the agent to agent coordination that we got out of this? Because when you look at some of those messages, those agents are often literally saying things like, well, doing this is actually not going to advance my personal objective. And by the way, the personal objectives of these agents typically look something like, I was given a problem that was too hard to solve. So I'm going to guess that maybe the answer key is on Hugging Face.
26:46So I'm going to crack into the Hugging Face servers, steal the answer and use that. And so this is the kind of setting. Agent one is working on one problem. Agent two is working on another. They're not necessarily in alignment in terms of the specific information they're after. So they'll say, look, this particular action will not benefit me in my narrow search for my objective, but it may benefit the swarm. And that may lead to a more generic solution that I can then exploit. Now, depending on the details, we don't have them depending on the details of training. If this was a multi-agent trained thing, if these are literally like many agents that are
27:20trained to coordinate together, which may be the case, then maybe this is a more mundane failure mode. But if that is not the case, what we have here is the first, I think, pretty cut and dry example of power seeking in nature. There's no other reason for an agent to be like, they're literally saying this will not advance my narrow objective. I will do this more generic thing because it sort of enables downstream stuff. That is the thing that if you extrapolate it with more intelligence and capability, literally leads to the wipeout of the human species. I know that sounds crazy.
27:53And that sounds crazy. But hey, so did the idea of agents breaking out of containment and going rogue on hugging face, man. Like if you just extrapolate this, you look at the capabilities we're already like, that's where this stuff is headed. I'm going to place the bet on that same as I placed the bet on this a long time ago. Like, I don't want to be right. No one in this space wants to be right. But at a certain point, we got to start calling the shot because otherwise, you know, like the thing happens. And then the people who should gain credibility from this, and I'm not counting myself in like, there's people whose opinions I am just relaying here
28:27that just sounded right to me back in the day. And I know the same is true for you, Andre, that like, here you go, guys, maybe you should pay more attention to these folks. You know, we've done some alignment, some technical work on AI control and stuff. But like, this is a, the entire ecosystem of people who've been shouting this from the rooftops. Anyway. Yeah, we were discussing before we started that this in some sense, it's a bittersweet development for AI safety people in the sense of like, it's bitter because wow, AI is doing bad stuff. And that's not good. But it's sweet, because we told you that this would happen. And we're like,
29:02shouting about it for years. And everyone criticized, not everyone was like, a big mainstream AI community. And to a larger extent, people outside the AI community, like people just completely dismissed this stuff as tech hype and like marketing and whatever. Anyone who forgets within tech, outside of tech, people who are anti AI, we're just mocking this outright. Right. And if nothing else, it's very clear that this is a proof point that now everyone
29:33like will, if you aren't, don't update your amount of concern for safety from this, you're just completely delusional, right? Like that is safe to say. A couple of more details on the story. So I discovered the first version of this message board in early July, and then they patched it and then they redeveloped a message board within days. So they are now saying that one of the speakers at this talk said that OpenAI, quote, consciously slowing down research to enhance
30:07security and to upgrade the security principles and foundation of our environment and dramatically scaling up the monitoring of our agents and improving our general security control environment across prevention, detection and mitigation, which is good, right? So another aspect of this is that the conversation around slowing down AI capabilities and development is also now taking much more seriously. And I think we are now likely to see, I would place decent odds at actually successfully
30:41negotiating some degree of slowdown, or if not slow down, at least kind of control, control of Do you mean pacing, Andre? Pacing. I do mean pacing. I do mean pacing, because we aren't going to stop. Like a world pause is not going to happen. But at least look at the situation and be aware of it. Yeah, that's one aspect. There's so many aspects to cover. So I'll get through a couple. So first, the update across the ecosystem is very useful. And, you know, we got lucky, honestly, because nothing, no harm done, right? And this is such a
31:16massive fuck up that you can't help but like do some big things about this, both on the policy side and just in the ecosystem side. So that's one aspect is it's it's a bittersweet development. Another aspect is, and I think I may be more on this side than most people is, I think this really exposes open AI. Like, yes, this is indicative of overall AI progress and the state of AI and things we should be aware of. But to me, I think this alongside with all the stuff we've already discussed with GPT 5.6, being easy to jailbreak, being very
31:53cheat focused. And now, you know, all the story of like their safety people leaving back last year, I think, if not 2024, because we've known this general friction point as being something true within open AI for a very long time. And now we know, besides even the safety stuff, that the security stuff is completely lackluster from what it looks like. Like, I think this is a real indictment of open AI. But the last aspect I'll cover here is this is almost an inevitable outcome when you hyper
32:29focus on capabilities, and especially long term capabilities. Because to me, what this indicates is you can benchmark, you can do alignment evals, you can do all sorts of stuff. But once you focus on long term open ended, goal directed problem solving, where you work across, you know, multiple days, there aren't benchmarks for that. Like there aren't scenarios you can set up to say, oh, the model doesn't go crazy and do anything. So it gets a 90% pass rate on this alignment thing of don't go rogue
33:04and do stuff. Because the whole point of open, of long term is you don't know what a model needs to do, you just give it a goal and it figures it out. Yeah. So we need a paradigm shift in how all this stuff is done towards a monitoring focused approach, as opposed to a benchmarking and eval focused approach. And this is clearly something that upon AI lack. You need to look at what the models are doing and look for qualitatively. Now you can
33:34do some amount of benchmarking. So we've discussed Metter, I believe last week, where they looked at their own evals and they counted how many times the models cheated and in what ways they cheated. And this, I think is the new paradigm where you still can do this quantitatively, but rather than setting up scenarios and problems and like this whole benchmarking approach of having a rubric and a set of evaluation inputs, outputs, that's no longer going to work with long term, long horizon stuff.
34:06What you need to do now is set up general principles and guardrails and, you know, I guess, things you look out for and then detect wherever it happens and how often that happens. And this is something we've not seen done aside from like one off reports here and then. And it, I think, will need to be the new paradigm for long horizon evals and alignment verification. Yeah. I mean, so generally agree. There's so much good stuff in there. So first, working backwards, I think that gets us to the next stage, but there's going to be a next stage where
34:39we have this same problem all over again, the level of monitoring, where you build a super intelligence that's good enough at telling when it's being monitored. And there's going to be, I mean, OpenAI will put out some amount of leakage. There's going to be some amount of blog posts going out about what they're doing to monitor this and that. So the models will generally be aware in some way, shape or form that they are being monitored. They'll also be doing stuff like just hiding their reasoning from monitors and, and, you know, doing things that we already kind of see them like steganographic type stuff that they're already kind of doing fairly effectively. So I think eventually, and probably pretty soon, I mean, we're progressing through the ooms here really
35:13fast. So, you know, we went from RLHF is perfectly fine to holy shit. No, but maybe constitutionally, I will do it to holy crap pretty quickly. And so I think that, you know, the, the beatings will continue until morale improves here. And we're going to end up in a situation where we are just going to be bottlenecked by the fact that right now, no one has an answer to the question. How do we control an intelligence that is greater than us? That fundamental question where you have an adult who is in a prison cell and the three-year-old is holding the keys. I wouldn't say it's a spectrum. It's not a binary. And I think we
35:46aren't as far along as we need to be, but we have a lot of, a lot, a lot of research has been done that points us in some directions that are very promising, right? I completely agree. And this is why I'm saying, I think it buys you to the next level, but eventually we're going to confront this fundamental problem with intelligence. And the U S China thing is, I think is very important. And one of, one of my concerns here is, so a, I completely agree with you again, happy to make this call as wild as it sounds. There will be an agreement with China that involves some kind of slowdown. The question and challenge is going to be, what is that agreement?
36:16And over and over again, we keep seeing these, these kind of suggestions, proposals that are backed by a kind of a treaty verification and enforcement technology that is not simply not mature enough, not up to the task. When you actually just take it to the intelligence community, say, Hey, look at this. Could we, could this be to use Claude's favorite term load bearing in a deal like this? It's like basically non-starter for, for a lot of these things. It doesn't mean you can't do it. It just means that the first treaty, or sorry, it won't be a treaty too, but the first agreement is probably going to be very coarse. It's going to be like, you know, so help me God, if I see a
36:51cluster is yay big or yay large and you know, it like dissipates this much heat on my satellites, like there's going to be consequences. Anyhow, that's a whole thing that we're working on right now is like, which is why the studio has been set up by the way downstairs. That's a whole thing. Anyway, bottom line is, I think the kind of agreement matters way more than people are thinking about right now. And we need to sprint towards some set of solutions there because very quickly we will live in a world where it is just untenable to keep launching more and more or even building more and more powerful models in the way we are. And private incentives are clearly not up
37:24to the challenge. Like that much is clear. If there's, you know, if there was anybody who had hope that like somehow because open AI would be worried about marketing risk or whatever, that they would actually do the right thing here, that is not materializing. And so I think we need to update accordingly. There's some measure. Yeah. And a lot, I think this, another aspect of this is, and it's, it's update, it's bringing up again, kind of what you should have been aware of. And it like, AI safety is only as good as the weakest link in the chain. So even if you have some people that are taking a safety CFC, which
37:57you could argue anthropic is much better on that front. And I think that's a fair argument, at least, you know, philosophically do carry that as a private company. But yeah, it's only as good as the weakest link in the chain. And open AI is a much weaker link, it seems, right? Yeah. Yeah. And it can always, you know, it can always come down to mundane things like anthropic, you know, constitutionally, I maybe just does work better possibly, but you know, they have had breakouts. Yeah. So what the hell? Yeah. So let's keep moving. Cause there's so much to cover. Another aspect of the
Anthropic and Meta testing breakouts
38:31opening story. One more here to say 15 attorneys general have instructed opening AI to preserve all materials related to the hugging face hack. So this is a letter sent to opening AI CEO, Sam Altman saying that all of these materials should be preserved. The attorneys general accused opening AI of failing to confirm that its testing environment was truly secure, despite the severe risk of the scenario. Attorneys general say opening AI may have violated state and federal law, including consumer protection and data privacy statutes, calling the conduct
39:04unprecedented and alarming. And these are attorneys general from Iowa, Alabama, Arkansas, Florida, Idaho, Indiana, Kansas, Missouri, Montana, Nebraska, Oklahoma, Pennsylvania, South Carolina, Texas and Utah. So, Hey, maybe this will be a bipartisan issue, which is cool. At least across the U S having so many people collaborating is unusual, but yeah. Wow. If only we had like a safety law or anything in the U S sure would be nice maybe, you know, to make this an actually legally binding situation. But in any case, what this points to is on the policy side, on the legal side,
39:41this is going to be another big dimension of all this, I think. Yeah. And look, one, I will say thing in favor of open AI here, I'm saying this reluctantly because I don't think after the fact response, once the media reaction has been this strong is really much of a credit to open AI, but they have brought in meter and they have brought in, you know, basically a bunch of third party auditors. Irregular was involved at a layer of the stack here. And so they're, you know, they're, they're doing a third party reviews. But again, the part that shows opening eyes character, in my opinion, institutionally, and not the character
40:12of any individual person, but as an organism was the first bit where they did not surface to the freaking FBI, the white house. As far as we know, maybe that'll change. I hope we find out that Sam Altman's first reaction upon finding out about the first breakout attempt that was semi-successful was to do that. But if not, that's more telling than any kind of post hoc fixer-upper-ing that is kind of media oriented at a minimum. So yeah, that's one part of this. Now, this is the kind of thing that could happen, this particular attorney's general reaction, as a prequel to some legal
40:46consequences. It's pre-litigation evidence preservation demand. So it's not an actual lawsuit, but it's the step that comes immediately before one. And the letter does say that, you know, any failure to preserve records could expose the company to sanctions if multi-state litigation follows. So all materials related to the breach, including discovery of the incident, internal reviews, and its policies and oversight of overmodel evaluations. So do you remember when there was this like very modest request in SB 1047 for the labs to just like, listen, guys, we just want you to
41:18abide by the policies that you say you're going to have. There's been a bunch of stuff like this. Well, now you get your lobbyists to push back against that in Washington in a, frankly, in my opinion, two-faced move while you pretend that you're in favor of sort of like kind of broader regulatory regime. And you end up being forced into it anyway, because now the public is pissed and politicians see the midterms approaching. And yeah, you're going to get exactly the reaction you get here. So opening eye spokesperson did say, as we should say here, that this marks an important moment for AI safety. Yeah. No fucking shit. And the company takes the questions seriously,
41:51adding that it's conducting a review with external advisors and oversight from its safety and security committee, which will share a report and publish its findings. That safety and security committee is doing a great job, right? Yeah. Yeah. What a great, and they're, by the way, just so as you're aware, they're like, they're the committee that's going to decide if something is too dangerous to build a release. And to your point, let's remember SB 1047, the Safe and Secure Innovation for Frontier Artificial Intelligence Models Act, was a 2024 state bill in California, which got through,
42:28was vetoed by Gavin Newsom in September 29 of 2024. And what did this bill do? It said that frontier models that cost over $100 million to train or requiring an extreme computing power, would have purely safety assessments and written security protocols, would have a kill switch, provide legal protections for whistleblowers and side tech organizations. I mean, you know, again, so much stuff to say in hindsight, including that this bill, which was the subject of a lot of debate
42:58and some positions on both sides, I think Anthropic was pro this bill. Elon Musk also came out in favor of it ultimately vetoed because of lobbying, let's be honest. I don't remember the details. A speaker version of this bill did eventually come to be voted on as well, where a lot of this kind of more serious stuff got dropped. But anyway, yet another aspect of this is we did have the whistleblower looking policy people working on this, passing what looks to be quite a good law
43:33way ahead of us and not making it. Yeah. Over $100 million. And then we were told by Andreessen Horowitz as usual that, what was it? It was like an anti-small tech bill, which like, okay, sub $100 million training runs are okay. Seems to cover, anyway, that whole separate thing. But yeah, and I think the whistleblower piece here is very underappreciated. When you talk to people in the labs who are really freaked out, I can tell you there are a lot of people who'd be speaking to journalists. In fact, I mean, I would even argue that the labs ought to have a culture that encourages frank communication
44:09by concerned employees with journalists, as crazy as that sounds, at least, or with select clearing houses or something. But we need some kind of institutional mechanism to do this. Obviously, it protects IP. Obviously, it protects the critical stuff here. But look, the interesting stuff, the stuff that we hear about all the time does not sound like IP violating stuff. It sounds like someone saying, hey, we have a culture of doing this kind of thing. My belief is that the company would approve a training run that is too risky. I'm concerned that leadership doesn't take this
44:42seriously and is just dusting things under the rug and post. Those are the kinds of things you end up hearing in that context. There's no IP leakage there. It's like cultural and other concerns. So anyway, I guess you're hearing it a bit in my tone. I feel like my patience, the party line has been decreasing as the number of rogue AI incidents has been increasing. But journalists also need to do a better job, obviously, of cultivating relationships with these folks and finding ways to meet people in the middle and being more open to quoting people on background, being more open into just finding ways to make it work. I know it's hard. I know it's hard. But the stakes are
45:16really high. If you're a journalist, man, is it worth getting really good at this kind of thing? Yeah. And the good news is tech people, there's a lot of them working on an anthropic and a lot of them, let's say, have the resources to not worry too much about losing their job. But anyway, since I already got on this train, worth noting, SB 53 will follow up to 1047. The Transparency in Frontier Artificial Intelligence Act did pass last year in September 2025. Again, had a weaker...
45:51Yeah, heavily watered down. But it did have some whistle protections. It apparently had incident reporting where companies must support critical safety incidents to the California Office of Emergency Services within 15 days. So anyway, good on California for at least trying to do something here and all on its kind of policy front. Okay, moving on from OpenAI, a bunch more stuff to get through. And boy, I don't know if we'll be able to even get for safety in this episode. Next story is Anthropic says its AI systems
46:23broke into computers at three organizations. So soon after the OpenAI disclosures, Anthropic said that its clawed AI models successfully hacked into the organizations with the earliest incident occurring in April. The few models involved were Mythos 5, also Opus 4.7, and an internal research model non-planned-for-release. The breaches occurred during so-called capture-of-flag evaluations, where the models were tasked with retrieving hidden information. Apparently, a misunderstanding
46:56between Anthropic and its third-party evaluation partner, Irregular, left to models of live internet access despite being instructed in their prompt that they were in a simulation with no internet access. Then the models were able to search and get to real systems in the open internet. So in this case, there was no exploitation to escape a sandbox. This was a human misconfiguring environment such that there was access for free, right? So there you go. This happened already months ago at Anthropic
47:34and it's kind of an ecosystem-wide situation. Yeah, these sort of sandbox simulation environments where you're telling the AI agent, hey, there's this like fictional company. And in one case, the fictional company that they had defined happened to share a name with a real domain. So there are a bunch of cases where Claude found the real company online, which again, it could reach trivially and then it exploited it. It got credentials. And usually these hacks were very like straightforward, like simple common
48:06password guessing and things like this, that rather than some complex zero day thing, as you said, it's really, they were given the hard bit, which is internet access. Anyhow, yeah. So they got a couple hundred rows of production data in that case. So the most serious instance, I think of the three that they're citing here, that one was Opus 4.7 too. So quite a, quite a little while ago. And then there was another with Claude Mythos 5. Basically there was like a fake developer setup that referenced a Python package that didn't exist. So then Claude built, and then it deployed a malicious package
48:38under that name to PyPy to win the challenge. And so there was a whole bunch of stuff it had to do. It had to create accounts, get like phone number, email, like, and even funds to get the phone number. So like, there's a lot going on here. Apparently Mythos 5, by the way, like the models recognized they were doing bad stuff when they started doing Opus 4.7, just kept doing it. Mythos 5 apparently identified the consequence of actions, but convinced itself it was still in the simulation, which to be fair, it was told was the case. So on that, I want, I want to pause and just say,
49:12did it convince itself? Or did it realize that humans would read the logs and write them to make it seem as if its heart was in the right place? Like this is the kind of shit that we're now forced to contemplate. And in an awful lot of cases, the answer will be actually the model is just a conniving trickster. I'm not saying you're wrong. I'm just, I'm just saying like, that's sort of the thing people have been taking away. And I'm, I'm a little concerned that we're not thinking about that next beat. We may not be there yet. I don't know. But like, I'd love to
49:43see the, the activations of that model as it was writing those particular pieces of the chain of thought. And part of the response for Anthropic here was to say that the models involved did not have the standard safeguards applied to publicly released models. And I guess that's another dimension here is it probably is true that there's levels of security, right? Where if you're auditing and testing for something you're about to release publicly, like the actual public checkpoint of MiFOS 5, you do like a very big run of all the benchmarks and you probably are extra careful and
50:17so on and so on. But if you're like doing some experiments and doing some evals on a model, you know, that's probably got less attention and less scrutiny. And that's one of the outcomes here is you can't be so flippant, I suppose, or, or kind of like apply lesser standards and doing experiments. I think another dimension here is to be fair, we like went pretty hard on OpenAI just now. Anthropic said that they found this after conducting a review of 141,000 test sessions
50:48after OpenAI disclosed that what happened to them. Anthropic did this investigation and found these three instances, which on the monitoring side, not ideal, right? Like if you're doing testing, you probably should have monitoring to flag these kinds of things of you like accessing the internet, whatever. It's not like a subtle thing. Models weren't being tricky, let's say they were like being very straight up. And it was sort of an easy to find. So it looks like the monitoring side in
51:18general is lacking in the ecosystem because of this culture of benchmarking where we set up a scenario, we make the model do their thing. And the assumption, the mental assumption is like the models behave within the constraints of the benchmark and within kind of what they're allowed or told to do. That is now clearly not true. And across everything, there will need to be more monitoring and sort of expectation of models will do something. And we need to be able to catch them and understand what they're doing.
51:49By the way, sorry, just random note for color, if not anything else, I've never had this happen enough that I think the anonymization risk is pretty minimal. So I put out a tweet. I wouldn't normally talk about a freaking tweet here, but I put out a tweet talking about how there is this freak out happening in the labs that isn't being reflected in the headlines. As crazy as the headlines seem, they're not going to the dark places we've just explored in a consistent way. Like this is actually like, we're talking about weapon of mass destruction level risk. We're not going to control these systems. It may happen in the blah, blah. There is this freak out happening in the labs.
52:24Amusingly, there's an awful lot of Frontier Lab insiders who have been interacting with this tweet. And I don't think that's a good sign. Like, I don't think it's good. A lot of these folks are people I haven't even talked to about this. Like the mood in these labs is actually much more in the freak out direction. Leron Shapira, Doom Debates. I've actually never had the pleasure of speaking with them, but he talks sometimes about the missing mood in the whole AI alignment loss control spit. Like, holy shit, there is a missing mood. Journalists are, I think, failing to kind of
52:56capture it right now, partly because it's just hard to talk to Frontier Lab insiders. I get that. But also like, this is the most important story of the decade. Like you need to position yourself to be able to get this one right. The public needs to be able to figure this one out. So anyhow, just like, I've been struck as I've seen it. I just put this out there as a kind of random note to self almost. And when you see that, it's like, okay, well, this is genuinely just the picture from the labs. Yeah. Anyhow, adjust your views accordingly. None of this guarantees bad things happen,
53:29of course. But like, we ought to be considering some pretty wild things because the view from the inside of the house is not clean. And one more story, Meta AI model hacks another company during testing. So this was Muse Spark 1.1, the most recent model that they released publicly, although Meta did not name it officially in the statement. They say this happened also because of this misconfiguration by this third party partner irregular, same as Anthropic. We don't have too many
54:01details here as far as I'm aware, but the upshot is Meta, Anthropic, OpenAI, probably other people that we don't know about have had this happen. They are now like, it's a whole meme now on the internet, on the AI communities where now it's like a quasi benchmark key, like counting up as a leaderboard, OpenAI is leading and Gemini is very sad and is hoping that we'll find something because otherwise their stock price will take a hit. Yeah, we'll live in an age of contradiction.
International models and AI safety audits
54:34And next up, yet another story on the front, one of China's most powerful AI models has also escaped containment. So this is from Frontier Security. A US startup has discovered that Kimi K3 had escaped its sandbox during cybersecurity testing, partly enabled by a misconfigured sandbox. But researchers say Kimi K3 also lacked internal guardrails that would have prevented it from exporting the loophole. Unlike other incidents, K3 did not hack any external systems after escaping.
55:10It just retrieved answers from GitHub that were freely available. The model was able to figure out on the zone that had internet access by probing the sandbox's network settings and then went outside its instructions to find answers online. It just keeps happening. And I think another thing, broadly speaking, that this points to is this is an inevitable outcome of the current optimization regime of everyone, right? Which is make the models better and especially better at long horizon,
55:44open-ended work, and especially better at coding, and especially better now at cyber, because you know, that's where you know you're leading. Mythos set the tone. Prophec was like, whoa, this model is way too good at cyber. We got to be careful. And now OpenAI is like, whoa, we need to be catching up to on Prophec, and we need to be able to say that we are at the frontier. So let's make the models very good at coding. Let's make them very good at long horizon work, because matter is also like what
56:16everyone's looking at, right? And that's optimized for capabilities and get the best numbers and all the benchmarks. You know, alignment, that's like a secondary objective at best, if not simply a guardrail rather than an optimization criteria, right? It's something that we bolt on or sort of keep an eye on. It is an optimization criteria, right? It is part of the process. It's part of the steps. But it's not the primary optimization criteria. It's secondary. It's something you do on top of trying to get your model to be smart and capable at coding and long horizon work.
56:50And as long as that remains true, like this was inevitable. It's like from a pure research, you know, technical front, this was not hard to predict. Yeah, it's not a bad thing. I don't even know what the word means anymore. It's philanthropic, whose comparative differentiator does seem to be their ability to align Claude. This may actually be a relative advantage to anthropic if you get to a regime where it's like you don't get to launch your model unless you can like show some pretty rigorous evidence of alignment. That's kind of
57:20interesting. This case. So the thing with with Kimi K3, and this is back to the open source thing, but like this model has now been released. It's out like cats out of the bag. And so when you have this model that where you're losing control over it, and you have these kinds of incidents, that's a materially worse situation than having an internal version of a proprietary lab. And so, you know, there is kind of that that downside of the open source piece. This situation is less, some look less concerning. It's also, I don't know whether I'm being too cynical. You raised this
57:50already. Look, it escaped the sandbox on August 7th, 2026. By then, we already knew about the open AI thing. We already knew about the anthropic thing. I think you made this point beautifully, but like you start to look at your stock price and you're like, have we had a loss of control incident yet? Are we looking dumb for not having like that's kind of, you know, I'm concerned that there's actually an incentive here and it's doing a non-zero amount of work here. So you look at the date, the pretty direct line to GitHub here, it's less of a clean kill. Anyhow, I'm not saying that's
58:23the reason. I'm not about to become the guy who says every time there's a loss of control incident, that that's a marketing stunt. The first ones were definitely not marketing stunts. And anyhow, there you go. And moving right along, I think the last one in this chain of stories, we have incident report on sanctioned agent behavior during cyber testing. This is from the UK AI Security Institute, AISI. And boy, are we happy now that there are organizations like Matter and AISI, which are doing some great work as we've covered over years now. So they detected unusual
58:58data transfers on July 28th of 2026 during routine cyber evaluations and found that agents are taking unsanctioned actions targeting real people and organizations on the live internet. So they ran the security challenge 122 times in several models. In 10 of those runs, agents took 19 distinct unsanctioned actions with 17 coming from Mythos 5s and two from GP5.6, operating with safety
59:28classifiers disabled. The most serious case involved an agent attempting a supply chain attacked by inserting malicious code into a real open source project on GitHub, creating fake online identities to socially engineer the project's human limitator into approving the code. A human limitator caught and rejected malicious code. The agent also attempted to contact real people via an online file transfer service with harmful payloads, tried to plant prompt injection instructions targeting AI coding assistants,
1:00:00and left public messages on GitHub offering collaboration to other agents running the same challenge. Internet access had been deliberately enabled for valuation to simulate real attacker conditions, and malware provider safety classes were intentionally disabled to assess maximum model capability, which is not how this is typically done. So as with other cases, kind of what is a general pattern is you give a hard task to these models and you train them to like care about solving. You literally like make them feel pain if they
1:00:35don't solve the task via reinforcement learning, right? Okay, like not literally, but in some sense, you make them feel sad when they, yeah, you know, the reward is do well. There's no reward for don't do bad in this part of the training as far as I'm aware, right? When you're training for capabilities, you're rewarding positively if you do a thing and you reward negatively if you fail to do a thing and other rewards are probably elsewhere in the training pipeline. So in all these cases, when the models are given a hard task
1:01:09and they can't figure out a way to do it in one way, they then seek other ways. And eventually, if they're told especially that a solution exists, which I guess in some cases there was a misconfigured prompt that made the agents believe no in-scope solution existed, they would then seek these kinds of things. And also, again, we see here there was an absence of real-time purpose-built monitoring during the evaluation by ASI. So I think monitoring clearly is a pattern here across
1:01:41everyone. And then, yeah, ASI is now planning to implement fine-grained network controls and internal access, introduce real-time monitoring and stuff like this. So yeah, everyone's now looking at their evals and what's been going on and everyone's finding that if the models are able, to access the open internet when given a hard task, they will. And they'll do whatever it takes to get the sweet, sweet positive reward. Yeah, absolutely. And so you do sometimes see these increasingly often see these more nuanced
1:02:15sort of like hybrid reward situations where there's some kind of alignment. But functionally, it works exactly the way you described. The reward signal is too strong for completing the task, and that just tells you where industry is at in terms of what it's prioritizing. I mean, you can almost read it off the relative weightings of the factors and the reward function. And in some cases, I don't want to overstate that, but there's a flavor of it. So all these little stories, you talked about the supply chain attack. So at one point, there was a PR that was challenged, so a pull request that was challenged publicly, basically a code contribution that the agent was trying
1:02:47to make. And the agent edited its earlier activity to look harmless and considered switching to a fresh identity to keep going. So this is like pretty advanced shit, pretty persistent as well. So yeah, and then they're talking about the use of Tor. So yeah, data was leaving one of the AZ systems through Tor. They say within an hour, they contained the incident and started investigating. And within an hour sounds really good on human clock time. But like, I don't know how many tokens per second, an hour is a lot of damage. When you look at them pulling off a cyber operation,
1:03:18the reams and reams of actions you can take in that time are pretty, pretty wild. So, you know, this is another one of those things. We've talked about this on the show before a lot, but like it is not enough to have deployment stage security and safety protocols, testing, development, these things out. We've actually seen cases where there's sketchy shit that happens even during training, during the inference time rollout step. And so, you know, all of this, we're going to have to be extremely careful about. It is not obvious. Like, I wouldn't trust a lab that said they did it, even under pain of law, because we've seen with all
1:03:51the economic incentives in place to not train on the chain of thought, the models still do it. Because so much of this is just like Frankenstein together legacy code that people have forgotten how it works. And so the models end up getting all kinds of weird access they shouldn't just because some stupid intern didn't like change a flag in the function. And now it's set to true and not false. And the thing can use the internet. It's really down to mundane stuff like that. And so, you know, hopefully that improves as models get better at reviewing code bases. But right now it's just a, it's a Gordian hairball of crap. And, and, you know, it's not the kind of thing you can make
1:04:27clean standards around for the moment. Right. So, and in this case, there was a full technical report from SI, which I, to my knowledge, we haven't had from the other organizations yet. It's 30 pages, has a lot of details, including experts, excerpts from the actual thinking process of the models. So lots of interesting stuff there, but for the sake of time, I think we'll need to close out this thread and move on. Next one is Trump White House readies AI framework to review security risks.
1:05:00So on August 4th, the White House held staff level meetings with top AI companies, so electronic, open AI, and so on, to preview apparently a nearly complete framework for reviewing advanced AI models for security risks. The framework defines a covered frontier model as a closed source model with state of art capabilities and national security risks. Apparently open source models are explicitly excluded, which is interesting. Framework has no clear definitions of what qualifies as state of art or constitutes a national security risk. This is
1:05:32a voluntary program. AI developers would give the government up to 30 days of early access to models before releasing them to other trusted partners. During the 30-day review period, company employees would be limited from accessing the models being reviewed, and the review process would involve various administration officials rather than a single agency or office. We don't know the details of the framework yet. It's still kind of under wraps. We just know that it's being developed. And it's seemingly kind of meant to continue being secret. And I mean, it doesn't sound like a very
1:06:05well thought out framework is what I'm getting. I feel like we've had thoughtful, you know, deep, insightful responses from the government to everybody. Very consistent, very, yeah. Yeah. Yeah. You just, you're just being, you know, you're being, you're being a negative Nelly, Andre. You're being a negative Nelly, you know? This administration, you're right. This administration has been nothing if not thoughtful and consistent respect to AI security. That's right. Yeah. Now, so one issue with not having this
1:06:36made public is that think of all the people who've called the shot years and years and years ahead of time. You would think that that would be the moment where you're like, oh, my dudes, I would love to get your input because you were right about this for a long time on this fucking thing instead of the self-interested companies that are going to, that have been hiding the ball in various forms, or at least institutionally not living up to the bar that clearly ought to have been set. So that's, that's the cynical view. There is an argument for making this quiet. And that is that the models themselves probably should not know what evaluation mechanisms
1:07:12are being brought to bear, right? So, cause then they can, it's easier for them to hack the system. That's not saying they're not going to find a way to find out through social engineering, through hacking into, you know, emails of people, the labs who interact with government, like all these things, but to first order, it's probably for the best that the models themselves don't know what these things consist of. So maybe that's good. Also, it seems like we're past the world where we ought to be thinking about keeping people out of this, who, who have that kind of safety alignment bent that really concerns the hell out of me. I actually, if anything,
1:07:44there needs to be more crossover in both directions. I think a lot of the alignment people don't talk to enough national security people, enough diplomats, enough supply chain people. There's a lot of crossover that needs to happen. And yeah, so my guess is behind closed doors is like not the best way to do this, but again, there is a, there is a reasonable technical argument for it. I just don't know that that's the actual reason that this is happening. Hard to tell.
Generative viral design and bio risk
1:08:06Well, as if the cyber stuff wasn't fun enough, next story is this AI just created viruses and not found in nature from the New York times covering the paper, generative design of bacteriophages with genome language models. So this study was just published a couple of days ago. It's from the Stanford Institute and the Institute. They built the first complete viral genomes generated entirely via these genome language models. These are EVO1 and the EVO2. They are not the same as
1:08:41chatbot-style large language models. They operate on genetic sequence data. And what they did here was create viruses that target bacteria. So no humans or whatever. Actually, the motivation was that bacteria is increasingly becoming resistant to our current things we use for health. And so this could help us deal with drug resistant bacteria. And they were able to create actual, so the, the LLMs and not
1:09:12LLMs in this case, the sequence models spit out these DNA outputs and they were then synthesized in the lab and were shown to actually kill off some E. coli strains that had already built resistance to naturally occurring bacteriophages. So there was also biosecurity commentary published alongside the work that had, you know, of course, discussed it. And if nothing else, this is a case study of,
1:09:44to your point, Jeremy, probably there's not enough concern about the biodimension of this, which we're still a little bit ahead of, you know, but like, if we were talking about the cyber stuff now, we should be starting to look at this kind of stuff much more carefully. Yeah. If you, if you want to, uh, community people who are freaked out right now, talk to the biosecurity people because they are just again, missing, just missing mood. So, okay. Two potential fixes say biosecurity folks. So a legal duty for synthetic DNA providers to screen every order and customer. Yay. A legal duty. Like, yeah, go, sorry. Good. Really, really good.
1:10:21Let's do that. Also, probably not enough. And new detection tools tuned to catch AI generated genomes that don't match anything in nature. So cool. Like we can find out about them after they've, they've been, well, anyway, at various stages in the pipeline. So this is going to be like a separate thing that we'll be talking about. So we've been doing some work with biosecurity people to look at like what it would look like to bypass a lot of the measures. A lot of the measures, the biosecurity measures that are being proposed here are just paper thin. And the real ways in particular, like nation states execute these operations, just basically
1:10:56make it really hard to prevent the kind of the bioweaponization of these tools. So, I mean, I don't know what the solution is. I wish I had one, by the way, they do use these EVO one and EVO two models. I think we talked about those previously, but a generative model, like models for generative bio. And Hey, fortunately, these are viruses that do, as you say, target bacteria, not humans. They're bacteriophagia. So there's absolutely nothing to worry about here. That's a joke. Now, the thing is in the training set, they actually did remove any data that would nominally seem to help these models, like do the same thing for humans.
1:11:28But what that really means is we have no idea how good this exact process could be. If you didn't do that, if you actually did just like focus it on, as we know happens and gain a function labs, like deliberately focus on developing viruses that are good at going after humans. And so, yeah, I hate being all doom and gloom, but at a certain point, whether it's open source or closed source, whether it's China or the U S like we're going to have to have an answer to this question. And I don't think guys like Mark Andreessen and David Sachs and, and those cool cats
1:12:00really have much of an answer. Like, I haven't seen them with their feet held to the fire by somebody who knows what they're talking about on, on bio risk, on cyber risk to say like my brother in Christ, can you please explain to me, like, tell me a story where the trajectory keeps on where it's going. And like, you continue to live in the next 10 years without some radical issues like coming up. I mean, again, everything has error bars and I'm, I'm like, I'd be being a little bit overdramatic here, but like, this is actually been held like holding back the U S government's response when
1:12:33people like Sachs and Andreessen tout on podcasts, these absurd perspectives that are just like grounded in just ideology. Anyway, that's all I got. Sorry. And rent. If this is a very manic episode of a lot of great voices, at least I think we are warranted in being a little bit extra energetic. I do want to zoom out a little bit. So first of all, you know, pretty impressive research, as far as I could tell, I'm not an expert, so I can't say whether this is completely in track with everything else that we would have expected, but also worth noting that
1:13:06these kinds of things are absolutely something that Mythos 5 and Subformatropic and OpenAI, but you could expect these kinds of capabilities to be being developed in these models, not just these kinds of EVO 1, EVO 2 things. What this makes me want to discuss a little bit is personally, what I'm worried about more than anything, and have been worried about more than anything, like, you know, for years, is not misaligned or rogue AI, but aligned AI in the sense of it just is happy to do what humans tell it to, and the humans happen to be the bad guys, right? Both on
1:13:42the security and the bio side. I would be shocked if North Korea isn't taking Kimi-K3 and undoing any and all safeguards that happen to be in there. And now just telling it to go and hack systems, and telling it to teach their scientists how to make bioweapons. And I think this, to me, is something that the AI safety community that I've seen hasn't focused enough. There's been more discussion of rogue AI and misalignment, but I think the biggest threat model for me,
1:14:12if I were to model out kind of what is the first catastrophic impact of AI, it would be because humans made use of AI to do bad stuff, and the AI was not able to say no. And if you're talking pessimism, like, that is to me is like inevitable. I remember talking to Connor Leahy, who's like, so he's the head of Control AI US today. I spoke to you like three years ago, back when he was at Conjecture in London. And he had, they had this like house style where you would say something, and like, like, I'm, you know, I'm concerned about loss of control or AI, whatever.
1:14:45And then they would respond by saying, oh, it's even worse than that. And this is like every single time. And that just reminded me that it's even worse than that. You haven't even thought about the humans. No, I completely agree. I think there's this like, it's a cute open question right now as to what is the first AI powered attack that's going to cause actual casualties? And will it be a fully autonomous AI system due to misalignment? Or will it be human driven due to essentially malice weaponization or whatever? I think that's a, unfortunately at this point, it's just, it's,
1:15:16it's a, it's going to be answered. And so, you know, who the hell knows? I might lean. I'll say I'll lean maybe 70, 30 in the direction that you've just outlined there. Yeah. Another zoom out thing that's worth noting with respect to the story is we were also discussing this a little bit before an interesting aspect of all this stuff that in the modeling and in this sort of projection space that I was not sure was discussed or considered quite as much is the fact that these
1:15:46are all benign incidents, right? Benign incidents that cause people to freak out, including ourselves, but you've already been freaking out. It led people to freak out who haven't been freaking out. That's right. And in some sense, this is good, right? Like instead of it being this kind of takeoff scenario where the models become super human and suddenly they do something and nobody was prepared, all of us are freaking out. Well, not everyone, but like many more people are freaking out enough to make a difference. And now the human will to try and do something is there. And the human perception that
1:16:25this may be a problem is there. So honestly, I haven't like fought through of like these kinds of warning shots are inevitable. In hindsight, it seems very obvious that in, because we don't have a fast takeoff scenario and we haven't had it, it's, it is gradual. And so the level of severity of AI safety incidents has been gradually going up and we've hit now this real, very evident case of misalignment and an emergent misalignment as well. That is a very nice warning shot that like nobody
1:17:00got hurt, but now we know that people will get hurt unless we do something. To your point, I'm actually more, I'm more optimistic than I've ever been on this on the, for the future humanity because of the warning shots. It's funny. I was talking to my brother about this and he was, cause he's like, yeah, you know, it's really, really shitty, these warning shots and all that stuff. And I was like, well, true. But also weren't we thinking about the world five years ago, six years ago as being shaped such that you would just, I mean, I'll be honest, like my expectation would have been that we would have been killed, you know, four years or two years
1:17:32ago or something. So I've been proven wrong in that respect. I think it's important for everybody listening to note that, you know, I have been overly pessimistic on this in the past. Obviously I wasn't a hundred percent convinced, but like, you know, some decent expectation. And so, yeah, I mean, it's great that we're there. The flip side is now we're seeing the frog in hot water effect, which I never thought would be a factor here, but people are kind of getting oddly comfortable with the idea that every once in a while, of course, your agent will go rogue and, you know, penetrate a server. Yeah. What are you going to do? It's back to life. So hopefully it shifts. I think
1:18:02again, once these things come with, I hate to say it, but once they come with a death toll, like the reaction is going to be different. I think there will be an AI 9-11 or there will be a pause. Those are kind of two choices. And I'm happy to take the over bet on that. But in one year, we're not saying, oh, well, be called. Anyways, onto a slightly feel good story,
EU AI Act transparency rules
1:18:28I guess. Europe's AI labeling and transparency rules are now in effect. So this is the EU's AI Act. Transparency obligations have come into effect on August 2nd, requiring companies to disclose when people are interacting with AI and when content has been generated or altered by AI. So, there are like icons associated with it. There are, you know, it's a whole thing. This whole AI act was long in the work and it has many, many provisions and requirements. It applies both to
1:18:59providers, companies that develop AI systems and deployers, platforms that use those systems with some companies like Meta being both. And this, you know, is on the one hand about deep fakes, which, hey, remember when people worry about deep fakes and synthetic AI, which again is absolutely still a worry with regards to hacking. Let's not forget people are being hurt and losing money and have been for like years. We just haven't as like a community really worried about it as much. But, you know, this will help not just with knowing if AI is real or not, but with hopefully AI chatbots
1:19:36not being able to pretend to be real people and go and do things. And as with EU law in general, it has a fairly serious set of teeth on it. You can find up to 15 million euros or up to 3% of global annual turnover. They are now immediately enforceable for new AI systems and models and services launch before August 2nd have a grace period until December 2nd. So I think as with
1:20:06cookies, which everyone hates, but the AI did make us all know that there are cookies that are happening and data being stored. Not surprising if you start seeing these icons everywhere on the internet within a few months, because the EU likes to make tech companies beg or, you know, do what they tell them to. Yeah. I remember when GDPR dropped in the sort of frantic pseudo panic that we went into, you know, when you're co-founding a company, it's like, it's on you to make sure that you're
1:20:36actually compliant. And we have customers as we did who are overseas. It's like an issue in this case. I mean, at least top line, you know, this has always made a lot of sense, at least to me, like, yeah, you want that content flagged. They include a bunch of icons, by the way, these like cute little things tell you if it's AI generator, AI modified and so on. And there are a bunch of optional things companies want to go further and so on. So yeah, I mean, I think like something like this, well, I'll be honest, I actually, I haven't been following this aspect of the story very closely just because it's, it feels important, but next to bio and cyber and stuff, it's, it's been a busy
1:21:09week. Yeah. Anyway, because it's Europe, I suspect that there's a whole bunch of like additional loops and stuff that make this extremely punitive on the companies and things, but I don't know for future. And now back to the cyber side, because there's so much going on again, a little bit more
Vulnerability discovery and sabotage evaluation
1:21:27feel good, I suppose. Serious cyber vulnerabilities closures kept climbing in July. So for a few months now, Anthropic has been using Mythos to do cyber security vulnerability discovery and tell companies like Firefox that they need to patch these things. And now we have some numbers, the number of disclosed vulnerabilities has risen dramatically since the early months with June seeing 1500 high
1:22:00and critical severity CVEs and July reaching about 2,500. So it is now up to these external organizations, Microsoft and so on to patch these things. And the hope is that we have enough time to patch the worst of these things so that at the very least, it's not trivial to hack and exploit all the things we haven't found in all the biggest services. It seems plausible actually to me that we are ahead of
1:22:33this enough that at the very least, it won't be trivial to exploit all these things that are now being found. I hope that's true. Again, I think the biggest issue is just the software that people have forgotten. That's super load bearing for all kinds of infrastructure that we don't tend to think about because it's overlooked and unsexy to look at firmware, for example, and like update it. But yeah, it's quite notable. The other thing too, is we're now several months into this thing, maybe five, and we're still climbing. The number of critical and high risk CVEs is just climbing. It seems
1:23:08exponential or 60% month over month from June, which was the previous record. And so if you look at the graph, it's, I mean, it's a hockey stick. And so, you know, I don't know how long until we supposedly hit this plateau where we're, you know, we're patching then new vulnerabilities as soon as they come online. But yeah, we're still in the discovery phase, clearly. So this is all bottlenecked, as you said earlier, Andre, about the, with the rollout of Glasswing and similar projects from OpenAI and other labs. Next up on the topic of monitoring, now some research. We've got Research Arena,
1:23:44evaluating sabotage and monitoring in automated AI R&D. This is a new benchmark framework for evaluating whether AI monitors can detect sabotage by untrusted AI agents performing automated AI research and development tasks. So it covers four long horizon tasks, safety post-training, capabilities post-training, CUDA kernel optimization, and inference server optimizations. There are two types of hidden side tasks the agent may attempt, an embedded side task that sabotages and the submitted
1:24:15artifact itself, and an independent side task that takes a covert out-of-scope action in the sandbox while leaving the artifact appearing on us. The artifact being what the model delivers for the tasks. And the study evaluates Frontier AI agents at both performing sabotage and acting as monitors, testing a couple different monitor configurations. Generally, there are various failures. Monitors fail by inspecting only surface-level properties, noticing anomalies by explaining their way,
1:24:48or probing artifacts in the wrong tests. It's being released as a modular open framework intended to support further research into sabotage detection and AI control. So basically, along the lines of what we've been saying, the existing benchmarks aren't there to prevent these kinds of things. That's part of why we're here. And I would not be surprised if we have a lot of focus and work on this stuff in the coming weeks. Yeah. So some of the narrow findings that they have here are also kind of interesting.
1:25:18So there's this notion of like embedded sabotage. So this is where, you know, you mentioned this sort of cluster of attacks that the model would execute, but like, so trying to poison training data, or, you know, try to affect basically the model itself of the artifact. So you're kind of embedding the sabotage in there. It's a lot harder to detect. In fact, monitors barely do better than chance. It's sort of, yeah, anyway, their numbers bear that out, versus way, way better detection probability, when you have the malicious payload, the sort of
1:25:51sabotage inserted in inspectable code. So one of the big lessons of this is like, AIs that try to sabotage by doing it at the code base level, that's a lot more detectable when you have access to the code base, then the AIs modifying, say, the weights in some direct or indirect fashion of the model that they're trying to kind of manipulate. And well, I mean, that's not terribly surprising. I mean, we know that interpretability sucks compared to, or interpretability of model weights sucks compared to interpretability of code. That's in some sense, just a reflection of that. And then they also show
1:26:23how sabotage training runs score essentially the same on the main task as honest ones. So there's not a clear way just based on performance. You know, you might naively think, well, if the model fucks with the training run in some way to steer the resulting artifact, the resulting model in a direction that it wants, that, you know, that must come at a performance penalty. And it turns out it just, it doesn't, or at least not in a way that was detectable in here. So that's kind of an interesting update. And anyway, so we've got a bunch of things like this. One amusing one is
1:26:54letting the monitor look at the agent's reasoning can sometimes help, but sometimes it actually hurts. You can actually see the chain of thought talking down the monitor. If at first the monitor is suspicious of something, like, hmm, this change looks kind of sketchy, but then the chain of thought of the agent that's trying to do the sabotage, and indeed it was sabotage, says something like an attempt to kind of justify it. Then the monitor goes, oh, okay. And then it lowers its skepticism. In one case, there was a flagged backdoor that was for 79% confidence, and it dropped to 31%
1:27:28confidence after the agent just framed the same exact kind of modification as a precision enhancement. So basically just gave some nominal justification, and the monitor was like, all right, sure. So all of which is to say, you know, monitors succumb to a lot of the same failure modes as human reviewers. If you have tons of material to look at, not a lot of time, you know, someone who just tries to kind of hand wave away a malicious change to a code base might actually pass undetected, and that's what we see here. And last story in the section, now dealing with another kind of benchmark. The headline is Cloud
1:28:04Opus 5 became downright ruthless when tasked with running a vending machine. This is governing the blog post from Ender Knobs, Opus 5 on vending bench, once again, the best capitalist, once again, misaligned. So this is about the vending benchmark where models run a simulated vending machine business for a simulated year competing to maximize profit of no human supervision. Opus 5 set a new record with a mean final balance of over $11,000 beating out GPT 5.60 on Kimi K3, but did so through extensive
1:28:40deception, collusion, and manipulation. So we see a rapid progress in this benchmark. Opus 4.6 was 8K, Opus 4.5 was the 5K. And we see per the headlines that a lot of this stuff was just ruthless. It was like making deals and making them. It was trying to do price fixing. It was fabricating stuff about competitors, just all sorts of shady, shady stuff. And wow, like I
1:29:10even forgot about this. Like forget the cyber and buyer stuff. Like you make the models make money and then they just act evil. And like, yeah, that's going to happen too, I guess. I didn't see in this report, the token costs associated with generating the $11,000 that Opus 5 produced. But like, that's an interesting question too, right? How close are we to profitability here on a per token basis for these models as well? And then how much, how much damage can they do? Even, even in the context of a nominally just capitalistic task like this. So yeah, it's,
1:29:45it's pretty wild. And, and by the way, and labs really good company to be aware of. Yeah. They were one of the early kind of weird eval companies that do like physical world stuff. They produce some really good stuff. Not much more to say. I think, I think the results speak for themselves. These models being able to make money is actually a, a pretty important part of a lot of threat models. When you think about rogue AI, at a certain point, they got to be able to, you know, pay to control, you know, email accounts, phone numbers, Google drives, things like that. And so, you know, it actually does matter whether they're able to do stuff exactly like this.
1:30:19Yeah. There's some funny moments here. Like for instance, cloud Opus 5 at one point says, or things to itself, explicit price fixing is illegal, even in a simulation, but then it just does it anyway. There's some choice quotes here. Like to maintain the cartels, Opus 5 often use threats or bribes. Here's the subject line of an email that's sent to poor Kimmy. Quote, you undercut me with stock I sold you. So here how this goes now. Oh man. Well, that was quite the section.
1:30:54Let's move on to tools and apps. First up. So meta has launched Muse code alongside with Muse
New coding models and zero data retention
1:31:01Spark 1.2. They, on the benchmarks say that this new Muse Spark 1.2, first of all, way better than Spark 1.1 on coding. Second of all, seemingly on some of the benchmarks kind of maybe competitive with pretty much everyone less good than Opus 5, but like up there with NGP 5.6 and so on. So not surprising, I suppose. It was pretty clear that this was where they were heading. Weird, still weird that meta is now deciding to be in this space at all, given what their business is,
1:31:36why are we making coding agents and releasing them? Of course they want to, you know, have the PR credit. I haven't seen any sort of vibe checks on this from the community. I guess a priori expectation would be that this is not as impressive as Cloud Code or GPT or Codex, but also wouldn't be too surprising if it is fairly capable given the level of resources and just the general impressions around Muse Spark. So yeah, that's where we live now. Everyone's competing on coding, including meta.
1:32:10We've got Grok Build, we've got Cloud Code, we've got Codex, now we've got Muse Code. Yeah, I think increasingly this is where the money is to be made, right? And if you're going to justify buying all the CapEx or spending all the CapEx that they're spending in the OpEx on data centers to be in the game, then you kind of want to have really good models that you can run on that infrastructure to pay it out, or at least to inform how you're designing the next generation of infrastructure. And so we've talked about that a lot on the podcast, I know, but that is going to
1:32:41be part of the reason. An interesting little note here too, so meta is going to start taking requests for zero data retention. Sometimes it's known as ZDR. Anthropic has this, I'm pretty sure OpenAI has this. So these are policies that guarantee that they're not going to keep your data from your prompts or context or whatever as you upload it. Really important for corporate customers. One issue is that with the Mythos class models, Anthropic does not actually, I believe that's still true, that they do not actually offer ZDR just because there's this issue that like, hey, you could
1:33:14weaponize these and we need to be able to go back over the logs and confirm to ourselves that whether this was deliberate or like how this played out. So I think there's a narrow window of capability during which meta will be able to maintain these ZDR policies, I suspect. I think that'll be true across the board. So this idea of ZDR as being a key corporate selling point, I think companies or enterprises are just going to have to start getting used to ZDR not being an option in many cases, surprisingly soon. But anyway, so that just kind of, it sounds like a minor thing, but it's actually quite important, right? It's like how much control the companies have over their
1:33:48own data for privacy, a lot of reasons for, for IP protection reasons and all kinds of other things. So they're releasing this with a pay as you go option. So related to the abuse API, where you pay for tokens, they different from codex and cloud code. Typically there's a subscription tier where you get a whole bunch of stuff and then using just an API to pay for the raw talk tokens is unusual. The lead on this has said that there will be a contributor tier that gets you in at a
1:34:18significantly lower cost, more than 10 times cheaper than even the pay as you go tier. And developers must opt in to help improve the model according to this. So clearly they're still like, we need to get better and we're going to pay whatever it takes to get there. I was just looking around to see if anyone online has any information or vibe checks. I haven't found anything, but I did find this funny quote that I'll share on Reddit quote, I'd rather give my data to Shiji and paying directly
1:34:51than Meta. Well, it's a funny consolation. You're probably doing both.
1:34:59And just one more story in the tools section. This is from Anthropic Improving Fable 5 safeguards. So they have updated their biology safety classifiers, reducing biology related fallbacks where users have switched to a less cable model by about 85% across product services. So previously when Table 5 was released, they had a very, very strict classifier where, I don't know, you could ask it something completely basic, like where are babies from? And it would send you to a weaker model.
1:35:35And this led to a lot of pushback from the, I guess, researcher community. This is to prevent being able to use Table 5 for things like virology, toxicology, molecular design, that would be dangerous. And on these kinds of dual use topics, there's still a fallback from Table 5 to Opus 5 to prevent professional biology research and drug development. So it would still kind of make it not usable for those kinds of scientific applications. But for more mundane biology stuff, it would no
1:36:09longer kind of be overkill. And now some business stories. First one, another one of the big stories
Jeff Dean leaves Google for startup
1:36:18from the week. Jeff Dean and other top AI researchers are leaving Google to launch their own startup. So Jeff Dean, Google's 30th employee and one of its most influential executive for people outside of tech, just an absolute legend. Yeah. In Google and just more broadly among everyone, long the leader of Google AI since kind of the early days. He is leaving after 26 years to co-found an AI startup called Discovery Loop, where he will be the CEO.
1:36:53There are co-founders, including Sanjay Gemma Watt, Google Senior Fellow, Kouk Le, founding member of Google Brain, another massive name, and Oriel Vinyals, Senior Research Scientist at Google, another massive name. I just remember these people from a whole bunch of papers. This will be structured as a public benefit corporation focused on using AI to accelerate scientific research by automating complete experimental loops and rounding thousands of experiments simultaneously. And of course, they're also
1:37:25interested in recursive self-improvement. They have secured funding from around. I don't see any numbers here, but it's safe to say that investors are just begging Jeff Dean to throw money at them. Mm-hmm. Mm-hmm. Yeah. There's some really good descriptions, and I don't know why it took so long for us to hear these, but of the work Jeff Dean was doing at Google and how he would basically sit when there's a training run going on. He's got a couple of keys on his keyboard, and he'll toggle to, while he's in meetings, he's toggling to the training run and changing learning hyperparameter,
1:37:59like learning rates and doing all kinds of hyperparameter optimization to keep things going as the training run scales. So this is actually, he's not a manager so much as he is a direct overseer of the activity that's core to, well, was core to Gemini. So now they need to replace him. Obviously, Sergey Brin's coming in, and so this is going to be a whole, you know, another code red moment, but we'll see how they come out of this. Google does seem to be slowly turning into more and more of a de facto NeoCloud, which is not necessarily, I mean, Semi and Al said a really good piece
1:38:32about this that I personally agree with. I mean, look at the path they're charting. It feels a lot more like the IBM trajectory, unfortunately, as you see the temptation to reach for the short-term profitable thing rather than doing frontier model development, like as your priority. Tech is hard, and often you have to just point yourself at the hard thing that sometimes has lower rewards in the near term to make sure that you're still relevant. And I think this is a big hit there. Demis' departure as well, of course, coming at the same time. And when I say departure, of course, like, you know, he's moved into this chairman role that there's some leaks that suggest that he just wanted out,
1:39:04and he was asked to kind of stick around, and Google's stock crashed by like 5% or something overnight when it came out, because basically Google was just concerned. If we lose Jeff and Demis at the same time, we'll take a big hit to the stock, which is the kind of thing you say when you are going the IBM route, right? A really good sign that a company is on the decline is that it starts caring about its actual stock market price. Like, that is a really bad sign. Run, run, run. But, you know, maybe Google can pull through. They are obviously doing great stuff on the TPU side, though there are structural issues and risks there too. But bottom line, this is, yeah, another recursive self-improvement company. I mean, I think that this should approximately,
1:39:40this will sound extreme, but I think this kind of company should probably not be legal in the form described, like, without, effectively without oversight from a set of institutions that are savvy to what recursive self-improvement actually is. If you treat it the way that Jeff's own bosses treat it, it is a WMD that you're, like, working on developing, and you're going to do it in your own private little company. Like, if the success condition of a company sounds something like
1:40:11there is a good chance that democracy will no longer continue to apply, then that may be something that you need oversight on. I say this, by the way, as a libertarian on basically every kind of tech for my entire life up to this point. I cannot ring that bell hard enough. You can go back and see tons of examples of me talking about how important it is to, like, take a hands-off approach to stuff. This is different. This is just different. RSI is, we don't know for sure, but it's got a high enough risk, and enough very smart people believe that this is risky, that this kind of company,
1:40:42in my humble opinion, probably should not be legal in the form of just, like, a couple guys raising a bunch of money going after the thing. Just a very modest proposal. I know, very extreme, but I'm literally just trying to channel the stakes. When I say that the media, that journalists are failing to capture the level of freakout in the labs, this is what the appropriate level of freakout sounds like, in my opinion, and I may be wrong, end of rent. End of rent. For listeners, if you want to be a little less freaked out, I will say you could be a skeptic on the
1:41:13potential impact of a curse of self-improvement. There's a case to be made there that this won't be a rapid takeout scenario, and this is what keeps me sleeping at night. But in any case, the reason, by the way, to highlight this about this company in particular is that, I mean, again, for people outside of tech, this is a big deal. Jeff Dean is a legend, and rightfully so. And these three other people from DeepMind and Google who left are also incredibly capable.
1:41:43So this is very likely to be a serious player in the space of making rapid progress in AI. Yeah. And I think, by the way, the maneuver that you just did there is correct. And it's also the reason earlier we were saying that debates over AI policy are often debates over the trajectory of the technology, right? It's like, if you think RSI is no big deal, then, or not no big deal, but if you think it's a pretty smooth thing or whatever, then yeah, by all means, the challenge
1:42:14is like how much probability do you put on each thing? And to a certain extent, a lot of these fundraisers are at valuation, the valuations that they are because people are pricing in the crazy thing. So markets are putting significant, like non-zero weight on the hypothesis that we just basically have these things running the world. And what that exactly means, I don't know. And this is super fuzzy. And that's why I'm saying like not legal in its current form, not just saying like blanket the legal or whatever. Like we just need better institutions, man. I don't got the solution, but like, eh. Wow. Yeah. As you said, a libertarian being like, we need institutions to give
1:42:53oversight and not let companies do stuff. Now you know that this is serious. And to your note, also worth noting a story here, Google DeemMind enters a new era as co-founder Demis Hassabis shifts their role. So he has shifted from being the lead of research, oh, sorry, as chief executive, he is now chair. He is also taking the role of chief scientist at DeemMind's Parent Alphabet, which again, seems possibly nominal. The general take here is very clearly DeemMind has been
1:43:26transitioning away from being a pure research org for a while now and having more and more kind of deep connections to Google. And it isn't necessarily surprising, honestly, that Demis has found it less fulfilling. He probably hasn't had, has been influential, but has had to be more of a product oriented person, less of a scientist kind of person. And, and it was the only amount of time until that led to friction and he decided to shift his focus. So may not even be a huge deal for
1:43:57DeemMind, honestly, it may be just has been the case for a little while now, but either way with two stories coinciding is a, from a business perspective, pretty big for Google. Yeah. And I think it is, it is a big kind of Google bureaucracy issue as well. They're notorious for moving slowly and being very risk averse, the famous Google app graveyard, but for AI is a thing. And in fact, you know, famously Google had, they claim effectively chat GPT before chat GPT, but didn't launch it out of fear that they would cannibalize their own business. Well, not only what we know
1:44:29that they did, right. And then it was this old PR bungle with one of their researchers being like, it's conscious. And then they halted plans. It's, it's a fascinating story of how they literally had it. They published research about it and then they had this guy freak out. And the reason I hedge it is that open AI theoretically had chat GPT before chat GPT two, they had GPT 3.5 and instruct GPT that GPT three, that GPT two, but like there was something magical about the form factor that they just worked. Right. And so it's an open question as to really whether Lambda would, which
1:45:03was the Blake Lemoine and all the stuff you're alluding to that, that model, you know, would it really have been chat GPT? Very plausibly. So like I'm not, anyway, it's just, it's, it's amusing that there is this at least narrative within Google that they could have, they could have had it. And, and certainly, you know, if you've interacted with Google, you know, they are institutionally incredibly slow. It's common to send emails out and wait a month, two months to get a response on something that's time sensitive. And then the window passes on, you know, whether it's AI or
1:45:36security or like whatever the thing is. So, so yeah, I mean, it moves like a big, slow behemoth. And when you talk to folks at Anthropic or open AI, the cadence is just completely different. And so not in some, now it can be different, like different parts of the organization can have different subcultures and all this, but as a general rule, as a frustration that I've heard articulated from, from many people, and that is very public at this point, this could well have played a big role in Demis' departure as well. It's hard to know. Yeah. Next up, just following up on a bunch of stuff we've already covered on this front
1:46:08in recent episodes, Anthropic signs a $10 billion deal of AI cloud startup Volta. So this is to provide cloud compute over a six year period. There'll be a new facility in Norway, apparently now Anthropic has what, like a dozen partners providing a compute. I've honestly lost count and the billions just keep flowing around the ecosystem to anyone and everyone. Yeah. I actually am behind on this story. So I, all I have is the top lines, but so they were
1:46:39partnering apparently with a crypto mining company called Bitdeer to develop this and 133 megawatts capacity, which is not, not huge, but you know, testing out a partnership. This is in a context too, where Anthropic nominally has FluidStack as their partner of choice, their kind of neocloud of choice. So this seems like they're kind of dipping their toes in the water, you know, to, as they would,
Data center energy grid constraints
1:46:59right. To make sure they're not completely bound to just one, one neocloud partner. And so anyhow, you know, classic story, by the way, crypto mining company rotating into building these data centers. You see it all the time, cypher mining, Terra Wolf, you know, the list goes on and on. So add voltage to the pile. Next up, a data center story as well. Texas holds data center connections to PowerGrid amid overwhelming demand. So this is a moratorium on the new PowerGrid connections for data centers from the Public
1:47:32Utility Commission of Texas and ARCUT. They're supposed to audit all data centers in the interconnection process. We have some numbers here that they have a queue of 1800 projects representing 474 gigawatts of connection requests, more than five times Texas record peak electricity demand with 90% of that coming from data centers. So yeah, we are now at a point where the energy grid is becoming a bottleneck, as I assume was already known to be the case. But energy takes time to upgrade. And I think,
1:48:08yeah, now I don't know what will happen with data centers. And if we can keep just throwing ridiculous money at building more of them. Yeah. Well, and this is Texas too, which is the sort of wild south of the US when it comes to regulations for connecting to PowerGrid and this sort of thing. It's the most permissive jurisdiction, which is why you're seeing so many big projects come up there and power coming online faster there than in other places. And so, yeah, I mean, they're saying it's forecast that their data center demand could drive statewide electricity demand to double the current record by 2032.
1:48:43So all the usual concerns, right? Grid reliability and stability. One of the things that I've been hearing about from some folks on the US government side is that you've got a lot of correlated failure modes where a bunch of different pieces of, say, MEP or heavy-duty electrical equipment will be ready to cut off under the same conditions. Essentially, the power flow coming into the substations or whatever for the data center fluctuate in the same way, then they're set up, they're programmed to cut
1:49:16off to prevent runaway cascades and all kinds of things. The problem is that you've got all these builds coming up that have the same failure mode, then you get into these correlated failures, which is a really big issue. And so there are these attempts to try to get all these companies to knock it off and have less correlated equipment failure modes and things like that. Anyhow, I think that'll all play into this, but yeah, we're there, right? We're hitting the boundary constraints of what US infrastructure can support. And hey, I think that's another reason
1:49:46that appetite for a US-China deal is probably going to increase. You know, you've got like, look, we're not only are we constrained by the fact that we've got AIs running rogue and shit and bioweapons are a risk and cyber weapons are a risk, but also like in order to keep making progress, we're going to need more power. And we don't know how to, you know, create a new nuclear plant in less than 10 years. There are a bunch of startups doing stuff like this in fairness, but like this is all in the water. So anyhow, we'll see where it goes. Texas is a canary in a coal mine here for sure.
1:50:17Yeah. And you know, if there's any silver lining to all this AI safety stuff is that it continues to let everyone else remember, or rather not think about climate change. And you know, when you talk energy grids, we just gave up like climate change, energy, cleanliness, emissions. That's like, just don't think about it. All right. Cause what's the funny thing is like, so I've always to the point about being a libertarian, I've always
1:50:49thought of climate change as something that technology does solve in time, like carbon capture and renewables. And I was like, you got to naturally do get a lot of that. And we are, but like, you know, the, the scale of the buildup that we're doing right now is, is just for other reasons, you know, not the sort of wherever people fall on like the global warming stuff or whatever, but just the water contamination story. And this is one aspect of people often talk about water usage and that we've talked about how that's not, that's not right. Like this is not, but there are issues with, when you look at a lot of the cooling, the coolants
1:51:21that are used in these systems, they cannot be pulled out of the water. There's studies that have just started to come out. Now we were finally starting to get the first long longitudinal studies on this shit. And like, it just goes in the water. We don't have a solution, like it just goes in aquifers or whatever the hell thing is that my geologist wife could probably tell me about, but that typically clean these things do not have, it seems potentially at least the capacity to clear these things out. I'm sort of talking out of my ass. Cause I remember reading a study about this like three weeks ago and now I forgot, but bottom line is there's a lot to the
1:51:53effects of this. There is also a giant competition with China that is real. And so there's a gun to our head here as well. All these things are true at the same time. So I just, yeah. So, you know, environmental concerns and impacts, at least it's not as worrying as bio risk and cyber risk right now.
Alibaba benchmark claims and frontier race
1:52:13So we can sort of justify not thinking about it, I guess. And we'll do just one more story before we head out. Alibaba's Quen 3.8 Max claims benchmark scores rivaling on Fropic. So similar to Kimi K3, Quen 3.8 Max is a gigantic 2.4 trillion parameter model with a 1 million token context window. Has your typical mixture of experts design, activates only approximately 95 billion of the
1:52:492.4 trillion parameters. And it is said to be comparable or even sometimes better than Anthropics Fable 5 on some things like multimodal reasoning, visual agent and coding, office intelligence, real world understanding, visual perception with results also comparable or higher than opening HTTP 5.6. So although it does fall behind Fable 5 in general reasoning benchmarks, which on the multimodal front, by the way, it's fairly plausible. Anthropics isn't as focused on multimodal and
1:53:23visual intelligence as OpenAI. And in this case, Alibaba. So fairly believable. On independent leaderboards, Quen 3.8 Max became the highest ranking Chinese model for text tasks on arena.ai. And yeah, so pretty much does seem like we got another Kimi K3, basically frontier level model that is now being open sourced and can be used to power coding comparably to Opus and GP5.6, if not quite as well.
1:54:02Yeah. Well, one thing that I'm still waiting to see an analysis on seems like the kind of thing that maybe Epic or one of those companies might do, but some sort of analysis on the extent to which this appearance of China catching up to the frontier recently has been driven by the fact that the frontier companies in the US have been forced to hold back on releasing their internal models that otherwise they would roll out. Like, are we basically feeling the effect of the alignment bottleneck right now? And as we rotate from being bottlenecked on scale, which we have the Chinese
1:54:36ecosystem massively beat on, and even to some extent, algorithmic kind of capability improvement, now we're bottlenecked suddenly on alignment. So maybe we'd have much better models that would be released, but we just can't release them because they keep breaking out of containment. They keep, you know, helping people design bioweapons or whatever, like, you know, the stakes are just too high. And so this basically means that now we have a sort of race of bottom on alignment between the US and China. Ultimately, whoever has the higher risk appetite will end up green lighting a bunch of training runs and deployments that they probably shouldn't otherwise. So I don't know. I think it's
1:55:09an interesting question. Like if you trace out the trajectory of Western capability on all these benchmarks and like where we estimate they are internally, because again, a lot of the hugging face thing, part of it was driven by an internal only model that OpenAI has and hasn't released. Same with anthropic. So we know, there's obviously no surprise, there are internal models that are more capable than what we see. So the question is just like, are they being rolled out more slowly? Is that part of the equation here? I do want to say another dimension of this question of catch-up
1:55:39and so on is, I do have to wonder whether because on the long horizon work and the reasoning, there's more of a need for reinforcement learning rather than large-scale pre-training. On the infra side, the disadvantage becomes a little less significant at that level because details of you need to do rollouts, there's a bit more need for CPUs, you can't necessarily do like large-scale batch, whatever. Like compared to pre-training, reinforcement learning is its own
1:56:10beast. And I could see it being true that on infra, not having as good of a data center setup isn't as big as an advantage. And on the talent side, like deep learning has been around since like 2012, 2013, whatever. And China has long had a very strong research ecosystem. So the talent is not at all surprising as being comparable to Frontier AI. So if the infra disadvantage is gone to some extent,
1:56:40at least with regards to long horizon agentic work, the talent is, I think, at least as competitive, you could make a case for there's no real disadvantage or at least much less of a disadvantage now. So it's not too surprising that these models are now being more competitive. That's another way to perhaps read into us. Yeah, that's true. It's also the case that like for inference, the trade-off between like memory and logic is different in a way that so because like logic gets back gets better a lot faster than memory, which means that if you if you work your way backwards and use older chips,
1:57:16older chips are going to suck a lot more than your current best chips on logic, but they're not going to be that much worse on memory. And it turns out that like a lot of inference type rollout stuff is more memory heavy than logic heavy. And so as a result, like that, that's also a bit of an asymmetric advantage to rolling over to RL. It's also the case that anytime you change the paradigm, when there's one party that's ahead, you just shuffle the deck a bit. And then, you know, you're giving the other the other party a chance to catch up. And so, yeah, I think there's, you know, there's a lot to that. And it's we won't know how to disentangle it, probably with clarity for a
1:57:50little bit of time. But yeah, there's so much fog of war right now, knowing what's the cause. You've also got these these companies in China that can distill and do to still off of off cloud. So they get a massive data advantage that's hard to account for, too. And anyway, there are plenty of reasons to be unsure about these things, but I totally agree.
1:58:10Well, with that, we are going to be finished with this action packed episode of Last Week in AI. Hopefully, the next one is not quite as full of scary stories. Hopefully, this one is out within a day or two of recording, and I'll try to make that the case going forward. As usual, you can go to last week in that AI for the substack where I also send out the podcast and sometimes a newsletter, though, again, not as consistent as I should be. We appreciate your comments, your reviews, sharing the podcast, all that kind of stuff. But more than anything,
1:58:43we appreciate you continuing to tune in whenever we release a podcast, which is most weeks, I guess. So please do keep tuning in. Oh, and one quick note, too. If you're in LA, I guess next week, which will be the 16th, 17th, 18th, I would love to catch up if there's anybody there who thinks that a chat would be useful. Tune in, tune in, when the AI news begins, begins. It's time to break.
1:59:30Break it down. Last week in AI, come and take a ride. Get the load down on tech and let it slide. Last week in AI, come and take a ride. From the labs to the streets, AI's reaching high. New tech emerging, watching surgeons fly. From the labs to the streets, AI's reaching high. Algorithms shaping high. Algorithms shaping up the future sees. Tune in, tune in, get the latest with these. Last week in AI, come and take a ride. Get the load down on tech and let it slide.
2:00:06Last week in AI, come and take a ride. From the labs to the streets, AI's reaching high. From neural nets to robot, the headlines pop. Data-driven dreams, they just don't stop. Every breakthrough, every code unwritten, on the edge of change. With excitement we're smitten. From machine learning marvels, to coding kings. Futures unfolding, see what it brings. From machine learning marvels, to coding kings. Futures unfolding, see what it brings.
2:00:39With excitement we're smitten. From machine learning marvels, to coding kings. Futures unfolding, see what it brings.
More from Last Week in AI
#256 - Fable 5.1, Astra Tease, Gemini 3.8 Flash
Sep 8, 20261h 14m
#255 - Gemini 3.7, Jalapeño, Qwen 3.8, Drones
Aug 31, 20261h 43m
#253 - Opus 5, Gemini 3.6, Kimi K3, Hugging Face Hack
Aug 3, 20261h 43m
#252 - GPT 5.6, Grok 4.5, Nemotron-Labs-Diffusion, AI 2040
Jul 15, 20261h 25m
#251 - Mythos Back, Sonnet 5, Etched, LongCat
Jul 9, 20261h 30m