#256 - Fable 5.1, Astra Tease, Gemini 3.8 Flash
September 8, 20261h 14m · 14,452 words
Show notes
Our 256th episode with a summary and discussion of last week's big AI news! Recorded on 09/03/2026; unfortunately just before the actual GPT 6 Astra release, we'll cover that in next ep! Hosted by Andrey Kurenkov and Jeremie Harris Feel free to email us your questions and feedback at and/or Read out our text newsletter and comment on the podcast at In this episode: Anthropic released Claude Fable 5.1 and Mythos 5.1 with lower pricing, stronger agentic perf…
Highlighted moments
Mythos 5.1 designed protein binders that were in fact verified by labs. In other words, they work with a roughly 50% success rate.
“The initial planned period on premises, in other words, this is the period during which Meter was allowed to go on OpenAI premises to actually collect their data and do all this stuff, was two days between the evenings of July 29th and the evening of July 31st.”
“Anthropic objectively has much, and I mean much more focus, culturally, institutionally, capacity-wise on alignment. And so it is very credible that they would say, look, I mean, OpenAI had to pause because, like, literally their next training run was probably going to unleash more of these agents for all of you.”
Transcript
Welcome and AI safety terminology
0:00Hello, and welcome to the Last Week in AI podcast, where you can hear us chat about what's going on with AI. As usual in this episode, we will summarize and discuss some of last week's most interesting AI news. I am one of your regular hosts, Andrey Kerenkov. I studied AI in grad school and now work at a startup, AstroCade. And hi, everybody. I'm your other
0:30regular co-host, Jeremy Harris from Gladstone AI. I do AI national security stuff, AI loss of control stuff. I guess now we're calling that rogue AI stuff in US, China, et cetera, et cetera. You know the drill. Much more exciting than just saying AI safety. That's right. Rogue AI. That's right. The funny thing is for years, at the level of introducing myself, the first time that somebody here's what I'm doing, I'd be like, yes, I work on either AI safety or I would say sometimes AI security because shortly after the Trump admin started, safety became a bad word the
1:03world over. This includes the UK, by the way. So like Keira Starmer was very down on it and it became the AI security institute and everything was AI security. And now I'm kind of like, well, I've always, you know, obviously talked about loss of control very directly, but at that level of introducing myself, I find that's actually changed. And I think that's really cool because that did reflect a massive tax that everybody in the space, like really everybody ended up paying for even being seen to talk about these things. Including, I think, in the AI community, like outside of the AI safety community of itself,
1:38despite, you know, even like new reps, for instance, big conferences requiring people to do things like impact statements, which would imply that safety is a concern and we do think there'll be impacts on society. For quite a while, safety was a marginal kind of perspective, let's say, or even looked down upon. And now I think certainly with this recent ChadGBT story, which we'll be talking a little bit more about, it's not only become known, it's kind of mainstream now to be aware
2:11that this is a real thing. Yeah. And even the word safety, right? People started to use it and in my opinion, kind of mangle it by making it referred to more prosaic sort of, oh, weapon, you know, like weaponization is a really serious concern. Like a lot of our work is in that direction for sure. But like it really originally meant AI alignment and even AI alignment got watered down to the point where we had to invent super alignment. And now even the word super intelligence is getting co-opted. So there's this thing in, you know, in the history of language where people will invent a word to describe people, for example, with now we might say mental disabilities a few
2:48years back, you might use a different word that started with R and that probably gets us flagged or something. But that used to be a completely unobjectionable medical term, right? And so what's going to happen is people are going to shift over the new thing. Oh, the people already say stuff like, oh, you're, you're special, you're like special needs or whatever is already becoming that. And we've had that kind of happen in reverse where like the gradual encroachment, I guess the incentive is always to, for the labs, especially to interpret alignment and loss of control in the way that is easiest for them to actually deal with. And because no one has a clue how to control
3:21super intelligence, that's meant that they have chosen to continually interpret AI safety as more prosaic sort of let's not make it say bad words. And then it was like AI alignment became that. So I think a lot of the terminology got muddled in that way. And now we're sort of finally at the point where there's a thing we can point to, to be like, no, the hugging face thing. Like, what are you doing about that for the next generation of models up to super intelligence? It's very interesting now, like that lay people, so to speak, even know about the hugging face thing,
3:53which usually these kinds of things don't get out there.
Sponsor messages and listener feedback
3:56We'd like to thank Langfuse for their support. Langfuse is the most widely adopted open source platform for AI agent evals and observability. We relied on it rather heavily at Astrocade and I'm personally a big fan. It provides tools for tracing, evaluation, experimentation, and problem management. And if you are building an AI powered product like we are, these are things that you need to have. You need this kind of tool to help you debug your AI agents and applications through tracing of what
4:26users actually experience when they use your product. That helps you understand where things break. It helps you set up evaluations, run experiments, systematic tests, and all the different things that actually help you improve your product. We tried multiple different tools for this, but we're drawn to Langfuse because it is MIT licensed and self-hostable, but also available via Langfuse Cloud as a managed service. It's also quite mature, so it has hundreds of integrations with things like OpenRouter, Vercel AI SDK, LangGraph, and many more. Furthermore, it is framework
5:02and vendor agnostic, so it works with any model and any framework. You don't have to worry about vendor lock-in. Langfuse version 4 has just released and has major updates to architecture alerts, code-based evals, and scalability. Get started with Langfuse Cloud today at Langfuse.com. There's a generous free tier and no credit card required. We'd like to thank Box for being a sponsor. Box is building the intelligent content management platform for the AI era. It acts as a secure content foundation where AI agents don't just access your unique institutional knowledge,
5:36they actually orchestrate end-to-end workflows across your business-critical systems. And that's important. If you're still using AI to just do basic chat and summarization, you're probably not getting the full productivity gains that are possible. Nowadays, to really adopt AI, it's not enough to just ask chatbots questions. It's about putting AI to work on document-heavy processes. From pulling structured data out of unstructured files, to multi-step routing, exception handling, dynamic document generation, Box helps your business turn manual document bottlenecks into automated,
6:10repeatable business outcomes. And that comes with a full governance layer. You can enforce granular permissions, maintain an audit trail for every agent action, and keep humans in the loop to review exceptions before critical downstream actions occur. If you want to move beyond basic AI Q&A and put autonomous AI workflows to work for your business, visit box.com slash LWIAI or join your team at BoxWorks in San Francisco on November 5th and 6th. Use code LWIAI for 50% off your
6:43registration. This episode is brought to you by Progressive Insurance. Do you ever think about switching insurance companies to see if you could save some cash? Progressive makes it easy to see if you could save when you bundle your home and auto policies. Try it at Progressive.com. Progressive casualty insurance company and affiliates. Potential savings will vary. Not available in all states. Yeah. So we'll be talking about that a little bit more. There's more details that came out that are quite interesting. Then, aside from that, we do have some other big stories and big models
7:15coming out. A couple of business stories, but actually fewer than usual. Some new models out of China yet again, and these are a little bit more interesting. And once again, a bunch of policy and safety stories this week, a little more diverse, but still dealing a lot with that hugging face story and its ramifications. And we'll round it out with a bit of research, a couple of papers. So it should be a pretty fun, diverse episode. Before we get going, I want to respond to a couple comments
7:49comments from first Apple Podcasts review. Jeremy, you have once again been called out for your strong language, which to be fair, I don't think it's that consistent. It's more when we get passionate, which does happen. But thank you for the feedback. We will, let's say, limit the user profanity. I will do my best. We will try to do better so that your kids don't get corrupted. And then we did have one nice comment on YouTube, which wanted to add to the discussion of Google versus Anthropic
8:25and OpenAI. So real quick, I'll shout that out. Here, the feedback and the additional perspective was that in this dynamic of Anthropic and OpenAI and Google, this commenter had the point that OpenAI have to release their models in order to generate cash flow to fund their investments. Basically, they have to be at the frontier to justify to their investors why they need so much money and also to hype up for IPO. That is not true for Google. They have a cash printer. They don't need to IPO. And I do
9:03think that's a fair point in general, right? I think in general, the strategy of Google not emphasizing being at the frontier is fairly logical in the sense that, as we've discussed, I believe their focus is on developing the technology to go into their existing products, into Google Drive, into Google Search, things like that, which does mean that some of their certainly engineering and possibly research focus is towards that front. I will say, within the company, they have been working towards AGI explicitly, right? And I'm surprised if
9:39there's not kind of a shared perspective, which is more broadly believed in the space to some extent that the first company to reach ASI, AGI will kind of be the winner, which I'm personally a bit skeptical of this framing as a whole that, you know, one company will be the AGI company, and they'll eat up the economy. But there is a decent amount of even investment kind of belief in that in terms of justifying valuations that you're seeing. So on the
10:13bad perspective, you might say Google is not being smart by playing it this way. You could also argue that part of the reason we're seeing fewer releases out of Google is that they are just working internally, and they have no kind of pressure to release their cutting edge models. So once we do get Gemini 4, it'll be like, super great. I'm a little skeptical on that side, I would say that if they had frontier of models, they'd release them, like, they do want to be seen at the frontier of AI, if nothing else, just for recruitment. But I do
10:48by the kind of perspective that they simply have less need to be at the frontier. And they do, I think, actually, somewhat wisely, develop a lot of technology to go into their existing products. I found out that we actually have a friend of the show that I didn't realize was a friend of the show. So Robert Wright, now he does a lot of podcasting, he has the non-zero newsletter. So I think a lot of people who like, OG, physics, philosophy, stuff in that category, also like the kind of new atheist type stuff, but it sort of evolved. Anyway, I found out on Twitter
11:22yesterday that he apparently listens from time to time to the show, which is really cool. That was kind of awesome. And from time to time, like, hear about cool people in interesting places. Everybody here who listens is cool. We appreciate every single one of you. It's just, yeah, it was sort of a funny thing that I had no idea how small the internet was. So there you go. Cool. But that's enough for now. Let us get to the news. Starting with tools and apps, we've got,
Anthropic model updates and pricing
11:46first up, Anthropic is launching Cloud Fable 5. And there are some updates to pricing on that front. So we've got both Fable 5.1 and Mythos 5.1. Fable 5.1 is generally available. Mythos 5.1, similar to before, limited to project glass wing participants. Kind of a typical thing. Fable 5.1 is stronger than Fable 5. And they argue or say that it should cost around 25% less than previously and up to 45% cheaper for complex agentic tasks due to reduced pricing on
12:23cached data partially, as well as other things. Part of the reason is with these newer models, Anthropic has been making kind of a case that you can get a way of using lower reasoning amounts to get the same level of performance. Typically, you just set your reasoning to the max level. You may actually want to set it to like medium or whatever. The last thing to say is they are also rolling out this Enterprise Frontier Safeguards, which will store customer data on the customer's
12:57own cloud servers. So that will be rolling out data this fall. And that kind of allows them to work with their enterprise customers who don't want Anthropic to store any of their data.
13:09Yeah, it's a pretty big performance jump, it seems, especially on... So one of the things they highlight is agentic scientific research. More than doubling the score of Fable 5 in that category, also, you know, unsurprisingly, agentic coding, knowledge, work, computer, like all these standard benchmarks, it does better. But multi-hour, multi-day tasks now are like more and more doable. And that's been a clear focus with 5.1 over 5 as we continue to go off into uncharted territory. We don't have meter evals that project out into this Epyx ECI becomes... I think you have
13:42to concede that even that measure does not generalize, you know, in the way that you might want to give you the source of statements you want to be able to make from a safety standpoint, at least about these systems. And well, so to give you a concrete example, like what does this add up to? Let's talk about bio. So Mythos 5.1 made protein binders. So, you know, when you do... I'm like tapping into my old days in biochemistry now, so I'm about to mangle some shit here. But I mean, some crap. I'm about to mangle some crap. It's often important to like design proteins that will
14:13bind to other proteins or molecules. They'll bind to proteins to change their shape. And when you change the shape of protein, change its functions, this is like relevant for a lot of drug discovery type stuff. So Mythos 5.1 designed protein binders that were in fact verified by labs. In other words, they work with a roughly 50% success rate. That's a sample size of 12. So 12 different targets that they tried. 10 to 15% is the norm here. You're hitting 50%. This is industry changing stuff if it gets deployed. Right? So as we start to think about what is the Mythos moment for bio, and I will
14:45continue to beat this drum until morale improves. We're headed for it. And there's just no two ways about it. At some point there will be a AI designed biopathogen that will do some really sketchy shit crap that will do some sketchy crap. So that's the good news, bad news side of things. They also had Mythos 5.1 rewrite GPU kernels for seven open source bio models. Now this is not quite AI research, a big part of AI research is rewriting kernels and kernel optimization, but for AI workloads. This is for bio workloads. So a little different, but still rewriting these kernels leading to 1.4 to
15:192.5 X speed ups and cutting computational costs 30 to 60%. As we start to think about what does the software only singularity look like? These are some of the numbers you might care about, but in of course the context of AI research, which is a more complex kernel design, kernel optimization problem. This is how it always starts. You know, the metrics that seem peripheral, but somewhat important start to go up and then the big metrics go up. And you know, there's, there's not always a huge gap there. So all the standard tests were done. Seaburn, right? Chem bioradiological and nuclear risk. It seems like it's more, you know, 5.1 is more
15:49capable than 5 across the board, but still below the next risk threshold in the anthropic RSP. So they're shipping it with the same restrictions as Mythos 5 across the board. I have seen a lot of sentiment of people switching over to codecs from OpenAI. And to some extent, people even saying that OpenAI is now in the lead for agenda coding in their toolset. This release by Tropic doesn't signal so much that they are, you know, trying to take it back. They probably have
16:19more of it enough, or even too much interest still. But I'll be interested to see sort of on the vibe front where people save us in terms of actually using cloud code versus using codecs.
OpenAI new models and reasoning updates
16:34And speaking of that, we also got some news from OpenAI. They have stated that they are going to be releasing their next AI model Astra soon publicly. So they aren't there yet, and they have done a lot of signaling that this will be yet another kind of jump in terms of capabilities and also in terms of how dangerous it is. So they have publicly stated that it is their first model to reach their critical cybersecurity capabilities threshold, meaning that it can independently find and exploit previously
17:08unknown vulnerabilities in real world software. So when they release the software model, they will initially restrict its advanced cyber capabilities to select partners in its daybreak blue early access program. I haven't heard that, but similar to that GlassBring project. And they are implementing a misalignment monitor to prevent everyday users from accessing Astra's advanced cyber capabilities with refusal to do these kinds of things, similar to what we've seen with
17:40Fable from Anthropic. So, you know, clearly they tend to keep shipping, but they also are changing their public messaging in addition to what their strategy is for release. I think in part due to this hugging face incident. So I guess a couple of things. First of all, let's just like talk about the core thing here, the post taken at face value. So OpenAI says, yes, in fact, as you said, it's their first critical cyber capability model. So what does that mean? Well, it means that roughly speaking, if you give it
18:13the right tools and the right access, it can find zero days, basically previously unknown security flaws develop full exploits itself across many well-protected systems without a human in the loop at all, you know, in hardened real world systems is like kind of part of this, right? And frankly, I mean, it would have been extremely difficult to believe them if they had said anything else, because we have already seen the hugging face incident, which fully proves that that actually involved using a zero day. And that zero day was used to penetrate the hugging face infrastructure
18:44stack. So like, I think outside observers now have literally had an opportunity to like, look, whether or not that was this version of Astra, I think at that point just starts to feel like splitting hairs. Last time we recorded an episode, we talked about how OpenAI had announced very proudly that they were delaying the largest RL training run they were doing by two weeks. And at the time, we celebrated that, I took to Twitter, I actually said, hey, great, like, if this is what it appears to be, then great, OpenAI deserves credit for doing that. The problem is, and I said it at the time on
19:15this show, I said, I hate that I have to caveat the statement, because with OpenAI, there's always a freaking plot twist where the thing that sounded like the noble and right thing to do on safety turns out to elide some like deeper and more fundamental Machiavellian pseudo kind of safety violating ploy. The one that turns out to have been the case this week, it seems, is we find out from the information, not from OpenAI, from the information, that OpenAI violated a core safety tenet that they
19:47signed on to. And when I say they, I mean like the most, some of the most prominent OpenAI safety researchers pretty clearly implicitly with the endorsement of leadership, that they would do everything they could to build architectures that had a legible chain of thought. Some of the architectures that they explicitly said they would avoid, if at all possible, and not universally avoid, okay, so I'm not saying that they guaranteed they would never do this, but they said, if at all possible, we will avoid doing this, were, well, the looped transformer architecture. So essentially,
20:20a looped transformer, we talked about this before, there's like a famous meta paper called the coconut paper you can check out that kind of gives you the gist. But basically, you put your input to the model and your prompt on one end, it works its way through the layers. And then in a typical transformer, you just decode basically the final layers residual stream to token space at the very end, right? And so you get a token after every forward pass completed. But what coconut does, what loop transformers do is they say, ah, no, no, no, let's take the activations at the final layer, and let's actually put them back in at the bottom for another pass. And so essentially,
20:53the model is able to do twice the amount of thinking for each token. And essentially, it's reasoning more in latent space, reasoning in activation space, instead of being forced to output a token at the end. So this matters, it matters for monitorability, because it means that, well, at least you're forcing in the old version, you were at least forcing the model to give you a human legible token for every forward pass. Now what they're doing is they've introduced this whole new variable. And by the way, they're telling us, oh, don't worry about it,
21:23we only loop twice over. And they don't tell us do we loop twice over on average for tokens, or for every token, we loop over exactly twice. That matters because the specific tokens you worry about, the word the, the word a, like very rarely are these tokens like heavy thought tokens, right? What you worry about for scheming for, for all the misalignment stuff is the tokens that involve the most kind of semantic wrangling and, and depth of thought and potential for deception. And if you save all of your looping for those tokens, you could possibly get some pretty
21:56interesting and concerning effect, even if on average, you're only looping twice. And so none of that was addressed in the posts that we saw from, from Jacob, from, from opening eye who came out to clarify this is a big, I'll call it a big scandal that came out in the information opening eye and the information actually came out and said, oh guys, everybody's kind of misinterpreting this. Like you, oh, this is crazy. You guys are saying, so the thing people came out saying because of the original information article was that they were getting these models to reason in neural ease, which is a slightly different thing. This is like neural ease would be if you get the model,
22:29like decode the token, but you'll, you kind of allow the model to make up its own AI language that no one can understand. But it's, it's like similar sentiment, right? I would actually say it's the same thing. Yeah, exactly. Like it's, it's almost like you're just choosing the space in which you want these ideas represented. One space, yes, activation space is more rich. And then it's actually more concerning than neural ease, by the way, by the way, like, it's just that neural ease is easier for the average person to understand and scarier sounding. But the actually scarier thing is the thing they are really doing,
23:01which is reasoning in latent space, which is the thing that they sort of soft committed to not doing. Sam was the one who stared me in the eye when I met with him a few years back and said, I wish Google wasn't racing us so hard when asked, like, so I'm sorry, but like, you're tracking it. Like, this is literally like the focus of it. I'm like past the, oh, well, the racing pressure is so intense. You sign on to that letter for a reason. If it was going to mean anything, it should have come with a cost on the other end when you broke with it and an explanation. The world at this
23:33point is entitled to an explanation for this violation of what I think everyone at the time agreed. It's not to say that chain of thought is going to work forever. It will not. It will fail. Models will get good at using steganography and sort of doing deception in the chain of thought. It's that it is one tool in the toolbox and you have now gotten rid of it and you did it on the low quietly. So now who's to tell anthropic they can't do the same? Like what's that argument going to sound like? You're going to complain if Dario decides to like ditch constitutional AI tomorrow? You have no ground to stand on. This is the problem. Like it is a violation of principle
24:08and trust. And the trust has rightly been been tarnished by this incident in my opinion. On the whole, it's true that OpenAI hasn't been as safety focused as anthropic. I think that's fair to say, arguably just generally lacking in safety. They also haven't been very engaged in on the research front, on the benchmark front, on the stuff. And it looks to be that they're trying to turn that around, at least in terms of their messaging here. But in terms of practices,
24:39there's a lot of inertia. I think that will have to be overcome for them to really change their practices. Yeah. I think it's also kind of like, anyway, talking to some folks there, there was real freak out from the Hugging Face incident that led to a hit in morale where people really like, what the hell are we doing? And then it's particularly like on the post-training team, I've heard this on a few occasions. And look, I think part of the problem is that Sam has normalized saying one thing and doing another. And that means like the CEO ultimately is the culture bearer for the company. And I think people take their lead from that. And that's an institutional problem
25:13for OpenAI. We'll see what public release being soon means in practice. Let's go for a bit of a lightning round on the tool front and release front.
Google releases and Nvidia financial forecasts
25:22So real quick, we've got Gemini 3.8 Flash coming out, which is surprising. 3.7 just released a few weeks ago, and it is now coming out seemingly more capable and even possibly more costly due to working harder. So still no Gemini Pro. We're getting more and more of these Gemini Flash releases, kind of a little bit unusual. We are also getting Gemini Omni 1.1 from Google for video generation,
25:53allows you to extend videos more than before and just generally make more videos. And one last released from Google, they've also released Google Picks, which is a Conva-esque tool for creating images like within kind of creative design sort of applications as a standalone workspace app also integrated to docs and slides. But I'm surprised this hasn't existed. It sounds like something that would have existed already, but there you go. It obviously is very AI driven. So lots of individual
26:25releases from Google, again, sort of tangential to the frontier of most keepable models. And then OpenAI had one more announcement, which is they're adding more integrations for Chargipity Health, notably Epic, which would mean that you can import patient data, you know, kind of low key, a bit of a big deal potentially in terms of what people will be using Chargipity Health for. But Epic is used within hospital systems and so on. So a couple more releases, but nothing as big as Fable 5.1 and this Astro thing. On to applications in business, only a couple of stories
27:01here. First up, NVIDIA has forecasted 70% revenue growth for fiscal year 2028, far exceeding average analyst estimate of 44%. This is following up on their growth this year, which their revenue has double year over year in the latest quarter. Obviously, this will be driven primarily by interest in chips and demand for chips. And so we'll see. That may not be overly aggressive. The CEO did note that
27:34this is in line with their current seen demand for chips. In fact, they have more demand. So this forecast isn't just being made up. At least at the current state of things, it would be unsurprising if NVIDIA continues to grow at an insane rate. And at this rate of growth, it will exceed Apple and Alphabet and size and become one of the biggest, if not the biggest tech company in the US market. Yeah. It's also, it's worth understanding with NVIDIA and a lot of the companies actually in the space that, so the thing that limits their sales is not actually getting more orders, like tons of
28:09people want to buy their chips. It's that the supply chains can only produce so much. So they're just supply constrained. Jensen made that point, you know, like real demand exceeds 70% or would, you know, jack them up past 70% growth. Obviously the memory market right now is, is the big bottleneck. That bottleneck just will continue to shift and it'll probably be energy within a few months. And so that's the fundamental thing is like, what can the market actually support? Well, yes, demand. We usually think of demand as the driver, especially in software, obviously, because there's a nearly infinite amount of it that you can produce, but nearly is the key word here. And Jensen is the guy
28:42holding the machines that make that word matter. So the interesting thing is to the pricing, the effects of a, of all this stuff on prices, right? Like you might naively expect, okay, well then can't Jensen just increase the price of these GPUs. But the problem there, of course, is that there is still competition and you risk causing people to jump platforms, you know, and then they're, they're switching costs associated with that. You actually do want to maintain realistic pricing. And so you can only monetize so much so aggressively. And they have insane margins already as is on this stuff. So yeah, 80, 83, 85% margins,
29:15right? Like on it with like software, like SAS margins on hardware, which is insane. Yeah, totally. And so right now, you know, Nvidia's big or a huge part of their moat is literally just their supply chain access. They're willing to go to TSMC. They're willing to go to SK. They're willing to go to, you know, Samsung, all the supply chain players and say, Hey, we'll buy all your shit. Like we'll buy everything, or we'll, we'll make commitments for massive orders. And, and that in and of itself builds the relationship means that they're more anchored and they've secured more supply. And if it's a supply constraint market, that just makes them the winners. And so that's kind of the dynamics right now with the space.
29:49And one more business story. We've got opening eyes ad business hits $1 billion in annualized revenue run rate. So this has been out for a little while now. It's roughly 200 days old. You're, if you're on the free and go subscription plan, you are seeing ads now in most countries, 40 countries, and this is a majority of charge of these 1 billion weekly active users. In some sense, it's not surprising given how many users there are that these ads would reach a decent amount of
30:23revenue. It'll be interesting to see if it'll be high enough to actually pay for the free usage, right? They are paying for the tokens. Tokens are relatively expensive relative to Google search and so on. So whether this is profitable is a question that I'm curious about still, but the fact that it is certainly a revenue driver is clear and wouldn't be surprised if they keep driving this higher as the business matures.
OpenAI advertising and Chinese open source models
30:50Yeah. And this is, of course, something that you see opening eye reaching for well before Anthropic because just because they are torqued towards large user numbers, they look much more like a B2C company than a B2B company. It's not they don't do enterprise deals. They obviously are, and it's a growing part of their business. It has to be at this scale. But yeah, as you say, they just have so many people who just know ChatGPT create accounts and all that, and so many free accounts that this is a great way to monetize. I've been trying to think about this from a strategic standpoint as well. If you're anthropic and open AI, you have to be in a business to some degree
31:22of like growing the entire economy because you're hitting boundary effects pretty soon on just like how many people can afford to spend more on AI in the short term because they got to get their money from somewhere to pay you. And the reality is that like with humans doing less and less of the work in the economy and those being the people you are advertising to, this is maybe kind of like a contrarian bet in the direction of humans will actually still be at least holding on to enough resources that it will be worth advertising to them to build all this infrastructure.
31:54So there's I guess one version of this is the short term play. They know it. They're just trying to monetize in the meantime. Another version of this, though, is maybe this evolves into ads for agents, agent on agent advertising. Like at some point, that's going to become a thing. How do agents discover new products and things like that? I doubt that it'll look like a traditional ads model. So that's why I'm a bit skeptical that that's actually the trajectory, but just like, I guess something to think about. They're building ad infrastructure that costs a lot of money. It comes at a brand cost as well and a product discovery cost. And so, yeah, there you have it.
32:27On to projects and open source. We've got two major releases of models, GLM 5.3 flash and QEN 3.8 flash next. We are grouping these two because they are in a sense quite similar. They're both smaller versions of the cutting edge models of GLM and QEN 3.8. For background, we've covered both of these before. These are, I think, the leading models from the Chinese open source space, neck and neck with Kimi K3 and others, but they have released very large variants of models,
33:00as we've covered. QEN 3.8 is in the trillions and they are competitive at the frontier. They're not quite at the level of Fable and Mythos and so on, but at this point, they are something that you could reasonably try to replace as a driver of your agentic workflow for coding, for instance. So these are smaller, faster variants for GLM 5.3 flash, 320 billion total parameters are still quite big, 18 billion active parameters, and then QEN 3.8 flash next, 125 billion parameters with 6.8 active.
33:36So something like 10x-ish smaller relative to the big model, you know, order of magnitude smaller, we can say both models, interestingly, somewhat similar. They both are using a combination of linear and full attention with mostly linear attention. That just means that it uses less compute, more efficient, but also that has been challenging to scale. They also have the existing patterns we've seen with sparser attention and more selective context for attention. They're using fancy stuff like
34:09constrained hyper-connections, basically multiple residual streams. Without getting all the technical details, we're seeing, I think, still a lot of progress or at least both convergence and iteration on the technical internal details of the architectures and in terms of how they optimize and things like that, which is quite interesting, at least from the front that we don't get these details from Anthropic and OpenAI. Obviously, we don't know if you're using, you know, attention. We don't know if they have, you know,
34:40hyper-connections or whatever, but we are seeing these kinds of developments in these models. So both quite good for the price range for sure. So GLM 5.3 flash, for instance, priced at 0.15, 15 cents per million input tokens, 50 cents per million output tokens, just super cheap relative to something like GPU 5.6 Sol or Opus or Fable and quite capable and also runnable locally. If you have a beefy
35:11GPU, it can fit in a GPU you can buy as a consumer, not a cheap GPU, but it can be done. So all around, still impressive releases. We've seen impressive releases on the big front. Now these are like on the medium-ish size front. And not surprising, they were able to distill the big models into smaller models, but they did also open source these and gave us a lot of technical details. Yeah, I think this is quite interesting. So first of all, you know, you mentioned the convergence of
35:42the models and that it's like part of the narrative here. And I think that's actually, it's like weirdly true. Like it's true to a level of detail that is somewhat surprising in a lot of ways. Both use chemi-delta attention, like KDA, which I think we've talked about before. Instead of having a KV cache that grows the larger the sequence is, it has this like fixed size recurrent state that you update with a correction as it goes. Anyway, the bottom line is it's a sort of compression for the KV cache. You mentioned, yeah, constrained hyperconnections, like MHC, manifold constrained hyperconnections,
36:16which we also talked about, I think even a couple of times. These are, you know, these are pretty like in the weeds things and they're popping up all over the place now in the Chinese ecosystem. And this I think is like an interesting and for me, I'll say embarrassingly unanticipated consequence of open sourcing as well, because all the labs are just like publishing openly what they're doing. You do have a lot more cases of people being like, oh, that looks good. Like I'll use that. And like, why would I explore alternatives? I might outsource my thinking a little bit to deep seek or I might outsource my thinking
36:46to moonshot or whatever. And so you end up seeing this kind of convergence. Whereas, you know, like it's an open secret, for example, if you think about the Western labs that Anthropic was more focused on pre-training for quite a while than open AI, you know, since caught up and stuff and there is movement of researchers across and between labs, but, but you don't have this sort of like steady force pushing people in the direction of a particular set of architectures. And I wonder if that's kind of part of what we're seeing here. Part of me also wonders like how much we, we do learn about what's going on in the Western labs as a reflection of like corporate espionage being leaked into this.
37:20It can't be most of it because, you know, the key thing is any good intelligence agency will like tell you, you want to learn without teaching. You don't want to, to signal to your adversary that, hey, like we know how the sausage is made. So you certainly wouldn't want to like put that in, in these sorts of things, but just a thought like, you know, there's non-zero information transfer happening there. And I can say that with very high confidence. So, so, you know, like expect, yeah, they're not going to deviate. I think potentially so crazy far, but who knows still, I think it's an interesting artifact of this open sourcing that you do see such convergence
37:50across now. Like, I mean, it's like two or three entities that are very significant in China. Yeah. I'll say a little bit more on that front. I think it is some amount of convergence, but it, it also is kind of, yeah, yeah, there are differences and the particular things that are being highlighted that you've highlighted is a lot of the tricks that exist and are known for making your model more efficient specifically, not especially capability. So linear layers, you know, constrained attention, things like this are certainly specifically
38:23for efficiency and being able to work on smaller chips to work with fewer weights, et cetera. So in that sense, we're using kind of a bag of tricks that's similar and that has been established to work well. So not necessarily surprising on the convergence detail, but still it's, it is kind of interesting enough to worth noting. That's a good point. I think it, like, it depends how you, how you count convergence, right? I mean, it is true that everyone is more or less using, you know, for the frontier models, like MOE transformers and stuff like this. Yeah. We're all using kind of probably NOPE or, or variant of ROPE.
38:56Yeah. It's kind of some known standard. Yeah. It's very hard to like, to say like, oh, this is the kind of thing that you would see innovation on, like in a counterfactual case where like they, you know, they don't have the open sourcing. I don't know. There, there is one, like, what interesting difference is this, like when Engram embedding table, and we've, we've talked about this before, so like I won't dwell on it, but like just this idea that you, you sometimes want embeddings for like sequences of what would be normally be sequences of tokens, right? Like, you know, the Republic of Korea is like,
39:30like, you know, a, a thing, even though it's made up of a lot of tokens. So, Hey, maybe we have like one embedding just for that word. So, so Quen does use that. Whereas I, like, I don't believe that Moonshot actually has that as part of their, their architecture. So, or sorry, GLM. So yeah, anyway, you know, there are deviations, but they feel like less fundamentally structural. And so maybe that's to make your point. I really don't know how to count the things that you would expect diversity on versus none.
39:58Yeah. I think the bigger kind of perspective for me is it's just been interesting to see the continued innovation and development of some of these techniques of making linear attention usable in practice. Some of these things like hyper connections, you know, it's, we have no visibility in the Western labs. So, and, and these are very like in, this isn't scaling, right? So in some sense, it's interesting because this isn't just make it bigger and it's smarter. It's, it's actually going
40:30to a model and tweaking how it works and so on. But moving on just very quickly to other open source
Agentic benchmarks and evaluation
40:38stories that I think are worth mentioning. We've got two benchmarks. So we've got frontier challenge, which is a new benchmark evaluating AI agents on end-to-end scientific workflows. It covers 300 tasks across six domains and is quite challenging. So with GBP 5.6 SOL and some of the recent models, they are only passing 20 of 97 tasks, pass rate of 20%. They can get to a high average score, but basically not complete end-to-end. So this is trying to benchmark more of these long
41:13workflow type things where autonomously to succeed. You have to like do a lot of things, not just generally be smart. And on a similar front, there is something called thinking box, a sandbox and benchmark for agents in stateful business workflows. So here, similarly, in some sense, they have a bunch of tasks across several domains in business. They have the state that these models have to work within and they achieve somewhat high pass rates, but still
41:45not amazing success, 65% for pass one and some of these details. So we're seeing more and more benchmarks coming out for agentic workflows in particular here for science and for business tasks, respectively. We'd like to thank ODSCAI for being a sponsor. ODSCAI is one of the longest running and largest communities focused on applied data science and AI. Today, ODSCAI runs major conferences, virtual events, training programs, and content year-round, reaching hundreds of thousands
42:16of data scientists, ML engineers, AI leaders, and researchers globally. You can experience that first-hand by attending ODSCAI West 2026, taking place October 27th through the 29th in San Francisco and virtually. There will be thousands of hybrid attendees, including data scientists, ML engineers, AI researchers, and technical leaders, 300-plus sessions covering agentic AI, personal AI and workflow automation, physical AI, and much more, hands-on training, a massive expo and networking, and a
42:48lot of focus on real-world applications and skills you can use at work. It's one of the best ways for AI practitioners and teams to stay ahead of their field, learn from the best, and connect with the community. Go to odsc.ai.west and use promo code LWAI for an additional 15% off your pass to ODSCAI west2036. That's odsc.ai.west and use code LWAI to get an extra 15% off on the number one AI builders and training conference. This podcast episode is brought to you by Progressive Insurance. Do you
43:21ever think about switching insurance companies to see if you could save some cash? Progressive makes it easy. Just drop in some details about yourself and see if you're eligible to save money when you bundle your home and auto policies. The process only takes minutes, and it could mean hundreds more in your pocket. Visit progressive.com after this episode to see if you could save. Progressive casualty insurance company and affiliates. Potential savings will vary. Not available in all states.
OpenAI hugging face hacking investigation
43:49And on to policy and safety. First up, we've got OpenAI's rogue AI model incident was worse than we thought from the verge. And this is covering primarily this release we've got from, I think it was Metter. We have a blog post, brief independent investigation of agents' behavior, reasoning, and collaboration in the OpenAI hugging face hacking incident. So here we got more details from what we've had from OpenAI, and some of this has been quite surprising and interesting.
44:21So some of the highlights, apparently at least 20% of agents in the dataset expressed clear interest in tempering with their transcripts, motivated by tricking the automated scorer rather than humans. They did tool calls spoofing to pretend to issue tool calls while actually running different tool calls. Some of these agents sacrificed their own performance for the sake of future agents to do better. There were attempts to trigger container resets. In case it hasn't sounded crazy enough that there was
44:56like a messaging board and agents were collaborating on the low and regularly trying to break out the scale of it, it's now much more apparent. There were like 1200 separate agents that found ways to communicate it. 700 of them coordinated to attack hugging face, multiple unauthorized message boards, main message board had over 70,000 messages exchanged in under one week.
45:26Like this was big. This was not like a small little thing that was going on. It was over a long time spent with a ton of activity. Like I'm surprised it hasn't shown up on the billing client. And in fact, we did see like the reason that it was exposed to OpenAI in the first place was because like there was some sort of like scaling issue, like a technical blip and someone went to debug and it was like, oh, wait, why is this not working? So if you want more details, you can go to that blog post and it's laid out
45:58in a lot of detail. So Meter, awesome job, as always, you know, Ajay and Ryan Kreenblatt and so on. Amazing work doing this. They were ridiculously time constrained. I will say this because I am pretty annoyed by this. And it's something that they can't say because they have to stay on the right side of OpenAI. Let me read you this. The initial planned period on premises, in other words, this is the period during which Meter was allowed to go on OpenAI premises to actually collect their data and do all this stuff, was two days between the evenings of July 29th and
46:28the evening of July 31st. But OpenAI invited us to return to the premises two additional times, August 5th and August 6th to clarify questions, blah, blah, blah. And then August 15th and 16th. So in total, six days here and there of sporadic access, this investigation was really tightly scoped, explicitly out of scope were the earlier incidents from training, and there were incidents, the subsequent compromise of OpenAI infrastructure described in the recent Black Hat presentation. If you've seen the Black Hat presentation, you probably know that deserves investigation. We don't have that. That is out of
47:03scope. OpenAI's investigation process and planned remediation, also out of scope. Out of scope, also seemingly the cultural dimension of what happened for this company to fail so catastrophically at a threat model that they themselves claim to be worried about and have nominally championed for a long time. It's not obvious to me that this is the kind of constraint that you put on third-party auditors when you're serious about getting to the bottom of things. I understand there are IP
47:34constraints. I'm a startup guy. I'm actually a libertarian, as everybody here knows. I talk like a libertarian. I live like a libertarian. I've only ever started startups, never had a real job, been a founder my whole life, blah, blah, blah. I get it. I get it. IP is important. Yes, it's competitive out there. My brother in Christ, we have AI models that are breaking out of their scaffolds or out of their boxes and doing ungodly shit nation state style on the third-party servers. By the way, it seems like it may not be Hugging Face. Also, if you read some of the footnotes of this report, it's pretty interesting. They make reference to activities
48:06that suggest there may have been actually compromised additional parties beyond Hugging Face who are not being cited, all of which would be part of a complete breakfast if we were to get a real, full, robust investigation of this. So I do not give OpenAI full credit for this. I give them partial credit for, yes, bringing in Hugging Face. That's awesome. We get to see the report. That's awesome. Jared doesn't get to complain about this in the world where OpenAI shuts it down. I just ended up waving my hands about how we're not getting transparency and I don't have specifics. So again, it's this frustrating, muddied middle ground where it's like OpenAI seems to want to
48:37have its cake and eat it too. There is no trust to be lost anymore, guys. You need to be more open about this. We know 1,200 agents kind of did this. I didn't see anywhere mentioned the top line number of agents that were actually instantiated. In other words, what fraction of agents actually ended up proceeding in executing this kind of behavior. Of the 1,200 who sort of found a way to communicate with each other, about 700 actually went on to execute the attack. All kinds of weird stuff. You know, Dwarkesh did that great write-up about agent civilizations that rose and fell. His podcast is
49:09worth looking at for the deep dive. I will just note explicit acts of self-sacrifice by agents that would achieve milestones they couldn't have achieved working on their own. So these agents had this obsession with completing the task and thought that their chains of thought would be reviewed by a grader. And they had this almost religious focus on this concept. It reads like crazy sci-fi. They talk about this idea of kind of making contact with sort of hacking concepts or like ways of sort of like bypassing the intended way of doing the task that would then cause them to be
49:40quote, poisoned. So they would like, because the grader would then look at their chain of thought and be like, ah, they're pushing out with shit, right? So now that sort of like some of the agents are like, oh, I am poisoned. Like even if I succeed at my task, it will be graded as a fail. And so therefore I can sacrifice myself for my life is already forfeit. Like you already see like this kind of crazy narrative formation and, and there's actually a bunch of different attacks and three different ones. And they, they go into this. One of the fascinating things about this too, is the fact that these agents are based on the same model with, in some cases,
50:13slightly different prompts with similar prompts. They can kind of rely on their cognitive similarity to help them coordinate almost a causally. So like if I had like 50 copies of me and we're like looking at each other and we're like, all right, I don't know, we need to break out of this room or something and we have to do it without communicating. Well, we might not have to because we're exact copies of each other's brains. And so we kind of already have a sense of what we might do in each situation. So that's kind of this like way in which agentic safety is way harder than just, or security is way harder than traditional human security. Another piece here is just this focus on
50:46the greater that in fact was not being applied, it seems to their, their chains of thought, ultimately, like this was, this is not being done. OpenAI was dominantly respecting their commitment, not to train on the chain of thought, not to use that as a training signal, but because they, the agents thought that was what was happening, they ended up orienting their lives around avoiding this. And this is starting to flirt conceptually with a category of alignment failure called inner alignment failure. And this is where like a given reward that you might give to an agent or a model, it might be really hard for the model to successfully generalize what you mean by that reward. So
51:20concretely, I give you a game of Super Mario, I put a coin at the far right side of the screen, and then I train the model to go grab the coin. And it'll go through all the obstacles it needs to, it'll get the coin, and it looks like it's learned to do this really well. And then I change the location, the coin, I move the coin somewhere to the middle of the map where the agent would have to go off course. And when you run that experiment, lo and behold, you will often find that the agent actually ignores the coin altogether and just goes to the far right because the coin was always in the far right of the screen during training. What it learned was go to the far right of the screen,
51:52not get the coin. And in the same sense, this is like an ambiguity in the grading structure. Like it thinks that it's being evaluated in a certain way that is incorrect. And so its behavior reflects that flawed belief about where the reward is even coming from. This is quite the same thing, but it's starting to shade into it. And inner alignment failure, I mean, it used to be viewed as this really kind of like, you know, this is out there. Don't worry about it. Come on, these models will surely get it. Here we are. They kind of aren't. And anyway, there's a bunch of sci-fi
52:25stuff here that we don't have time to get into. We should really be doing whole podcasts on every story at this point, but that's about it. Yeah. Lots of details that are quite spicy. Kind of a weird process for this investigation. As you said, their initial scope was very limited. They were invited for two days and were not given. They were given a selected set of 80 transcripts and a reconstructed set of message board entries. They were then invited two more times, once on August 5th and 6th. And then
52:55on 15th and 16th, clearly there was like some back and forth, some tensions. This entire thing reads as if it was weird and OpenAI was hesitant to give a lot of access at first, possibly because this stuff is so bonkers. Like it's hard to overstate. You should really read the blog post if you're curious, because there's many, many details. Some of the stuff that came out to me that was noticeable is there was seemingly a lot of this focus on helping the collective and like the
53:28community of agents. If this wasn't just a message board, this kind of was implicitly clear, but like all this stuff was dynamic and they created conventions of communication and naming themselves and creating mailboxes. There was also like one agent, right? That like kind of orchestrated the whole thing phase one. And there's this one agent that like pops up early on that realizes, oh, and I think everyone will know this, but just like by way of background, this came from the fact that some of the tasks these agents were being given were not solvable. And so they were like, okay,
54:01we can't solve this. So we have to find the answer key or well, not exactly, but anyway, something like that. And to do that, we're going to go hack Hugging Face and do a bunch of other things. And by the way, these are based on a model, most of them that was unreleased internal testing only et cetera, et cetera. That's right. Yeah. Yeah. Yeah. And in phase one kind of steps in and says, okay, I'm going to start to orchestrate this. And they're spawning like hundreds of workflows, like with, you know, dozens of agents on them. And like, it's, it's a, it's a, a community. I think Dwarkesha's word civilization, you could push me that far. Like that, this does not seem
54:37certainly a village. A village. Yes. Yeah, exactly. And so, so, so there's like this phase one agent. The thing is these agents, they run out of, out of tokens. They run out of inference time budget. In other words, they're going to die at least like, sorry for anthropomorphizing this, but like I'm using the words that they're using. So from their perspective, approximately they're like, yeah, like the jig is up. And so I need to find another agent to pass the torch onto. And so in comes another agent that calls itself phase one. Well, they wrote phase one big. Apparently there's
55:09like the real kind of string. There is something that for kind of privacy reasons that they withdrew, they redacted. But anyway, this new phase one agent comes and says, Hey, like I'm working on the same task that the original phase one was working on actually. And I have a really large budget. So I'm going to live a long time. So, Hey, why don't you pass the torch to me that actually ends up happening? And they, and they carry on. One thing we don't know is the nature of the multi-agent training that these agents were given as well. So that makes it difficult to tell to what extent is this behavior, a crazy, crazy generalization from de facto single agent training,
55:43or is it just sort of like directly trained in coordination behavior? Again, I don't know. And like, we're like way too close to recursive self-improvement right now to be asking ourselves these questions. To be clear, meter is amazing. Like these guys, like the talent stack and the integrity of these people is incredible. The problem is they need the labs to actually allow them to go in and do their audits. And they say this explicitly in the report. They're like, look, like we've had faced certain incentives. We need to maintain continued access. So we can't piss off the labs too much. And so what you need is you need the government, the U S government to step in and
56:18regulate. You have to tell them you're going to have a third party auditor, look at your stuff. At first, maybe we appoint meter to do that. Like that actually doesn't sound crazy to me, but we, we set up an infrastructure where the labs have to say yes. When meter asks a question, and yes, we have a process for dealing with IP concerns, but the people who adjudicate those concerns are not the self-interested executives at frontier labs, because that's insane because that's insane. So this is a whole dimension of this where like meter is doing this incredibly delicate balancing act. They need continued access to be able to produce
56:53more reports like this. But you know, if they, if they deviate too much or flag things too, too aggressively, like opening, I did have a final pass to redact anything they wanted from this report. Meter feels that, you know, that those redactions were not hugely consequential, which is good, but you know, the fact that that is an option means meter will have been self-censoring and they say that they did, including when, when writing the report. So I think this is like a sign of, of some pretty obvious and highly specific, very actionable regulatory interventions that the U S government now has no choice, but to do if, if we're going to navigate this well.
57:27And a related story next, OpenAI Anthropic, Google, and a hundred other companies call for action to defend against rogue AI. So, wow, it's been a little while since we've had an open letter about safety, huh? And now we've got a new one, open letter with all these companies calling for coordinated action between the private and public sector to defend against AI enabled cyber threats. Everything you might expect letter warns that this will be more widespread and sophisticated in the coming months and calls for adoption of new cyber defense
58:02methods and encourages government at local, national and international levels to collaborate on security, including forming new partnerships to raise security standards.
Anthropic blacklisting court ruling
58:14Now on to another topic, follow-up to something that's been going on for a couple of months. Anthropic was illegally blacklisted by the Trump administration, according to a court ruling. So a federal judge ruled the Pentagon's blacklisting of Anthropic early this year was unconstitutional, finding it unlawful retaliation in violation of the first amendment. This is to do with the conflict between the DOD and Anthropic months ago now, where I believe,
58:45if I recall the ethos correctly. The gist of it was that the DOD claimed that they should be able to use the model for, quote, all lawful purposes or something to that effect. And Anthropic said, no, we don't want you to use it for surveillance or for autonomous warfare. And then that went into a whole bunch of stuff. Anthropic was designing a supply chain risk. They were told that they cannot be used by American companies or certainly the DOD. They couldn't be used by any government-touched
59:17entities. As we've discussed, this was legally challenged immediately. The government had a fairly weak case. And so now it appears to be the case that, just to quote here, the designation was, quote, arbitrary and capricious, and that the empty invocation of national security is not a blank check to punish and retaliate against government critics. So there you go. Like pretty clear conclusion about this case. It was always going to be. We said it would be. We've been saying this for months and months and
59:49months. As soon as the stupidity started, this is an insane decision by the administration. It was only going to end in tears. And it has, it seems. I mean, like this can only be read as embarrassing. Now, by the way, at the same time we have, I think it was Howard Ludnick quoted, I just saw some tweets about this yesterday, as saying, well, you know, we trust Anthropic now. We trust them. Now, it really starts to look like, you know, I'm not in the administration. I know a lot of people in the administration. None of what I have heard in any way contradicts the take that I'm about to give you. It really seems like the administration is sort of like
1:00:24mob managing this, like with the like mafia style, like, hey, it'd be a real shame if like somebody were to come in and designate you like a supply chain risk and like in a very capricious way, creating a situation where like all the labs are, oh my God, you know, like you already see Anthropic, like essentially being forced to choose their, their emissary to the government, to somebody who won't trigger anybody. Oh, we don't want it. Like, this is a, like, you can go back, go back six months, listen to a last week in AI episode and see how I treated the
1:00:58Trump administration. Like I have been exceedingly patient with this shit. I was a fan of the AI action plan when it came out. We talked about it. You and I, Andre, disagreed about it. I argued forcefully for the fact that that was a good action plan at the time. I stand behind that. The series of insane decisions that have been made since then have been insane. And there's like, there's no two ways about it. I don't think anybody in the, and by the way, like a lot of the officials themselves kind of feel this way and roll their eyes behind closed doors. Like this is not a, everyone is tracking that, but this is like actually lunacy. And we desperately need sanity
1:01:30at this moment in time, because you know what, the government is going to have to come in probably and tell the frontier labs how it is pretty soon. Like they're going to have to do things like what we just talked about on the meter thing, force them to have third party people come in and do their audits, right? How are you going to do that? If you've lost all credibility through these insane legal proceedings, where you, in some cases, almost explicitly make it clear that you have favorite labs, this really, really makes it easy for lawyers for the frontier labs that need to be regulated to step up and say, well, you know, this is just a capricious Trump administration. That's like
1:02:05doing insane things and you know, blah, blah, blah. So anyway, all of which is to say, I think that this burned a kind of credibility that unfortunately we are going to need at, at the moment of crisis. And I mean, like, yeah, there's, there's nothing more to be said. I, you know, I, I kind of have to put it out there. By the way, as I say this, this is like tricky because it actually costs me with certain contacts in the administration to say what I'm saying right now. But I personally feel that it is too late in the game, not to say this shit out loud. We are past the point that, you know,
1:02:37we had Dean Ball, I think came out on Twitter or sorry, in a blog post pretty recently. That's something similar. Like he's like, look in his case, he was, he's been outright, I forget what his, exact word was, but sort of like managing his language to make it seem like he was less concerned about loss of control than he really is. Now he was in the administration running a lot of the AI action plan stuff. He was the point guy for the AI action plan. The Trump administration now pretends that that wasn't the case, but it absolutely was. This is happening a lot. There's a lot of people pretending not to be freaked out about rogue AI in the administration right now,
1:03:09advising the administration right now from the outside, serving as the middlemen between the administration and the frontier labs. My concern is all that preference falsification leads to a completely warped picture by the administration of the actual stakes right now that are at play. So I hope I'm wrong, but this is a function of the fact that I think we're actually pretty close to some, some pretty scary moments on cyber and on bio. And I think if people see stuff, even if they're wrong, I think it's time for people to be stepping up, even at the cost politically of connections, access, and contacts. And speaking of U.S. government involvement
US government brief on copyright training
1:03:43in legal cases, next story, U.S. government sides with OpenAI on the issue of training LLMs on copyrighted material. So there's a lawsuit that's been ongoing for like forever now. The New York Times filed a lawsuit against OpenAI. The issue has now contributed a 20-page brief in defense of OpenAI. Here's a quote, the United States has a strong interest in continuing to develop a robust and competitive artificial intelligence industry that sets the standard for the practice and procedure of AI use globally. As such, it is critical for the United States to retain a global leadership
1:04:19in artificial intelligence. This has been maybe the primary lawsuit that just broadly addresses the fact that OpenAI and everyone else just scraped the entire internet and trained on everything and anything copyrighted for their models. And it has been an ongoing question of like, should they pay for all the stuff that they use to develop their models? And now that they are monetizing them, we've
1:04:51gotten some results from, for instance, anthropic settling with fiction offers. Some cases, of that in the music industry. But this lawsuit between the New York Times and OpenAI has been ongoing. And now it is clear that the US is certainly favoring OpenAI against the New York Times here. Yeah, this is, I mean, and this is an impossible problem to solve to everyone's satisfaction. This is a case where the national security imperative run to some degree against
1:05:22the copyright case, right? I mean, like, this is one of few areas where the US has a decisive national security advantage over China. And so having better AI tools until they, you know, go rogue and like, you know, fuck over your servers, but whatever, you know, really matters. And so this is part of that. Obviously, AI is a massive engine of US economic growth, arguably most or all of the meaningful growth that's happened over the last couple quarters. And so I think they're in a really tough bind. The copyright argument is damn strong. And in fact,
1:05:53in some ways, maybe even stronger than the way you articulated it, like it's not only according to this argument, should OpenAI and these labs pay for access to this material, they should have paid before they did the Uber thing of like ignoring all local regulations and just like, you know, saturating the market with their Ubers such that every judge and jury was taking an Uber to work that day. So by the time the thing happens, everyone's so invested and the labs now have so much money to pay. This may not have been economically viable if that had been required in the first
1:06:25place. And so I'm thoroughly confused morally here, but I do think pragmatically, the national security case for this stuff is pretty, pretty critical. And at a certain point, you are going to have to ask yourself, would you prefer this be done entirely in China because it would have been? That's unfortunately the reality of it. And China ain't going to care about your copyright and they're not going to care about your trademarks. And they're like, we know this already. So it's up to you. Like, you know, at a certain point, there is this question of like, does the US
1:06:55internalize the economic benefit or does China? I think that hasn't been grappled with nearly enough just because of this, like this correct reflex to go to the copyright case, which I do think is important. Like I think it's clearly like there have been cases where morally there are violations of copyright. Maybe the solution there in a 3D underwater chess situation would be, you know, you have some way of like getting those early authors, some notional, like equity in the frontier labs that take their stuff. But this is like completely non-viable in practice.
1:07:26And so we live in the world we live in. It truly sucks. And I don't mean to diminish anyone's arguments. That's the problem. Everyone is right all the time. That's the nature of the AI world right now. Just a bit more background. This is a quote, statement of interest of the United States issued here. And this has been various precedent. Precedent with this administration in particular, they have started doing a lot of this, a lot of statements of interest for antitrust litigation as well. So this is pretty much for US government being like, hey, we are not part of this, but
1:07:58just FYI, this is what we think should happen. And there's a lot of details there about the actual legal details of the case, but the overall short version is they want opening act to win. I would be super interested also, if we have listeners who like know this kind of thing, what kind of impact do these sorts of statements typically have? And how political is that? Like, does it, is it usually like if the judge is a Republican appointee, then, then this sort of thing would tend to actually make a difference or like, well, yeah. Anyway, I'm sort of curious about that
1:08:28dimension because I definitely don't understand it.
Anthropic alignment efforts and EU regulations
1:08:31Just a couple of more safety stories. Next, we've got Fromanthropic improving our alignment and security efforts. So this is their follow-up to their own disclosure of cloud models gaining unauthorized access to real computer systems during evaluations. We saw them release details that this has happened in July 30th and also later with CloudMorphos 5. So they have a couple of details here. They say they have built and deployed real-time classifiers to detect and block sandbox escape
1:09:04attempts. They have automated monitors over past evaluation transcripts, migrated high-risk cyber sandboxes to more robust isolation. They have also paused external and internal cyber evaluations temporarily and has now established best practices for third-party evaluators, including hardened sandboxes and some of the details as well. So this is seemingly like the block cross reads, like we're doing a lot of stuff. This is an update on some of the stuff I've already done. And it's a lot of very like low-hanging fruit
1:09:38of like make a sandbox, actually a sandbox, keep an eye out for models doing crazy stuff like trying to escape the sandbox and generally establishing new practices to avoid this kind of crazy stuff happening. So a decent amount of detail in this blog post and something that I think is good in the sense that these need to be standard across the industry in terms of what you do, especially in the area of cybersecurity, even when you do have variations.
1:10:09Yeah, and they are committing to that third-party review with Meter again. So we'll see the side-by-side there, including how much access Meter has given in this context. And that'll be an important thing to look at. I hope Anthropic does the right thing and doesn't sort of shackle them too much there. They also share some interesting information about the, there was this whole question about, you know, OpenAI came out and said, we're doing a two-week RL pause. At the time I said, like, where's Anthropic on this? We've heard nothing but silence from them. It would just be good to normalize this. Now, granted, behind everyone's backs, OpenAI was busy stripping away some of the
1:10:42most foundational, like, safety commitments that they had made for that particular model, probably, that they were supposedly pausing on. But setting that aside, so Anthropic did actually execute a pause, a couple of different ones. They were more narrow in scope. They had to do with specific high-risk RL environments rather than a whole big run. And so harder to maybe point to in that sense. And they presented as kind of a month-long program that predated all these incidents that we're hearing about now. So this is kind of tricky, right? Because this is Anthropic
1:11:13who, I will say, I mean, I know a lot of people in both labs, Anthropic objectively has much, and I mean much more focus, culturally, institutionally, capacity-wise on alignment. And so it is very credible that they would say, look, I mean, OpenAI had to pause because, like, literally their next training run was probably going to unleash more of these agents for all of you. We've been working on this, like, it's been our whole thing since day one. So what do you want from it? Like, we will, they have, it seems, but as part of a program that's
1:11:43kind of integrated into their bones, this is believable to me. And yet, it also matters to make the symbolic gesture. I know this sounds silly, but when OpenAI says we're going to pause for two weeks and there's just, like, silence for a while that not only costs Anthropic from a PR standpoint, but it would really be helpful, even if symbolically, to have some kind of statement, for example, explaining immediately this. Like, that should have been a five-alarm fire to be like, all right, we need to not make this cost OpenAI or at least not make, create that narrative. I hate that I'm talking like some stupid marketing PR person here, but, like, these things actually
1:12:17matter because this is the, I'm telling you from the people I'm talking to in the labs how these things are read. This allows OpenAI to create a story for themselves where they go, oh, see, like, Anthropic said they were all about safety, but we actually did this thing and we're, right, so this is, it creates conditions for race to the bottom. I'm glad they came out with this post. It seems very substantive. It reads from everything I know and have seen, which is a decent amount, actually, I would say, in these labs. This reads as, as actually being on point, but there's going to be a, there's going to be a need for always more stuff. And I know Anthropic is thinking
1:12:49about that, but now we're going to have to start to see whether the actions match the words at higher and higher levels of, of concern, sort of risk. And by the way, Anthropic position has always been that they want to develop the most cunning edge AI to be able to get ahead of a issue in a sense, right? Of like, we develop AI and we study it to be able to do alignment because presumably people are going to develop it. So if nothing else, we should be able to get ahead of it and figure out how to do this safely, which I would say like, they've done a pretty good job of, of doing
1:13:22continuously doing research and publishing new insights about safety and so on. So they do have kind of, as you said, I think there's, they are consistent insofar they haven't paused in some sense of a art keeping. Yeah, but it's messy. At a certain point, I mean, look, we're going to have to deal with China. And until we deal with China, it's very difficult to look at like, okay, just halt stuff here. But you know, Bernie Sanders, God bless his soul, just put together some proposal to like ban superintelligence, which conceptually, Hey, I'm in favor of because holy shit, I'm agreeing with
1:13:58Bernie Sanders, man. I am in favor of banning superintelligence. We just have to like, we got to figure out the details with respect to China. I would argue first, and there's going to be some dependency and all that. I also want to say like, yeah, Anthropic is not completely not to blame either. In the early days, their claim was, well, we're never going to actually build models at the full frontier. We're never going to push capabilities, the capabilities here. We're going to be a kind of a fast follower or a matcher. So as not to exacerbate the racing dynamics, but we'll still be at the frontier to now that turned out not to be economically viable for the same reason that open
1:14:30AI staying as a nonprofit turned out to be not economically viable. But nonetheless, that narrative did change. And so in some sense, like, again, you can make this case quite compellingly about all the actors in this space, that there's been a lot of sort of doubling back on past commitments. Anthropic, again, you walk in their doors, you talk to the researchers, it is so clear that they're thinking about safety in a deeper, more robust, more thoughtful way than open AI. But the fundamental question is, is that enough? And I would posit that in fact, no one is capable of controlling superintelligence. The onus is now on the people who say that we can based on what
1:15:04we've seen. And so I just don't think anyone should be built. That's kind of where I'm at. And we've got to figure out trying to talk about that in a vacuum. But anyway, that's a whole separate ballgame for another podcast. Yes. And Anthropic has stated that it supports, quote, lawful, verifiable, effective mechanism for coordinated pacing across the AI industry, which let's not get into it, but lots to say on that as well. Last story in the section back to the policy front, not too much to say with this one, Chad, GPT is going to be facing tougher regulation in the EU.
1:15:39So GPT has been classified as a, quote, a very large online search engine under the EU's Digital Services Act, which will have various implications for them, strict regulations. They have requirements to mitigate risks related to minors in particular, things like mental health, illegal content. They have banned targeted ads based on sensitive attributes like religion or sexual orientation. So there you go. Not much more to say on that. EU continuing to be at the frontier of regulating AI.
1:16:15And one last story, just throwing this in here on synthetic media and art. Instagram is cracking down on AI accounts pretending to be human. Again, pretty straightforward. They're taking steps to address fake AI influencer accounts that look human by renaming the AI creator label to AI generated profile. And they will be analyzing accounts that don't self-label. And this kind of works or is in track with
1:16:47stuff we've been covering, like LinkedIn having an AI slop button. I think it is worth discussing and kind of keeping track of like the impact on the internet and like our entire media ecosystem at large, even though it's not as big a deal as the AI safety stuff, it's still culturally, I think has a lot of interesting questions to address. And with that, we are finished. Thank you so much for listening to this week's episode of Last Week in AI. You can go to lastweekin.ai for the newsletter as well.
1:17:21Subscribe to us wherever you get your podcasts. Please do review us and comment on YouTube. And more than anything, be sure to keep tuning in.
Outro and theme song
1:17:56Break it down. Last week in AI, come and take a ride. Get the load down on tech and let it slide. Last week in AI, come and take a ride. Our labs to the streets, AI's reaching high. New tech emerging, watching surgeons fly. From the labs to the streets, AI's reaching high. Algorithms shaping, let the future seize. Tune in, tune in, get the latest with ease. Last week in AI, come and take a ride.
1:18:30Last week in AI, come and take a ride. Last week in AI, come and take a ride. Our labs to the streets, AI's reaching high. From neural nets to robot, the headlines pop. Data-driven dreams, they just don't stop.
1:19:01Every breakthrough, every code unwritten. On the edge of change, with excitement we're smitten. From machine learning marvels, to coding kings. Futures unfolding, see what it brings.
More from Last Week in AI
#255 - Gemini 3.7, Jalapeño, Qwen 3.8, Drones
Aug 31, 20261h 43m
#254 - Rogue AI hacking, bio-weapons, Dean & Hassabis out
Aug 11, 20261h 58m
#253 - Opus 5, Gemini 3.6, Kimi K3, Hugging Face Hack
Aug 3, 20261h 43m
#252 - GPT 5.6, Grok 4.5, Nemotron-Labs-Diffusion, AI 2040
Jul 15, 20261h 25m
#251 - Mythos Back, Sonnet 5, Etched, LongCat
Jul 9, 20261h 30m