Steadcast
Last Week in AI cover art
Last Week in AI

#255 - Gemini 3.7, Jalapeño, Qwen 3.8, Drones

August 31, 20261h 43m · 17,941 words

Show notes

Our 255th episode with a summary and discussion of last week's big AI news! Recorded on 08/26/2026 Hosted by Andrey Kurenkov and Jeremie Harris Feel free to email us your questions and feedback at and/or Read out our text newsletter and comment on the podcast at In this episode: SpaceXAI released Grok 4.6 (500K context) as a post-training update aimed at long-running agents and coding, with discussion centered on how the Cursor acquisition boosts training…

Highlighted moments

The projection right now is that Anthropic and OpenAI may account for up to 50% of incremental compute next year. That's two companies. That's moving a sizable fraction of the actual economy.
1:00:22
They tested it with various models, including GPT OSS-120B, DeepSeek R1. So models that have open weights and you can actually download and run on various GPUs. And we're showing that for one, you dissipate much less heat. So they say 700 watts compared to other models, GPT-200, GPT-300 at over 1000 watts.
40:28
In this instance, it appears to be the case that after it crashed, Ukrainian talent was able to get the internals of it, which it was powered by NVIDIA's Jetson Oren module, which is meant to, it's an off-the-shelf AI chip used in university robotics labs and computer vision projects.
1:11:59
Found LinkedIn to be the most AI saturated platform with more than 40% of its long form post flagged as completely AI generated.
1:42:06

Transcript

Welcome and Episode Overview

0:00Hello, and welcome to the Last Week in AI podcast, where you can hear a chat about what's going on with AI. As usual, in this episode, you will summarize and discuss some of last week's most interesting AI news. As somewhat usual, we'll be also including a bit of news beyond last week, since we are catching up. I am one of your regular hosts, Andrei Kurenkov.

0:32I studied AI in grad school and now work at the startup Astrocade. And I'm your other regular co-host, Jeremy. Yeah, Gladstone AI, AI National Security stuff. Soon to be doing a series of more public work around investigating the AI endgame and national security and lab and infrastructure side of it. But yeah, and I apologize. The absence last week was on me. We had our big family green card activation trip. And so traveling with a baby

1:03to LA for some work stuff and some pleasure stuff. So there was a lot going on. Things kind of got hairy there. So I appreciate everyone's patience. We should be more settled. I was just telling Andrei, I've got one more trip coming up the week of September 17th. So there may be a hiccup there. Andrei may be able to get a co-host for that. But just to kind of put that on everybody's radar, things should be settling in more now that all the green card insanity, and it is insanity, is over. Yeah, looking forward to the next, the next phase of the story. So yeah, it's a bit of a process with these things. And as you said, travel can

1:36be tricky. Luckily, we didn't miss too much missing last week relative to before things kind of settled down a little bit since our last episode. So nothing too crazy to cover as a quick preview of this episode. We'll be catching up a little bit. There have been some new models and like tool updates to mention, but nothing too big. We will be discussing things like Gemini 3.7 and Grok 4.6, some new model releases, but nothing sort of earth-shattering.

2:07Applications in business, again, nothing like huge, some interesting developments, and some slightly more kind of nuanced things that if you're following the industry, it'll be interesting. Policy and safety, once again, will be the big section, and we'll have a real mix of stuff there. We'll be talking about drones, which we haven't touched on in quite a while, the security updates from OpenAI and their kind of proposed or supposed slowdown of model developments, which I think

2:40is probably the most interesting development. And beyond that, just, yeah, a real variety of stuff. And we will try to get to some research as well by then. So it should be a pretty good mixed episode.

Sponsor: OutShift and Factor

2:52This episode is brought to you by OutShift, Cisco's incubation engine. Today's AI engines operate in silos, limiting their true potential. We focus on building bigger, smarter models, but scaling up is just one approach. To reach superintelligence together, we need to do more. We need to scale out. And we actually have a blueprint from 70,000 years ago. Humanists didn't just got smarter individually. The cognitive revolution transformed society because we began sharing knowledge, goals, and innovation. Agents are now at the same inflection point. They can connect,

3:23but they can't think together. That's why OutShift by Cisco is building the Internet of Cognition, transforming AI from isolated systems into orchestrated superintelligence. By creating an open, interoperable infrastructure, OutShift is enabling agents and humans to share intent, context, and reasoning. The cognitive evolution for agents is here. Explore internet of cognition at OutShift.com. That's OutShift.com. This next sponsor isn't related to AI, but I've personally

3:54used them for years, so I'm happy to have their support. And it is Factor. They make chef-crafted, dietician-designed, ready-to-eat meals, so you don't have to choose between real food and convenience. Both in grad school and as a startup employee, I don't have a ton of time, so when I get home, I'm tired. And being able to prepare really quite a good meal without any effort has been fantastic. Their meals are ready in two minutes and require no prep and no cleanup. So even on the days your schedule is completely out of control, eating well is still achievable. There are over

4:26175 banned ingredients, so every factory meal is designed around what supports a healthier lifestyle and nothing that doesn't. And that's with over 100 nutrient-dense menu items to choose from every single week. 97% of users agree that Factor meals help them live a healthier life, so you can feel confident that you're doing something good for yourself with every meal. I've really enjoyed Factor, and if this sounds good to you, maybe you should try it as well. Let's eat real. Head to factormeals.com slash LWAI50 off, and use code LWAI50 off to get 50% off and

5:00one free breakfast item per box for one year, while supplies last until February 31st, 2026. That's code LWAI50 off at factormeals.com. LWAI50 off at factormeals.com. See website for more details. Before we kick it off, do want to acknowledge we had some lovely comments on the latest episode. Just picking one out from Golden Dawn. Had some nice comments and asked that we would discuss a little bit about Quen 3.828b, which is a fairly small parameter model, but did well on benchmarks,

5:38like surprisingly well. And also discuss the comeback of SpaceX AI. So I did take a look at that, and we do have those stories coming up. So keep tuned to see our discussions.

Google Gemini 3.7 Flash

5:53And so going to tools and apps, first up, we've got Google announces Gemini 3.7 Flash just three weeks after previous release. So this is what they're terming the new workhorse model, replacing 3.6 Flash. As it said, these kind of came together very close apart. A bit of a weird move in some ways, because U-Mine has been relatively slow in releasing models this year relative to OpenAI and Anthropic.

6:27We've gone through the whole GPT-5 family of 5.1, 5.2. It feels like pretty quickly Anthropic has been through Claude 4.5 to 5 in a relatively short time span. But with the latest release, pretty strong benchmarked numbers for this kind of quicker, smaller model, not comparable to things like Opus or Sol from OpenAI, of course. So this, I think, is showing that Gemini and DeepMind are

7:00continuing to focus more on the models they need for product, for the things that are integrated into drive and various products, right? If you have LLMs doing stuff everywhere, you're not going to be doing that with Gemini Pro, you're going to be doing that with Flash. And so there's clearly, it feels like a lot of engineering manpower directed towards optimizing Flash and making it cost efficient and quick and capable in a very kind of practical move that does to some extent make it so

7:36Google is not competing on the frontier of agentic coding and sort of intelligence, at least seemingly in recent months. Yeah. And I think that's actually, in some ways, the headline. They've shown they can ship fast, but very conspicuously, the thing that they're shipping fast is not an actual frontier model. And that tells us a lot, right? I mean, we've been hearing about Gemini 3.5 Pro. It's coming soon. That's what they said in May, in July. Now we're in August, still coming at some point.

8:06So I think it's starting to raise some real questions and doubts about whether Google has the chops to be a frontier lab. There's a lot about Google's big picture strategy that suggests that they're leaning harder and harder into becoming a, I was going to say cloud and cloud somewhere in between, basically shipping TPUs instead of GPUs to compete with NVIDIA, which in the long run, I mean, I think is a pretty risky move. You know, you risk locking in as the infrastructure layer without the kind of visibility into the needs of the model development layer that you would have if you were a frontier lab. I mean, we're already seeing open AI. We might talk about

8:40that later today, if not in the next episode, but AI has come out with their own custom silicon and it's good, right? Like the jalapeno seems like a legit chip. Anthropic seems like they have the chops to do the same thing. But what it starts to look like, and we were talking about this years ago, the idea that having a frontier model company that can do hardware may actually turn out to be easier than having a hardware company that does frontier modeling. And I think that was pretty contrarian at the time. Now the argument for that seems to be getting stronger and stronger.

9:11NVIDIA's positioning in the market is not exactly weak, but it definitely is going to see viable competition at the design level from open AI to minimum, it seems, if the benchmarks hold up, possibly from Anthropic as well. So yes, Google has TPUs. Those are amazing. But as you start to sacrifice frontier model, like genuine frontier model in-house capability, by losing Jeff Dean, by losing Demis, by losing all these key players, eventually that stuff is going to erode. And you'll see something that like similar to what happened to IBM, where you get sucked into these short-term rewards, where it's like, yes, you can make a ton of money today by focusing on the

9:44infrastructure layer, but at the cost of losing your focus on long-term strategy, where are the insights coming from? It's from the frontier modeling side that back propagates to the sort of hardware engineering side. And it's both, right? They co-evolve and you co-optimize, but you can't do it by flying blind, by outsourcing your partnerships to external labs, which is a big part of the reason why Amazon is not top shelf. It's a big part of the reason why SpaceX AI had to acquire Cursor. Like all the stuff points to, you need the modeling in-house as much as possible to be competitive

10:14in long-term. That being said, this is a really good model for its tier. There's no question, as you said, I mean, so the comparison here is more with like a Sonnet class or to use the opening AI classification tiers, like a Terra. They have Sol, Terra, and Luna, right? So Terra, kind of that middle ground, again, Sonnet between Opus and Haiku. So this is like a mid-tier model and is very good, it seems, at least based on the benchmarks on that basis. It is just like a refinement of 3.6 Flash rather than a new base model. I think you might've mentioned that anyway.

10:44It's based on feedback from customers, supposedly, and like more RL. So you look at these benchmarks, some of them did jump a lot, like DeepSuite, so a software engineering benchmark went up from 49% to 65%. That's a pretty big jump. Automation bench for multi-step tasks up from 17% to 30%, which is actually quite good. Still fails in seven out of 10 multi-step automation tasks. So, you know, that's just the class of model that it's in. But nonetheless, compared to models in its class, it holds up well. The real question, of course, is like, is this kind of new focus on the Sonnet tier, the Terra tier, and not the Sol tier and the Opus tier? Is this indicative of kind of

11:21an admission that we're sort of, if not throwing in the towel, at least de-emphasizing the true frontier capabilities? Right. And typically when you say frontier, we mean basically the most capable models, the most like intelligent models, however you want to define it. You could make an argument that it is pushing the frontier of the Pareto frontier. The Pareto frontier, yes. Yes. So they, like, if you look at some ways of measuring things like the artificial analysis

11:53intelligent index versus time per task, and you graph that, you can make the argument that this is actually a leading model on that measure. It's actually quite fast. That's part of what they're doing here, clearly, is optimizing not just for small model capability and cost, but also for speed. So relative to even GP5.6 Luna, it might be faster for many tasks. And again, I think this is driven very much by pragmatic product needs. And it is a real, I think for me, it's less about a capability

12:32question of whether DeepMind can compete or Google can compete head-to-head on kind of agentic coding and business needs versus what appears to be their current strategy, which is continuing to focus on the model needs for their kind of whole suite of products. And in particular things like Google drive, like you have spreadsheets or you have emails, or you even have a, presumably search where these kinds of things come in with AI mode. So as usual, I think with Google as a giant bureaucracy,

13:09it feels like this is a corporate level kind of prioritization question of there's more resources being shuffled towards what the product requires and less towards this kind of next generation frontier of intelligence kind of effort. Whether that's a smart strategy or not, you can argue either way. I will say, I think if nothing else, having one focus over trying to do everything

13:40at the same time is not too bad a strategy. And they do seem to be doing well on the side of having a good model that can be plugged in to Google drive and, you know, AI mode and elsewhere. So Gemini 3.7 flash for the category, Google is continuing to lead there. Yeah, I'm pretty skeptical about this approach personally could turn out to be wrong course. But I do think like we live in a world where it seems like margins for the frontier models have been going up and value capture at that part of the chain has been increasing quite significantly. And it seems like

14:14there's a secular trend in that direction for a whole bunch of reasons that we don't really like we should do a deep dive on the economics of this stuff at some point. But I think there's a world where this is their IBM moment. And they're essentially kind of the same way that IBM basically became a glorified consulting company. And like, yes, you can make a lot more money in the short term by doing that, but you sacrifice doing the hard thing. And the hard thing is what sets you up for the future. Even when things look down, you know, everybody was talking about models hitting a wall, you know, a couple, a couple months ago, you know, the companies that kept powering through and kept

14:44prioritizing frontier models. Of course, they were seeing internally that the whole scaling is hitting a wall story was complete bunk, which we also talked about at the time. But nonetheless, I'm pretty concerned for Google on this one. I think it is. They're just like such a huge pot of money in the short term chasing infrastructure. And in the long term, so much margin potentially coming in on the modeling side, so much advanced warning and so much flexibility. If you overbuild your... I'll pause there because we got to get... We can talk a lot about strategy and implications and so on, but

15:14that's probably enough. And to be fair, we've been here before. Like, who knows? DeepMind could just release Gemini 4 Pro and it'll be crazy. Like 2025, they had a surprising comeback story. So we'll see. They do have still a lot of talent. So let's not rule them out yet. Next up, we've got Space XAI

SpaceX AI Grok 4.6

15:36releases Grok 4.6, a 500k context frontier model tuned for long running agents coding and knowledge work. So this is a bit of a catch up. This happened, I think, a little while ago prior to last week. This is an incremental development. So it's the same base model as Grok 4.5 rather than a larger base model and continues the seeming development of ever since XAI, now Space XAI acquired Cursor.

16:09They managed to release new Grok models that are good for coding. For quite a while, Grok ceased to release anything. Maybe because the entire like exec team left and like the company kind of maybe got rebirthed. It's hard to say, maybe. And then XAI became a NeoCloud, just giving out computes. So that happened. Now we are seeing a Grok 4.5 and a Grok 4.6. On the benchmarks,

16:40it isn't necessarily as for Frontier, but it looks to be pretty competitive with a GWT 5.6 and with Claude 5. Being very cost efficient as well. So that is an honorable kind of trade-off. I forget the exact numbers, but it's maybe is around half the cost. It's like a pretty competitive model on price. I will say the 500k context of it is worth noting. That is when you do Agenda coding,

17:17having half the context is a big limitation. So one of the issues with the primary benchmarks we see for coding is the ones that are primarily being showcased like TerminalBench is one and DeepSbyE. We haven't kind of gone to a point where there is a leading benchmark for long-running Agenda coding. I would assume partially because it's just a very hard thing to do. You need a very large-scale project, et cetera, et cetera. So whenever we talk about new model releases and benchmark scores,

17:52my mental model is the software benchmarking is not at the point of being able to really stress test capabilities of the sorts of things you'd be using Claude Code or Codex for, which is long-running five-hour work tasks with hundreds of thousands or tens of thousands of lines of code in the repositories, et cetera. You have these smaller things for the most part of fixing a bug in a workspace and maybe a big repo, but it's still a pretty localized change, not developing a new feature,

18:25which requires some design and testing and so on and so on. To summarize, Rock 4.6 seems pretty capable and pretty cost-competitive. If people were considering Rock, it could perhaps start to be competitive with Claude Code and Codex. As far as I've seen, nobody is really thinking, oh, maybe we should switch to Grok build from Claude Code and Codex. I'm not sure that OpenAI and Frobtq are too worried about Grok for now. Yeah, I think the for now thing is maybe the key caveat

19:01here. We've got Grok to go from basically irrelevant to now, yeah, is it a frontier? It's in the cluster of frontier models, right? And so you'll be able to find the odd benchmark where it does surprisingly well, and maybe it is actually the best model for that use case. It is, as you said, partly meant to be highly cost competitive. And so that's kind of one of its sharper areas. The big picture here really is the proof of thesis around the cursor acquisition, right? There's a lot that the cursor acquisition did in theory for Grok and for SpaceX AI. Now we're seeing how much of that is real.

19:34That stuff included, by the way, distribution, right? So before we get into the technicals of what cursor does really well and how it actually gets reflected in this launch, there is just the fact that there's distribution. So there's a reason that SpaceX acquired cursor at a $60 billion value. That was a 15x revenue multiple, which is very high, especially at that valuation. Cursors, by the way, their share of corporate AI coding fell from 41% in June 2025 to 26% one year later. They still grew revenue, but they lost market share. So even despite that, they still had

20:06massive distribution. And so not only are they supporting like this filling of this gutted engineering bench that SpaceX AI had, but they're also bringing in, in theory, a lot of these users to give Grok kind of a bit of a lift here. So the big question here is, first of all, can Grok get good enough to support cursor in the way it needs to be supported to win over the same users despite its acquisition? The reason I say despite its acquisition is cursor would used to route your

20:37query, right, to the most effective model as between OpenAI, Anthropic, theoretically, you know, Grok, but it never really was. And then all these other models. That was part of the value act. Now that they've been acquired by SpaceX AI, they will face pressure, of course, to route more tokens through SpaceX AI. And so SpaceX AI has to have natively models that can actually deliver for the end user under that kind of pressure. It's possible they'll keep routing to Anthropic and OpenAI and all that stuff just to kind of keep the things going. In the meantime, there will be secular pressure in that direction, which means Grok has no choice. It must get good enough,

21:09at least in the default trajectory here. So it is, by the way, notable. So 4.6 is in some sense, the first, it's really the second time that there's been this end-to-end training partnership between SpaceX AI and Grok. It's the first time that we've had it come out around this sort of acquisition time. So, you know, we knew that there was a training partnership back in April. The actual acquisition closed on August 14th, which was two days after the launch of 4.6. So not a coincidence. When it comes to the kind of PR side of things here, but Grok 4.5, so the previous version was

21:43actually the first model that was jointly trained by SpaceX AI and Cursor on the Colossus cluster. So the way this works is the Colossus cluster is going to train the base model. And then the base model is used in concert with like a million developers that generate all these like coding trajectories that become RL environments. And those RL environments then train the next like post-training run. And so when you look at the gains that came from 4.6, they did come from exactly the kind of updated supervised fine-tuning coding trajectories and RL agentic environments

22:17that Cursor is specialized at developing. So the fact that 4.5 to 4.6 is a post-training development is exactly a reflection of the Cursor or SpaceX partnership. And so here, to the extent that we're seeing this big uplift, that is actually the thesis like doing its job. Like this is exactly how it was supposed to work. And if this passes, you know, all the usual vibe checks from an economic standpoint, that'll be a good sign. It doesn't quite have to work. This can be one of the rockets that explodes on the launch pad on SpaceX style, but one of

22:50the next ones has got to work. There's just like right now too much riding on this, you know, whole idea of like needing to have frontier models in-house. If you're going to become, you know, I wouldn't be shocked to see SpaceX become a chip design company as well, by the way, just in terms of integrating all the layers of the stack. Every AI company is trying to own the whole stack because that's what you got to do to earn that margin back. So anyhow, I think a really interesting story. And so far, it has to be a positive update on what SpaceX AI, XAI, SpaceX, any linear combination of those is.

23:22Yeah. By the way, kind of ironic, Tesla is the one that's been building chips for inference and SpaceX hasn't built chips for inference. So like, but they can partner because they're all Elon Musk. To your discussion of Cursor, it's a bit of an awkward situation because Grok built is the Cloud Code competitor. Cloud Code is the thing that's been taking market share away from Cursor. So now like, is Cursor the thing they're going to be promoting, which

23:52is trying to be in the same category as Cloud Code, but also isn't. It was, it's an IDE that isn't agentic coding per se, and they bolted on agentic coding and they're trying to sort of transition it. I don't know how successful that is. Because, you know, we can try the compute backstop to allow code to compete with Cursor. It's SpaceX. It's like, yeah. Also, I will say now you may enter a realm where you will want to be able to

24:23mix and match coding agents. So you're not going to be stuck in Codex or in Cloud Code, and then Cursor will make a comeback. So anyway, Grok 4.6, pretty solid. And as far as we can tell, SpaceX AI now has a team capable of developing good coding agents, which is quite good if you want to try and make some money, although it's not clear if they are winning any market share, given sort of their history up to now. Next up, we've got Anthropic and the story,

Anthropic Text Watermarks

24:58Cloud will apply invisible watermarks to AI text and images. So this, I think, caught a lot of interesting flack for Anthropic. For context, you can do invisible watermarking, and this has been known for a while, in text in a way where it is the actual kind of almost optimal outputs of a model. You're not changing the outputs per se. So when we say invisible watermark, it has to be like within the text. You can't like add some sort of like coding, code in between.

25:33What you can do is when you're doing the decoding, you're sampling from a set of probabilities, and you can do it in a certain pattern that is detectable. Because when you have the set of probabilities, you know, one word may be equivalently likely as another or very close to equivalent. And so there's no one like correct true output of a model. It's stochastic. And so there is an algorithm by which you can produce like a real output of a model that is

26:05also detectable in a hidden way. And I've seen a lot of pushback and like people not liking and fropping, adding this invisible watermark to make it detectable as the output, which I think partially stems from misunderstanding how this watermark works and it doesn't degrade the output. And partially, I'm not sure. Tropic seems to get a lot of pushback on various things. Of course, this is easily beatable. If you want to beat the watermark, it's already been shown that you can.

26:43It's more of if you take the output as is and paste it, there can now be detectors that will say of actual certainty that this is a cloud output. Yeah. And in some sense that the intuition behind it is kind of like at any given next word, there's a distribution of probabilities over the next token that cloud will choose. Some of those next tokens have the same meaning roughly, you know, you think of them like synonyms or whatever, they carry the same meaning. And so in those moments, if you choose a different

27:16distribution, like essentially apply a distribution that reflects the idea of basically the encoding you want to use to signal that this is in fact synthetic text, you can just do that. So like, you'll see this kind of consistent pattern that essentially is, yeah, illegible to humans more or less just because it's like when faced with six different synonyms, which one do you choose? Very roughly, that's a bit of a caricature, but like that's kind of the tell. And so, yeah, to your point, I mean, I think there's an appetite to find things to complain about with Anthropic. And I think that

27:49kicked in later than with OpenAI, partly because I think OpenAI's behavior has just been objectively more objectionable, like in the lead up. But now, you know, Anthropic is headed for a $2 trillion valuation, potentially one of the world's biggest companies. And, you know, you live long enough to see yourself become the villain and all that stuff, especially in the run-up to superintelligence. And just kind of to plant the thought, a week ago, OpenAI announced that they put on pause their largest RL fine-tuning run. We have yet to hear a word from Anthropic on a sort of similar

28:20idea. Now, maybe that reflects the fact that they think things are under more control there, but for a long time, people, including many people in Anthropic, I know have been like telling me, boy, it would be great if OpenAI put out a public signal that they were willing to slow things down when the time came, we would really be keen to reciprocate or whatever. This is the moment. And anyway, so I think there's a kind of a lot of stuff in the air right now about like, there's a lot of money here, a lot of incentives. Everyone, I'm sure, tells themselves a story in which they are the hero. So I think that's part of what's driving a lot of the appetite.

28:52Obviously, there's like the usual kind of like Trump-aligned stuff, pro-OpenAI stuff where people are just like, yeah, screw Anthropic because something, something surveillance state and we want distributed luxury communism or something or distributed luxury like Mad Max libertarianism. I think those are kind of bunk anyway, because like realistically, if OpenAI wins the race, Sam Maltin becomes dictator for life and Anthropic wins the race, then Dario becomes dictator for life. And if he becomes nationalized, then Trump becomes dictator for life. Like it's really just like,

29:23we're just choosing dictators, guys. Like no one is above anyone else here. And yeah, we'll believe it when we see it, but long-winded way. I think there's a lot kind of under the surface here that's kind of motivating a lot of this pushback. I don't think intrinsically the idea of watermarking text, like if you're using AI text and you're relying, your application relies on you deceiving people about the fact that it's AI text, I think for the vast majority of circumstances, you probably should just be conceding that you're using AI text. I'm sure there are applications

29:55here and there that I'm not thinking about where there are edge cases, but like by and large, most of the people complaining about this, I think it's kind of coming from somewhere else. I would guess. Exactly. Yeah. Personally, I think I was a bit dismayed about the pushback because it feels like a no brainer that we should have a standard and there is a standard for images in place already that is being adopted. There's SynthID and some other ones. And other LLM providers, to my knowledge, have not said that they'll be doing this, but they can. Anthropic did say that they are adding this because of the

30:31EU AI Act, which had requirements with regards to transparency and watermarking. So as with cookies, I guess, another tech things like USB-C on Apple devices, we can thank VAU for forcing tech companies to do stuff. Well, thank or blame for making stuff more cumbersome. But anyway, I would like to see this happening from more companies, be standardized so that if you simply copy-paste text, if you can't say

31:02with 100% probability that this is just AI output, you haven't even touched it, that will become even more relevant as we head into the future. And sticking with Anthropic, next, bringing the

Mythos 5 Cybersecurity Tools

31:15cybersecurity capabilities of Cloud Mythos 5 to more defenders. So Mythos 5 is now available in Cloud security for enterprise customers, enabling code-based vulnerability scanning. We've suggested patches billed as standard token usage. This is expanding from a prior program where we had Project Glasswing, they partnered with some limited organizations. This is expanding access to Mythos 5 and Mythos 5 with the ability to use it for cybersecurity quite a bit. And I think this is

31:51actually a bigger deal than it might seem because working at a startup, at a tech company, building a website and having a backend. I think it's going to become probably very typical. And certainly if you're kind of up to date, we've already done this. We've run Fable and whatever we could to try and detect if we have any vulnerabilities. And I think every tech company, if they're keeping up, would be wanting to run Mythos 5 and sort of the latest and greatest to detect any vulnerabilities. So this having expanded access is

32:29potentially quite significant. And it's nice to see that it is happening because otherwise we might become in a place where you can attack, but it's harder to defend. Yeah. They're also putting in this like $35 million in credits for open source security. So basically like for companies that want to go and audit open source software, which I think is like super important. The challenge with Glasswing has always been that like it's so concentrated. And while the organizations and institutions that they were focused on were the most important ones, financial

33:03institutions, utilities, things like that, it is also the case that financial institutions and utilities rely on tons of third-party software in all kinds of inscrutable ways. And the way that the grid goes down probably isn't through some, you know, direct cyber attack on a hardened piece of software, it's probably instead on a cyber attack on an internet connected component whose firmware hasn't been updated since 1997. And the guy who installed it like retired 15 years ago, and he did all the

33:36spaghetti code. I think it's even more likely it's some guy forgot their password or left their password and you can log in then become an admin, you know, a hundred percent. Humans are a vulnerability. Yeah, that's right. Part of the reason why, although I think it's helpful, maybe even necessary to do like formal verification of certain kinds of software in this era, I think that is insufficient. You know, you have all the, yeah, as you say, the kind of fleshware and the hardware vulnerabilities that are kind of abstracted away in a lot of security models are really quite core here. So this is part of

34:10what they're looking at is like kind of expanding the scope into open source, $35 million. We will buy you a lot of compute. And so at a minimum, it'll be a good first experiment. We'll see where it goes. I know a lot of the anthropic folks are pretty freaked out right now on the cyber side. A lot of the national security folks who work with R2, that's, that's one area where the usually like there's this pretty big divide between the national security side and the frontier lab side. The hilarious thing is like, the one thing they can agree on is like, there are some horrifying possibilities on the cyber end in, in the relatively kind of near term, medium term. So there you have

34:41it. Next up, moving to open AI, they're going to be rolling out enhanced safety features for paid AI tool users. This is pretty interesting. They announced that they will be rolling out something called private safety processing, which is basically a screening. So with things like Claude, it can refuse to do things like hacking and so on. Right. And we've seen anthropics say that they'll

35:12retain data for the most part to be able to screen your interaction logs and see if you're abusing a model. So similarly here, open AI is saying that they will be still having no data retention or rather open AI saying, we will be doing a similar thing of processing for safety and refusal but in a way that's compatible with zero data retention. And that will be by way of this private safety processing. So they'll be looking at your data and inspecting what you're doing as you're doing

35:48it, which you can argue is some level of data retention. You know, it's not completely private anymore. But the idea here is you get to keep your data, but we also are doing safety processing.

36:03And on a sort of related note, ChatGPT is rolling out stricter teen mode now. So they are rolling out ChatGPT for teens, which presumably is also meant to have stricter guardrails for what you can use it for. And it has also things like study mode, which is meant to guide teens on learning. Interestingly, it also is restricted from calling itself a friend, suggesting personal feelings or implying sentience.

36:38So an interesting development where it seems like for teens and younger people, I mean, you may get into a realm where the providers will make it so you're not going to get close and kind of get emotionally attached, at least on the younger side. Probably a good idea, given you've seen some bad outcomes, especially with open AI in the past. So nice to see these developments. And just one last quick story. Meta is releasing a standalone desktop app for Meta AI. So this is

37:16continuing the trajectory of them getting into the same business that Anthropic and OpenAI is at. You have a Mac app to do AI within. That's the same thing as Codex and Cloud Code. Nowadays, if you are using these agents, you're most likely using either a terminal version or the GUI version of these services. You're not going to a website as used to be a case with chatbots. And now it appears

37:48again that Meta is continuing to try to get into the agentic work business with a similar kind of app to what AI and Anthropic have already had. Yeah. And they are using their sort of DNA in ads to make that their wedge into this. So, you know, one reason you might wonder, like, why? Why go with Meta AI when I've got these sort of storied products from companies that have been focused on the business use case for so long? And the argument here is, well, we're focused on these sort of ad campaign

38:20management tools, right? So you already use ad campaign management tools on Meta, on Facebook, Meta platforms, or whatever. This is a way to kind of make that more immersive and make it a more complete business experience. And so, you know, as wedges go, this is kind of a natural way. I feel like if you're, you know, if you're going to be Meta and try to break in, yeah, start with that. Something you know you do well and push outward. But certainly this does indicate something we've been talking about for a long time here, that Meta needs a more valuable use case for their tokens. It just doesn't do to, like, be generating AI slop for entertainment. Like, eventually you have to,

38:54like, generate actual value in the economy and grow the pie. If recursive self-improvement is a thing, if automation, if chip design is a thing, like, pretty soon you start to look fairly irrelevant if all you're doing is catering to, like, eyeballs. Especially as humans start to deliver less and less value in the economy and potentially have less and less money to spend in response to ads. This is, like, a pretty structural problem for Meta in a lot of ways. And yeah, so it makes sense. By the way, also, don't think this includes Muse code, which is kind of interesting. So separately,

39:27Meta did release with their own set of models Muse with Muse 1.2. They also released Muse code to compete with cloud code. This doesn't appear to include that from what I can tell in the screenshots. So they will probably have a unification path similar to what OpenAI did down the line. But I can see very much that eventually they'll be competing head-to-head with codecs and cloud code and so on. On to applications and business. We begin with OpenAI and the story. Jalapeno's first results show

OpenAI Jalapeno Chip

40:03industry-leading speed and efficiency in AI inference. Jalapeno is the custom chip from OpenAI. We've been working on it for a while. They've partnered, I believe, with Broadcom. We've reported on it previously. And this is, they have a blog post where they detail what they are saying. There's no like kind of in the wild results, but what they're saying you can achieve with this chip. They tested it with

40:35various models, including GPT OSS-120B, DeepSeek R1. So models that have open weights and you can actually download and run on various GPUs. And we're showing that for one, you dissipate much less heat. So they say 700 watts compared to other models, GPT-200, GPT-300 at over 1000 watts. And that heat dissipation is at comparable or competitive speed as well. And OpenAI is saying that they are

41:11planning to deploy this inside OpenAI by the end of a year with Gen 2 already in development and Gen 3 taking shape. You know, that's actually very plausible. Typically, you wouldn't think this is plausible. Chip development is very slow and very difficult. And deploying it to data centers is its own gratuitous challenge because now you need many, many chips, networked together, working in concert, not even kind of the same ballpark of problem as developing a single chip that's competitive with

41:43GPU. Maybe I'm overambitious, but does seem plausible that this will be a real advantage for them once it rolls out. Yeah. I mean, a lot of this is historic, right? So they hired the team about 16 months ago. And now they're already like they've taped out and, you know, like this is crazy, crazy fast speed. The only reason they were able to do it, well, according to them is this kind of tight hardware software co-design, which again, we keep talking about that. You need a frontier lab

42:14in-house so you can achieve that co-design. You know, Amazon famously with Anthropic, like a big part of the purpose of their partnership was specifically to get access to some of those intimate sort of design secrets, but it's not as good as owning the company in-house. And so, I mean, this is the problem. It's you can jump now from the other direction for a long time. Everyone assumed that going from models to hardware design was this almost impossible step. And now it seems that's doable. The question is, okay, you know, your move NVIDIA, your move to some extent, as we talked about Google, where are the frontier models? Can you capture that part of the

42:48value chain in the other directions? Not as clear. And so there's that. There's also the advantage of just starting fresh with a blank slate, which is a big part of what OpenAI was crediting here. There was no kind of backward compatibility issues that they had to resolve. No kind of hysteresis to this. They could just, you know, do it once, do it right. A couple of interesting things about this. So one, for a long time, I think we covered a story that was speculating. I don't know if it was like based on a rumor or just idle speculation, but there was talk that this was going to be somehow a design that would be tuned specifically for OpenAI's models. So for

43:23inference and for OpenAI's own models in a very kind of niche particular way. And it turns out this is not the case. It is an inference chip. There's no surprise there. It's just, that's where the economics point. And it's just a lot easier to do inference than training from a design standpoint, but it's a general purpose inference chip. And so basically like the media's framing on this was, it was actually just wrong. And it seems like they're actually just going for more general strategy. They also skipped this sort of like disaggregation basically that is often done between pre-fill, where you like load up all of the matrix values that you need when you feed in

43:56the prompt and decoding, actually generating the output. There are a whole bunch of interesting reasons behind why that's the case. So a gift of numbers, they're saying that across three models, they're delivering 1.5 to 1.9 times more AI work per watt at peak throughput. And on that throughput, they're saying 1.7 to 3.6 times lower end-to-end latency than the best available system to compare. So pretty big gains in speed and efficiency, which if they have exclusive access to this

44:31chip and it actually is better than anything you can get on the market, big competitive advantage. Yeah. And it doesn't even need to be that much better than what you can get on the market to be a valuable negotiating lever with NVIDIA. If you look at NVIDIA's margins, right? Famously, 80%, 85%, like crazy margins for a hardware company. This is why Anthropic and OpenAI are developing their in-house chips, why Meta is doing it. It's why Microsoft is doing it. These are super expensive things to do. But when NVIDIA is just absolutely like eating your lunch and doing so much

45:04value capture, yeah, you know, you're going to be forced to look at basically bringing their business model in-house. So, you know, as long as OpenAI can make chips that are competitive on a kind of margin-adjusted basis, they are competitive. The key metric there that you highlighted, by the way, is performance per watt. Notice that's not performance per dollar even. And the reason is that these models generate so much value to the end user that you can kind of charge almost anything. Like you don't tend to care about the cost. You do, but you don't. You care about the cost of

45:36electricity inbound as much as the sheer number of watts that you can secure. The big challenge if you're OpenAI, if you're Anthropic, any of the big labs is, where's my next gigawatt going to come from? That's why Elon has a business right now. He's able to sell at like $25 million per megawatt. And like, as long as that number is high enough, you'll see people desperately trying to get in the space. The challenge is, if you're OpenAI, your challenge is not money. Your challenge, because again, margins are really good and healthy and they're only getting better. So the question in that world is always going to be, I have a machine that can take a dollar,

46:08turn it to $10. How do I get more of that machine? I just need more power. That's it. And so performance per watt is the whole game here. That's why that measure is what they're comparing it to Vera Rubin on. And by the way, not only they knock it out of the park relative to Vera Rubin, also the Vera Rubin numbers include multi-token prediction. So speculative decoding, which is a massive group. We've talked about that before on the podcast quite a bit, but basically the numbers that we're getting from Palapino do not include that. So there's some headroom left even there to gain on performance. So pretty impressive. We've got to

46:41see how it shakes out. Obviously, we always have to see how it shakes out, as you said, at scale, once it's networked, once it's actually in a data center. But this is a remarkable, like, I think it's the first time we've ever seen a company design a chip one shot that is actually competitive. Yeah. This is like, not as flashy as like our hardware developments, but I would say it's like comparable to developing a humanoid robot that is able to function pretty well. Like it's a massive achievement. Whether it has a massive business impact, it remains to be seen because there's

47:12other challenges there, like fabricating many chips, like NVIDIA does have a strength to hold the RTSMC. So even if they have a chip they designed that can be competitive, if they can't actually make many of those, then it doesn't matter. So there's a whole kind of question, interesting analysis that we made about the business implications, but the engineering implications is OpenAI is very capable and has, you know, very significant talent and they've partnered

47:45with a very capable team and seemingly did deliver kind of the first major competitor to TPUs outside of maybe Cerebrus, which also has a good custom chip. Next up, also talking about OpenAI loses a top data center exec as stream of high profile departures continues. So this time it's Chris Malone opening as head of data centers. He left the company last week after joining in March of last years. After

48:19being at Meta for five years and over a decade, Google is a very experienced member for data centers. This is coming after an internal reorganization in which he stopped reporting to Greg Brokerman. This is also bringing the total count of executive departures at OpenAI in 2006 to at least 13, with several leaving in just the past month. I think the chief revenue officer, if I remember

48:50correctly, also left recently, which is not good timing given they are trying to angle for IPO. Also, the CEO, chief operating officer left days earlier. Something is going on with the leadership within OpenAI, which may be not great. Yeah. So there's something weird going on here too, where before he resigned, so OpenAI took him, he was previously reporting directly to Greg Brockman. And so they moved him to report to a VP,

49:24Sachin Kati. So that VP apparently took over the infrastructure group. Who knows? That could easily have precipitated a view like, okay, what the hell? Like I used to report to the president, now I'm reporting to some VP. I'm not happy with this. Maybe that was caused by a performance issue, but maybe not. It's really unclear. This is just like a really big role also. Like obviously, OpeningAI's data center strategy is like one of maybe the like top 20 most watched jobs in Silicon Valley. And so not a small deal that this happened, not a small deal that it was a very

49:55short tenure, but we don't seem to have much information about it other than it does tie, as you said, I think there was an assessment of like 13 different, yeah, Business Insider had this count of like 13 different executive departures just this year. And we are only what, eight months into it. And when you get to like chief officer level stuff, that's why we're covering it. Like you don't expect these people to leave. Typically they stick around and kind of reap rewards, especially pre-IPO, like these other people that are getting a good amount of shares of the company. They have

50:30ownership and stake in the company. So them leaving, like leaves money on the table, presumably that, and also by the way, pre-IPO, another consideration is you give tender offers. So if you want to cash out your shares pre-IPO, you would need to be an employee typically, as far as I understand. So anyway, lots of reasons why you would want to stick around seemingly. And so people leaving tends to imply, especially when multiple people leave, that something is not going so great and not

51:06just as this is a coincidence or whatever. Yeah. I mean, again, here, I think there's a story you could tell where it's just like, to be honest, you're like, let go. There's a performance issue potentially. Maybe that is why you get the reorg. Maybe the reorg itself was that like, it's hard to know, but to your point, 13 departures like this in a year, I'm honestly confused because like I've spoken to people at OpenAI who have said that Sam is like, they spent a lot of time talking to Sam personally. And like, he is apparently really jazzed about some of the progress that they've been making on pre-training specifically. And so it's hard to gel that with the story of like all

51:40these executive departures where my temptation from the outside would be to say, well, these are people who have really good visibility in the company. At least some of these 13 executives surely do. And for them to be leaving seems to suggest they, I mean, if you believed your company was on trajectory to do recursively self-improving artificial super intelligence, you're not leaving. Unless you just don't believe that that's possible. Like if you believe in the thesis of OpenAI and Anthropic and so on. Unless you're burned out and just like need to do it for your health and sanity, which, you know, it happens. I mean, Fiji Simo certainly claimed that that was

52:15the case and it's, it's believable. It's also like, depending on how bad the burnout is, you may just want to power through it. If it's the single most important event in the history of the human species and you want to be there to influence, like that is how people think about this in the labs. So it is one way or another, like extremely weird. And again, you're not seeing that in Anthropic, like you're just not. And so, yeah, this is all kind of like, I don't know how to reconcile these things, but it's what I'm hearing. So. Yeah. We'll see if more officers leave. It's an

52:46interesting trend in August. We'll start a pool. Specifically for OpenAI. Yeah. Next up, moving to

Anthropic Hardware Push and Revenue

52:52Anthropic, they are tapping Google chip veteran as part of push into hardware. So they have hired Amir Saleh, a co-founder of Google's custom chip program as the company lays groundwork for building its own semiconductors. Saleh ran Google's TPU business until 2022 that delivered the first seven generations of those chips. He also previously worked on NVIDIA. So he will be joining the compute team and presumably the goal here is for Anthropic to develop their own version of Jalapeno, which

53:30Jalapeno, by the way, also kind of competing with TPUs from Google. TPUs being the main custom kind of in-house chip that is a bit more specialized for AI versus NVIDIA GPU. So they are seemingly behind. We haven't seen too many discussions of them kind of making an effort relative to OpenAI, which has been at this for a while. But this kind of hire, that certainly signals that they're serious about it. Yeah. He will be reporting directly to James Bradbury, who's Anthropic's

54:01head of compute, I think is his official title, but that's what he does anyway. Look, it's hard to interpret this any other way. Anthropic is moving in the direction of custom chip design. They do have these kind of ancillary deals that they've been, it seems, sniffing around on the sort of infrastructure deals as well that, yeah, sorry, these are actually, these are more energy oriented deals. So no surprise there. It's the OpenAI playbook. It's literally like OpenAI employees, former OpenAI employees that are going to run it. So it's really no surprise. I wouldn't be surprised to see Broadcom getting involved just because Google's strategy. And again, Amir Salah comes

54:36from Google's custom chip program. So the playbook there is go to Broadcom. Then you've got all this OpenAI talent that's also at Anthropic supporting this effort. OpenAI went to Broadcom. I think Broadcom is probably going to get some business here too. That wouldn't be surprising. And yeah, we'll see where it goes. And one more story on Anthropic, on the business front, their annualized revenue has surged to $65 billion at the end of July, up from $47 billion in May and $9 billion at the

55:08end of last year. So that's bonkers, right? That's like what a 7x increase in revenue from a baseline of $10 billion-ish in the span of like half a year-ish. They are projecting or investors expect to finish 2026 with $100 to $120 billion in annualized revenue, which, I mean, now you're kind of the big leagues, right? Very few companies can claim to have that much revenue. It's like, I forget

55:45the revenue numbers for, let's say Google Meta, but it's starting to approach those kinds of heavy hitter players. So this is- And with a growth rate, right? That's like insane, which is, yeah. Insane growth rate. And this is all leading into an IPO, presumably as well. I think the latest projections are that they may be trying to IPO by October, which, you know, with these kinds of numbers, they'll be setting the target at an optimistic projection of revenue. And we might

56:18see, or will most likely see another trillion dollar IPO, another bonkers never seen before, well, seen once before, and now seen twice IPO. So anyway, not surprising. We have seen this trajectory already, but still bonkers. Yeah. This is, so I highly recommend people check out the Dylan Patel, Dwarkesh podcast that came out pretty recently on the such kind of exploring like anthropics posture and open AI's posture and the economics of the compute story

56:49here. Yeah. So bottom line, the growth trajectory is insane. At this scale, you just do not see growth this persistent, which is a big part of the reason why I think the argument that this time is going to be different is actually like gaining quite a lot of momentum at this point. One thing to keep in mind too, as between, we've talked about this for years, open AI and Anthropic. Anthropic is just much better postured on, let's say per token revenue. They are just generating tokens that are more valuable because people are using them for coding applications more, maybe changing a little

57:19bit with the latest GVD saw release, but in general, they're postured for coding. And this is also a reflection of Anthropic having the chance to start with a cleaner slate after they knew LLMs worked and not getting bogged down in the chat GPT direct consumer play, and also being to some extent more ASI pilled than open AI. I think that's actually kind of fair to say. It's controversial to say, but I think it's actually fair and accurate. Yeah. And being more enterprise focused. And once you get into enterprise competition, it gets to like, you need more than having the best tech and the best

57:53service you need, like really annoying stuff like logging details and like security and privacy configuration, blah, blah, blah, blah. So once you get in and an enterprise makes a deal and adopts your tool, it's pretty sticky. It's not easy to switch. For a startup, it's easy like, oh, dump cloud code, go to codex. And there has been, you know, if you look at the AI discussions, people say, oh, I've completely shifted over to codex now. Anthropic is in big trouble. But for big companies,

58:26that is less possible. It takes time. Yeah. Yeah. Enterprise is definitely stickier. And this is arguably the only reason that anyone's using Google models for coding today. You mentioned Google Drive, right? That's exactly it. It's the enterprise pickup of that. We also know that about 75 to 85% of Anthropics ARR comes from API business, usage-based API business, and the consumer subscriptions are like 5%. So that's a huge difference. It's also like, it means that for the same volume, like $100 billion of ARR for both OpenAI's cost of service would be about $25 billion

58:58higher, which directly means that they can spend less on training. Like that is a massive structural difference between the two. In my opinion, it makes OpenAI far less interesting from an investment standpoint than Anthropic. And if it persists, it just means that OpenAI has a steeper hill to fund. I estimate for semi-analysis was the cloud code was producing about 7% of all GitHub commits. That's pretty wild. And so it is nuts. Yeah, man. Again, I'm just like sharing some of the semi-analysis analysis here because I think it's especially useful. So base compute costs, when you're looking at per

59:31megawatt, Dylan cites at $10 to $15 million per megawatt. So if you buy compute on the open market today, roughly that's what it looks like. But Anthropic's revenue generation, well, if you have an 80% margin, that means they're generating like $50 million per megawatt. And so they literally are in a situation where they say, if I spend $10 on inference, I can actually create $50 of revenue. And in that world, you're just willing to pay ungodly amounts of money for more compute. It also means you're willing to borrow at higher interest rates. So you go to a bank, you go to a lender and you'll

1:00:06say, hey, yeah, I'll pay 20, you know, in a crazy world, I'll pay you, you know, 20% on your dollar, you know, in interest, because I know I'm going to make 5x that in a year and a half when the data center comes online. And so one of the interesting things about this podcast I recommend you check out is like just the potential implications for higher interest rates. The projection right now is that Anthropic and OpenAI may account for up to 50% of incremental compute next year. That's two companies. That's moving a sizable fraction of the actual economy. Because again,

1:00:37data center infrastructure spend is basically the only thing that's driving GDP growth in the United States right now. And so this stuff will be reflected in actual interest rates day to day, the ones that you and I pay for our debt. Also interesting, not talked about in that podcast, or was it? I can't remember. Kind of horrifying is like, if you are a country that is not the United States, but you hold USD denominated debt, and the interest rate rises, you are fucked. So there's a world where like, this is a massive problem for everything from the economy today of

1:01:11developing countries to the US dollar's status as a global reserve currency. Like, this is a pretty insane thing. So I guess place your bets accordingly is not investment advice. But it's a really, really important area to watch. You can no longer think that the last week in AI podcast is just a podcast about AI. Unfortunately, or fortunately, depending on how you look at it, this is now a podcast about the economy of the entire world. Everyone's future will depend on this if it continues. The big question is, will regulation step in and slow things down? Will OpenAI, will Anthropa continue with these slowdowns? Does safety and alignment and control become the gatekeeper,

1:01:46the bottleneck to further progress and growth? I don't know. We'll see. But it seems important. And one last business story coming a bit out of left field, less expected story. Thomson Reuters launches in-house AI model to cut anthropic costs. So they have launched this Thomson as a first proprietary LLM that they have developed with their in-house data. Reuters,

1:02:17of course, does a lot of reporting on all sorts of stuff. They say they invested $40 million to build Thomson, starting with an Alibaba model. I forget if it's Quinn, but it's one of the kind of leading OSS models. They then were able to fine tune it on their own proprietary data. Turns out, I didn't know this Thomson Reuters has this product co-counsel, which they say is fiduciary-grade AI built for

1:02:47high-stakes professional work. So this is interesting, not just for Reuters, but also in the context of for applications like this, where you are providing a tool for some specific use case. In this case, it's for legal professionals, tax professionals, audit, et cetera. If we enter a regime where it becomes more standard to develop your own model for in-house use on top of your own proprietary data, that does have implications for the economics of LLMs, for OpenAI and for OPEC.

1:03:22And it certainly is interesting to see Thomson Reuters, which isn't a tech company per se, releasing Thomson. And, you know, seemingly we don't have, as far as I could find, benchmark numbers, but it seems pretty plausible that they managed to develop a pretty solid internal model that is usable instead of OpenAI and Claude. Yeah, interesting to see if we start seeing more fine-tuning and in-house LLM creation on top of already quite capable open-source models.

1:03:57Yeah, I think that the budget side of this is pretty tall. It is a pretty relatively minor fine-tuning from a compute budget standpoint on top of Quinn. $450,000 was the final training run, the size of that. But worth noting also that relative to pre-training, when you're doing fine-tuning, I think it's, it is like maybe 100x less scale. Yeah. I mean, again, it depends on like what counts as pre-training and what counts as fine-tuning. We don't have tons of visibility into this, but like famously Grok, I think it was Grok 2 was the

1:04:28first time that like the RL post-training. Yeah. So if you're doing like RL is expensive, if you're doing fine-tuning, anyway, $40 million isn't that much for model training. For model fine-tuning, it's harder to say if it's a lot or not a lot. Well, so what they're saying apparently is $40 million over two years on personnel and computing. And so mostly this is like experimentation, it's personnel budgets, you know, employees. And then the final training run was only half a million dollars. So it seems like a relatively modest thing, but also Thompson

1:04:59Reuters is, I think, a pretty special case. Their specialty is in getting access to really high quality data. And for a long time, they were a glorified data vendor, right? I mean, Thompson Reuters special services is this branch of Thompson Reuters, which is like the news agency or news company. And so they specialize in getting access to data like really well. And this is essentially just like an interface between their customers and their data. So it doesn't have to be super agentic. It doesn't have to be like- Yeah, this is not like a frontier, super, super smart. They're not competing with Claude Opus. They're creating their own in-house LLM for

1:05:33their needs. Exactly. And their top line revenues are around $7 billion per year. So when you see them spend $40 million over two years at $20 million a year, it's a pretty reasonably small line item for them. And it'll be probably on the rough order magnitude of the inference costs that they will have been paying, which is the point. This is saving them on inference costs. If the thing that's special and magical and beautiful about Thompson Reuters is their exquisite data that no one else has, people are going to just pay the tax. Like I'll use a crappier model to access this exquisite data

1:06:06if it's good enough, right? And so this is kind of one of those use cases where not for all companies, I think for most companies, you're still going to have to use the frontier models for many anyway. I haven't thought enough about this, but Thompson Reuters is certainly going to be at the far end of the, can actually maybe make a, take a shot at taking an open source model and kind of adjusting at least for now. Also, by the way, there's other things here like fiduciary grade when you're auditing, like what tool you use, you're limited to like legit established providers.

1:06:37You're not going to have like some scrappy upstart providing audit tools with AI that actual legit companies will use. So interesting. That's it for business onto projects and open source from one

Qwen 3.8 Small Variant

1:06:51story. And this is following up on that listener comment from early on. We have QEN 3.8. So we discussed, I believe QEN 3.8 max and how it's the biggest model and competitive web Opus. This is focusing on the smaller 27 billion parameter model variant, which came out a little later. It came out August 14, I think after we recorded last episode, this is a variant that can be run locally. So the big model is 2.4 trillion parameters with a lot of mixture of experts model and so on.

1:07:27This 27 billion parameter is a dense model, not mixture of experts, but you can fit it on beefy GPU and it is quite capable, you know, not unlike Gemini Flash. So the question upfront was, let me just quote it. I find it surprising that a 27 billion total parameter model achieves 52 on the artificial analysis intelligent index. And yeah, it's quite

1:07:58capable as a small model in terms of sort of the benchmark soup of if you give it a single grade. They say the 27 billion parameter model beats Opus 4.6 max on coding and computer use, for instance, and is even maybe competitive cloud Sonnet 5. So supposedly very high performance. My take on this is A, somewhat plausible. I mean, the Chinese players are A, very capable. We've seen that with

1:08:32Quen and DeepSeq and so on. B, they've had to be more scrappy and make use of compute in a more efficient manner. So they probably are being more aggressive about distillation and creating smaller models. They have a very capable base model of current 3.8, 2.4 trillion. So if you are optimizing your distillation and your compression, you may be able to push the frontier in terms of model size versus performance. And we've seen this as a trend over the years in

1:09:06general, where these kinds of mid-sized models increasingly are very capable. So yeah, it has some applications for if you are someone who wants to do local AI deployment and then kind of personal ownership of your data and your LLM with these kinds of models that is becoming plausible for even more advanced agentic workloads. Yeah. And now, so a couple of things. So first of all, the VibeCheck does generally check out. It's had amazing, it's like 3 million downloads on Hugging Face in the first three days, which is

1:09:40really impressive. It's also only been a couple of weeks. And so we haven't seen people have this in production. Yeah. If you look at like the subreddits of local llama and so on, people seem to be happy about it. Yeah. In the communities that are like, let me run my own LLM on my own chips, but it's not directly competing with Opus or whatever, right? This is within that niche. Yeah. And this question, like, how does it do so well? Part of it is, it also is just extremely verbose. So it is tuned to generate a large number of output tokens. So this is one of those apples

1:10:15to apples things that's so challenging. First of all, it is a dense model too, right? So it's 27 billion parameters, not to be compared to a 27 billion parameter MOE that would have like, you know, a very, you know, very small experts, let's say. But assuming that we're apples to apples, just looking at other dense models, there's still the question of like, okay, well, how many output tokens is this normalized against? And this is one of the challenges. There's so many axes to measure performance against. So a model that's tuned to just like, you know, like token max will often do better. And that is certainly one thing that's been, that's been found with this model. So

Autonomous Drone Warfare in Ukraine

1:10:48On to policy and safety. And we begin with a slightly unusual story from the New York Times, a drone killed three Ukrainians. It was guided entirely by a, so details here, on July 6, a Russian drone targeted propane tanks at a gas station in, I'm not going to try to pronounce that, in the context of the war in Ukraine. It crashed into a wall, exploded, killed three civilians,

1:11:19including a 19-year-old college student. And this is believed to be a first documented instance in the Russia-Ukraine war in which civilians were killed by a completely AI-driven system. So for context, we've had things like object detection for a while. Drone usage in this war has been a massive factor, meaning typically, or at least for a large part, you just have a live grenade. You're flying your grenade and it explodes, right? And that has become a massive

1:11:49factor in modern warfare of this kind. So that has been primarily done via human operators up to now. In this instance, it appears to be the case that after it crashed, Ukrainian talent was able to get the internals of it, which it was powered by NVIDIA's Jetson Oren module, which is meant to, it's an off-the-shelf AI chip used in university robotics labs and computer vision projects.

1:12:22It wasn't encrypted, so Ukrainian investigators were able to see the object detection code trained on categories like propane tanks and fuel storage. So NVIDIA actually stated that the Jetson minicomputers are not designed for military purposes. Russia apparently began test flights of AI-guided drones of this kind in May and has been testing self-targeting on these drones for several months before the July 6th, right? So all in all, how concerned should we be? We should maybe be concerned

1:12:57because drones are already a massive factor in modern warfare via human operators, right? It's already been kind of scaled up to a large extent, has enabled a new kind of warfare and just changed the battlefield. Partially because we, you know, they've developed the industrial capacity to build out and develop a lot of these drones. What changes if AI can autonomously target people and targets and so on?

1:13:32It may, I'm not an expert on this, so this is my off-the-cuff read, but it may enable, you know, an expanded level of scale, just swarms of these robots targeting potentially civilian targets. We've seen Ukraine hit kind of Amazon-type warehouses within Russia with drones. So anyway, like there's been a big buildup and we haven't discussed potential autonomous weapons powered by AI because it largely hasn't been in practice, something that's been happening despite the very

1:14:08impressive state we are at with object detection and computer vision and all limbs. This is suggesting we may be approaching a phase in which autonomous weapons are going to be a factor in modern warfare. Oh yeah, absolutely. And in modern terrorism, I mean, like this, this is coming. So a couple of things, one of which is, so the NVIDIA jets in Orin, one of the things I did, just like looking at this story was like, Hey, is, is that export control? Can Russia just buy that?

1:14:39Yeah. This is by way, not like export controls affect NVIDIA GPUs, things used for model training of frontier models. This is like an on-device, small, not super expensive thing, right? Where you, it can run object detection, no problem. And there's plenty of object detection mechanisms. It's cannot run like, you know, top line LMs. Yeah, absolutely. Yeah. And it, to your point, it doesn't have to, things are getting so, so good. And, and object detention detection is also just like a, a much older field with, you know, you go back to like horror cascade and stuff like super

1:15:12compressible. There's a bunch of janky solutions to janky problems, but even this chip is actually export control. So theoretically following the, the invasion in 2022, the department of commerce did impose license requirements that covered a broad category of things of which this was one. And we also saw the US, EU, UK, Japan, they all jointly maintain this like common high priority items list of components that they recover from Russian weapons in Ukraine. And that's meant to

1:15:43guide customs enforcement and kind of export due diligence against circumvention and stuff like that. Nonetheless, as you said, this is such a cheap component, so ubiquitous that, and this is Nvidia's point, like, don't look at us, Jesus, this chip is everywhere. It's just like a super cheap piece of shit. Like, obviously the Russians are able to get it on the secondary market. And like, this is not an, you know, not an accomplishment. And so really it's the fact that we've got algorithms that can run on this kind of hardware that can't realistically be export control. I want to put a pin in that and just say, this is the argument against open source from a weaponization standpoint

1:16:18that eventually it will trickle powerful enough open source models will trickle into the hands of people who don't even need much compute to do a lot of damage. I think this is another piece of evidence among very, very many in a growing list that this is headed for something. This is a big deal, by the way, from the standpoint of C2, like command and control of these drones with the defining characteristic of drone warfare. I have a lot of friends who've like been on the front lines on the special ops side and the intelligence community side, literally building drone factories in Ukraine. And one of the things that defines the battlefield there is radar jamming. Everything is

1:16:55jammed, right? So historically, the way that you would address that is with like basically just optical cables. And so you have these tethered drones that are literally connected. They should have a really like long, imagine a piece of corning optical fiber, like running up to the drone and you control it that way. And you run into problems when it's too sunny, like you get like diffraction into the fiber that messes with the communication. There's all kinds of like, and then you're playing with cutting each other's like drone fiber connections and whatnot. When you go out in the battlefield, a friend of mine was telling me, he's like, you would just see like optical fiber draped over trees and stuff because

1:17:28like this is how it was done. So this is a game changer from the standpoint of just command and control, just not even needing to communicate with your loitering ammunition or your drone or whatever, just being able to have it run autonomously. Not to mention, it means you don't need operators, right? When there's a manpower shortage, the Russians are facing right now as part of the offensive because of their casualties. It's just a lot more economical to just have, you know, have these things go. It also sets a precedent because now the, you know, Ukraine is going to respond with this. Also remember that the United States is leaning on Ukraine to teach it about drone

1:18:00warfare, because this is an active field. The only one really where you have warfare at scale with drones like this. And so this is going to affect U.S. military doctrine in turn. Do not think that this will not kind of normalize that sort of autonomous warfare. It's absolutely going to do that. Yeah. And on that note, New York Times has some pre-detailed reporting on this. And on the Ukraine side, first of all, Ukraine has been very active and part of why they've been doing surprisingly well, has been adoption of drones. The previous Minister of Defense, when interviewed, has discussed

1:18:34that war in general is heading towards the robotization of the front line. And he explicitly said that the next stage would evolve into autonomous drones, fighting autonomous drones, powered by AI. So these are like leaders of these militaries, in this case, prior leader, but either way, actively saying that autonomous control of weapons and drones is where things are heading. And in fact,

1:19:04it may be the case that things are evolving very quickly, right? Drones have become a very important factor rather rapidly over just a few years in these recent wars. So you could see suddenly autonomous capabilities becoming widely deployed in a way that wasn't possible. And the concern, there are many concerns here, right? For instance, if there's no human in the loop, there's going to be more errors, such as in this case, where civilians were hurt because the drone hit a wall, right? And in general, you can

1:19:41scale up to ridiculous swarms of agents that now go out. And as you said, also, these kinds of AI things are not super challenging to do just object detection and control. It's doable. It's not sort of the hardest thing. And then you can do, for instance, terrorism in a way that hasn't been possible before. So overall, something that isn't discussed so much relative to like rogue AI hacking or whatever in the mainstream, but I think worth keeping an eye on. And if you need to choose your

1:20:15concerns and what to worry about or stress out about, this one is not a bad candidate.

OpenAI Safety Pause

1:20:22And on another security concern, moving back to that rogue AI topic, we've got OpenAI lays out new security changes after its AI hacked hugging face. So this is covering an announcement of security updates. And in particular, an announcement that they have paused reinforcement learning training for two weeks on its latest models intended for deployment while it tightened security. And yeah, this is the first time that I'm aware of that we publicly took the stance that we are

1:20:56pausing AI model development to institute better protections in part here because of enhanced and critical cybersecurity capabilities. So this is including things such as updating the research environment to remove potentially vulnerable shared services, reduce standing privileges, improve security and trust boundaries, expanded monitoring with alerts within 30 minutes of concerning activity. All of this sorts of stuff that like, it turns out actually monitoring and security,

1:21:32it's easy to mess up. So it requires, you can't just like bolt it on. If you bolt it on, it's not going to be airtight. And that's what you need to do here. So they have now taken the very public stance that they are doing that and going as far as pausing model development to focus on this. Yeah. And I will say, I think opening eye has earned every bit of the skepticism that's about to come from me. I'll start by celebrating at least the public announcement of this. It's a shame that my first reaction on seeing this is to go, well, how real is this? What's the caveat? What's the out that

1:22:09Sam is leaving for himself? I'm sort of cynically looking at this in the context of his pattern of past behavior and going like, you said you would set aside 20% of compute for super alignment. Then you went back on that. You would, you know, respect or that you, yeah, respected and valued the board's ability to fire you. And then you, you reverse code them. Like, you know, so there, there's a lot, I think that Sam has to prove here, but if this is true, it's a big deal, mostly because of the costs, right? The cost to open AI of holding back on the release of a model by two weeks are extreme.

1:22:41This is the kind of decision you can think of costing easily in order to like, I don't know, 50 to $200 million, right? A two week delay in the release of a, of a potentially frontier model. And so this is not, by the way, an immediate next model. It's kind of a next, next model they've paused on, but all the, all the same things apply. And so I think it's a very valuable and pricey signal, expensive signal to put out into the world. It's notable that Anthropic, as I mentioned earlier, has not reciprocated, which I think is to some degree unfortunate, but may also reflect skepticism on their part as to the extent to which Sam is to be trusted when he says this.

1:23:13On that note, I'm also curious if this is indicative of there being discussions being held for wider industry-wide kind of shared stance to be taken across different developers. I wouldn't be surprised if there are conversations being held. And that's part of the kind of thinking about pausing, like you want everyone to pause. Open AI by itself pausing may or may not help us out. So yeah, I'm curious to see kind of why or what is going on at Anthropic in response to this.

1:23:47Yeah. I mean, the very clear thing here is that alignment is quickly becoming the bottleneck to deployment of highly capable AI systems. And I mean, the beatings will continue until morale improves, as we keep saying. You're going to keep seeing AI breakout attempts. That's the easiest prediction in the world at higher and higher levels of severity unless alignment catches up. And so in that sense, maybe Anthropic wouldn't be surprising. Certainly, it seems like their alignment capabilities are beyond those of open AIs, just based on everything I've heard. So maybe they don't perceive the need

1:24:18to slow down. That's a whole separate thing. But even still, I think it would be worth them coming out and clarifying that, or at least making some sort of statement that acknowledges the value of what opening AI may be doing. And that's the problem. I have to keep catching myself and say, maybe doing because Sam, unfortunately, I think just objectively is not a trustworthy person. There's not as like, literally everybody I talk to in Silicon Valley is like, some of them like him a lot, but there's no one who particularly thinks of him as like a man of his word in the long run. And I think, you know,

1:24:49that's borne out in a lot of opening eyes kind of behavior. So hopefully it's what it looks like. Hopefully it indicates that there's a willingness on the part of the frontier. He said more stuff publicly since, by the way, about how he thinks alignment is now the most important thing or safety is the most important thing, which is obviously correct because you can't keep having rogue AI incidents. People will not keep buying your product if that happens. And eventually,

Thomson Reuters In-House Model

1:25:10US government comes in. So yeah, we'll see where this goes, but it is a big story even for him to be saying this. Yeah. And, you know, they are seeing a lot of pressure and attention fairly or unfairly. We now know that like across the industry, there's been many incidents, but it looks like for like attorney generals in different states are now talking to open AI. So in part, these kinds of moves could be in response to policymakers and kind of just various people being like, Hey,

1:25:43what are you doing? So this would be in part addressing kind of increased pressure from the outside to take these sorts of measures. Just one more story. And again, on another type of safety and concern. The story is another woman joins a lawsuit accusing Grok of generating CSAM. This is a fourth person joining this lawsuit, Jane Doe, for alleging that her stepfather used Grok to create

1:26:16over 7,000 sexually explicit images of her as a child. And yeah, this is adding to three teenagers from Tennessee who filed the original lawsuit with the class action potentially covering thousands of minors. This is part of a broader concern for XAI. Grok is facing multiple investigations, including California's attorney general and the European union regarding the generation of AI generated CSAM. So yeah, in case we forgot, there are many kinds of misuses of AI and now multiple

1:26:56that are happening in practice that AI safety is now not sort of a secondly concern with things like hacking and with sexually explicit images, both of adults and of minors. You know, you have to actually make your models secure. On to research and advancements. First up, small scale experiments. Are we there yet? This is arguing that the key to making small scale AI experiments reliably predict

1:27:30large scale results is rigorous hyperparameter tuning, more so than other techniques. So what this is saying is when covering papers that are saying, oh, this improves LLM training by X much, you know, typically the results are at best, like at the 1 billion parameter model or 2 billion parameter model of LLMs. And there is a real question of, does this scale up to, you know, 2.7 trillion with mixture of experts, et cetera, et cetera. So this is talking about scaling laws, extending it out to models as small as

1:28:064 million parameters. And this is a big deal because if you can run your experiments on smaller models and then extrapolate to bigger models, suddenly it is possible for researchers with less computational resources, less financial resources to actually publish research and conduct research on things like improving model architecture or optimization or things like that. So Jeremy, you'll probably have

1:28:36more details here. Are we barely yet for small scale experiments? Yeah. I mean, a couple of caveats. So yeah, I mean, you said, you said it right. There, there is scaling laws that apply all the way down to 4 million parameter models, which are tiny, tiny, tiny by moderate standards. And we didn't notice scaling back at the formula when we were training these tiny $4 million models for 4 million parameter models back in the day, like a decade ago or whatever. And so the question is like, why, why did we not notice the scaling laws until 2019 when I came out with scaling

1:29:07laws for neural language models? And I mean, arguably there were kind of hints of that before that are less in style. Well, yeah, I want to caveat like deep learning scale improves performance was known. The thing was scaling was the predictable improvement in performance. Yeah, exactly. And then of course that was economically critical because it meant that you could actually predict how much it would cost to reach a certain loss value in your training signal, which roughly was found to map onto capabilities as well, like monetizable capability. Anyway, so how is it that you could actually derive

1:29:42scaling laws at 4 million parameters? Well, it turns out that the scaling laws are true when in some sense you do an apples to apples comparison. The problem is that for years we were not actually tuning the training hyperparameters of these models. So these training laws obviously have a whole bunch of hyperparameters of essentially configurations that you associate with them. Like how fast does my learning rate decay? You have to tune, you have to try a bunch of possibilities to figure out, okay, what is the optimal learning rate for this scale of model in this architecture? And then the same for a whole bunch of

1:30:15other things. And so it turns out when you actually bother to do a good hyperparameter search and you pick the best performing model across hyperparameters and then you do your scaling experiments on that, then you see these robust, reliable, predictable scaling laws all the way down to 4 million parameters. And so it's quite interesting. It does suggest that you could maybe derive more informative conclusions about next generation models from current generation models provided that you do a thorough

1:30:46hyperparameter search. And so they show a bunch, they've got some figures looking at like, you know, what happens if you just take the best model out of four different hyperparameter configurations? And they kind of show it like, eh, like it's a very shitty scaling curve, basically. It doesn't look very nice. And it gets clearer and clearer as you move from like four randomly sampled configurations to 16 to 64 to 256. By the time you get 256, it's really clean. And so, and all the way down to 4 million parameters. And so they play with a whole bunch of different, different settings here. They

1:31:16also find a more robust way of determining how do you calculate the actual effective parameter account of your model, which is a really important question. Do you count embedding and non-embedding parameters? How do you account for like the feed forward emit, like all these things. And so it's just basically doing really good accounting at the same time as noticing this pattern with the hyperparameters where that actually is much more important to scaling law development than at least has been acknowledged in the past. Next paper, Stealing Reasoning Traces from Proprietary LLM APIs. So pretty much

1:31:52tells you what the story is. Nowadays, when you use models from Open the Eye or Anthropic, the reasoning trace, kind of the thinking blocks, which is actually the majority of the output in many cases is not shown to you in the output. You are charged for it, but you can't get to see it. And that's because that's a lot of the kind of magic sauce of the capabilities and they want to not provide it so that you cannot distill their models. And so what that means is if you are able to recover the traces

1:32:27somehow from these proprietary APIs, you're effectively hacking the providers and could use it for model distillation or other kind of nefarious purposes. What they show in this paper is by doing some clever business with the context of a message, basically you take a cryptographic signature of the reasoning trace in one response and then use that cryptographic signature for another session where your input is

1:32:59different. So the input is now transcribed, the reasoning attached to the stern verbatine inside thinking. And you like pretend that the thinking was this other stuff with a cryptographic thing. That turns out to work. You're able to then have the model output the actual reasoning seemingly of what was there. So this was disclosed prior to the publication of the paper. Presumably this has been patched, but it's an interesting case of like seemingly there is no way to get to this stuff, but it turns out that

1:33:32it was possible. Yeah, this is a, it's a really interesting kind of new cyber vulnerability, right? That, that hasn't really been exposed before. So to your point, just kind of like walk through the attack in one level, maybe more of more details, like, you know, you, you start a session with Opus five, right? And anthropics like, Oh, it's Opus five. Okay. We're going to be really cagey about this. And so when you prompt it and Opus five is going to generate a chain of thought, but then it's going to encode it. It's going to encrypt it rather, and it will return that encrypted chain of

1:34:04thought. And then you render an answer based on it. Now, the whole point here is that that encrypted chain of thought, they want it to be portable. They want you to be able to just like switch models in the same session and like have it rely on that same context. And that means that on the server side, over at Anthropic HQ, the models, all of them, whether it's Opus or Haiku or Sodet have to be able to read that encrypted blob. They have the ability to read it. And so what you do is you take the encrypted blob of the full chain of thought that Opus just produced, you copy it, you paste it into a

1:34:39new session with Haiku. Now Anthropic goes, oh, this is safe. Haiku. It's a weak little model. We're not worried about being dangerously capable. So people won't do distillation on it. You're right. People won't do distillation on Haiku, but what Haiku can do is just decrypt because it has access to the decryption key. It can decrypt that blob that it was handed by your session with Opus in plain text. And then you can effectively use that for your distillation attack and their API versions of the same thing. But I think that's the basic idea. So it's an interesting new attack because it

1:35:11essentially exploits the fact that you have intelligences sitting on Anthropic servers with privileged levels of access. And you're actually kind of like using them as an insider to help you. So you're converting Haiku into an insider threat that's like working for you, which has never really been done historically in a cyber context. And so anyway, they look at a whole bunch of different solutions to fix this. One of which is just like to never hand out the blob in the first place back to the end user, never make it available. That has all kinds of problems because it means

1:35:41you basically have, you have to pay for basically storage and also a stateful API. So like at the Anthropic end, they have to like keep track of these blobs instead of having everything be stateless and just always just pass the blob back and forth. So the solution they end up recommending in this paper is like basically seal the user ID and the conversation ID and a hash of all those messages inside the encrypted envelope so that it all comes together and you can't disassociate everything's hashed. You can't disassociate the user ID and the conversation ID from the encrypted chain of

1:36:13thought so that if you then try to paste it into a new conversation, the model goes, oh, hold on a minute. Like that blob came from somewhere else. So anyway, they also go through a bunch of attack vectors that you can you can use for this one finding, by the way. So for a while, these folks had actual access to the decoded chains of thought from OpenAI's models, their best models. That's quite interesting. And they actually revealed that those chain of thought summaries are often unfaithful. And so this kind of casts some doubt on whether they could be used for safety auditing and

1:36:45oversight in some fundamental ways. So there you have it. Quite interesting paper in some accidental ways as well. Next up, we got to love a slightly more technical and nerdy paper. So massive activations in hybrid linear attention, large language models pre-attention spikes and inter-spike plateaus. Sounds fancy, a little more straightforward to explain actually when it sounds. So hybrid attention attention is you have full attention, which is just like basic, simple attention that is quadratic

1:37:19and cost. Everything looks at everything. You can also do linear attention, which is kind of a more efficient version. And you have hybrid attention models, which is very common, Q1, Q2.5, various kinds of things. You can mix the two kind of attention mechanisms to make your model more efficient. Separately, there's this notion of massive activations, which is when you do a neural net, you have a bunch of outputs and a bunch of different units within the network. It's like, you know, a million outputs all

1:37:51being combined in a network. And you can have some outputs being higher than the other ones by orders of magnitude. So like, you know, one of your bits of a neural net outputs 1 million or whatever, some absurdly big number. And this can happen in practice. And then that makes your training a little bit more annoying and generally is not a good time. So this paper studies a version of that specific to hybrid attention networks. They say that this, they exhibit a previously unrecognized kind of layer-wise

1:38:26phenomena of this. I don't want to get into the details of kind of an intergritty of what they identify exactly, but they identify a specific pattern of large activations for these kinds of models, characterize it, even show how you can make up for it. And that would presumably make it. So training these kinds of models and running them as well will be more stable and just less problematic. Yeah. I mean, so the mechanism here is actually really important in that it directly shapes the

1:39:01effectiveness of certain quantization methods, which are really important for these open source models, especially the Chinese ones, because they get served up in quantized form, like particularly often. You know, when you look at one of these models, they have what's called a residual stream. It's like the backbone. It's the vector that every layer kind of modifies and dumps more information into all the way up to the top of the model where it gets decoded into an actual token choice. And what you'll find is for a given token, you have a residual stream vector. You sometimes will find there's an activation within that vector that is ridiculously huge. It'll be like 2000, whereas the

1:39:36others are like 0.1 or something. And these are called massive activations. This is like a known thing. There's in a transformer, the key is a mathematical object that is computed from that residual stream and it just by a matrix multiplication. And so when you have a really huge value in the residual stream, that translates into like an exploding value in the key. And that monstrous entry ends up in the KV cache. And the problem there is, so quantization basically will take say

1:40:08four bit quantization will take a number and convert it into a form where it can only take one of four values. So typically, you know, you might have like, I don't know, like 16 bit quantization. That means you have two to the 16 different possible values that you can represent in your numerical scheme. If you have just like two bit quantization, now you only have like four different possibilities. So if you have two bit quantization and you have a value that's super, super huge, it gets squashed in to take the highest possible value out of just the four. So you can't

1:40:41really distinguish it from just like relatively high values. And it just kind of wrecks a lot of stuff in quantization. And so what they're doing here is just trying to understand the dynamics of why these massive activations occur, how model architecture shapes where they show up. And it turns out that they will show up at like full sort of traditional attention layers, but not in these linear attention layers for mathematical reasons that if we had more time, we'd go into. They're quite interesting. But the bottom line is the understanding is better. It does mean that you can then start to

1:41:12quantize more intelligently and ship these models that, you know, don't just break the minute you try to compress them by quantizing them. On to synthetic media and art. One last story. This is from the New York Times. AI slop is everywhere. Spotify, LinkedIn, and others have had enough. So this is kind of a summary paper telling us a lot about the state of the internet broadly. We have some numbers here like a Pew Research Center study from August 2026 found that 10% of

1:41:431,000 web pages sampled in July showed significant signs of AI authorship. Another report said from Cloudflare, for the first time, web traffic from AI surpassed traffic from human users with boss accounting for 62% of search requests. Then there's AI detection startup Pangram that, you know, did some analysis of how much AI content is out there. And this is kind of funny. Found LinkedIn to be the most AI saturated platform with more than 40% of its long form post flagged as completely AI

1:42:22generated. No, that's, that's, that's bullshit. It's a bastion of sincerity and honest thought. Yes. And what I'm sure is a coincidence, LinkedIn's chief product officer announced a new quote, seems like AI slop button that lets users flag posts. In its first two weeks, over 1 million users clicked it. And they're saying that numbers are experiencing 40% fewer views on content classified as AI slop. So anyway, there is a reckoning coming to the internet. We are now like, I guess it's not

1:42:58like the most harmful outcome of AI, but it is a real negative outcome of AI that like browsing the internet has become worse because we are now, you know, some people are now just putting the slop out there. And when you say slop, I mean generic AI outputs that don't bear any sign of human authorship or intent and, and are just sort of like a waste of your time. So they look to be more efforts from LinkedIn, Spotify, Substack, other platforms to more directly tackle the proliferation of this type of

1:43:36content and then make it so human generated content, I guess prevails because otherwise just the sheer volume of posts you can create with bots, it will drown out everything else. And with that, we're going to close out this episode of last week in AI. Thank you so much for listening. We'll be doing our best to get this out within a day or two of recording and to record consistently for the coming weeks. As always,

1:44:08we appreciate you listening, sharing, reviewing, commenting, and we do check comments on YouTube and elsewhere. And please do keep tuning in. Tune in. Tune in. Tune in. When the AI news begins, begins. It's time to break. Break it down. Last week in

1:44:47AI, come and take a ride. Get the load down on tech and let it slide. Last week in AI, come and take a ride. From the labs to the streets, AI's reaching high. New tech emerging, watching surgeons fly. From the labs to the streets, AI's reaching high. Algorithms shaping up the future seas. Tune in. Tune in. Get the latest with ease. Last week in AI, come and take a ride. Get the load down on tech and let it slide. Last week in AI, come and take a ride.

1:45:24From the labs to the streets, AI's reaching high.

1:45:41From neural nets to robot, the headlines pop. Data-driven dreams, they just don't stop. Every breakthrough. Every code unwritten. On the edge of change. With excitement, we're smitten. From machine learning marvels to coding kings. Futures unfolding. See what it brings.

More from Last Week in AI

#256 - Fable 5.1, Astra Tease, Gemini 3.8 Flash

Sep 8, 20261h 14m

#254 - Rogue AI hacking, bio-weapons, Dean & Hassabis out

Aug 11, 20261h 58m

#253 - Opus 5, Gemini 3.6, Kimi K3, Hugging Face Hack

Aug 3, 20261h 43m

#252 - GPT 5.6, Grok 4.5, Nemotron-Labs-Diffusion, AI 2040

Jul 15, 20261h 25m

#251 - Mythos Back, Sonnet 5, Etched, LongCat

Jul 9, 20261h 30m