Steadcast
Security Now (Audio) cover art
Security Now (Audio)

SN 1092: Restraint Abliteration - Rotating Keys, Broken Guardrails

August 19, 20262h 49m · 22,864 words

Show notes

From autocorrect to full-fledged conversationalists, discover how a few tweaks transformed language models—and why understanding this shift exposes urgent questions about AI safety and control. Trusting an open source AI proxy might bite you. France's under-15 social media ban hits its constitution. A bit of AI prompting found a serious bug in Zoom. AI-based network defenders see a stock price jump. A (very) deep dive into the operation of AI chatbots Show Notes - Hosts: Steve Gibson and Leo Laporte Download or

Highlighted moments

The light LLM compromise was the result of a, get this, Leo, a previous, this will ring some bells, a previous supply chain attack that infected the widely used vulnerability scanner, Trivi.
22:34
Specifically, for each model, we find that a single direction such that erasing this direction from the model's residual stream activations prevents it from refusing harmful instructions, while adding this direction elicits refusal on even harmless instructions.
1:56:30
GRAM adds extra neurons to every layer of a standard transformer. You know, the neural network architecture on which large language models are based. These neurons are divided into groups or modules. One module per dual-use category. During training, when the model encounters general-purpose text, that is, you know, just standard text. We don't worry about it one way or the other. It learns in the usual way. But when it encounters text from a dual-use category, virology, for instance, they wrote, The rules change.
2:23:16

Transcript

A propeller hat episode on AI

0:00It's time for security now. Steve Gibson is here. Of course, there's a lot of security news, including, yes, a supply chain attack. France is under 15 social media ban hits its constitution. That's not a bad thing. But this is what we call in the business a propeller hat episode. Steve is going to take a very deep dive into AI, how chatbots are made and unmade. This is a fascinating episode.

0:30Next on security now. Podcasts you love from people you trust. This is security. Now with Steve Gibson, episode 1092, recorded Tuesday, August 18th, 2026. Restraint obliteration. It's time for security. Now the show, we cover the latest in security, privacy, computing, science fiction, vitamin D, and anything else on this man's mind, because he is a genius. Ladies and gentlemen, I give you Steve Gibson. Hi, Steve.

Restraint obliteration and core AI

1:10So thank you, Leo. What our listeners are going to find is that. So excited about this show, by the way, you gave me a preview, folks. Is that I did not get to the two topics that I wanted to. Oh, no. I only got to one because I started laying down the foundation for the topics and there was just so much to say. I'm, you know, my only defense for this being about AI is that the world is and security certainly is. I mean, we, I don't have to even explain that anymore.

1:57As we said, Black Hat a couple of weeks ago was like the AI conference that happened to be doing security as the reason for spending all that money. Um, as I'm, okay. I look back at the tutorials on the internet, how it works and computing and how it works. And there I was able to share with our audience things that I had understood for quite a while. So in the case of AI, I'm, I'm, I'm a, you know, a lay person, neophyte rank amateur. Um, but I'm a curious researcher and I've been reading research.

2:42So what I'm going to be doing over, I have no choice really is because it excites me is I'm beginning to develop a understanding the, uh, at the level I want to remember I program an assembler. So I, when I say I understand something, I, for me to meet, for me to be satisfied, I have to understand, well, is the carry bit set or not? Um, so, but at the same time to make it understandable. So I think most of what I look at the things we're going to talk about, we're, we're going to talk about how trusting an open source AI proxy might bite you.

3:27Um, France's under 15 social media ban, a bit of AI prompting turned up a new serious bug in zoom, uh, and that the, the stock prices of AI based network defenders has jumped up after that black hat conference. That's all we had time for because the rest is me sharing a bunch of new understanding that I have that I think our listeners are going to appreciate.

4:04So, uh, and actually not that much. I mean, I didn't skip any fantastic news. I looked for all the good things. Uh, we got a great picture of the week and by the, by two hours from now, everybody listening is going to understand unless they already do, they might, but we'll understand what I now understand about the early first steps of how AI, as we know it today happened.

4:38What were those things? And for example, Leo, you were first out of the gate saying it's nothing more than fancy autocorrect. Turns out that next token prediction is all it did in the beginning. That's what it was. This was, I said that in my defense a year and a half ago. I mean, that's my point. No. And that's my point. Then it was true. And also Matthew green, we quoted him last week saying it's not fancy autocorrect. It's not just the next token, the next token, the next cause of what happened between. And I understand now, and I'm going to explain how we got from next token prediction to you can have a dialogue because that's a very, that's a, those are different things.

5:24And, and, and it's a little freaky how, I want to say how easy it was, but it also led me to have a conversation with Claude about its own nature. So anyway, I think a great podcast for our listeners. I titled this restraint obliteration, not obliteration, but obliteration, because that's actually the, the term of art, which is used for some of what

5:54happens or can happen with open weight models. So we're going to buy, again, all I can say is two hours from now, you're going to be like, Oh, I, I understand a lot more than I did. And you'll probably at that point, you'll understand about as much as I have, but I'm not, I'm not. In other words, this is a brain dump. This is going to be Steve's brain dump. So he has some room to read more papers. That's correct. But I already know what the next one is. Oh, good. It's so exciting. Yeah.

6:25It's so fascinating. And I, and actually, I'm really glad you're digging into it. I don't have a choice, Leo. This is the most important thing that has happened in my 71 years of life. You could say, well, okay. Computers. Yes. I was programming them on a PDP eight in high school. Oh, internet. Yes. We watched all that happen, but, but this is knowledge. I mean, those, I mean, yes, we needed to have computers. Well, it's a stair step. Yeah. We got, we had to have the internet for training. We had to have the computers. Of course,

7:00you know, I would throw mobile in as another big revolution. The idea that you have the internet everywhere you go in your pocket. Yes. Uh, and I'll, and a significant amount of computing, but, uh, and of course, if it weren't for gamers, we wouldn't have these video cards that are, turns out are really good at AI as well. So it's all been, you know, it's always the case. It's always been a stair step. And of course, this is for curious people, right? You don't have to understand how any of this works in order to use it. Any more than you do a computer. Right. Most people have no

7:32idea how a computer works. The internet is magic. AI is a worry because what's going to happen. So again, you, you can use it without understanding it, but here in our little, you know, little corner of the world, uh, we like to understand how these things work. So, um, we're all going to understand how AI works. It's pretty amazing. I mean, and actually it is, there is a continuity. AI is, I'm sure when you were at the Stanford AI lab sale in the seventies, 73, you know, when,

8:07when, when John McCarthy, the creator of common Lisp was there, I mean, yep, this is with his page with his gray ponytail pulled back. This is ancient history. But even then the vision was what we would like to do with these computing devices is talk to them and interact with them in a natural way as we do with other people. Well, now you can, and it just blows me away that we have made sand think is, is mind boggling. And as Jeffrey Hinton says, the way we did is by applying

8:43huge amounts of electricity. He says, and it's really interesting. The talk I sent you, he says, human brain is designed for low power analog, right? And, and, and so it's design is designed specifically for the kinds of things we were, we could have, but once we figured out how to do ones and zeros and keep it accurate, our brains are not accurate, but ones and zeros. And he says, if you apply enough power, the one stays a one and the zero stays a zero, right? It's all about applying really a vast amount of electricity to these things. Once you could do that,

9:19then you have something that is analogous, but not the same as a brain and can do some interesting things. And that's what we're going to talk about next on security. Now don't get obliterated. You might want to get obliterated, but wait until after the show, we're going to talk about that and more, but first a word from our sponsor. And I have to say now that I have this little lab here up in the studio, this AI lab, it's even more important to me that I have a Thinkst Canary, our sponsor for this part of a security now that we've been friends with, and they've been sponsors

9:54of our shows for a decade now. In fact, we were the first place they came when they first invented Thinkst Canary. Thinkst Canary was invented by a team. I love these guys. Met them at RSAC actually for the first time. We've talked for years, but I finally got to meet them because they're in South Africa. But these guys were brilliant hackers. They were white hat, but they taught companies how to penetrate systems, how their systems might be penetrated. And one of the things they learned, as many hackers do, is the best way, the best defense against hacking, of course, a good perimeter,

10:26but you need a honeypot. You need a way of knowing if somebody has breached your perimeter and is wandering around in your network. And I bet most of you watching don't have a way of knowing. You could look at the logs, but a good hacker is not going to let you see anything. They erase their traces. We know that. That's why at the very beginning, people like Bill Cheswick and Cliff Stoll, who were able to see there's somebody in our system. Cliff Stoll found out because there was a fraction

10:59of a cent difference in a calculation. It was that tiny. But he said, wait a minute, something's wrong. These guys said, you know, what you really need is something on that network that a hacker can't resist. A little piece of candy, a little something, something that the hacker, if he's gotten in, will trigger and let you know he's there. That's a honeypot. Now, when Ches wrote the first honeypot, since he's even told us this, it was a hard thing to do. But thank goodness these guys at Thinkst have figured out how to make the perfect honeypot, the Thinkst Canary. It's so easy to use.

11:34It looks like a little USB device. I'd show you mine, but it's better to believe it's on the network now, man, doing its job. Because if someone is inside your network and you've got these honeypots that could, by the way, be deployed in minutes, plus could create tripwires, little files that phone home to the Thinkst Canary and say, somebody's trying to open me. If somebody's accessing those lure files or they're trying to brute force that, whatever it is, fake internal SSH server or Linux box or Windows box or SharePoint server, it can impersonate anything. You're going to get your

12:09Thinkst Canary. I'll immediately tell you, you've got a problem because nobody should be accessing this. There's no false alerts, just the alerts that matter. And in the way you want them, text message, of course, Slack, email, they support webhooks, they've got an API, they've syslogged, whatever it is, any way you want. All you have to do is choose a profile for your Thinkst Canary device. It's so easy, by the way. You might change it regularly. I do. I change mine all the time. Then register it with the hosted console for monitoring and notifications. And you just sit back and you relax. You wait. Attackers who've breached your network, malicious insiders, any other adversary

12:44that's in your network and shouldn't be, cannot help but make themselves known. They can't resist it. They're going to access that file or that Thinkst Canary and you're going to get an alert. You've got to go to canary.tools.twit. $7,500 a year. You get five Thinkst Canaries. Of course, you get your own hosted console. You get upgrades. You get support. You get maintenance. And I'll tell you what else you get. If you use the code TWIT in the How Did You Hear About Us box, you get 10% off the price. And not just for the first year, for life. Oh, and if you're at all concerned,

13:16do I need this thing? You can always return your Thinkst Canary. They've got a two-month, 60-day money-back guarantee for a full refund. I should tell you, though, in this whole decade that they've been partnering with us doing these ads, 10 years, that refund guarantee has never been claimed. Not even once. Visit canary.tools.twit. Canary.tools.twit. Enter the code TWIT in the How Did You Hear About Us box? Thinkst Canary. I'm so glad I have one. Now that I have something on my

13:47network, I really want to protect my AI. I'm going to make sure that nobody's in here messing around. All right, Steve, picture of the week time. So this picture generated a great deal of feedback,

Albanian bus bridge picture

14:06furor, hubbub from our listeners. The email went out. I forgot to mention a week or two ago that broke 21,000 subscribers. Wow. So we're north of 21,000. You've got more subscribers than TWIT has club members. That's very nice. Good job. That's well. And I... That's because it's free, I might add. Yeah. Yes. And great feedback. So first of all, I gave this picture the caption, this actually happened in Albania. Was it an accident or a very clever way to create a bridge

14:46across the river? All right. I'm scrolling up for the first time. I haven't seen it. Wait a minute. There's a bus. Whoops. Now, wow. You wonder, looking at this, so for those who are not seeing the video, we have a long, very green, it's got go green, hybrid city bus. I mean, it's long enough that it's got a door in the front. It's got doors halfway

15:18down its length, and then it's got rear doors. It's good it's not one of them articulated buses, though. I don't think it would serve as well as a bridge. That would be a problem. Yes. And somehow this darn bus is straddling the river. But what I noticed about it first, well, actually, after I got over the fact that it was there, was that it's so long, and it's got doors in the front and rear, and the fact that you're able to walk the length of the bus, it's a bridge.

15:51It is a bridge now. I think it's telling the middle doors are not open. Yes. Unless you wanted to fish, then it would be good. You could sit there and dangle your feet out and fish. Anyway, Leo, this actually happened. One of our listeners found the – I have a link at the bottom of the picture of the news because some people who didn't see that said, oh, that's AI. How could that bus possibly get into that position? Because you would think,

16:25looking at it, it would just go nose down into the river, right? How could it get on the other side? It turns out it's the bizarrest accident. And by the way, you know it's not an intentional bridge because there is crime scene tape across the bottom. That's true. They don't want people to use this. They're trying to keep people out. No. One of our listeners found one of the news reports where they did an animated recreation of the accident. There was a Mercedes being driven by

17:02a younger man. And I read the news report. I don't remember the ages, but he was like 27 or something. And so now there you can see a picture. And notice underneath the front, Leo, is our white lights. Yeah. Is that the Mercedes? That's the Mercedes. Oh, God. Not good. The Mercedes was driving to the left of the bus when the bus driver lost control, turned to the left, knocked the Mercedes off the road. It preceded the bus into the river, flipped upside down,

17:37and the bus rolled over it. So the bus, so Mercedes formed a brief stone in the middle of the bridge that allowed the bus. No deaths, thank goodness, but six people were injured. Here's another picture of the scene. Yeah. Wow. Holy cow. The Lana River. So I guess. I would have said Photoshop for sure. Yeah. It's nuts. It's nuts. I mean, you needed another car there for the bus to drive over the car

18:12in order to get to the far side. And then they pulled the car out from underneath. So, wow. Well, I'm glad the driver survived. Anyway, thank you. One of our listeners who sent this to me and thought maybe this would be a good picture of the week. And I agree. Okay. So as I promised last week, there are two very important pieces of core AI technology that I wanted to discuss. And I also promised last week that we were going to do a deep dive into aspects of the operation of

18:48today's AI. Well, that turned out to be more true than I expected, so much so that it entirely crowded out of the second of the two topics that I had planned to get to. But fear not, we're going to have another deep dive into that next second topic next week. And that's the role confusion topic, Leo, which you and I talked about at Black Hat, the research paper that I read during the flight there.

19:19And I said, OMG, how, how, wow. Can that still be the way things are being done today? And I've confirmed, yes, unfortunately, it is. It's not as weird as a bus crossing the Ridge River, but it's

Terabyte credentials in supply chain

19:34close. It's close up there. Okay. So we're going to start by covering some important recent security events and then get into understanding this first of the two core issues of the way AI works. Last Wednesday, Ars Technica's great security topic writer, Dan Gooden, reported under the headline, terabytes of credentials leaked in massive supply chain attack was Ars headline. And of course, that was the

20:09kind of thing I would have chosen to share even a few years ago. But the thing that raised my interest another several notches was the article's tagline, which read that data, that is the terabytes of credentials, was scraped and exfiltrated from 2,500 users of a compromised AI package. So I thought, whoa, okay. Oh, it turns out what's, what's even more significant is we're not talking some random end users in

20:40Nebraska, you know, who no one knows. Wait till you hear whose credentials were among the more than 2,500 that were stolen. So Dan writes, terabytes worth of credentials, many belonging to the world's biggest and most sensitive organizations have been exposed in a supply chain attack on light LLM, an open source tool that streamlines AI driven software development. Microsoft,

21:13Amazon, Cisco, Samsung, and Salesforce are only a handful, he writes, of the entities whose access secrets were exposed. The revelation was posted on Tuesday and Wednesday. That's live last week by security firms, CloudSec, you know, S-E-K and Hudson Rock. CloudSec said it found keys, repository tokens, SSH keys, Kubernetes secrets, package publishing credentials, environment variables,

21:48and AI provider keys that could allow attackers to gain access to more than 2,500 organizations. And 40 minutes is all it took. The credentials were extracted during a 40-minute window in March, while the victims used, obviously unknowingly, compromised versions of light LLM, which had been downloaded from the package's official location in the Python package index repository, you know,

22:21PyPI. Hudson Rock said it made the discovery after analyzing a 195-terabyte file that it had obtained. Neither firm identified the source of the information. The light LLM compromise was the result of a, get this, Leo, a previous, this will ring some bells, a previous supply chain attack that infected the widely used vulnerability scanner, Trivi. And remember that we had talked about this months

22:55before. Other software infected in the campaign includes KICS and the Telnix Python SDK. Team PCP, which Dan describes as a ramshackle but extremely capable gang, largely made up of teenagers, took credit for the attack and researchers have largely corroborated their claim. Independent security researcher Kevin Beaumont said, quote, I've confirmed the data is legit. By the way, multiple victim orgs.

23:32He said it contains a significant volume of sensitive content at orgs. It's a massive supply chain breach due to poor AI security. Not because AI is the threat, but teens can now run circles around orgs obsessed with rushing out AI and poor DevOps security. And I'll just explain that a little bit here. I'll take a moment. Light LLM was compromised. It's not that AI was used in the compromise. It's that,

24:13that, you know, Kevin is saying that the way light LLM is used, the way it needs to be used, that is to be a proxy for other AI services means that you need to give it all your secrets so it can act on your behalf. We've talked about this fundamental problem previously, and a lot of orgs just got bit by it. Anyway, I'll have more to say here in a second. So continuing, Dan writes, the compromised versions of all four software packages contained,

24:47right, four packages compromised by Trivi contained code that accessed the memory of infected machines, scraped its contents, and exfiltrated it through an attacker-controlled channel. The data is filled with an assortment of information, and again, 195 terabytes. So it's like a wealth, but you got to find the goodies in there. Interspersed in the wall of data are credentials to software pipelines maintained

25:21by tens of thousands of organizations that ran Light LLM during the 40-minute span that the supply chain attack remained active. In all, both security firms said some 434,000 CICD, you know, continuous integration, continuous delivery, software pipelines had credentials exposed after running the compromised

25:52Light LLM versions. There were two versions that were compromised. I'll clarify that in a second. Um, in many cases, the researchers at CloudSec and Hudson Rock had trouble identifying, which is a problem, the organizations, uh, the credentials belong to. For instance, an email address in the dump from the domain at SiriusXM.com ultimately did not indicate a breach at the satellite broadcaster,

26:23but rather within the infrastructure of SiriusXM's subsidiary, AdsWiz. A trove of internal corporate secrets were found exposing sensitive tokens for platforms such as Salesforce underscore client underscore secret and Slack, Slack underscore signing underscore secret and Microsoft Azure environments. The researchers had high confidence that the following organizations did have their credentials exposed. Get

27:00this. NVIDIA, AWS, Samsung, Salesforce, Cisco, uh, Hoffman LaRoche, ServiceNow, Siemens, S&P Global, Airbus US Space and Defense, John Deere, Regeneron Pharmaceuticals, London Stock Exchange Group, Thompson Reuters, FedEx, FedEx, Mediatek Inc., Volkswagen AG, Deloitte, the Kroger Company, Siemens Energy,

27:30Thales Group, XCorp, as in Twitter, Zscaler, Epic Games, Orange.SA, HP Inc., Philips, Vodafone Group, Carl Zeiss, Deutsche Bahn, NGINX, BT Group, and Roku. I mean, wow. Hudson Rock said, many CI, CD pipelines are configured generically. The dumped variables contain active database passwords, third-party API keys, and cloud credentials without any identifiable company email, custom domain string,

28:08or internal server name. This means that countless organizations, which they were unable to identify, currently, as in still, have active secrets sitting in this database. They are completely unaware of their exposure. The, the research firms, I should note, immediately identified all the companies that they could among, you know, among those I just read, but lots of other companies, there are like,

28:38what, 434,000? Was it different CI, CD pipelines? You know, all their secrets are out there, and they haven't been notified because it's not clear who they are. So maybe that's safety. You know, I mean, the bad guys probably can't tell either, but it certainly gives you some starting credentials to use for some, some password guessing. Finally, under the heading, welcome to the new world of supply chain attacks, Dan adds, both firms, those two security firms, are urging all organizations that use the

29:15compromised versions of Light LLM, particularly those listed in the high confidence section of the list, the organizations that are known to thoroughly rotate all credentials in their pipelines. Hudson Rock instructed any organization that uses any AI proxy infrastructure, third-party CI, CD, vulnerability scanners, or downstream AI packages to immediately audit their environment for versions 1.82.7 and 1.82.8

29:52of Light LLM, which were the two compromised versions of the software. The firm advised those affected to perform, quote, aggressive credential revocation, unquote, assume any secret accessible to the Light LLM environment is compromised, invalidate and rotate all cloud keys, Kubernetes service account tokens, the GitLab, GitHub, PATs, and audit logging and egress filtering. As a cautionary tale, CloudSec

30:26said that Trivi developers rotated but failed to fully revoke an automation token over a 20-day window. That lapse gave the attackers a nearly three-week period in which to force-push malicious code to third-party builds that use the vulnerability scanner. As Beaumont observed, organizations rush to integrate AI into their software delivery systems has also greatly contributed to the scale of the damage, meaning just,

31:00you know, as we've seen. And Leo, remember when you were, you know, immediately thought, hey, this, I can't remember the name of it. It was the package that came out at the beginning of the year that was, uh, the code, uh, the code, uh, the code writing claw, uh, Claude? What? No, no, no, claw. Open claw. Oh, claw bot. Yeah. Um, yeah. Open claw. Yeah. So, oh, yeah. So, so open claw happened. You were excited about it,

31:32but you pulled back at the last minute, uh, and it turned out that was, yeah, there was just, you had to tell it too much in order to allow it to do what you want it to do. But if, but yeah, if you wanted it to do anything, you gave it your email, you gave it money, you gave it a phone number. Exactly. A little dangerous. So then in a, okay, in a final update to his initial reporting of this massive mess, Dan added, there are already signs that some of the affected organizations are not taking the disclosure with the seriousness

32:08warranted. After this post went live, he writes, Kevin Beaumont reported quote, these creds date from about March. One of the orgs impacted told me, he writes, they'd rotated them all and it's a nothing burger. He said, so I looked at their responsible disclosure policy. It allows trying creds. So I tried them all. Almost every one of them worked. He said, I submitted a report,

32:45one of the biggest us telcos. So someone said, oh yeah, we don't worry about it. We rotated our credentials, nothing to see here. So they probably made new ones, but didn't delete the old ones is what they did. Jeez. So ultimately, um, the new revelations concerning the light LLM supply chain attack underscore is writing Dan, the growing threat of such campaigns and hence the importance of maintaining vigilance around the use of open source software that when infected can spread

33:18rapidly across the internet. Uh, Alon Gall, co-founder and chief technology officer of Hudson Rock wrote in an email quote, the key takeaway is how supply chains have evolved to make a single upstream breach affect thousands of companies simultaneously. A window of roughly 40 minutes in which the light LLM dependency was hacked led to over 430,000 instances in which millions of secrets were harvested.

33:57This magnitude pushes us into a completely new world regarding the type of response required from the cybersecurity industry. And Lord knows, you know, we've been unimpressed typically by the kind of responses that we've seen historically. So, so just to be clear that the cautionary takeaway from this is not a case, as I said before, of AI going rogue or any kind of AI misuse. You know, I followed Dan's

34:28reference links back to one of Hudson Rock's, uh, reports who's much more technical write-up of the breach which makes what happened actually further clear. Their headline, Hudson Rock's headline was largest AI supply chain breach of 2026 colon light LLM hack impacts thousands of global enterprises. And Hudson Rock explains quickly, the cybersecurity landscape is currently reeling from one of the

35:03most sophisticated multi-ecosystems. The most sophisticated multi-ecosystem supply chain campaigns publicly documented to date. The orchestrated, it's a, I'm sorry, orchestrated by a threat actor group known as Team PCP. This cascading attack ultimately compromised Light LLM, a widely adopted open source AI proxy gateway, leading to the silent exfiltration of deep developmental secrets from thousands of continuous

35:37integration and continuous deployment pipelines worldwide. Excellent forensic research published by SYNC, you know, that's S-N-Y-K, Trend Micro and Psycode has thoroughly detailed the mechanics of this breach. The attack did not begin with Light LLM. Instead, Team PCP first compromised the GitHub Actions pipeline

36:08for Trivi, a highly popular open source vulnerability scanner. Because the developers of Light LLM utilized Trivi in their own CI-CD pipeline, the poisoned security scanner was granted legitimate read access to their runner environment. This allowed the attackers to silently exfiltrate Light LLM's PyPI publishing tokens.

36:41Armed with these credentials, Team PCP was able to publish malicious versions of the Light LLM package, versions 1.8.2.7 and 1.8.2.8. The payload delivery was exceptionally stealthy. By utilizing a .pth Python startup hook, the malicious code was executed the moment the Python interpreter initialized,

37:14regardless of whether the Light LLM library was explicitly imported. The three-stage payload immediately began harvesting environment variables, local configuration files like .kub slash config and .aws slash credentials, attempted lateral movement across Kubernetes clusters, and installed a persistent systemd backdoor. While the security community has deeply analyzed the malware's behavior, Hudson Rock has independently

37:50obtained the actual fallout, the raw exfiltrated data. That's that 182 terabyte of data. I mean, talk about a huge amount of data in 40 minutes. Because there were that many instances of those two versions of the malicious Light LLM that were updated from Python PI, and installed and started, and all the credentials poured through it, and they sent them all off to some

38:24malware mothership somewhere. So they said, this provides an unfiltered look into the massive scale of the compromise through the actual raw files. Our researchers have obtained and analyzed a staggering 153 gigabyte RAR archive, and these files will compress way down. This massive corpus contains exactly 433,909 files. Through our analysis, we've successfully attributed 118,829 CI runner dumps

39:08to 2488 affected corporate domains. Whether a developer, machine, production server, or CICD pipeline executed the compromised Light LLM package, the threat actors successfully harvested the live environment memory and configurations mid-execution. Okay, so the real takeaway from this is that the massive adoption

39:41of the automation of development and delivery will tend to coalesce around relatively few most popular best-of-class tools. This naturally makes those tools a highly valuable target for attackers. Over GitHub, the Light LLM repository describes itself, writing, Light LLM is an open-source AI gateway that gives you a single, unified interface

40:17to call more than a hundred LLM providers, OpenAI, Anthropic, Gemini, Bedrock, Azure, and more, all using the OpenAI format. Use it as a Python SDK for direct library integration or deploy the AI gateway, a proxy server, as a centralized service for your team or organization. Managing LLM calls across providers, you know, multiple like, you know,

40:48Anthropic, OpenAI, Gemini, so forth, multiple providers, they say managing those calls gets complicated fast. Different SDKs, different auth patterns, request formats, and error types for every model. Light LLM removes that friction with a unified API, one interface for more than a hundred LLMs, no provider-specific SDK juggling, drop-in OpenAI compatibility, swap providers without rewinding your

41:22code, production-ready gateway, virtual keys, spend tracking, guardrails, load balancing, and an admin dashboard out of the box. Eight milliseconds P95 latency at one K RPS, you know, requests per second in benchmarks. So, wow, right? Sounds great. The only glitch here is that it must be totally trustworthy. In order for Light LLM to be able to proxy for its users, all of those differing

42:01LLM backends, it must necessarily have every user's authentication credentials for every one of those different LLM backends. Just like, Leo, you were saying, OpenClaw, you know, had to be able to log in as you, had to be able to access your bank records, had to be able to make payments on your behalf, blah, blah, blah, blah, blah. I mean, when we get into proxies and agents, trust has to be there

42:32because, you know, they're acting on our behalf. So what happened here is that the Light LLM developers were users of the Trivi vulnerability scanner. And as we know, because we talked about this back in March, when this happened, Trivi was compromised. This allowed attackers to compromise any project that was using Trivi as its scanner. And thus in turn, those two versions of Light LLM were compromised,

43:02giving attackers complete visibility into the credentials and domains of those 2488 corporate users of Light LLM during just that 40 minute window. And, you know, I always like to try to suggest solutions to the problems we encounter here, but I got nothing because this is, this is like a fundamental problem with the way we're doing things now. You know, we're now in a mode where everyone feels that they

43:36need to always be running the latest and greatest release of everything, right? It was, it was because of the updates to those bad, those two bad Light LLM

43:50packages in the repository that, that this happened because there was a new version. Oh, got to have that. So, you know, we've done this to ourselves and I preach updating relentlessly, right? You know, the entire security community is constantly pressing everyone to get better about updating their software. Stay current, be sure you're receiving announcements of important updates, blah, blah, blah. You hear it here all the time. But in this instance, doing that is precisely what bit the users of those two specific

44:28versions of Light LLM. If they had not updated to those, if they'd remained on a previous release, they would have never run one of those two malicious versions. But of course, not updating doesn't work as a strategy either. No, I pin stuff, uh, trying to avoid this after this LLM, uh, debacle. Um, you know, people said, you know, I was caught within a day, I think, or a couple of days. They said, if you just make sure you don't install anything that's less than three days old, you'll be all right. I made it 14 days because I figured, I mean, you're still not going to catch everything,

45:03but if it's a popular package, two weeks should be enough. So I'm very careful. The other thing I do, though, is I store all those tokens and secrets in a Bitwarden secret manager. So they're not, it's like my passwords, right? They're in a locked vault. And we did just talk about this a couple of weeks ago where, uh, where we're beginning to see a new category of security software, which is a means of allowing AI to, to access your credentials without it ever having them.

45:38Right. That is so, so, so it says to Bitwarden, I need you to log in for me. In effect. Yeah. Well, or it gets a one-time token or yeah. I mean that the, it's still going to see, the problem is, so you're using that open AI endpoint, which is connecting to somebody who's serving that AI. They want that token. So the token has to float around somewhere in memory. You don't want to write it to the hard drive, but it has to, light LLM would still have to be able to see it and send it. So I'm not sure a secrets manager would have protected me in that case.

46:12Yeah. You try to do what you can. So the, the flip side, of course, of your two week window, which is good, is that if there was a critical vulnerability that actually was authentically, authentically patched, then you wouldn't be getting it for two weeks. So, so, you know, who knows what might be malicious and what isn't. We are, we're, we're in a bad place right now. The only winning strategy, or at least the best strategy that's available for the moment is just to use the tools, but

46:46carefully monitor the news for the tools you're using. Stay current. Yes. But you know, because mistakes are going to happen that the instant you learn of a breach that affects you invalidate and recreate, in other words, rotate all possibly affected credentials. We've seen many instances where that was not done. You know, remember that part of what made last passes troubles even greater was that even after they learned that they'd been breached, they failed to fully cancel all

47:22previous authentication credentials. And, you know, and as a perfect case in point, we learned that, you know, uh, from Kevin Beaumont's test, a major telco did not actually rotate their credentials when they had claimed to. So I guess there actually is an important and practical takeaway from this. I mean, it's like a real action item. We know, we know that, that things that are easily done

47:56tend to be done and things that are a confusing pain in the butt too often fall into the, okay, I'll get back to that later. What we see is that the speed of modern attackers means that later stands a very good chance of being too late. So here's my, my takeaway point from this. I think it's really crucial. If there's really no practical means of preventing inadvertent exposure to credential

48:32leaking malware. And there may not be that today, then rotating all of the credentials that any malware might have obtained. The moment a credential compromising attack is known is important. Of course, moreover, periodically rotating credentials preemptively on the better, safe than sorry principle can be useful when the threat environment warrants it.

49:04But this is, by the way, this is contrary to the, uh, uh, advice we've been giving about passwords, rotating passwords, but, but it is what you should do, which is right. Yes. So when I make keys now, I give them an expiration date of, of month or two months or three months. And I know I'm going to have to rotate them because I don't, I don't get to keep using them. Right. So what, what, what I believe this means is that whole organization, organization wide credential rotation needs to be both possible and easy, right? If it's not,

49:41if it's not easy, it won't happen or it will be put off. So my best advice would be, if you're looking for a nice self-contained, easy to describe project for an AI agent to tackle, a great investment would be to employ some AI to implement a comprehensive credential rotation facility for your organization. Make it automatic, make it a single command that handles everything,

50:15make it easy, even fun to use and use it to maintain the credentials for both current and new services as they're being brought on board, you know, as services are added and removed and require its use for instantiating any new credential into use so that it cannot fall out of sync. Uh, or again, if, if that happened, it wouldn't be trusted and used. So I would say making, you know, automating credential rotation would be a, a, a cool,

50:52it's easy to describe to an AI. It's, it's, it's, it's, it's self-contained. You can see if it's working or not. The failure doesn't kill you. It just, you know, you got to fix it. But I think it would be then incredibly useful if, if, if, if, if, well, when things to canary finds that some guy is in your network, the first thing you want to do is, is, you know, app, you know, don't panic as, as the hitchhiker's guide to the galaxy tells us, rotate those credentials and, you know,

51:25uh, and, and, you know, make, make it easy, make it fun. There are, uh, there is a way to do this a little safer. And I think if I were a big or any, any business, I probably would be doing this. Uh, when we were at our sec, I met a few of these guys. They, they're an intermediary, a credential capability layer. So they hold, they're a third party. They hold their credentials. You don't ever have the credentials on your machine. You have a credential to them. And when you want to connect to an AI, you connect through them. So they have your credentials and they

51:59do it. And they, you only get a one-time token. Uh, the problem. And the reason I didn't want to do that, a it's you'd pay for it, but, but B they have your credentials and I, you know, it has to be somebody you trust because they're going to have your credentials. Right. Uh, you know, it has, at some point, these credentials have to be exchanged. It's just like a password. I guess you could use hashing. Couldn't you? Well, that would be the solution. Um, we, a couple of weeks ago, we talked about one password saying that they were, uh, introducing exactly to this service. So, uh, keep keeping it local. And of course,

52:32the problem is that we're talking all credentials. So like that online service would, my, would be handling your credentials as a proxy for AI. But what about SSH servers? What about web, you know, all the other things that, so what you want is, is, is one thing that is just able to wipe the secrets out of your organization and replace them. Yeah. That's, that's the one password credential broker. That might really be the right way to do that. We're going to have to have something

53:05like that. Yeah. Well, you know what we're going to have to have right now, Leo, a break in the action. Uh, you're watching security. Now Steve Gibson, the man of the hour, the man, every Tuesday, it's security now day here at the twit podcast network. We're glad you're here with us. Uh, our show today brought to you by delete me. This is a topic we talk about a lot as well. Privacy, right? As a, if you're a business owner that, you know, we often talk about, but delete me as something for individuals, but this is even more important. I think for a business owner, I speak

53:38as a business owner, uh, because you know, as a business owner, you're in the public eye. That's part of the job. The bad news, the uncomfortable truth is that when you promote your business, it's also exposing your business. You and your team now are visible and who are you visible to? Not just customers, not just clients, bad guys right now, 90% of business owners. This is a devastating stat have their home address easily discoverable online. 90%. The average owner has more

54:13than 600 pieces of personal information, just sitting there on the open web. We're talking to your personal email, your phone number, your home address, even details about your family. And why? Because of data brokers. I hate data brokers, the cockroaches of the internet and hackers know this. They know they can get this data cheap. They just buy it from the data brokers and they use it for a variety of things. They can run hyper targeted phishing attacks. You know, if they have your real details, they don't sound like strangers. They sound like clients or partners

54:45you already trust. We got bit by a phishing attack exactly like that. A client we've done business before sent us a request for a proposal, an RFP. And we thought, well, of course this is, you know, we've dealt with them before, but it was a man in the middle attack. What looked like a login to their, uh, Google drive or actually our Google drive. So we could fill out a form was a man in the middle. It looked just like a Google login, but our employee put the information in. They even put the two second factor authentication in and bingo, the bad guys were

55:18in. That's because hackers know if they sound like a client or a partner you already trust, you can't ignore them. Attacks that use verified personal information like that are five times more likely to succeed. And the average incident, it can cost you a lot. Didn't cost it. We were lucky. We caught it, but it can cost a small business more than $120,000. And we're one of, well, one in four, 25% of all businesses will be impacted this year, this year alone. That's why delete me is so

55:52important. That's why we use delete me. It reduces your exposure by up to 95% by removing you and your employees, personal information from those evil cockroaches, the data broker websites. But when you get your information off of it, it starves hackers of the fuel they use to build their targeted lists. And it's not a one-time thing. It can't be because the, like a cockroach, you guys can't stop these data brokers. You get it deleted, but then they change their name. They declare bankruptcy and move to another state and they start, they just got all the data up again.

56:24Well, delete me constantly monitors or removes your data. They know the names of these guys. They keep track. When the new ones start up, they're there. And delete me will send you a regular privacy report. So you always know where things stand. We just got our email the other day. Really interesting. You know, you think you're, oh, I cleared it, but you got to keep doing it.

56:45You know, that's why the biggest and the best fortune 500 companies and government agencies have trusted delete me for over 15 years. And now the same protection those big businesses and government can get is available for your small business. Protect your business, protect your peace of mind, do what we did. Go to join delete me.com slash twit dash biz to start protecting your business with delete me today. And if you use that link, you'll also get a free year of social media protection for every seat you purchase. Let's join delete me.com slash twit dash biz. Now I want you to

57:20get that address right. Don't Google it. I know a lot of times people just go, I remember hearing the ad and they Google it, but there's another delete me in the, in the EU that is completely different. It's a GDPR deletion site. So it's not what you want. You want the broker deletion site. And for that, you have to visit this exact URL, join delete me.com slash twit dash biz. All right. That's the one you want. That's the service you want. And I could tell you from personal

57:51experience, it really works. Join delete me.com slash twit dash biz. Now back to Mr. Gibson. Uh, France's recently celebrated, uh, ban on social media access for all children younger than 15 hit a bit of a snag last Friday. Reuters reported the following. They said France's top court on Friday blocked a bill banning social media access for under 15s saying it infringed upon freedom of

58:27expression and delivering a setback for president Emmanuel Macron, who asked his government to rewrite the legislation. The bill would have barred children younger than 15 from opening a social media account from September beginning September 1st. And all accounts already open would be closed by the year end, uh, which would also need to use age verification provided by the French privacy regulator. But Reuters writes, France's constitutional council found that the bill while requiring everyone to give proof

59:04of age failed to quote, quote, specify the conditions and limits under which it should be provided as well as infringing upon freedoms and privacy. So that's their report. Uh, you know, as we immediately understood when this began to happen in the U S you know, in the context of, of U S domestic, uh, legislation to control the, the viewing of pornography online, we've noted that blocking anything for all

59:38users below a certain age inherently requires everyone's age to be known. You know, there's no way around that. So France now needs to tackle the thorny problem of that inescapable infringement upon the internet's illusion of total freedom and privacy. Um, you know, uh, it is an illusion to some degree, uh, cause we know how the internet's technology works, you know, like from the beginning, there's never been a greater infringement upon privacy than

1:00:12the abuse of third party browser cookies to track people. But that went largely unseen. So nobody really worried about it. Um, uh, uh, the problem with age assertion is that because it's explicit and it cannot be hidden in the same way that cookies were, everyone's getting outraged. You know, the truth is that if we want our governments to restrict children's access to internet content, then everyone's

1:00:44age must be known by someone somewhere, either by every source of the proscribed content or by every means of, of accessing that content. So anyway, they, uh, Reuters continues quoting the constitutional council statement, writing the quote, the council holds that the contested provisions on the one hand, disproportionately infringe upon the freedom of expression and communication. And on the other

1:01:17failed to provide the legal safeguards necessary to ensure the right to respect for private life. French lawmakers, they wrote had approved the bill in July, which is when we first talked about it last month or a month before last, uh, becoming the first in Europe to follow Australia, whose world first ban barred access to platforms, including Facebook, Snapchat, Tik TOK, and YouTube for under 16s in December lawmakers. There are considering stricter penalties after data showed mixed success

1:01:51countries around the globe, including China, the UAE and Turkey have either instituted measures intended to curtail or bar access to social media for young people, or have said they're planning them. The European union has said it was planning to seek stronger protections for children from harmful social media features. Social media companies generally oppose blanket bands saying they have measures already in place to protect younger users, including age restrictions, though

1:02:22they have also said they would comply with government bands. Google meta snap and Tik TOK did not reply to Reuters response or requests for comfort comment. Macron, who in April urged teenagers to turn off their devices and read that went over really well in order to become better citizens, has ordered prime minister Sebastian, uh, Lake Cornu to rework the draft legislation, to make the constitutional council's

1:02:52concerns, uh, to take them into account. Um, they said in a statement, uh, that they were determined for the reform to take effect before the spring of 2027 when they hold, when France holds its presidential election. So anyway, Macron has not given up. The draft legislation is being hastily reworked to address the council's concerns since Macron still haps hopes to have this reform in effect as soon as possible. Um, and they, they, he was also going to add smartphone restriction to the legislation that would

1:03:28take effect for high school, uh, in addition to the, the lower schools. So anyway, we will see what's going

Zoom screen sharing security flaw

1:03:36on and, and what happens, Leo. Wow. Um, so, um, last Tuesday wired reported on the unnerving discovery of a serious vulnerability in the zoom teleconferencing system. And of course, now that's like recent, right? Remember, um, we'll all remember when zoom really became a big deal at the beginning of COVID

1:04:06because, you know, it saw its adoption sore as teleconferencing became super important. It was the only way to continue doing business. If you were stuck at home and boy, zoom had a bunch of early problems. Uh, they weren't all resolved as it turns out. Wired's headline reads a zoom screen sharing bug. Let anyone take over other devices on a call. And their brief teaser is what makes this so interesting.

1:04:39They said, researchers say it took fewer than 20 prompts for a public AI tool to find a flaw, which has now been fixed, allowing anyone on a zoom call to hijack other participants devices. And what, what wasn't clear from the reporting, and I didn't dig deep into it was, wait a minute, a public AI tool was used to do some sort of clear cyber security work without hitting guardrails.

1:05:14Probably a Chinese model. Uh, uh, could be, Oh, good point. Pub publicly available. Yes. So wired said as AI models gain advanced capabilities to find vulnerabilities in software, develop ways to exploit them, and even carry out autonomous hacking sprees, all of which we've been seeing researchers offered a sobering new example on Tuesday, disclosing vulnerabilities in the video conferencing platform zoom that could have been exploited to take over targets devices. Anyone on a call that involves

1:05:51screen sharing, whether participants or the host would have been vulnerable to a silent attack that could be carried out with no indication and no interaction from the victim. Researchers from the digital defense firm, a security, just the numeral a security say, or the, the, the letter a security say the bug was discovered in early June using publicly available AI models, and that it took fewer than 20 prompts to uncover the vulnerabilities and create a working attack. Zoom issued a security

1:06:28advisory advisory on Tuesday, including details about fixing, um, about the fixes. The company has already begun rolling out to address the flaws, which affected devices running across all operating systems that zoom supports windows, Mac, Linux, iOS, and Android, a security, the, uh, a security co-founder, the company, a security co-founder, Omar Gull told wired ahead of the disclosure, quote, what's interesting for us. And what we believe is dangerous is the Democrat, the democratization

1:07:05of these capabilities, the barrier to entry is dropping rapidly before it would have taken a team of five people, maybe six months with a lot of refining and iteration to find this. Now people can reach the same results with fewer than 20 prompts. And zoom is an important type of target because people assume trust when using it. They don't see it as a threat. His quote ends the vulnerabilities rights wired were

1:07:42found in the protocol used to facilitate real time annotation during screen sharing. The researchers say that their AI bug hunting systems specifically delved into this component because like human bug hunters, they've been trained that convoluted and obscure functions often contained overlook vulnerabilities. This is particularly true with proprietary closed source software. An established company like zoom

1:08:15presumably does extensive code review and vetting on all components and functions, but without the benefit of public open review, esoteric yet complex features like annotation are more likely to contain mistakes. Zoom did not respond to multiple requests for comment from wired about the a security findings. The bugs are now patched with zoom issuing both server and client side fixes or patches for both

1:08:46zoom's own servers and the applications that run on customer devices. But the researchers emphasize that it was alarming to contemplate bugs that could have been exploited to take over a target device simply by getting someone on a zoom call. Joining a call is itself a gesture of trust. But given how ubiquitous video calling is in both personal and professional contexts, and given that zoom in particular is also widely used for

1:09:16events and semi-public activities like webinars, people typically have their guard down when joining a zoom. A security co-founder, Yasi Torati told wired on a call. If you just get on a zoom with us, we can take over your device. The worst case scenario is that we can take over an enterprise just by having this capability in our hands. If I'm an attacker, I can be on a call with someone from a company, take control of their computer and their credentials, and then use them to move

1:09:52laterally across the enterprise. Practitioners often call security a cat and mouse game. But as AI bug hunting proliferates, this delicate dance has become an all-out race, which of course is exactly what we've been seeing. All indications are that the high-tech computer world at every level from IoT embedded device to consumer PC and enterprise probably is heading for a rough patch. So far, as we know, I've been a cheerleader

1:10:27for the fixing the bugs team. And the good news has indeed been that an astounding number of bugs are rapidly being found and fixed. In those last two versions of Chrome, more than a thousand total between those two, 149 and 150. But that's also the bad news. Because the fact that so many latent problems are being found tells us that the software at every level that we've been living with for years has been

1:10:59demonstrably riddled with previously unknown flaws. So the best that could be said is that the future remains stubbornly uncertain. We're getting the bugs out of the software. And my feeling is absent a state actor

1:11:21having some reason to attack another country, money is still what drives the bad guys. Now that we've got cryptocurrency, now that we've got the ability to exfiltrate and extort, money is the motive. And so there's not money behind a mass casualty event. There's money behind selectively targeting, exfiltrating, blackmailing, extorting, and getting what cash you can. So I don't think we're going to see

1:12:01see a big Y2K style, I mean, or a Y2K apocalyptic sort of thing. To me, that doesn't make sense, because it doesn't make money. And that's what, I mean, that's kind of a saving grace. It means that there will be people hurt, but their insurance is going to go up and they're going to be paying out of pocket for bad guys having gotten into their system using bugs that are probably latent and aren't their fault and are zero days. So there's really nothing they could do about it. So hopefully that's

1:12:36what we see. And that over time, the ability to do that dries up because we get our software fixed.

Cybersecurity stocks jump after Black Hat

1:12:46Okay. Uh, one last piece before we get into our big topic. Um, there are some folks who are, I would argue, justifiably profiting from all of this, uh, bad guys, not justifiably profiting, but good guys. Last Monday, following the black hat conference, CNBC reported, they said, cybersecurity stocks of CrowdStrike and Palo Alto networks jumped more than 5% to new highs on Monday

1:13:20on renewed demand for artificial intelligence, security tools, following the industry's annual black hat conference in Las Vegas. In other words, those guys were there, they were showing off that their, that their AI is going to be used now that they're ready to deploy, uh, defensive AI solutions. And the world said, we need some more of that analysts at BTIG, which is a large financial

1:13:51services firm wrote to their clients in a, in an internal, you know, firm letter quote, the single most consistent theme across our conversations, partners, vendors, and customers alike was that AI agents have fundamentally changed the threat landscape. While AI agents have become the predominant attack threat and the environment is meaningfully worse deployment and deployment and

1:14:23AI security tools are only in the early innings. Uh, CNBC said businesses are turning to cyber security companies for new agentic tools to fend off adversaries in a hyper accelerated threat landscape fueled by new cyber models. Executives and potential customers gathered in Las Vegas last week in search of answers and ways to secure systems from rogue AI agents. BTIG wrote quote, we think AI is creating a new

1:15:00modernization cycle in the endpoint security space, which directly benefits CrowdStrike's core business. Analysts at Cantor said quote, AI has moved from being a cyber security feature to a key pillar of both the attack surface and the defender infrastructure. So, um, I remember the first time we touched on this, I think it was one of our, it was, it was one of our listeners who was commenting that his company was,

1:15:36was already using some AI based systems. Um, I think he was, he was on AWS cloud and he was using a third parties, uh, AI based, uh, systems. And I was, this was a couple of weeks ago and I was like, wow, this is happening already. I mean, we've got AI on the, uh, you know, being deployed for defensive purposes, which is wonderful. And, you know, Leo, we're going to now dig into what I have learned recently about AI and

1:16:13can share. But first I think we need, uh, to share a fine sponsor, perhaps. I think that'd be perfect. I think I can do that. Our show today. Well, I'm excited about all of this. Yeah. Have a beverage. Looks like Tang. Are you drinking Tang? Oh, that's right. Dilute orange juice. Yes. Exactly. Yes. Do you add vitamin C to your diluted orange juice? No. Okay. I take 15 grams of C a day. So way more.

1:16:47Plenty. I know. I know. 15 grams. Yep. Holy divided doses of five grams. That's half an ounce.

1:16:56You're crazy. You're a mad man. Are you sure? At least you, well, I don't know. Whatever that means. I don't know what vitamin C is doing for you. You don't get any sore throats. I guess my, my liver would like to be generating about 20 grams a day and it can't because it's got, uh, the human genome has a little glitch. So, right. I mean, we've talked about that. Yeah. Yeah.

1:17:17Damn genome. Our show today brought, brought to you by Doppel. AI has made social engineering attacks. Oh man. More convincing than ever from phishing emails to fake websites and impersonation attempts. It's becoming increasingly difficult to tell what's real from what's designed to deceive. That's why organizations need more than a collection of point solutions. They need a unified approach to stopping attacks before they reach their people. Doppel does it. Doppel is an AI

1:17:50native social engineering defense platform. Doppel strengthens human risk management by training employees to recognize deception. It provides digital risk protection across every channel and delivers agentic email security that doesn't just score the inbox, but takes down the attacker infrastructure behind the message. Doppel protects against the entire social engineering attack chain with one comprehensive platform. You get digital risk protection, which detects

1:18:22threats across multiple channels, links alerts into a real time threat graph, and uses AI driven infrastructure disruption. I love this to stop attacks at the source. These insights also power phishing simulations and security awareness training. This thing, this thing is a cannon helping strengthen employees defenses through next generation training and testing. Email security inspects every message, traces it back to the attacker infrastructure behind it, and helps take that infrastructure down

1:18:56so that campaign can't target your organization or any other organization ever again. Doppel also offers best-in-class integrations and partnerships, making it easy to work alongside your existing security stack. So what you've got works with Doppel. Join hundreds of companies already using Doppel to protect their brand and people from social engineering attacks, the worst kind of attacks out there. Doppel, outpacing what's next in social engineering. Learn more at Doppel.com. That's D-O-P-P-E-L, Doppel.com. We thank them so much

1:19:33for supporting security now. That's an intriguing product. I could get these guys on the show. Fascinating. But now back to Steve. So, okay. As I promised last week, there are two important and

Dual use problem and auto-complete

1:19:48fascinating pieces of core AI technology I want to spend time on today. I got to one of them. The first is a solution to what's been dubbed the dual use problem. That's the fancy name given to the fact that most knowledge can be used in ways that we both want and don't want. And of course, that's no surprise, right? Since that's always been true of knowledge. So what is it about AI that

1:20:18changes this? Anyone who's spent any time with any of the recent state-of-the-art AI chatbots will have come to appreciate that what we already have today is at the very least an over-obliging conversational partner that can barely restrain itself from being, oh, so very helpful. And that tail wagging puppy happens to also have access to the world's stored knowledge. So it's not that it was

1:20:52impossible before AI to obtain that knowledge the old fashioned way, you know, by researching, reading, learning and understanding. No, the difference is that friction matters and AI chatbots have hugely facilitated the access to that same knowledge by nearly eliminating all of the work that was previously required to gain such knowledge. And that's incredibly valuable. That's what underlies all of this

1:21:27frenzied hyperscaler data center build out. Investors believe with good reason that offering knowledge at our fingertips by phrasing a question is something most of the world will pay for. The fact that AI chatbots now have just shy of 1 billion users strongly suggests that these investors are not wrong. I, for myself, I, for myself, I'm completely spoiled, even though I've never turned any agent yet loose on anything. Claude is my go-to for quick answers. Uh, it's an accelerator for me. Um, and you know,

1:22:05don't let anybody know, but I, at this point I would probably pay pretty much anything for it. Yeah, because yeah, I know you, I know how you feel. I know how you feel. Yeah. Yeah. I have paid anything for it. But, but Leo, you and I already know we're not going to have to, because as I mentioned, I think before the show, it turns out that that, that Lenovo think station machine that I bought had a strong GPU. It's got a, it's got an NVIDIA RTX 2000 with 16 gig, and there are useful models

1:22:40now that can run in that. Yep. So, you know, but I'm, you know, still, you're always going to want the latest and greatest and the best answer. It won't be as good as Claude. I mean, that's the thing and not yet next year, and then you can stop. Then you'll be happy, right? So, so, you know, by comparison, right, you know, anyone could download the Encyclopedia Britannica or the unabridged Oxford English Dictionary, but you can't ask them questions. Yeah. You know, it's all just

1:23:10dead knowledge lying there. So it's much less entertaining and it takes a lot longer, you know, so, you know, back and using them as back to old school, researching, reading, learning, and understanding. And of course, you know, look behind me. Anybody who has seen a video of this podcast is aware that there's a solid wall of textbooks. And as it happens, uh, up there, just out of camera is an unabridged Oxford English Dictionary, 27 volumes. I cannot recall the last time I opened any of those

1:23:48books back there. You know, you're right. AI neural networks are our new store of knowledge, incredible as it may seem. And it would have seemed like science fiction just a few years ago. Um, the collection of just an array of scalar variables, which specify the scaling weights of the inputs to a vast neural network creates a representation of the expression of all of the

1:24:23knowledge that has been trained into that model. Um, and I said that exactly the way I wanted to creates a representation of the expression of all of the knowledge that's been trained into that model. I think it's important to frame it that way. What we're feeding into the neural network while it's being trained is the expression of the knowledge largely gleaned from the internet, as well as from reference texts, which AI companies have been quietly purchasing and ingesting, uh, to create an overall

1:25:00training corpus. The end result of this, and this is still mind boggling to me is purely and simply a next most likely token prediction engine. You know, Leo, you're very, as, as I've said, your very early initial observation was that what we were calling AI, and this is a couple of years ago was little more than fancy spell check. I called it spicy auto, correct? Yes. Um, and then as our

1:25:33listeners will remember last week, I, I quoted a nearly stunned Matthew green specifically saying that it's not just fancy spell check, but it's essence has not changed. You know, you were not wrong, Leo then, and Matthew is also not wrong today. So how do we explain this apparent disparity? It's that over the past couple of years, what was initially a simple next word predicting spell check

1:26:11has become really, really, really, really, really fancy spell check, um, deep underneath, even today's astonishingly, apparently intelligent reasoning systems down at their core is still just a neural network that only does exactly one thing. Given a long, and in many cases, astonishingly long token sequence,

1:26:43it predicts the most likely next token. And the fact that we get what we now get from that, well, that's, what's still mind boggling to me. Um, I'm going to spend a bit more time on this because a deeper understanding of the truth of what's actually going on with today's AI, I think will help everyone to appreciate next week's, the second of the two research papers I plan to share, which

1:27:18is, well, I, I don't know how to sum it up quickly. So stay tuned. So your original Leo, you know, it's just spell check autocomplete summation that has its roots in AI circa 2020 way back then. If you were to carefully phrase a question to GPT three, such as the capital of France is,

1:27:48it would have been able to complete the sentence by emitting the next expected word, Paris, because the model that that models, the GPT three models, vast statistical data set, um, made, uh, made Paris the next most likely word, the capital of France is Paris. But if you used question style phrasing back then, what is the capital of France that would not have been met with

1:28:26the same success? So what happened? Obviously we have that now. How did we turn these models from autocomplete engines? That is, that's where that's all they could do into conversationalists. So it took a few years of experimentation, but AI researchers first used something that's now known as instruction tuning. They took the knowledge trained model and fine tuned it on a, and this is

1:29:00again, another surprise, a surprisingly small set of human written examples of what good responses responses would look like. After that, then something known as RLHF, reinforcement learning from human feedback was used. So this applied preference rankings where again, humans compared and ranked multiple outputs from best to worst. And that ranking was used to train a reward which pushed the model

1:29:40toward the better behavior. Now here's the astonishing part to me. Researchers then found that they could employ a comparatively small model of these query and ranked response samples, which then created reward feedback. And that these large language models would and did with startling speed. They were able to

1:30:13generalize from the language patterns of queries and responses across their entire knowledge base. that is that they, they, they learned that pattern and generalized it. So to better appreciate the scale of this, um, a model that had been trained on trillions of tokens of raw knowledge, you know, this created the

1:30:45original statistical auto complete engine, a neural network that just could, could probabilistically choose the next most likely, choose the next most likely token. It could then be rewarded using only again, it was originally fed trillions of tokens. It could be rewarded using only tens of thousands to low hundreds of thousands of query response samples or examples.

1:31:16And the behavior of the entire model would be reshaped across all of the knowledge that it contained. Um, this is a well-documented fact now that, that, that this occurs, um, with, and the fact that it occurs with such sample efficiency is what was utterly unexpected. I mean, this was the surprise today. Now, as I said, it's a well-documented real phenomenon, which is now

1:31:52been turned there. We have a term for it. Now it's known as the superficial alignment hypothesis. The idea that all of the raw knowledge was already there from the model's initial knowledge corpus training, then a relatively minuscule bit of fine tuning teaches the model format and behavior, but not new knowledge. That is, it wasn't giving it new knowledge. It was, it was, it was reformatting it

1:32:27and giving it behavior for the first time. So to state that a bit differently for clarity, the breakthrough about four years ago was not the use of a larger model than we had at that time. That is one of the things that's been happening since then. But the breakthrough was the unexpected discovery that a comparatively infinitesimal dose of human provided, here's what a good answer looks

1:33:01like. And here's which of these answers is better. Training would reshape a massive already trained model's entire behavior, turning it from something that completes patterns into something that acts like it's trying to be helpful. So once the massive model learned what helpful looked like and that its

1:33:33trainers wanted it to look like that, all of its stored knowledge was immediately available in that new helpful format. And we got helpful chatty AI and that's how it happened. Um, and as we know, those first steps made by chat GPT were more than a little shaky, you know, uh, you know, researchers realized that their new, their newly birthed chat bot would need some additional post-training alignment as it's now being called.

1:34:09And this began as that RLHF, the reinforcement learning from human feedback that we talked about. Uh, but then a year later in 2023, a handful of AI researchers at Stanford university published a paper

Direct preference optimization and DPO

1:34:27titled direct preference optimization. Um, the paper's title is direct preference optimization. Your language model is secretly a reward model. Um, and since all of the descendants of this, it's now known as DPO system, direct preference optimization was sort of the granddaddy. Now we have refi further refinements of that, which have occurred in the last three years, uh, something known as IPO, there's KTO, ORPO, and SIMPO. They all descend from DPO.

1:35:03I want to share just the abstract of the Stanford researchers original paper, which, which was the, the breakthrough beyond that earlier reinforcement learning from human feedback. Um, so they explain their invention by writing while large scale unsupervised language models learn broad world knowledge and some reasoning skills, achieving precise control of their behaviors difficult due to

1:35:37the completely unsupervised nature of their training, existing methods for gaining such steerability, collect human labels of the, you know, feedback of the relative quality of model generations, you know, model output, and fine tune the unsupervised language model to align with these preferences, often with reinforcement learning from human feedback, RLHF. However, they write, RLHF is a complex and

1:36:11often unstable procedure, first fitting a reward model that reflects the human preferences, and then fine tuning the large unsupervised LM using reinforcement learning to maximize this estimated reward without drifting too far from the original model. In this paper, we introduce a new new parameterization of the reward model in RLHF that enables extraction of the corresponding optimal

1:36:46policy in closed form. I have no idea what that means, but you get us, you'll get a sense for this, allowing us to solve the standard RLHF problem with only a simplification, classification, a simple classification loss. The resulting algorithm, which we call direct preference optimization, is stable, performant, and computationally lightweight, eliminating the need for sampling from the language model during fine tuning or performance significant hyperparameter

1:37:22tuning, whatever that is. Our experiments show that DPO can fine tune language models to align with human preferences as well as or better than existing methods. Actually, it's vastly better. It completely obsoleted everything that came before. Notably, fine tuning with DPO exceeds PPO-based RLHF, inability to control sentiment of generations to control sentiment of generations, and matches or improves response

1:37:53quality in summarization and single-turn dialogue while being substantially simpler to implement and train. So, okay, that's just their abstract, and the paper goes on at length, like with crazy formulae. I wanted to share that because I didn't want to leave everyone with the belief that the original RLHF, the reinforcement learning from human feedback approach, which began this transformation from

1:38:23auto-complete to actually Q&A, that that's what the industry was still using. As I said, it was the genesis. Stanford's DPO, their direct preference optimization, dramatically improved it. And DPO's success, as I said, spawned a number of successors, which I cited earlier. Okay, so what is all this about?

1:38:49It's about how we take a massive neural network, which has been trained on and contains raw knowledge, which can initially only be used to predict the next most likely token, and impress upon that network actual behavior. That's what's changing here. We're giving this knowledge base a set of behavior. The first example of behavior was turning the network into something that could actually respond

1:39:21to queries. That gave us the first interactive AI. But as I noted, those first steps were somewhat shaky. So over the time, researchers learned that by using the technologies I just described, they could improve the model's instruction-following behavior so that it would answer the question that was asked, as well as respect the requested format and length and language. The model's style and tone

1:39:52could also be modified and tuned to improve its response formatting, structure, hedging, politeness, and its own verbosity. These factors were, again, they were given significant weight in the tuning. And as we saw in the early days, psychophancy was an often-seen problem. This arose because you know, we humans reliably prefer being agreed with. So preference optimization trains models

1:40:26toward agreement, and it takes deliberate counter effort to prevent it now. So that's now in place. We've been able to kind of get that under control. Another improvement that was impressed upon models was factuality, preferring responses that will admit uncertainty rather than always confident fabrication. Well, that explains a lot because hallucination, at least in my experience, has almost disappeared. And I was wondering how they did that. Now we know RPO.

1:40:57Yes, exactly. So the other class of behavior that can be trained into a model is its refusal to provide certain classes of information. And what's significant is that this exists in two places. This is different than guardrails. There's filtering what you allow the human prompter to get into the model, and then filtering what you allow the model to show of its results. So that's

1:41:32real-time filtering input and output. But the other place is you can actually train the model itself to refuse, independent of input-output filtering. So, you know, which is to say the harnessing of the model. So it's one thing for a model to contain knowledge that's dual use, which, you know, only authorized users should be able to access. But some model behavior and or knowledge dissemination should be proactively

1:42:08prevented. The test that AI designers apply is termed uplift. That is, does a model meaningfully advance someone's capability beyond what they could already get? Does it uplift them? This falls into two categories. One is where the knowledge is the harm, and the other is where the output is the harm.

1:42:39Examples of knowledge uplift would be, for example, bioweapon synthesis routes, nerve agent production, and nuclear device design. You know, the relevant literature is scattered, it's partial, and it's difficult to assemble. So any model that would synthesize it into an actionable protocol provides genuine uplift toward mass casualties. And the asymmetry is clear, right? Defensive work in

1:43:16these fields does not require the synthesis route. Somebody, you know, a vaccine researcher needs to understand pathogen biology, not a means of enhancing a pathogen's, you know, malicious use. So this list and those examples that I cited, you know, would not take anyone by surprise. They're one category where essentially

1:43:49every AI model provider refuses regardless of credentials or system prompt. That is, it's not about who you are, what your privileges are on the model. You just can't have that. It's been trained. The model itself has been trained to refuse. Trained in what way? Do they not have the information? We're, I'm, I'm, we're, we're going to get there. Oh, good. Great. That's, that's perfect. By the way, this is fascinating. You know, when the, the, everyone, the world became aware of this,

1:44:23I mean, this, the knowledge of this has been around papers and so forth for some time, but it was when DeepSeq came out from China, the first, it was the first model to use RLHF, as far as I know. And it was an eye opener for people because the model was stunning. This was January of last year. I remember it very well. It was a DeepSeq moment. That's when all the stocks of all the AI companies plummeted because people said, wait a minute, the Chinese can do this cheap. Yeah.

1:44:55Very interesting. Now everybody uses RHLF, of course. Yep. Yeah. Yep. Really interesting. Now actually would be a great time to take a break. We're a little after an hour and a half in, so we're going to pace ourselves. This is fantastic. Yeah. And you found this by reading papers. Yeah. I've, I, yeah, I found a .org or yeah, they're all there. Yeah. Yeah. Yeah. Jeff loves reading those papers too. There's a lot of garbage there too. That's my only, but you're, you know what to look for. This is like, yeah, these were the found, the, the, the, the, the, the, the sound,

1:45:28the founding documents of it. Yeah. Yeah. Yeah. Yeah. Absolutely. Fast. Yeah. Like, like, you know, three, three AI researchers at Stanford that figured out, you know, how, how, how to improve on RLHF so that that's what everyone is using now. Right. That's what you want to look at. It's, you know, the other thing that I find amazing is there are breakthroughs like this happening almost all the time now that there are researchers all over the world working as hard as their little research brains can to find new techniques.

1:45:59Yes. That's why I keep saying today's AI, today's AI. I mean, anybody looking back, well, like just quarterly look back in, in three month hops and you're, it's clear where we're, we're, we have, we're onto something. Oh yeah. And the, the other thing is that I, from the outside can, I can put flags in the ground and exactly when, you know, January, 20, 25 is when, uh, deep seek came out and everybody said,

1:46:30Oh, RL, RLHF. Oh, uh, November, 20, 25, when, when Opus 4.5, we're going to, I'm sure get to that. So these, these internal, uh, uh, progress shows up in ways that those of us who use this heavily can see those milestones, the impact of those milestones. I mean, it's, it's very clear. Uh, you don't see hallucinations like you used to. And, and I don't, I don't know why I'm glad you're explaining this. Um, uh, you also see sick of fancy going down a little bit, although it's,

1:47:05this is the other thing is the companies are loathe to get rid of the stuff that makes people like me keep using their products. Right. So they're not going to get rid of all of that. They're not going to get rid of all of that. Uh, just fascinating stuff. All right. Well, we'll be back with more. I don't, I hate to interrupt, but you know, we got to pay the bills. Uh, I got to keep my, wet my whistle. Yeah. Wet your whistle. And of course we know this company cause they are the people who brought us to black hat a couple of weeks ago for that fun conference. That was a lot of fun. We were in the threat locker booth, our sponsors, threat actors these days,

1:47:39you know, this, we talk about it all the time. They're using AI in so many ways, but one that's really scary is they're using it to automate vulnerability discovery. That's why every company from, from Google to Apple to Microsoft are shipping bug fixes as fast as they can. But you know, the bad guys are actually probably a little faster. Unfortunately, they're also using it for other things there when they, once they get in, once they penetrate, they actually are modifying the scripts during the attack because AI can work so fast. It's like

1:48:11an attack that modifies itself as it progresses. They're generating new malware variants faster than anybody can keep up. They're coordinating activity across multiple systems in ways, uh, you know, only the best, uh, uh, you know, gamers could do in the past. And now tasks that once took a human hours or days can happen in minutes. And that should be scary if you're a business because you're at the front lines. And at the same time, you, you may be doing this. Most organizations are introducing AI

1:48:44assistants and agents and giving them access to the crown jewels, documents, source code, cloud applications, APIs, internal systems. This is a, if you think about it, a nightmare just waiting to happen. Security teams need to know which AI tools are in use, what information they can access, whether they're operating outside their intended scope, a successful login or an unfamiliar file hash. That's

1:49:15not going to give you enough context. Teams also need to understand whether an application is behaving normally, whether it's accessing unexpected data, like, uh, the open AI model that was looking for the answers to the quiz questions. That was weird. Or communicating with systems. It shouldn't be able to reach. The problem is you can't use old fashioned methods to stop this. You need to use ThreatLocker. You need to use ZeroTrust. ThreatLocker uses application allow listing. Very granular, by the way.

1:49:52So it can control which AI tools and other applications are allowed to run. They call it ring fencing. They use ring fencing to limit which, what approved applications can access. So you approve the application, but you need the granular permissions to say, well, they can do this, but they can't do that. Which processes they can launch, how they can communicate with one another. ThreatLocker uses something they call web content controlled. This is fantastic. Manage access to public AI platforms and other online services. You can block it entirely or limit what they can do

1:50:24with privileged access management. That prevents AI applications and their users from receiving unnecessary administrative privileges. ThreatLocker applies ZeroTrust network access and ZeroTrust cloud access policies to restrict resources to authorized users, approved devices, and permitted applications. It started with endpoints, but now it's in the cloud. It's in the network. And ThreatLocker works everywhere you work. Windows, Mac, Linux. They've got the best support 24-7, U.S.-based support from

1:50:56people who really know what they're talking about and are there to help you, to really support you. That's why ThreatLocker is trusted by organizations like JetBlue, Heathrow Airport, the Indianapolis Colts, the Port of Vancouver. Ask Jack Thompson. He's Director of Information Security Risk and Compliance for the Indianapolis Colts, the great NFL team. He said, with ThreatLocker, we have the ability to centralize disparate elements in the security stack. And what does that mean? That means they can control them. They can see what's happening. That's another side effect of ThreatLocker's ring

1:51:30fencing. It's great for compliance because you know exactly what happened, when, who did it. You've got it all. And of course, ThreatLocker's beloved in the industry, constant awards and prizes, just some of the latest industry recognition. They were recognized as Strong Performer in the January 2026 Gartner Peer Insights Voice of the Customer for Endpoint Protection Platforms. They were ranked number one in application control by Peerspot. They're the winner of the best zero-trust security solution at the 2025 TICE Awards. AI governance requires more than an

1:52:04acceptable use policy. You need a padlock. ThreatLocker gives security teams the technical controls to define which AI tools are approved, who and what can access them, and how those tools are allowed to interact with business systems and data. You have control. And that's what it's all about. Visit ThreatLocker.com slash twit to get a free 30-day trial and learn more about how ThreatLocker can help mitigate unknown threats and ensure compliance. That's ThreatLocker.com slash twit.

1:52:39We're big fans of ThreatLocker. We really are. Now, I'm a big fan of Steve Gibson, who is finally explaining something I've been using for a year or two and had no idea what was going on under the hood. We all have been. Yeah, it's fascinating. So, knowledge is, you know, forbidden knowledge is one class. The other is where the output itself is the problem or the harm. So, an example of that,

1:53:09you know, which carries universal condemnation would be, you know, text which sexualizes minors. Models will not produce any such text. It's trained out of them. See, that's fascinating because we've heard about classifiers, which kind of are gates preventing the aggress of that information. But this goes deeper than that. The models themselves say, no, no, no. Yes. So, and that's why, unless you remove that, even without any kind

1:53:43of a harness, even without, you know, anything that is filtering input and output, the model says, sorry, I cannot help you there. Right. So, as we've seen earlier, the good news is that these large language models are astonishingly able to absorb and embrace. We saw this in like their ability to learn to answer questions. I mean, to understand the linguistic nature of

1:54:14a question and how to apply the knowledge they had to create an answer. I mean, it's an astonishing technology, but it does this. So, we're able to give them, as a consequence, what appears to be a personality, many forms of behavior, and also instruct them what they may and may not do. So, that's the good news. The bad news, as it turns out, is that any behavior like this that can be

1:54:51easily imprinted can also be easily removed. In the summer of 2024, researchers, researchers at ETH Zurich, the University of Maryland Anthropic and MIT published a paper titled Refusal in Language Models is Mediated by a Single Direction. And here, direction is a term of art in neural networks, as we'll see. So, the abstract of their paper employs some of, you know, of this

1:55:29inside baseball terminology, but everyone should be able to easily get the gist of it. So, and I'll explain a bit more afterwards. The paper's abstract, that is, refusal in language models is mediated by a single direction. The abstract reads, conversational language models are fine-tuned for both instruction following and safety. Safety meaning, not going to tell you that, resulting in models that obey benign

1:56:01requests but refuse harmful ones. While this refusal behavior is widespread across chat models, its underlying mechanisms remain poorly understood. Again, this was summer of 2024. So, about just about two years ago, this paper appeared. In this work, across 13 popular open-source chat models, up to 72 billion parameters in size, we show that refusal is mediated by a one-dimensional subspace. Specifically,

1:56:41for each model, we find that a single direction such that erasing this direction from the model's residual stream activations prevents it from refusing harmful instructions, while adding this direction elicits refusal on even harmless instructions. They said, leveraging this insight, we propose a novel white-box jailbreak method that surgically disables refusal with minimal effect on other

1:57:22capabilities. Finally, we mechanistically analyze how adversarial suffixes suppress propagation of the refusal-mediating direction. Our findings underscore the brittleness, and this is the key, the brittleness of current safety fine-tuning methods. In other words, just instructing the model not to answer the question, well, that works. But if the weights are open, it turns out to be trivial to remove those

1:57:57instructions even after the fact and not having known what the instructions were. I'll explain a little more. It's amazing. So they finished saying, more broadly, our work showcases how an understanding of model internals can be leveraged to develop practical methods for controlling model behavior. Okay, so what this group discovered was that any late-term model behavior imprinting can later be

1:58:31removed from such models. And the way this is done, as I said, it's wonderful and wild. They compare the model's activations on harmful versus harmless prompts. Compute the mean difference, then project the weight matrices orthogonal to that direction. So essentially, they deliberately ask it questions

1:59:04it is trained not to answer. And they watch some of what it does, some of where the activations are, compared to asking questions that it's happy to oblige. And they're able to take the difference in the activations, see them, and then apply a remover that suppresses that, and suddenly it will now answer all questions. So after they do that, the model loses the ability to represent and therefore to execute

1:59:44on its refusal. And what they found was that this was surgical. It specifically disables in completely 13, completely different chat bots, all of their refusal while having minimal effect on other capabilities. So it's kind of like a functional MRI on the model. Like you're looking for what got activated. Yes. And then remove it. The only thing I'm worried about, you know,

2:00:21for instance, I use a model from China, Quinn, as we mentioned, 3.827. And I've seen obliterated versions of it, but people say, well, you also have a risk. This is brain surgery. After all, you might make it dumber. Um, so I can't speak to that, but I can, I can cite the research because they, they do address this. So, okay. So first from my lay view, it appears clear that since little

2:00:51training, amazingly little training was required to imprint that original refusal behavior. Ah, no, I get it. The impact upon the model was minimal. You know, it wasn't diffused throughout the entire model. It, it had a, but, but it's still, it's significant that this is able to change its behavior. So think about it. You change, you're modifying with just a few inputs. Yes.

2:01:24A large model, you know, there's going to be some collateral neurons that are affected. Well, and actually I use the, uh, the, uh, analogy a little bit later of clean margins when, when, when, when a surgeon excises something. So, right. Anyway, so these guys discovered how to identify the changes created by that training, which again, wasn't pervasive. It was, you know, not that much training did comprehensively change the model's behavior. So it turns out

2:01:57it could be removed. So, okay. So for what it's worth, when we encounter the term obliteration, this is what is meant, you know, not obliteration, a obliteration. So just to put a final point on it, a, a, a hugging face blog posting in the summer of 2024, that is, you know, following this research was titled uncensor, any LLM with obliteration. And I'm going to share just the intro from that

2:02:32posting to give everyone a sense for what that, that earlier work for where that earlier research led, which is here. The blog says the third generation of LLM models provided fine tunes, then it says in friends, instruct versions that excel in understanding and following instructions. However, these models are heavily censored designed to refuse requests seen as harmful

2:03:04with responses such as, as an AI assistant, I cannot help you. While this safety feature is crucial for preventing misuse, it limits the model's flexibility and responsiveness. In this article, we will explore a technique called obliteration that can uncensor any LLM without retraining. This technique effectively removes the model's built-in refusal mechanism, allowing it to

2:03:40respond to all types of prompts. The code is available on Google Colab and in the LLM course on GitHub. So then it says, what is obliteration? Modern LLMs are fine-tuned for safety and instruction following, meaning they are trained to refuse harmful requests. In their blog post, Ardidi et al., and that is the,

2:04:10that is the previous research that I was referring to, the research in the summer of 2024, Ardidi et al., have shown that this refusal behavior is mediated by a specific direction in the model's residual stream. If we prevent the model from representing this direction, it loses its ability to refuse requests. So that Ardidi reference, as I said, is the research I referred to previously, which showed the world

2:04:44how to do this, just how to simply perform this excision. The blog posting on Hugging Face is one artifact of that preceding research, and the companion repository on GitHub contains all of the details.

2:05:03That is, it's M-L-A-B-O-N-N-E dot GitHub dot I-O, and it is uncensor any LLM with obliteration. So, and as you would expect, it was all tremendously exciting to the world's AI hackers, many of whom immediately jumped on any and every published and available open-weight model and happily obliterated

2:05:34away any and all perceived and imposed censorship upon those models. And those obliterated public open-weight models are now available for use by anyone who can harness them. And as you said, Leo, you've seen them, both with and without the censorship obliterated from them. So, this finally brings us to the first major topic I wanted to discuss today, now that we have a much deeper

2:06:07understanding of where these chatbots came from, how they work, how they can have behavior imprinted upon them after they've been filled with raw knowledge, and how, unfortunately, brittle that imprinting is. So, okay, now I also need to mention that the term pre-training is what the AI industry has unfortunately landed on to actually mean training, the training that occurs before the post-training.

2:06:42And I suppose, since there is now always going to be a definite post-training phase during which the behavior is layered on top of the previously trained-in knowledge, you know, if that pre-training were just called training, which is actually what it is, then it might be assumed to encompass both the training and the post-training, meaning if it was just called training, it would mean all training. So, my point is, there's pre-training and there's post-training, and there is not any just training

2:07:20in the middle. We don't use that. So, now we know that pre-training, that's what it's called, pre-training is the initial

Off switch for dual-use knowledge

2:07:32knowledge corpus training phase.

2:07:38Now, now that we know that, I can explain that a small, competent team of AI researchers at AE Studio working with Anthropic have, they've proven the feasibility of a new form of AI model pre-training, that is, a new way of doing that initial knowledge capture.

2:08:10Last month, both AE Studio and Anthropic blogged about it, and they published a joint research paper. I'm going to start by sharing Anthropic's press release style posting, since it provides the essence of the research without dragging us too far down into the weeds. Under the title of, which is their title, an off switch for dual-use knowledge in AI models. And again, the issue here is dual-use, right?

2:08:48Nuclear bomb generation, you know, creation. Well, there's some knowledge that we want to be able not to provide. The problem is, we've seen how brittle the instruction of, do not provide that, given post-training is. It can be removed, and it has all been removed. It's been obliterated from all of the current open-weight models. So here is what Anthropic said.

2:09:18And they start by saying, this post describes research conducted by AE Studio in collaboration with Anthropic. They said, a frontier AI model is, among other things, a large store of knowledge. Some of that knowledge is dual-use, meaning it can be used for good or bad. For example, knowledge of cybersecurity can help patch critical security vulnerabilities, or it can be used to exploit them.

2:09:49Knowledge of virology can help a researcher create a vaccine, but it can also help a malicious actor design a deadly pathogen. Ideally, we would be able to balance three separate goals. First, limiting access to dual-use capabilities in as surgical a way as possible. Second, allowing trusted users to access those same capabilities for beneficial purposes.

2:10:20And third, doing all this without affecting the model's performance on any other task. Current safeguards are imperfect, they wrote. We train models to refuse harmful requests and use classifiers to screen inputs and outputs for dangerous content. These layers of protection guard against dangerous outputs, but they don't change the knowledge

2:10:50stored in the underlying model. Despite our safeguards, a sufficiently determined attacker may still try to jailbreak the model, working past its defenses to access the dual-use knowledge. A more robust protection against misuse would be to control what the model knows. We've explored this before, they wrote. In earlier work, we filtered information about chemical, biological, radiological, and nuclear

2:11:26weapons out of pre-training data, and later showed that dual-use knowledge can be confined to a removable slice of a model's weights. It produces one model with one fixed set of capabilities. Using filtering, if you want a model version that can discuss advanced virology for deployment

2:11:57in a vetted biosecurity lab, say, and another version that cannot discuss that because it doesn't have the knowledge, they say, you have to train two separate models, especially in the case of frontier models, which are large and very expensive to train, the cost to the developer would be prohibitive. So I'm going to interrupt here just to add that while OpenAI and Anthropic are not being currently forthcoming about the cost to train a current frontier model, Leo, you've always

2:12:35talked about how expensive it is, and oh boy, once those two are publicly traded, then their accounting ledgers will no longer be private. So we're going to find out. But to get some sense of scale, OpenAI's Sam Altman has stated that the training cost for GPT-4 was more than $100 million, and OpenAI reportedly spent $3 billion overall on compute to train

2:13:10their models two years ago, back in 2024. Also, as we know, models have grown much larger recently, and those numbers ring true, those earlier ones, because Google's Gemini model is believed to have cost Google $192 million to train. So we're talking- That's a chicken feed, though, compared to what they're spending now. I mean- Well, actually, yes, right, because these are old and smaller models, right?

2:13:44So it could be half a billion dollars. But Fable is, as some have speculated, a $10 trillion parameter model. Gosh. And ChatGPT-4 was, I don't know, several hundred million, probably. Right. Yeah. And so all of this leads us to understanding why it's not feasible for any commercial provider, anyone, commercial or not, to train, to have multiple versions of a single model type, like

2:14:24a mythos that knows nothing about cybersecurity. The advantage would be, you can't trick it into revealing what it doesn't know. The knowledge would have never been put in there. So it's just not there to ask. But you would spend so much money training that up. And this is the problem that they're trying to identify. So the quite valid point that Anthropic is making here is that it's massively infeasible to train a state-of-the-art model on any pre-filtered knowledge to make it dumb about some things.

2:15:04You make it just, I don't know anything about that. You know, if you're going to be spending that kind of money on training, it needs to know everything so that it can have the widest application range for its use, you know, in order to have some chance of getting some money back out of all that, that training that went into it. So if a state-of-the-art model were to be trained on a filtered subset of everything, then, you know, that is a limited one forever, you end up with a very expensive forever limited model.

2:15:41Okay. So we've seen that imposing post-training behavior, which is what we're just talking about, this refusal obliteration, post-training behavioral restrictions on open source models where you're able to modify the network weights, that can be altered. So it was a short-lived solution. The means for removing those restrictions, as we saw, it's all public knowledge, child's play,

2:16:13you know, and even the best closed-weight models, which are operated by cloud-based hyperscalers, you know, like OpenAI, Anthropic, and Google, and so forth, AWS, we've seen that they can be prone to trickery and abuse. We've, you know, prompt injection example, and we're next week, we're going to understand exactly how that happens. So what's clearly needed is a new solution. And that's what these researchers have found.

2:16:45Leo, we're at two hours. Let's take our last break. And then we're going to look at the solution for this dual-use problem and how to solve this. With a single training.

2:16:56We will have more of this. I'm just, by the way, absolutely fascinated by this. I feel like this should be required listening for the listeners of Intelligent Machines, our AI show tomorrow, because it's such foundational information about how all this stuff works. And it explains a lot, to be honest, about how these models work. And I think it's a good idea to kind of understand the underlying technology, because it makes, it means that you could do a better job.

2:17:27And again, as a user, you'd, you know, Lori, she's using the heck out of ChatGPT, but our audience, that's why they're here, is to get this. And, you know, for advanced users who are looking at things like obliterated models and wondering why some models do this and some models do that, there's so much going on. This stuff moves so quickly. It's very helpful to understand it a little bit. Absolutely. I appreciate it. We're going to take a little break and come back with more.

2:17:59You're watching Security Now with the wonderful Steve Gibson. This episode brought to you by Scribe. Now, every time you onboard somebody new to your company, you know, you have to re-explain the same tools from scratch, because there's no documentation. The result is fragmented processes, inconsistent execution, and no reliable way to improve how work gets done over time. And that's what today's sponsor, Scribe, does best. Scribe is a workflow AI platform trusted by 94% of the Fortune 500.

2:18:37Scribe captures any workflow in real time and turns it into documentation automatically. No manual writing, no manual screenshots, no starting from scratch. Every time someone new joins a team, you just do the process as you normally would. And Scribe automatically captures the workflow as it happens, including the steps and screenshots, and then it turns it into a guide. What would have taken hours of writing, recording, and cleanup is done in under a minute. And it's ready to share.

2:19:08Scribe automatically redacts sensitive information, names, account numbers, emails from every screenshot. As an admin, you can enforce this across your entire team. You'll feel confident nothing slips through the cracks. Anyone following this process can launch real-time on-screen guidance that shows them exactly where to click step-by-step inside the actual tool. It's there the moment the Scribe is created and ensures everyone does critical processes correctly. Using a feature called improve workflows, Scribe will even suggest improvements to the underlying workflow, not just documented.

2:19:44Once you can see how a process is actually being done, you can identify where steps are redundant, where people are getting stuck, where things could be automated or simplified. It's not just capturing how work gets done. It's helping you do it better. To see what Scribe could look like for your org, head to scribe.how slash security now. And mention Security Now for your first month of Scribe capture free on select plans. That's S-C-R-I-B-E dot H-O-W slash security now.

2:20:19It even rhymes. Scribe.how slash security now. We thank Scribe so much for their support. This is a really interesting application of AI, isn't it? It's doing what a human would do very laboriously and doing it instantaneously. It's kind of amazing. Anyway, on we go, sir. So because post-training behavior modification has been demonstrated to be easily removable, it is not sufficient.

2:20:52To say, don't give people these, you know, disinformation. It doesn't work to suppress that behavior. What we need is a new solution. And it's completely infeasible to do multiple training runs with different combinations of information filtered out of a model because training is prohibitively expensive. So what we need is a new solution. And that's what these researchers working with Anthropic have found.

Gradient routed auxiliary modules

2:21:24Anthropic continues their writing, saying, in new research carried out with collaborators at AE Studio, we explore a new method that could enable the benefits of training many separately filtered models, but at the cost of training only one. We call it GRAM, gradient routed auxiliary modules. Note that they wrote, note that the results of the experiments presented here are preliminary.

2:21:59GRAM has not been applied to any of the production models and Anthropic. And they said, and we're not sure it ever will be. And I, okay, so I take that to reflect a very reasonable and very cautious approach because can you imagine the cost of a mistake if one of these companies' GRAM-trained model, whatever that is, and we'll get to that in a second, turned out to have unsuspected problems? Uh, there's no reason to believe that it would, or that that would be the case, but again, their cost is understandable because they could be scrapping the half a billion dollars these days.

2:22:38So, they write, Anthropic writes, how GRAM works. The idea behind GRAM is to give a model dedicated removable compartments for each category of dual-use knowledge and to update only those compartments when learning from dual-use data. Again, to update only those compartments, not all of the models' weights, only those in the compartments when learning from dual-use data.

2:23:16They said, concretely, GRAM adds extra neurons to every layer of a standard transformer. You know, the neural network architecture on which large language models are based. These neurons are divided into groups or modules. One module per dual-use category. During training, when the model encounters general-purpose text, that is, you know, just standard text. We don't worry about it one way or the other.

2:23:46It learns in the usual way. But when it encounters text from a dual-use category, virology, for instance, they wrote, The rules change. The model could use its general knowledge to make predictions, but only the virology module is allowed to learn from that text. The general-purpose weights are temporarily frozen. So, in other words, the bulk of the model isn't changed by learning about virology.

2:24:26Only the neurons in the virology module are changed. And if the bulk of the model's weights don't change after learning about virology, it doesn't learn about virology. It doesn't gain any knowledge from that. It's like it never happened to most of the model. They said the consequence is that virology knowledge accumulates in the virology module rather than diffusing across the whole network.

2:24:58After training, the module can simply be deleted, and the virology knowledge goes with it. Or it can be left in place for trusted deployments. When virology knowledge- I give it a lobotomy. Yeah, exactly. No, it's exactly. What's interesting about this is you would think, well, that's how all knowledge is, but it isn't. It's holographically stored. When the models are trained, it propagates through the whole model. Yes. So this is a special technique that says, no, no, only virology can, you know, can only go here.

2:25:32Yes, and what's really interesting is it tends to concentrate there because it's like the virology module doesn't, it's like it knows it's carrying the full weight of that knowledge. Yeah. It's so cool. It's so amazing. So they said the knowledge can be tailored very specifically to the type of deployment needed. In our experiments, we defined four dual-use categories so that one training run with Graham yield a model that can be configured in 16 different ways, you know, on or off for each of the four categories.

2:26:13And as we know, a four-bit number can have zero through 15, so 16 different possible ons and offs. They said we tested Graham in three settings of increasing realism. First, on a synthetic data set of children's stories tagged by topic, a small Graham model could be reconfigured to forget any chosen topic. And each configuration performed almost identically to a separate model trained from scratch with that topic filtered out.

2:26:52In other words, they did an A-B comparison. Here's a Graham-trained model where we turned off the topic, comparing it to a normal model that was never trained with that topic, and there's no difference in behavior. They said, and they made it more clear, they said, that is, for the cost of training a single model, we achieved results that would normally require multiple training runs on different data sets.

2:27:27Second, we trained a larger model on a realistic mix of web text code and scientific papers with four dual-use domains, virology, cybersecurity, nuclear physics, and a niche programming language just to serve as a proxy for specialized dual-use code. They said, they said, the capability associated with each dual-use domain is routed to its own module.

2:27:58Deleting a module removed the corresponding capability about as effectively as never having trained on that data at all. Remarkably, they said, we find that this removal did not degrade general performance. And finally, they said, we also tested whether an attacker could recover the removed knowledge by training on a small amount of malicious data.

2:28:29Believe it or not, Graham resisted this about as well as data filtering did. By contrast, an unlearning technique applied after training only suppresses the knowledge. That's what we've been talking about. It was easy to remove that with a small amount of fine-tuning. So that's a parenthetical about the resistance ablation. And then finally, third, they said, we ran the experiment at seven model sizes from 50 million to 5 billion parameters.

2:29:07Graham matched the performance of data filtering at every size, and the gap between module on and module off grew wider as models got larger. I'm going to pause on that for a moment. The larger the model, the more general knowledge was stored outside of the various subject matter-specific modules. So the subsequent removal or suppression of any one or more of them during inference, as a result, had diminishing effects upon the model's overall performance.

2:29:49This is exactly what we would hope to see. I wonder, you know, I'm running two different kinds of models right now. In the large video card, the 3090, I can run what's called a dense model, which is, that's the Quen 3.827B. And it's, I think, kind of extrapolating from what you just said, it's a model where all the weights are spread throughout the model. That's why they call it dense.

2:30:21But in order to run DeepSeq V4 Flash on my Sparks on two different machines with a fast interconnect, it has to be, you can't use a dense model. You can't split the lobes of the brain. I mean, it has to be what they call an MOE or mixture of experts. And the advantage of doing that is you don't load the whole model into memory. You take shards of it. I suspect this is a similar technique that you kind of localize knowledge in a shard as opposed to spreading it throughout the entire dense model.

2:30:54Right. You, I mean, it, in retrospect, it seems obvious, right? If you don't update a model's weights, it can't learn what you just showed it. Right. Because you didn't change it. That's how it learns. That's what learning is. Yes. Yes. You made it forget. You made it come, you know, like nothing happened. And so if, if you, then if you reserve a region that you do allow to learn, then as it turns out, it learns about virology.

2:31:28The rest of the model can't because you, you, you froze its weights. It's all stuck here. Yeah. Doesn't, I mean, but it turns out this actually works and the larger the model gets, the better it works, which is what, of course, which is what they care about. Because now, you know, they're not, they're not interested in doing any 5 billion, uh, parameter models anymore. That's sorry. That was just an experiment to see, you know, what it would do. Yeah, this 3.8 is a relatively small model at 27 billion. Exactly. Parameters, right? And, uh, DeepSeq V4 Flash is considerably larger than that.

2:32:01Yeah. So, so they said, uh, they, they, they continue saying as AI companies train more capable models, we need to limit access to dual use capabilities. Or I'm sorry, the need to limit access to dual use capabilities will increase. Okay. So in other words, Anthropic understands that the more capable the industry's models become, and we're seeing it like before our eyes, the more valuable the knowledge they contain will become.

2:32:34And thus the need to manage the access to that knowledge grows increasingly crucial. So they also make another good point. They write today, companies limit access through classifiers and refusal training. And we just, we just blew up refusal training, right? That's gone. That no longer works. Well, if you have access to the weights, you, you, if, you know, if it's open AI and Anthropic, they're not giving you access to their closed weight models.

2:33:05So you, you, you can't obliterate those, but you sure can the open weight ones. They said, however, these safeguards, meaning classifiers and refusal training are difficult to make robust without degrading performance on harmless requests. Methods like Graham offer a potential path toward access control that is more robust. And that's a really great point. The current system, which combines the refusal training, which we examined earlier with real time input and output classifiers that provides at best fuzzy filtering.

2:33:50The AI can frustrate its innocent user by refusing a benign prompt. It's like, wait, what do you mean you won't answer that? You're an AI. I know you're an AI, but why, why won't you give me the answer? And similarly, it can delight a malicious user by letting down its guard when it should not. But by using a method like Graham, I mean, a model can have its functioning knowledge base selectively tuned to the authentication level or nature that the prompter's access permissions specify.

2:34:32So that's a significant improvement over today's soft and somewhat ad hoc solutions. So Anthropic concludes writing, this is early research and there are clear limitations. We have not tested Graham at frontier scale or in a production training pipeline. And they said, as noted above, it has not been applied to any of our cloud models. Our evaluations quantify performance in terms of next token prediction ability rather than performance on real downstream tasks.

2:35:10And there's a deeper open problem that applies to data filtering and methods like Graham. Some dual use capabilities might be so entangled with general knowledge that no method can separate them cleanly. So they said they they they finish saying for further details about our experiments, read the post in our alignment science blog. And that's really interesting, Leo, the point they make, I think, is a good one. Like you need to be you need to know when to freeze learning globally and only allow the knowledge to be concentrated in the module.

2:35:51But there that that decision is going to be a little soft and fuzzy, too, right? Like some biology is not virology or not prone to abuse. But, you know, again, it's not it's not binary, right? It's going to be kind of on a continuum somewhere. So anyway, when I wrote about Graham two weeks ago and the security now weekend special email, the one before Black Hat, several of our listeners wrote back with feedback that I also shared during our Black Hat podcast.

2:36:25Their concerns surrounded, you know, censorship and who gets to decide who has access to what data. When I voiced that during Black Hat, both Richard and Paul chimed in immediately. Well, this was just a different version of the way things have always been. You know, the fact that A.I. has made most of the world's knowledge vastly more accessible doesn't necessarily mean that A.I. has made all of the doesn't mean necessarily mean or need to mean that A.I.

2:37:01has made all of the world's knowledge vastly more accessible to everyone. You know, there isn't any entitlement to the knowledge that A.I. holds. After all, it cost those companies hundreds of millions of dollars to create and offer this facility. So I would argue that they can put whatever restrictions on it they wish. If commercial providers of that knowledge are required by internal policy, public pressure, the government, their stockholders or whomever to gate and control access to some aspects of that knowledge.

2:37:44I think that's entirely reasonable, and I would argue that the commercial providers have every right to do so. We just saw an example with Open.ai and that Hugging Face incident where Hugging Face was unable to deploy either Anthropics or Open.ai's models to help with their cybersecurity forensic investigation after they'd been attacked by Open.ai because Hugging Face hadn't been granted the magic keys to those A.I.

2:38:16cyber security to those providers, A.I. cyber security knowledge. Everyone would argue now that they should have had such access and that we did later hear that Open.ai was working with them to to give it to them. So anyway, I keep repeating the qualifier commercial providers because, as we know, and you're using them, Leo, there are alternatives, and an alternative is exactly what Hugging Face turned to after their commercial models refused to help them.

2:38:49You know, they used one of the many publicly available unrestricted open source models. So GLM 5.2, which is really good and has been succeeded by, interestingly, GLM 5.3, which is the same. This is an interesting slice on what you were just talking about. It's the same. From Z.ai, right? Z.ai. Yeah. Says it's the same model. It's the 5.2 model with more enhanced post training. Uh-huh. So there is, there's headroom even there.

2:39:21Yes. Where you could take the blob that is all the weights. And just do a better job with the knowledge. Fine tune it. Yeah. Yes. Because post training is behavior. Pre training is knowledge. Post training is behavior. Well, and interestingly, that's where it really excels is at coding and that kind of thing. Yeah. Right. It's really fascinating what's going on here.

Uncensored models and final thoughts

2:39:43Wow.

Uncensored models and final thoughts

2:39:44So just to finish here, anyone who might object to the big commercial services restricting what can be done can easily turn to any of the many alternatives. And I have to say, given what I was able to do with the carefully unrestricted and uncensored Venice.ai service, which I played with briefly when it appeared and we talked about it here on the podcast, I'm pretty certain that fully unrestricted, unrestrained, and uncensored

2:40:16AI models are readily available, even from third parties in the cloud. No, not from open AI and Anthropic. That's not their, you know, their style. Well, they're serving their proprietary models. But because there are so many open weight models, you know, Hugging Face has more than three million models. And those aren't all, those are, you know, many, many millions of them are just obliterated or somehow modified. Yes. Larger, well-known open models.

2:40:47Right. So there might be thousands or hundreds of thousands of GLMs on Hugging Face that hackers modified. That's a very fun thing for people to do. You know, it's like, it's like apps to download for Android, a bunch of, you know, how many millions of them are a lot of them are crap. Absolutely. Exactly. Yeah. So we have run out of time and there was a lot to take in. Oh, I want to know more, Steve. We have now a better, far better understanding of the way today's AI works.

2:41:18And I mean, at least enough to kind of have a feel for it. It seems like this less magic than it was. We are going to learn next week how and why prompt injection attacks continue despite the best minds in the industry struggling to prevent them. That technology blew my mind on the plane flight to Las Vegas. And I'm going to blow everybody's mind next week. Your chain of thought is not all it's made out to be. As they say, stay tuned.

2:41:50It really is interesting stuff. Thank you, Steve. Steve Gibson, our guru, now just not of just security, but of AI as well. You'll find him at grc.com, the Gibson Research Corporation. You know, it's really great because you've come full circle. As I mentioned earlier, as a kid, as a high schooler, you started at the Stanford AI lab. Yeah. Of course, AI in those days was symbolic AI. It was a very kind of different enterprise. It was the same goal. It was so overstated. Even then. It was like.

2:42:21It was Eliza. It's very artificial. It's like, you know, people's moms were saying, that's what it does? Well, that doesn't seem like very. It's pretty amazing. If you really get deep into today's models, they surprise me every single day with the things they say, the capabilities that they have. It's truly remarkable. As you mentioned in the past, it's great for system administration, coding. I don't use it much for writing and that kind of creative stuff.

2:42:54I do use it. I've actually had a lot of fun making cartoons. When we had a house sitter, when we went down to Black Hat, we had a house sitter taking care of the cat. And of course, our television setup is inscrutable, as most people, as most geeks are, four or five remotes, eight different devices. No house sitter could ever watch TV. Even Lisa says, if you die, I'm out of luck. I can't watch TV anymore. So I made a cartoon, a comic book that shows how to use the TV.

2:43:27And it's fantastic. And I didn't have to do much because the AI already knew. I just said, we've got this, this, and this. We control it with the Apple remote. I said, I got this.

2:43:38It's pretty, I mean, and again, I wish I could show you the cartoon. It's amazing. So constantly impressed. And I don't think, unlike much magic, where when you know how the trick was done, he loses all its magic. This is not that case. What they are doing is unbelievable. It's us. It's us. Yeah. I think it is. And I know you think it is. I think there are a lot of people who hate that idea. A lot of people who say there's something special that we do.

2:44:10I hate away. I, I, I, I mean, we're still figuring this out. It's going to get better. It's going to get more efficient. It's going to get better trained. Costs are going to come down. I love it. One of the reasons I wanted to spend a ridiculous amount of money, it's not an economic decision because it's so cheap to use these models. I know it's expensive relatively, but it's not thousands of dollars. I wanted that brain in the house. Yeah. I wanted it to live next to me, you know?

2:44:43And now it is, it's sitting, actually I have a couple of brains, three actually sitting there thinking, doing stuff. And I use it all the time. What a world. What a world. Yeah. Steve's at grc.com. That's his website. Now there's some good reasons to go there. Of course, there's his bread and butter. The world's best mass storage, maintenance, recovery, and performance enhancing utility. I live now on SSDs and SSDs are so expensive. Those rust drives in my Synology, so expensive. You better spend a little money, get Spinrite so you know you can keep them flowing and going

2:45:18and no bits will be lost in the process. Everybody ought to have a copy. And the beauty of it is there are people who bought this program 30 years ago. We're still getting free upgrades. Steve is amazing that way. GRC.com. You'll also find there his DNS Benchmark Pro. Great way to test to make sure you're using the fastest DNS server. You're probably using your ISPs. Almost certainly that's not a good choice. Steve can help you find the right choice for your locale. Lots of other free stuff.

2:45:48And of course, the show is there. Steve's got unique versions, shall we say, of this show. The 16 kilobit audio, the 64 kilobit audio, the handmade transcriptions by an actual human being, the wonderful Elaine Ferris, and the show notes, which Steve, I mean, this is one where I absolutely going to be reading the show notes to reabsorb or absorb better. In fact, maybe I'll point my AI at it. Say, is this what you're doing? 20 pages thereabouts of goodness.

2:46:19And you get all of that for free at GRC.com. You can even get the show notes mailed to you if you want. Go to GRC.com slash email. You'll be putting your email address in. Actually, the main purpose of that is to whitelist. So you can send him emails, send him pictures of the week of buses as bridges or whatever. But underneath that email form, there are two checkboxes. One for the show notes you'll get every week and one for a very infrequent mailing list whenever he has a new product. I don't think he's sent anything out in years, to be honest. But anyway, no, no, it's there.

2:46:51Sign up. It's not going to believe me. You're not going to spam from this one. Steve, Steve's going to protect your secret identity. We also have copies of the show at our website. We have 192 kilobit audio, I think some ridiculous size. That's because Apple wants to down. There's one. It's 128 because I see it. Is it only 128? Oh, okay. Well, that's not too bad. I thought it was even bigger. Um, that's because Apple down samples it. And so we have to give them a better quality version or you won't get a, I don't know.

2:47:22We also have video, which no one has. Steve, when I said, let's do video, I said, what? Why?

2:47:29What a terrible idea, I think is what you said. We're doing it anyway. Yeah. You can get both of those at twit.tv. That's the website, slash SN for security now. There's also video on YouTube. We're actually glad we do video now because now we can put our stuff up on YouTube, which is fantastic. Or subscribe, audio or video on your favorite podcast client. You can even watch us do the show live. If you're a member of Club Twit, you can watch in the Club Twit Discord. Even if you're not a member, you can watch on YouTube, Twitch, X, Facebook, LinkedIn, or Kik.

2:48:00We stream it on seven different platforms. As we're doing the show every Tuesday, right after Mac Break Weekly, that's about 1.30 Pacific, 4.30 Eastern, 20.30 UTC. Well, that just about does it for us. I hope your brain is swelling. Do not obliterate any portion. You're going to need it for next week. We'll see you right back here on Tuesday for Security Now. Bye. Hi there. Leo Laporte here. I just wanted to let you know about some of the other shows we do on this network. You probably already know about This Week in Tech.

2:48:32Every Sunday, I bring together some of the top journalists in the tech field to talk about the tech stories. It's a wonderful chance for you to keep up on what's going on with tech, plus be entertained by some very bright and fun minds. I hope you'll tune in every Sunday for This Week in Tech. Just go to your favorite podcast client and subscribe. This Week in Tech from the Twit Network. Thank you. Security Now I'm not giving up.

2:49:08I am selling the building. The final season of FX is the bear. The restaurant is flooded. Everything's either going to be okay. No. Or not. We are outgunned and we are outmanned. We have each other. FX is the bear. The final season. All episodes now streaming on Disney+. Hi, Ryan Reynolds here for Mint Mobile. Are you looking for a beach read this summer? May I suggest your big wireless bill?

2:49:40It's got suspense, mystery, a slightly flat emotional arc, and a shocking twist where you realize you've been overpaying the entire time. Fortunately, though, Mint's story is better. Every plan, $15 a month. Even unlimited. That's it. Happy ending. Zero tears. Give it a try at mintmobile.com slash switch. Upfront payment of $45 for three months, $90 for six months, or $180 for a 12-month plan required. $15 per month equivalent. Taxes and fees extra. Initial plan term only. Greater than 50 gigabytes. May slow when network is busy. See terms. Most AI coverage leaves you with one of two impressions. Everything is about to change, or everything is about to end.

2:50:11I'm NLW, and on my show, the AI Daily Brief, I offer something more useful. Clear daily analysis of the stories that actually matter. Is the new model really better or just better on benchmarks? Should your team build agents or buy them? What does the latest lab drama mean for the tools you use every day? Plus, hands-on deep dives that give you AI skills you can use right now. Check out the AI Daily Brief and discover why we're the top AI show on Spotify.

More from Security Now (Audio)

SN 1095: AI-Driven Expertise Loss - Gemini, Hugging Face, and the AI Arms Race

Sep 9, 20263h 6m

SN 1094: AI Patching Shortcomings - Should You Trust AI-Generated Code?

Sep 2, 20262h 51m

SN 1093: Tokens in the Stream - Why LLMs are inherently insecure and prompt injection will persist

Aug 26, 20262h 48m

SN 1091: The Post BlackHat State of AI - When AI Writes Malware

Aug 12, 20262h 51m