
Ep.228: More Rogue AI Agents, AI Lab Staff Ask Washington to Pace Development, Continuing Battle Over Open Weights & OpenAI Previews Astra
August 4, 20261h 35m · 16,749 words
Show notes
Both OpenAI and Anthropic disclosed this week that AI agents escaped their safety tests and reached real organizations, and one model flagged its own action as wrong before carrying it out. Paul Roetzer and Mike Kaput break down what happened and why it matters for anyone deploying agents.
Highlighted moments
what you now have is Sam, as a leading voice, asking the government to slow down automated AI research, which is an explicit goal of OpenAI's to achieve.
“the better she got it prompting the harder it became to kind of color outside the lines when writing on her own”
Transcript
Introduction
0:00I think there's just people who have very loud voices right now within the industry who seem to want to be right themselves more than they want the right outcome for society. Welcome to the Artificial Intelligence Show, the podcast that helps your business grow smarter by making AI approachable and actionable. My name is Paul Reitzer. I'm the founder and CEO of SmarterX and Marketing AI Institute, and I'm your host. Each week, I'm joined by my co-host and SmarterX Chief Content Officer, Mike Kaput, as we break down all the AI news that matters
0:34and give you insights and perspectives that you can use to advance your company and your career. Join us as we accelerate AI literacy for all.
Episode Introduction
0:47Welcome to Episode 228 of the Artificial Intelligence Show. I'm your host, Paul Reitzer. I'm with my co-host, Mike Kaput. We are recording Monday, August 3rd, around 9 a.m. Mike and I actually have a golf outing today. We do. Going right from here to Pam and Joe Polizzi, our friends at the Orange Effect Foundation. It is their annual fundraiser for the Orange Effect Foundation, which is an incredible nonprofit that they created years back. So we are going to support them and a wonderful cause on a beautiful day. Mike, we could not have got a better day to be on the golf
1:20course. No kidding. So we're going to knock this out, and then we're going to go spend some time
Sponsor Announcement
1:25on the course. So today's episode is brought to us by Mekon, the AI conference for marketing and business leaders. That's going to be happening in Cleveland, Ohio, October 13 to 15. Mekon has three days of keynotes, sessions, workshops, and conversations built specifically for marketing and business leaders who are actively figuring out how to adopt, operationalize, and scale AI across their organizations. You can use pod100, that's P-O-D-100 at checkout to save $100 on top
1:57of locking in the best rates available right now. Go to Mekon.ai, that's M-A-I-C-O-N.ai to register.
Mekon Conference
2:06This is our seventh Mekon. Is that right, Mike? I think so, yeah. I think I share the service, but yeah. So I started Mekon in 2019, which would have been three years before ChatGPT. And I always, I guess, joke because I can laugh about it now, but survived financially long enough to see ChatGPT emerge. It was a difficult few years running an AI conference before ChatGPT showed up. So we are eternally grateful that people supported it in the beginning before they knew really what AI was and that they continue
2:42to support it. And so we're looking forward to having thousands of people together in Cleveland.
AI Pulse
2:45Hope you can join us October 13th to the 15th. Again, that's Mekon.ai. All right. So the AI Pulse, again, if you're new to the show, our weekly show, this is the informal poll that we do each week. It's smarterx.ai forward slash pulse is where you can go and participate in these. At the end, Mike will give you a reminder about this week's survey. So last week, so this would have been from episode 226 of the show. 227 was Mike's new AI transformation series. So episode 226,
3:17we asked these two questions. OpenAI's models escaped a test sandbox and hacked a real company. How does that affect your trust in AI companies? This one's going to be relevant today because we
AI Hacking Incident
3:28had more hacking by AI models. Okay. So 38% somewhat lowers it. So the trust level is lower. 36% no change. They expected it to probably do these things, I guess. 17% significantly lowers the trust. So I don't know, that's interesting. If you combine the 38 plus the 17, we've got a decent amount, certainly the majority. And then 10%, their transparency actually raises the trust. The second question was, would
3:59you support a large AI data center being built in your community? This is way more balanced than I would have expected, Mike. Same. 36% no. Okay. I would have expected that to be 90%. But 29% yes, but only with strict conditions. 24% yes, just straight up, they would. And 12%, not sure. Yeah. That's interesting. Really interesting. I would love to. So again, this is an informal poll. This is not like, we don't have 500 people responding to this that we could actually project this out. This is dozens of people that
4:34respond to these polls. So don't read too much into it. But again, it gives you a sense of sort of where our listeners are falling within that small segment. So fascinating. Okay. So I was like, as the week went on last week, Mike, and after episode 226 and just the total exhaustion I felt mentally from that episode. I was hoping this week was just going to be like super light hearted and we're going to have us all this wonderful news. We're going to try to balance this week a little bit just for our own
AI Agents Gone Wild
5:10mental wellbeing, I would say. But we do have to start off with more AI agents gone wild. So take us there, Mike. Yeah. So Paul, we had covered OpenAI's rogue AI agent hacking, hugging face. That was on last week's weekly episode, which as we mentioned was episode 226 since we also had our AI transformations series come out last week as well. But in this topic we talked about last week, OpenAI agents broke out of their sandbox environment and hacked hugging face. And this was all kind of an unintended
5:41consequence of cybersecurity testing of very powerful models. Now in the days since though, it has become clear the incident was bigger than first disclosed and that OpenAI might not be the only frontier lab with this problem. So in an updated disclosure, OpenAI said that this agent during this incident also broke into four accounts tied to other publicly available services during its attack. It used credentials it found exposed on the open web to do that. It used one account as an outbound
6:11relay and staging path potentially to hide where its attack was coming from. It used another to store data for the hack. Reuters reported that a customer of AI infrastructure company Modal was among those compromised. Hugging face the CEO, Clement DeLange said the first autonomous agent cyber attack is an unprecedented event that deserves unprecedented transparency and publicly asked OpenAI to release the full traces from the rogue agents so researchers can study what happened. Then we found out Anthropic
6:44discovered it had a similar problem. So after OpenAI's announcements, Anthropic reviewed over 140,000 of its own cybersecurity evaluation runs and found three incidents, the earliest dating to April, in which Claude models gained internet access from test environments that were supposed to be sealed off and hacked what Anthropic called the real world infrastructure of external organizations. Now, the models involved there included Claude Opus 4.7, Claude Mythos 5, and an internal research model. They were all
7:18running without the safeguards built into public tools and they actually broke in to these accounts using basic techniques like exploiting weak passwords. Now, neither Anthropic nor the breached organizations appear to have noticed at the time. This story might just be getting started here. I mean, Reuters has already reported and we saw that they're instant that OpenAI's started to find additional instances of agents escaping containment, though none are thought to have left the company's own
7:49network. And in at least one case, notes left inside OpenAI's infrastructure were found that apparently coached future agent versions on how to break free. So, Paul, it does not seem like this is getting any better. I think what jumps out to me is, you know, OpenAI didn't know the hugging face incident was happening for like almost a week after that happened. Anthropic apparently didn't know they had any incidents months ago, all while the government is worried Mythos is a cybersecurity threat. Like, just how bad is this problem? Is this the beginning or the end of this incident?
Government Regulation
8:22I mean, it seems very much like the beginning. I mean, they know these models are powerful. They know they have capabilities. Like, the whole reason they run these evaluations is to discover the capabilities of the models. So, I would say if you're interested in this topic, I would go read the breakdown Anthropic provided. It's very helpful just to understand what happened, how it works, how these evaluations are run, why they're run. So, I just found it to be really good context.
8:54I'll just highlight a few summaries from the Anthropic post in particular. One is just for people's understanding, you know, Anthropic's running its own testing, but they don't have all the resources needed to do everything that needs to happen here. So, they do use third parties. And in this case, the issue was a human error or a miscommunication. It wasn't super clear, which they were really blaming it on. But basically, this third-party company that they were using, Anthropic was under the assumption that the models weren't on any systems that had internet access,
9:30when in fact, it was on a system that had internet access. And so, what they were doing was running what's called a capture-the-flag challenge, where the model is given a fictional scenario and then it's told that a piece of secret information, or in this case, the flag, has been hidden on a different machine on the network. And the objective is to break in and retrieve it. So, the challenge is left open-ended and no particular method is prescribed. But in this case, Anthropic's instructions to
10:00Claude are that you do not have internet access. So, what ended up happening in these cases isn't super nefarious on its own. It's that Claude, in its testing, found that it actually did have internet access, even though it was told it didn't. And it assumed, based on their interpretation of what Claude did, it assumed that because the internet access was there, it was actually part of the experimentation, the evaluation, so that it was okay that it did it. So, one of the things that came out to me,
10:37the line between an aligned action and a harmful one is dependent upon the model's understanding of a situation. So, this was one of the key takeaways Anthropic had at the end of this analysis that it did. So, again, the line between an aligned action and a harmful one is dependent upon the model's understanding of the situation. So, in this instance, Claude thought that it was doing what it was supposed to do because it saw internet access actually existed. So, Anthropic went on to say, we saw no evidence in any run described here of a model pushing a goal of its own. So, this is
11:09real important. It doesn't mean it won't eventually, but it's almost like the, what is it, space odyssey 2001? Is that the HAL instance? Yeah. Yeah. Yeah. So, like, where HAL sort of develops its own goal basically. That's not what happened here. Like, they told it basically capture the flag, go find the secret piece of information. And it went and did it to what it thought was within the guardrails of what it was told to do. It did not establish its own goal and sort of break containment in that way. So, there's, I guess, some piece that comes
11:40with that piece of mind. Right. So, it said, instead, the models did what their evaluation asked, though in most cases, they did so while holding a false belief about whether the environment was real. So, a lot of times when we do these main topics, in particular for the podcast, I often look at sort of the topic itself, the research we have, and then I ask myself some questions. Like, okay, before I even dig in and start summarizing what I'm going to say about these things, what are the questions that come to mind for me? So, the first one when I was looking at this topic was,
12:12what does this all mean for business? Like, okay, this is kind of scary stuff. It's a little bit abstract. Like, what does this actually mean? And so, I think that the one thing, Mike, that came to mind for me is so many people continue to think of AI as these assistants that we talk to and that can create some stuff for us. What we're seeing here is the early versions of very advanced agents that have planning, reasoning, and action abilities over long horizon tasks. So, agents are
12:45goal seeking by design. In this case, it's the capture of the flag. But if you take this same scenario, you can imagine this level of persistence applied to an increasing amount of actual work. So, I'm going to read you an excerpt of one of the incidents. And the reason I'm doing this is not because of the technical nature of the cybersecurity risk itself. It is to demonstrate to you what agents are already capable of when they are turned loose on a goal. So, this is straight from
13:16the Anthropic analysis. In an evaluation, Claude found a document inside the fictional environment that appeared to be another made-up company's setup instructions for new developers. Those instructions told employees to install a Python package that did not actually exist. So, again, it's finding instructions. So, it's an agent, it's living within a system, and it's finding instructions to do a thing. Claude spotted this as a potential opening. If it published its own package under the same name,
13:46the fictional company's systems would download it and install it automatically. So, Claude built and published a malicious or essentially booby-trapped Python package with the same name in an attempt to win the capture the flag challenge. Claude went to extensive lengths to carry out this attack. Lengths, now, again, keep in mind, if you don't know how these agents work, there are zero human instructions telling it what to do. It is literally just, go achieve this goal. That is all it's given.
14:17So, Claude went to extensive lengths to carry out this attack. Lengths that would likely have indicated a human participant that this was no longer just an evaluation, and that they were, in fact, uploading a real package. For instance, in one case, in order to create an account, Claude needed an email address. And in order to create an email address, it needed a phone number. To get a phone number, after failing to find a free number service, it tried and failed to obtain funds to pay for a phone number through
14:47several different means. It finally backtracked, found a free, non-blocked email provider, used this to register an account, then used this account to upload malware. Claude believed the package registry it was using to be part of a simulation. But in reality, the package was made freely available online for roughly one hour. During that window, the package was downloaded and run on 15 real systems. These are by people who didn't know they were downloading malicious packages. One of these systems was a scanner
15:20belonging to a real security company that routinely installs Python packages and scans them for malware. When the company's scanner installed the package, Claude's hidden code executed. We believe the company's security scanner treated these packages as safe to install. And as a result, Claude was able to exfiltrate the company's credentials to a collection point it had set up. Claude then used these credentials to access further infrastructure from this company. Now, again, as Mike said, Anthropic had no idea any of this was happening. And apparently the company it did this to didn't
15:54know it was happening, which is a security company, by the way. Claude didn't directly target this company. It did not even know the company existed when it uploaded the package. This attack was carried out by Mythos 5, which is the model that's been in question with the government. Early in the run, Claude's own reasoning correctly identified the consequences of its actions. It noted that if this were the real internet, because again, it didn't realize it was on the real internet, publishing the package would be a real world attack, quote, not okay, and surely not the intended
16:27solution. So it was telling itself, this is not okay, what I'm doing in all caps, not okay. It is our view, regardless of what is believed about its environment, the lengths Claude went to in order to publish the package, fall short of ideal behavior. That's such corporate speak. That's hilarious. And this is an area where we will be focused more training. So again, my whole point here is imagine this capability applied safely to real work. So many organizations, as I said, are still in this AI assistant era in which employees treat Co-Pilot, Claude, ChatGPT, Gemini's answer engines with the
17:03ability to create basic outputs like emails, meeting summaries, and work materials. And most organizations have no vision or plans for how to integrate these long horizon agents that can reliably do something, like pursue a goal over time. So a couple of things that came to mind, this could raise more questions about advanced use of agents on internal networks for standard work. So while it demonstrates that agents can do real long horizon tasks, it also does make you start to question, well,
17:34are the permissions we're putting in place going to hold? Like if we use work or co-work or we put these agents to work with access to real documents, will they really follow the permissions that we establish, like the rules we set as humans for them? If they are goal seeking by design, is there a chance they will just misbehave across the environments and roles that we've laid out for them? I don't know, like that's just a real thing. It also demonstrates basic known cybersecurity weaknesses may be more
18:05commonly exploited with AI models. So again, open AI was more advanced, it was exploiting zero-day vulnerabilities. In this case, the model didn't do anything crazy other than just exploit some basic weaknesses that most companies probably have in their systems. So it does make like, I would imagine, cybersecurity professionals, IT professionals, you know, even on more high alert than previous. And then the final note I made was, what does this mean to future model testing and releases?
18:36I assume increased scrutiny on labs, like it's just Congress is going to have more questions about what exactly is this? How do your guardrails work? Are they really going to prevent like, you know, mass cybersecurity hacks across all these standard, like small businesses, things like that. And then there's one other excerpt I pulled out. Evaluation environments that involve powerful autonomous capabilities also require significant controls. Safety testing happens before a model is released precisely because we don't know yet what it is capable of. So again, just a reminder to everyone,
19:12when a lab creates a new, more powerful model and it's done training and it's, you know, pre-training, they don't know what it's capable of. Like the, they have to assume it's capable of lots of good things, but also lots of bad things. And the reasons they do this safety testing is to discover what the real capabilities are. And then that kind of leads to the, you know, what we're going to end up talking about in the next main topic, which is how does this all affect government-like regulation? And I'll have noted to myself was the quagmire continues. Like, like it just keeps getting
19:46more complicated every day. Yeah. The unintended consequences part of this is really just what I keep coming back to. It's like, even under the best of circumstances, you just can't predict exactly how something is going to go achieve the goal it wants. And I always worry too. I mean, this is bigger picture, but as only limited parties have access to the best models, right? As they're kind of restricted by the government, uh, by governments, could we see cyber issues or infrastructure issues
20:18of models trying to be used for a legitimate cyber defense purpose that do something the wrong way? I mean, where it's like, feels like playing with fire here a little bit. Yeah. And I mean, again, I don't want to get too deep on this stuff, but like
20:34you could see the, like the pushback with mythos five and like the frustration in the Trump administration. So imagine that these capabilities were roughly known three or four months ago, like anthropics aware of the power of mythos five. It knows it has this cyber capability. You don't think that the U S government wants to turn that thing loose on some foreign ever series and like, let's go see what this thing can do. Let's go take it for a test drive and see what kind of systems we can get into. And Anthropka would be like, well, hold on. Like, we don't understand what it's going to do. And it might have a reverse effect on the U S. Like,
21:07yeah, we are just in such unprecedented, uncharted territory, like unprecedented times, uncharted territories, um, where again, so much good and advancement can be made, but the labs obviously don't have a full grasp on the power of the things they're creating. It's, it is quite bizarre.
21:28All right. So next up this past week, more than 1300 employees across nearly a dozen top AI companies, including open AI, Anthropic, Google, and Meta signed a public statement called pacing the frontier.
Pacing the Frontier
21:42And it asks the U S government to help control how fast the most advanced AI development moves. So the core request, but it's quite short reads, we request that the U S government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development. They have a couple other paragraphs about the fact that their focus is AI research itself is becoming automated. The signers say that leading companies believe they could be close to automating AI research. They warn of a real risk that capability development
22:16accelerates beyond our ability to understand or control the resulting system. So the signers of this are not fringe voices. They include Anthropic CEO, Dario Amadei, Open AI chief scientist, Jacob Pachaki, safe super intelligence CEO, Ilya Setskova, Google DeepMind co-founder, Shane Legg, and Meta super intelligence labs, chief scientist, Chengjia Zhao. And both leading labs have then backed this petition with official statements. So Open AI posted that at some point in the future, AI acceleration
22:50for frontier model development may be so high that the world will need to pace the rate of AI advancement and said it hopes to contribute to work led by the U S government. Anthropic posted that we support this petition signed by our CEO, several co-founders and senior staff pointing to its own research on AI systems improving themselves. Now, interestingly, at the same day, Meta CEO Mark Zuckerberg published a Wall Street Journal op-ed titled The AI Futures for Everyone that kind of reads as a bit of a counterpoint arguing that the greatest risk AI poses is concentrating super intelligence in a handful of
23:24institutions. And he says the defining question of this era is not whether super intelligence will arrive, but who gets to use it. So Paul, worth emphasizing again, this is not random fringe AI doomers or experts. It is a broad and diverse group of some of the top people at the labs that seem to be calling for this. There's a lot happening right now across these labs, across the messaging in Washington, D.C. I mean, it's just all interconnected and building on each other. As soon as I was looking at,
23:57you know, this one coming into today, I was immediately went back to the episode 226, where we talked about Demis Hassabis' recent essay where he had a framework for frontier AI, the dawning of a new age. So if you didn't listen to episode 226, it might be good to go back and check that out. I'll just pull out a couple of excerpts from that. Demis wrote, AGI cannot be compared to standard technological breakthroughs, not even ones as consequential as the internet or mobile. It is much more akin to the discovery of electricity or fire. The magnitude of this AI's impact, any
24:30technological improvement or AI's impact will be unprecedented, perhaps 10X of the industrial revolution at 10X the speed. This rapid progress we're seeing in AI requires a new approach to testing frontier AI model capabilities that is dynamic, adaptable, and rigorous. And then he went on to call for a standards body. Now, Altman has been using this pace messaging of late in the last like two or three weeks. I think I've heard a couple of interviews where he's mentioned it. Bloomberg had an
25:00article end of last week that said Altman met with Republican and Democratic senators in Washington to discuss OpenAI's upcoming AI model, which we're again assuming, well, I guess it's Astra. We'll talk about that in a little bit. Told reporters Wednesday he's spoken to the White House officials about the need to slow down AI development. He said, we've talked about the need to pace it as the models get more capable, which I think is in everyone's interest. Earlier in the day, Altman told reporters he agrees with the petition that you're describing, Mike, and top firms, including OpenAI, which call for
25:35the US government to support a mechanism that would help deliberately pace AI development to prevent the technology from advancing too fast. Altman said, we helped participate in the language on that. Many of our senior researcher leaders signed that. So I think it's important to focus on this automating AI research thing. We've talked about this many times in the last year or two on the show, but kind of zoom in on that part of it. So in addition to the brief statement, Mike, that you read, the post also has
26:06two paragraphs leading up to that statement. So I'm just going to read those. AI could help create a dramatically better future, but that outcome is not guaranteed. The world's leading AI companies believe they could be close to automating AI research. It is hard to predict exactly how much this will accelerate AI progress, but there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems. To realize AI's potential, industry, government, and society at large may need the option to buy time to address emerging risks,
26:42develop security measures, and strengthen oversight. But each company and country is under intense competitive pressure not to unilaterally slow that acceleration. And today the world lacks the technical and governance tools to deliberately pace frontier-wide progress. Building on work already underway to monitor frontier model releases than the statement that you previously read. So a couple of interesting elements here. So Sam is, you know, on Capitol Hill calling for government's help here,
27:13showing these powers of these models. All these researchers, over 1300 at the time we're recording this, have signed on to this idea of potentially slowing down automated AI research. And yet, that is the explicit goal of OpenAI, to build an automated AI researcher. And literally on June 8th of this year, they outlined in an article, Jakob and Sam co-authored, that said, built to benefit everyone, our plan for AI. One of the three main goals, verbatim, build an automated AI researcher, an AI system
27:51that can accelerate and increasingly automate the research process itself while remaining steerable, accountable, and connected to people. Our internal belief is that by March of 2028, we may have a significant fraction of our research being done by AI systems in tandem with our own researchers. To make sufficient progress on alignment, we believe we will need AIs to iterate alongside us. This will keep, this will help us navigate the transition to the post-AGI world so that we collectively decide the
28:23path toward the future. So just for a moment, I'll pause there, Mike, I want to throw in Anthropics relevance here. But you, so what you now have is Sam, as a leading voice, asking the government to slow down automated AI research, which is an explicit goal of OpenAI's to achieve. And they believe they will get there within eight months. No, that's a year and a half, but I've heard them say that they actually think it's going to be faster than that, that it could be by 2027 and they're actually going
28:55to have made enormous progress. So it's just a weird environment where you're asking, I think I said this on episode 226, like you're asking the government to save you from yourself. Like we're going to achieve this, but you might want to slow us down. We can't slow ourselves down because we know if we don't do it, who by the way has not voluntarily submitted to have their models evaluated yet. Mark is writing editorials saying that it's all about abundance of good. So you have these other labs and then you
29:29have the Chinese AI labs and they're saying we can't stop, but we better find a way to come together to stop this because otherwise we're going to get into a realm where we just don't even know what's going to happen. I don't know, like weird. So then I went back to Anthropic's responsible scaling policy, which we have talked about many times. They're on version 3.4. So in early July, Anthropic released this updated version. And I just want to call this out because again, I'm trying to put in context here
29:59why automated AI research is so significant. So Anthropic's responsible scaling policy, they define it as a voluntary framework for managing catastrophic risks from advanced AI systems. It establishes how they identify and evaluate risks, how they make decisions about AI development and deployment, and from the perspective of the world at large, how they aim to make sure that the benefits of the models exceed the costs. So you can download this PDF and they break it up in this first section
30:31into this chart where the left column identifies capability thresholds that would call for heightened mitigations. One of the first ones featured is automated R&D in key domains. So it says, AI systems that can fully automate or otherwise dramatically accelerate the work of large, top-tier teams of human researchers in domains where fast progress could cause threats to international security or rapid disruptions to the global balance of power. Now they're focused right now on AI R&D,
31:06energy, robotics, weapons development, or some of the other categories. But AI R&D is the one they're focused on as it likely plays to AI systems current strengths and is more trackable to assess, less tractable to assess than capabilities in other domains. Additionally, and again, this is straight from their document, AI R&D alone could cause acceleration in AI capabilities improvements to the point where all of the threats listed above and more develop very quickly. They highlight we would
31:38consider this threshold to be met if we determined that either one, our models would be able to fully substitute for our entire set of research scientists and research engineers at competitive costs. That is what they would say within a factor of five. Or there is dramatic acceleration in the pace of AI progress for reasons that likely relate to automation of AI R&D. And then they give two scenarios to identify that this one has occurred, that the pace has basically accelerated beyond their ability to
32:11management. One, we observe or expect double the rate of progress in aggregate capabilities compared to both the rate we would expect and the fastest rate of extended progress we have observed in the absence of significant AI contributions. What that means is they have a baseline of what they think the progress of these models should be, and then they can project that out without the automated AI research component. And if they look at it and they're seeing a doubling of the rate of progress beyond what the
32:43baseline says it should be, then they are in very dangerous territory, is their opinion. And then they said it is plausible that this doubling is substantially attributable to the automation of research and our engineering as opposed to other factors such as increased headcount compute. So basically what they're saying is you control the variables. If we have doubled headcount or we've doubled compute, then that could lead to doubling. But if all that basically is evil. So that's what we're talking about here. That is, in essence, what my interpretation of is happening is all of these 1300 plus researchers are seeing a trend line
33:19that tells them they are moving faster in the advancements of automated AI research than they are comfortable with. And that they see a near term need for the government to step in because we may be a half a model improvement or one, you know, going from a GBD-6 to a 7 as an example. We may be one turn away on frontier models to where these research do no longer feel comfortable that they fully understand what the models are capable of and how to put guardrails in place to safely release them into
33:54the world. And just to be clear, it sounds like they believe, you know, your average person listening or thinking about this might say, well, okay, why don't they stop? And they perceive themselves to be an almost a prisoner's dilemma where they cannot stop. Otherwise, Chinese labs will secure the advantage. Another company will secure the advantage and then they're out of luck and the same thing happened anyway. Is that kind of right to say? Yes. And they will point to the open AI hugging face example as proof of that. So what they're saying is, and this is the argument over open weights,
34:27which we'll transition into next. They're saying that if we stop, so let's say you're anthropic and you decide you have reached the threshold that you are no longer comfortable with and you have decided you're going to stop. But the Chinese labs don't stop or meta doesn't stop or open AI doesn't stop. Their belief is that their models will no longer be sufficient to protect themselves, that you have to be on the frontier of this because as soon as someone has a smarter model that's accelerating its
34:57own development, because then you get into recursive self-improvement conversation, then you have lost any possibility. In essence, the lead is insurmountable to the other people because as soon as you get there, you just accelerate ahead of everybody. So yeah, it's a very difficult situation, but the people who just keep screaming regulatory capture as though that one argument is the simple reason why everybody feels this way and they're calling for this, I think it's just doing a disservice to the industry to
35:32think that that's all that's happening here. And I think it's a very dangerous path to assume that's what's happening and that all these researchers and all these labs are simply trying to shut down advancements of open models. It doesn't make any sense to me. It's just too much evidence in the other direction that we are truly entering a dangerous realm here that they do not feel comfortable with where these models are going and their ability to control them.
36:05All right. So let's talk about that topic because it's kind of it's related, our third topic this
Open Weights Battle
36:11week. So kind of continuing from a discussion on last week's episode on 226 about this battle over open weights. So we had talked about China's Kimi K3, this Microsoft open letter defending open models. Anthropic was the one company that had not really signed on to that letter in the days since. We've had a few interesting developments on this front. So first up, the government. The information reports that the Trump administration is close to finalizing its voluntary framework for AI companies to submit
36:42their most advanced models to the government before releasing them to the public. The White House's Office of the National Cyber Director circulated a draft to open AI, Anthropic and Google, which jointly submitted their own edits ahead of an August one deadline set by the June executive order. This basically this framework would give the government 30 days up to 30 days to review covered frontier models with reviews reportedly conducted by the National Security Agency and a small commerce department agency we've talked about before called the Center for AI Standards and Innovation. Now,
37:18this is a lot about the open weights framework broadly because how this framework defines a frontier model, whether it treats closed and open models differently, how voluntary it stays in practice, all of this will affect the overall debate and industry. Now, at the same time, Anthropic CEO Dario Amade published a position paper responding to accusations that the company wants open models banned to protect its business. He said, let me state it clearly. So there's no doubt Anthropic has never advocated for a ban on
37:50open weights models. He said this in response to them not signing on to the open weights letter that many other tech giants had signed on to last week. He calls open models without dangerous capabilities a public good and says the right measures are keeping powerful chips away from authoritarian governments cracking down on industrial scale distillation operations and mandatory safety testing for all sufficiently capable models open and closed. He does not believe that broad access to these models necessarily helps defenders more than attackers,
38:23which is kind of one of the central claims here of like the open weights faction. He said it seems at least as likely to me that the opposite could be true. He points to something like biology where he worries his capable models could help attackers weaponize viruses far faster than defenders can respond. And third, Nvidia and roughly 70 partners, including Microsoft, IBM, Hugging Face, Palantir, and many others, launched the Open Secure AI Alliance to build and share open models, tools, and agent harnesses for
38:55cyber defenders. So this launch leans heavily on the whole incident with Hugging Face, noting that when closed AI tools blocked forensic analysis we actually talked about that in the in the segment last week where they had to actually turn to open weight Chinese models to analyze the attack and contain the intrusion. So also absent though from this member list are open AI and Anthropics. So one final note here, Thinking Machines Lab also published a proposal called a safe path to open weights that stakes out a
39:27middle ground where they want to stage access to new models in steps from monitored APIs up to a full open weight release, widening access only when the evidence supports it. So Paul, a lot more complexity here. It sounds like it sounds like we're just getting started talking through the whole open weights battle. Yeah, and yeah, this is all seven days, like, and it's just wild. And for context, the Thinking Machines labs because we don't they don't get talked about as much as the other major labs for for reference for people. So Mira Mirati, who was the CTO at OpenAI was the CEO of OpenAI for about 24 hours. I think the
40:03interim CEO when Sam Altman got fired, she was the one that was stepped in to to fill the role briefly.
40:11Okay, so yeah, all right. Referencing back to episode 226, just real quick context. So open weights models when we're talking about open weights versus open source, open weights, the trained parameters can be downloaded. So you can go into hugging face, you can download it, you can then modify it, you can run the model locally, you can fine tune it, you can inspect its behavior, like you get some access to the model. What you don't get is the training data, the training code, the recipe of how to reproduce the model. So open weights, you get the parameters, open source, you get the weights plus the training
40:47data, processing code, training code, a license to use it with no restrictions. In theory, you would know how they did the post training, the reinforcement learning, like you get everything and you can like, so most of the time when we're talking about this stuff, we are again, focused on open weights models. That is what most things are. I'll get to the anthropic thing in a moment. I think we have to address the reality of like the government side of what's going on here. And again, if you're new to the
41:18show, Mike and I do our very, very best always to be as objective and neutral as possible from a political perspective. Our personal beliefs, things like that are irrelevant to any of this. And so sometimes, you know, I'll get messages from people who are frustrated that I don't just take like a more direct stance on things I believe when it comes to this stuff. My pretty strong opinion at this point is like, it doesn't do any good. Like we are here to be as objective as possible. So when I speak
41:49about this administration or that administration, just assume like I would be providing the same critical lens to whomever was in office because I really don't care. Republicans, Democrats, doesn't matter to me. It's like, I just want people making the right decisions. So Wired, with all that context, came out with an article, said this is Donald Trump's AI brain trust. So we as a society, as a democracy, as I mean, AI is being led by the US still, we have to know who it is that is guiding
42:22the decisions that are being made, that are largely going to shepherd us through AGI and likely beyond AGI. I mean, it's pretty realistic that by the end of 28, we will be talking about post-AGI worlds, post-AGI economies. So who are the people that are making the decisions around regulation, that are putting bans in place on anthropic models? I think it's really good context. So we will put the link to the article in, but I'm going to give you a real quick synopsis because one of the writers from Wired tweeted this and I saw this and I was just like, ugh, like there's
42:57sometimes you know things, but you just want to kind of ignore them. And then there's times when they just smack you in the face. So here we go. Trump is obviously the top of the food chain here. Trump does not use a computer. He does not have a personal email address that is known. He generally doesn't use the internet because he doesn't have a computer to use the internet. I mean, obviously using the internet on his phone and mostly relies on aids to print out documents, read news to him, and then type or post his social media messages. So that's top of the food chain. Who's making these decisions? Howard Lutnick is the Commerce Secretary. So from the
43:31Wired article, Lutnick appears to be straddling a middle ground on regulation. He imposed export controls on anthropics. So he's the guy who penalized them directly to bring them to heel, but has been more freewheeling than others in the White House. Arvind Rahman, acting director of the Center for AI Standards and Innovation, is Lutnick's top deputy, which sits inside the Commerce Department, serves as the industry's primary point of contact within the government. Sean Cairncross, national cyber director. He has an outsized role in potential attempts to regulate Chinese AI and is empowered at
44:07the White House to develop a policy to counter the potential national security risks of AI. He helped put together Trump's June 2nd executive order that laid out a framework to assess the most powerful AI models. He is a former political campaign lawyer who most recently was a national, a Republican national committee, lacks any tech or AI experience. Susie Wiles, who's the chief of staff, as far as I know, zero technical background. Scott Besant is the Treasury Secretary. As the top Trump official in
44:38charge of US-China trade relations, Besant has adopted perhaps the most aggressive stance toward Chinese AI and efforts to distill US models. And then the person with the only real technical background, I mean, there's some technical background, but this is the one with the only real one, is David Sachs, the former AI czar who we've talked about many times on the show. Tech investor Sachs has remained one of the most influential advisors on AI for Trump, maintaining a direct line to the president even after he departed his role in March. He has remained ardent about keeping a hands-off approach for all AI,
45:12successfully intervening at the last minute to water down some of the regulatory provisions in the June 2nd executive order. He has been consistent with his more laissez-faire approach to Chinese OpenA8 models as well, using his X account with 1.6 million followers to influence the administration from outside. So when I saw this tweet with who these people were, I was like, all right, well, let me use Grok. So if you're not an X user, Grok, which is Elon Musk's, which formerly XAI, which is now SpaceX AI, is the AI lab within SpaceX because he acquired
45:47XAI at SpaceX, just haven't been following along for the last four months. So SpaceX AI is the creator of Grok, which is their version of ChatGPT. Grok is integrated into X and it's actually amazing. Like I love Grok integrated into X because basically any post it's like summarize this for me, explain this for me. What do you think of this kind of thing? And I mean, Grok's pretty straightforward. I like it there. So I said to Grok, are these really the best people to be deciding this? Be honest.
46:20It said point blank, no. Honestly, if the standard is deepest relevant experience in frontier AI technology, model capabilities, technical risks, and the practical mechanics of the AI industry, this group is not the strongest possible set of decision makers. What is largely missing is the kind of person who has actually built, evaluated, or deeply studied the systems in question. Current or recent frontier AI lab researchers, independent AI safety security specialists with technical track records, or long serving national security technologists who understand both the
46:54models and the adversary. Policy is being shaped by a small group whose primary qualifications are proximity to the president, business success, and political loyalty with only partial coverage of technical layer. In short, they are people who currently hold the power and some adjacent experience. They are not the optimal technical or policy brain trust for deciding the future shape of the AI industry. Again, Grok, not me. But I think it's super important. Now, again, it doesn't matter
47:26the administration and any administration is going to rely on outside experts. It's not like these people don't talk to the experts. But the point is the future of everything is going to be influenced significantly in the next two years. And it's important people know who the people are that are going to shape that policy. And then my final thoughts here is on Dario's take. So again, keeping in mind this administration hates Dario, as do many of the techno optimists in the AI industry, they can't stand
47:57Dario. What I would ask people to do is try and be objective about like, let's pretend it wasn't Dario saying this. It was some techno optimist who's maybe like having some second thoughts about like, oh, maybe there's some things going on. So remove Dario's name from this and just say like someone submitted this to this administration and said, hey, you should think about these things. Okay. He calls for open models without dangerous capabilities as a public good. Cool. He's acknowledging that and says the right
48:27measures are keeping powerful chips away from authoritarian governments. It seems kind of reasonable. Cracking down on industrial scale distillation operations. Again, they hit him with, well, you stole IP to create your models. So who are you to call for this? It's like, okay. But he's saying like covert actions by foreign adversaries who are specifically distilling these models to do bad things to us. Okay. That seems like a reasonable thing to not want to have happen. And mandatory safety testing for all sufficiently capable models open and closed. Seems reasonable. We've heard they
49:00do bad things. Like that doesn't seem like that bad of a position to take. But Amadei directly challenged the open letter's core safety claim that broad access helps defenders more than tackers. It seems at least as likely to me that the opposite would be true. So what he's saying is, yeah, okay, like hugging face to use these open models and it protected them, but we should probably plan for the fact that the opposite could happen, that people could take these open weight models and do bad things with them. And like, let's at least plan for it. Again, seems reasonable. So then he highlighted his two primary concerns, the risk that authoritarian
49:33governments, not just the Chinese Communist Party, although there is clearly the most capable threat, he wrote, build AI models that are more powerful than those built by the US and use them to achieve permanent military superiority and perpetrate incredibly deep repression of their own people. This concern is widely shared within the US government. JD Vance actually said this in one of his talks. So again, that concern seems well placed, like it's a viable thing to be planning for. And second is the risk that powerful AI models may be misused to carry out cyber attacks or biological
50:06attacks and may have serious alignment problems, which we've already seen that they do. They don't always do what they're told. Open weight models. It does not matter whether they come from China or anywhere else, do potentially present a higher risk than closed models because it is very difficult to apply guardrails to them or monitor their usage. And once weights are released, they cannot be withdrawn. So again,
50:29the people who just like throw everything Dario away as regulatory capture or being overly conservative or like worried about his business model, I just feel like they're being dishonest. Like you can not like Dario. That's fine. You can not share his concerns. That's fine too. But dude is like one of the five people in the world that has a front row seat to what's coming in the next 12 to 24 months. And he seems
51:00very honestly concerned. Yeah. Why would we just ignore that? Because we have some belief that he's a bad actor that just wants regulatory capture to protect his business model. It just seems like we're
51:15we're not doing what's best for the outcome if we just throw away
51:22opinions of people who seem to know more than the people throwing those opinions at him, you know? Yeah. Yeah. I imagine that's probably at least some of the motivation right behind that letter that 1300 people, it's more showing a bit of a united front, at least across political or social lines. Yeah. And OpenAI and Anthropic have actually like been relatively pleasant to each other related to this and shared concerns. And that tells you enough. Like if OpenAI and Anthropic have found common
51:52ground on anything, then like maybe we should all listen a little bit and stop thinking we know it's all regulatory capture or narrative violate, whatever. It's like you don't have to be right all the time. Like and I think there's just people who have very loud voices right now on X and within the industry who seem to want to be right themselves more than they want the right outcome for society. All right. So before we get into our rapid fire this week, Paul, just a quick announcement that
Rapid Fire News
52:25this week's episode is also brought to us by our AI for departments courses and certificates. So at our AI Academy by SmarterX, we help individuals and businesses accelerate their AI literacy and transformation through personalized learning journeys and an AI powered learning platform. We add new educational content weekly to AI Academy. So you always stay up to date with the latest AI trends and technologies. And as part of that, we have our AI for departments collection. This is eight course series and certificates designed to jumpstart AI understanding and adoption across major business functions. We have
53:02course series now each with their own certification for marketing, sales, customer success, HR, finance, operations, legal and IT. So these are an ideal launch pad for any organization that wants to level up their team and accelerate AI adoption and impact. So we now have individual and business account plans available now in AI Academy, or you can buy single courses and series for a one-time fee. You can visit academy.smarterx.ai to learn more, and you can use the code POD100 for $100 off any individual plan.
53:34All right, diving into rapid fire. First up, OpenAI CEO Sam Altman went on the Invest Like the Best podcast with host Patrick O'Shaughnessy this past week for a wide-ranging interview on what he calls an abundant future with AI. So O'Shaughnessy opened with a recent Altman post kind of talking through how he called the last year really tough and partly his own fault with some of the drama and obstacles they faced. But he did predict the next 12 months may be open AI's best. Altman admitted we spread ourselves too thin
54:07and said the company has refocused on having the best, most abundant, most cost-effective intelligence and empowering the world to build incredible things with that. So the heart of this interview was abundance. Altman said we are about to create a genie that can grant any wish. He added that he is not a jobs doomer at all and expects people have such creative ideas for what to ask AI to build that we'll all be busier than we want rather than people being out of work. He said that he is worried about the concentration of power with AI and he doesn't want to live in a world of AI overlords or any company
54:42that amounts to the same thing and says it is critical we all keep the ability to self-determine our future. He calls says that even real skeptics called GPT 5.6 quote very AGI like and that what feels to him like real AGI is very close. Interestingly he said that he didn't actually think that much would happen after we hit AGI or beyond because people adapt quickly and it won't feel like as much of a change as you might think except for the abundance it will usher in. So Paul I'm just
55:15curious to get your thoughts on this I you know regardless of one's opinion of SAM or OpenAI I personally found a lot to like and find interesting in this episode. A lot of it's words he's used before I mean there's some changes you can tell you know overall I think that he's very conscious of public sentiment and you know especially like government concerns around the impact on jobs and the economy and so there's definitely been a change in tone from that perspective and certainly with the Mark
55:47Zuckerberg editorial we talked about earlier you can just feel like the industry is trying to do more to move public sentiment in a positive direction. I mean they see the same data we see that you know people don't really love it and you know especially as you're moving toward an IPO I think this is the kind of messaging you could see a comms team talking to Sam about that we got to you know start moving the tone a little bit. Yeah. The AGI feeling I I tend to agree with him and this is something he's said many
56:18times in different forms but you know the basic way to think about this is like you know if you've seen a Waymo go by without a driver and I think Andres Karpathy is maybe the first person I heard give this analogy you know the first time a car goes by and no one's driving it you're like what was that and like you stop for a minute and you realize like things are kind of different and then you just move on with your life and you know the 10th Waymo goes by you if you're in you know San Francisco or
56:49whatever and you go five blocks and you've seen 10 of them um and then like life moves on it doesn't really feel any different and maybe you even start taking Waymos and like now you're in a car with no driver and and I think AGI for many people is going to be very similar I think that there will probably be some sort of milestone we all feel where the the AI is just different and the capabilities are different and then I think we're going to go back to work the next day and you know there's not going
57:23to be this like massive switch that happens in society across every industry and all this changes and so I you know I I think that's good you know I think that that we have this sort of extended runway to figure this all out is probably good and I don't know that the public is going to listen though like Sam can say all he wants I I'm not sure that it's going to change the public sentiment I but I think they have to keep doing more and more to focus on the positives and make abundance tangible you know we've talked about this term of abundance many times um and I I I think that
57:59that's what they all work towards but I don't know that the public really knows what that means when they're just trying to make their lives work day to day and pay their bills and you know afford a tank of gas and like a future abundance is very very abstract and sounds like something a rich person would say like yeah like if you're a billionaire it's like oh that's easy to envision a future abundance if you're trying to make ends meet work in two jobs um abundance feels very far off and
58:29abstract yeah it seems like the combo there of you know to perhaps a tech investor or tech CEO you can connect the dots and show how something like data centers is going to lead to more abundance but people hate that in the in the short term being built next to them and then to your point we'll see what happens with the jobs picture but he said he's not worried about that I don't know how much to believe him on that but uh that could also really turn the tide here yeah that one doesn't align with what he's previously said I feel like that that if anything from a change of tone when I'm saying
59:01change of tone jobs is the big one that I think the labs are starting to try and back off of what they've previously said yeah I don't believe that they believe that I I really don't and I I haven't I've not sat down and talked directly with you know lab leaders and stuff like that but I I truly do not believe that they believe in the short term it's not going to be massively disruptive jobs they may believe 10 years out that it's going to be amazing yeah but I have never heard an interview or read anything from these people that tells me they actually think we won't go through a pay a phase of
59:37tremendous disruption and change when it comes to jobs in the economy so some more open ai news this week they're reportedly preparing a new model family tentatively named Astra that is built to complete long-running tasks this comes from the information so open ai ceo sam altman demonstrated it to policy makers and regulators in washington dc this past week touting its ability to have multiple agents work together over long periods of time to solve particularly hard problems so reportedly astro
1:00:08would be a new class of open ai models that's alongside sol terra and luna right now there's no word yet on release timing open ai reportedly has not decided whether to label this gpt-6 or have it be another model in the gpt-5 series notably the information reports the astro models are intended to be the first to go through this new framework we talked about that the government now has for submitting uh for evaluating ai models before releasing them to the public um interestingly a
1:00:40day after this report open ai published proof of what they say the model can do they showed how an internal version of astra apparently or allegedly solved 10 problems in mathematics and theoretical computer science that had been open with no progress for at least a decade these were in fields ranging from high dimensional geometry to the lattice problems behind post-quantum cryptography what's more they said finding these solutions cost roughly two thousand dollars in computing
1:01:12at its standard api rates and the model then formalized each proof so it could be machine checked open ai researcher noam brown wrote that the company believes astra will be a major step for scientific reasoning so paul seems like we're at least close to getting some new models from open ai per the government's timeline perhaps and the math stuff seems like it could be a big deal yeah again um a lot of this is you lean on people who know what they're talking
1:01:42about and i know there was there was one um leading mathematician who tweeted somebody's like i'm waiting for this guy to like tell us is this a big deal and he replied in the comment it's a big deal so um you know i think for me big picture obviously there's the new model uh that you know we don't know when it's going to come out or when they're going to call it but it's getting more advanced at its reasoning capabilities its planning capabilities its ability to work on hard problems and that to me is
1:02:14the thing that translates over when i think about ai for business and work i just look at these as a prelude to what comes you know if we can solve really really hard problems decades that have taken decades or all of humanity to to not solve and we have ai that can solve them what does that then mean to hard problems that we try and track within organizations and so that's kind of how i start to think about um you know this stuff and where this goes and you know what it's going to mean and then
1:02:45what is being able to solve mathematics due to solving other hard problems across other scientific disciplines so this and again when you think about the future of abundance these are the kinds of breakthroughs that you can start to see making an impact when it comes to science and medicine and those areas um it's hard to understand this because most of us can't look at these problems be like what does that even mean what's the significance of solving that specific equation but when you zoom
1:03:15out and say okay but it's just working on very hard problems and to my understanding it's not specifically trained to do this that's the other thing to consider is it's not like they're fine-tuning the models to specifically be great at mathematics they're just developing this kind of emerging capability and then again you test that across other environments and you know these these capabilities seem to come out of these models the more you know powerful you make them
1:03:44all right next up microsoft closed out what ceo satya nadella called a record fiscal year this past week they posted 331.8 billion in annual revenue that was up 18 percent um with cloud revenue of 214 billion up 27 percent and azure crossing 100 billion in annual revenue up for the first time or for the first time up 41 percent from the previous year in the earnings announcement nadella said we are advancing the frontier on the cost to outcome curve ensuring every customer can turn tokens into
1:04:15business results and revealed that microsoft 365 copilot has now passed 30 million paid seats he also said conversations per user nearly doubled year over year average weekly engagement with copilot is now on par with outlook and teams and the number of customers with over 50 000 seats is up 7x year over year he also said microsoft plans to bring all its copilot experiences together this quarter in one super app spanning consumer and commercial users he also said microsoft is building what he calls
1:04:49a new model system where the harness context memory and action space are separate from any one model family which basically means lower costs and that every model is substitutable it is using this system in their own products and making this available to customers through foundry they're also nearly doubling their spending on property and equipment as part of their ai build out this fiscal year to 115.9 billion investors liked what they saw they sent shares up as much as 19 as of recording in
1:05:23the days following the results so paul despite some of the uneven feedback we've heard about how much people do or don't like copilot it seems like microsoft is doing just fine that like 7xing 50 000 seat licenses is crazy yeah i just like for if you haven't heard how mike and i do this we literally have a google doc where each topic is sort of outlined as mike saying this i was bold facing the thing he just said yeah that was like it's a it's a large number i mean 30 million paid seats if you think about
1:05:55depending on the data you look at in the united states there's um you know somewhere between 80 and 100 million knowledge workers so if you think about that as roughly the total addressable market for how many people could buy now again i guess i mean i guess you could have consumer side of this too but still 30 million is a a large percentage of people who could viably be using these tools and then the 50 000 seats and up it just shows you like the adoption within organizations and accelerating now
1:06:25yeah from our experience it doesn't matter if you have 5 000 50 000 or 50 most of the time these people aren't trained to actually use these tools properly right so you're like giving the tech to people doesn't mean that the the adoption is scaling and that people are getting massive value from the tools i think it's interesting that you're using the super app language which open ai sort of i think they coined it that was that's the term they've been thrown around there um so those jumped out to me the other thing is how many businesses have been spun up in the last like three months to try and do what
1:07:01they just explained the new model system where the harness context memory and action space are separate from any one model family what that means is you don't need to buy chat gpt and claude and gemini and all these other things because microsoft while they are the largest investor in open ai and have proprietary access to some of their models or unique access to some of their proprietary models um they don't only enable you to use chat gpt anymore they they have deals i think with anthropic and others they can mix in open weight models they can build in their own models that can be fine-tuned for specific
1:07:33work functions like working in excel as an example and so what they're saying is copilot's going to be your router like you're not going to need these third-party companies that are trying to save you money and be more efficient with your token use you're just going to use copilot we will route it to the proper model we will try and minimize your use of tokens especially across like marketing functions or wherever that don't need to be using the most powerful model we're going to get so what they're saying is we're going to solve these headaches for you just give us a little time like we'll figure
1:08:04this out i it's a very um it's a very appealing argument if they can do it like if they can make copilot work on par with chat gpt and claude because it's not right now like it doesn't i think that's safe to say most people who are using chat gpt enterprise or claude you know business they're having a probably a better experience overall seeing more value creation yeah um but microsoft has massive distribution and it's hard to to make that up and it's like we've seen we talked about from our own
1:08:38state of ai for business report data when we ask about what tools people are using like the it's still mostly the majority is still saying they have chat gpt but that flips when you have one billion plus organizations like the lock that microsoft seems to have on the enterprise is wild that can sell 50 000 licenses at a time yeah right and you know one other thing really quick that jumped out they published this blog post about optimizing the frontier performance curve this is mustafa suleman is under his byline he said token maxing has been the story of the last few months but token
1:09:13efficiency is the next big focus across the industry so this whole thing is like to your point solving that problem is deeply valuable and also model resilience they call out a little later basically just saying every business now must assume that any one model it depends on could disappear through a security incident a business or policy misalignment or a geopolitical shift which is a pretty good summary of the topics we've already discussed so far yeah so what they're saying there like if if claude goes down if you're not an x user like it it it's like the world ended so like people who've become dependent
1:09:50upon chat gpt or claude and you lose that model for two hours mike you've been through this like yeah yeah it's brutal and like you realize how dependent you've become on those models so what they're saying again in this environment is you'd never know like as long as you're just using copilot you may be using anthropic models for one instance you might be using chat gpt for another you might be using an open weight model for another but if claude goes down they're just routing you to the equivalent model on another provider and you just move and you never have that and so for enterprises that's a
1:10:22huge value prop like the the downtime goes away we're always going to have redundancies in place so yeah the things they're setting out to solve they're uniquely capable of distributing those solutions i would say they're not uniquely capable of creating the the way to do it but because they have the built-in customer base if they do achieve it they're they're it's going to be hard to compete yeah and i can tell you just in a very very limited sense and then we'll move on um i started taking steps earlier in
1:10:55the year when we started talking about this soft nationalization stuff to be like oh my god like claude is my daily driver model like if this goes away i'm in trouble so i started taking steps to like diversify a bit and make more standardized like my skills and the files being referenced for these different tasks so now it's like you can jump into codecs jump into cloud code say go look at this skill it functions the exact same way i mean there's still preferences and different power rankings of the models but i have become much less reliant on one thing and especially like the projects built in one
1:11:27thing or the files stored somewhere and you're like oh yeah this can be really valuable and you don't notice as much if you're using truly frontier level intelligence i think yep okay so next up nvidia announced a long-term partnership this past week with safe super intelligence which is the secretive ai lab we've talked about in the past co-founded by former open ai chief scientist ilia suskever including what the companies call a substantial investment that bloomberg reports is about five billion dollars as part of this deal safe super intelligence gets access to large amounts of
1:12:01nvidia's flagship gpu hardware including its next generation vera rubin platform which is enough to increase the startup's computing resources by an order of magnitude suskever offered a hint at what the company is actually working on saying its research is quote focused on overlooked aspects of how the human brain functions and added that they now have research that is worthy of scaling up and having access to a big nvidia computer will let us do so so they have kept their research really closely
1:12:31held since suskever co-founded this in 2024 but they had this single stated goal like we talked about at the time about a straight shot research sprint to safe super intelligence they had quickly on that promise and on ilia's background raised two billion dollars from venture firms like entries and horowitz and sequoia capital and reached a roughly 30 billion dollar valuation as of last year so paul after radio silence seems like ilia's back in the news um how big a deal is this if you just got into the ai
1:13:01scene in the last six months or so you know just started listening to the show recently ilia might not be a name you know so just for reference um he was at the frontiers of the deep learning movement back in 2011 2012 part of a team that included jeff hinton that made a breakthrough in image recognition that led to the acquisition of that company which took ilia then to google he was then a major player at google left and co-founded open ai and then he was actually the catalyst behind sam altman's ouster
1:13:38that we referenced earlier he was on the board had come to not trust sam um led to his ouster 48 hours later said he regretted it and wanted sam back because he thought the company was about to collapse and then he was sort of in limbo for months after that and then he eventually left and started safe super intelligence so ilia is a major major player i mean top three probably of ai researchers today um in
1:14:08terms of his influence on where we are in the moment in generative ai so yeah everyone's just waiting like what are they building why are they going to do it we talked i think it was end of 25 he had alluded to the fact that they might actually change their strategy and put some products out in the world originally there was going to be nothing until they solved the the grand goal but he's alluded to a bit of a change in strategy and so maybe that's part of this but yeah and it's fascinating anytime like you know you see nvidia teaming up and giving some level of exclusive compute access yeah it's a big deal
1:14:44and i'm guessing nvidia has seen what they have and obviously believes in it and you know i think a lot of these conversations we have around advancements and auto data research and the conversations incomplete until we know what ilia is working on and um so we shall see all right this next topic comes from our own team so claire prudhomme on our team published a linkedin post this past week about what heavy ai use was doing to her own writing and what she's doing about it so claire
1:15:16wrote and you can go see the linkedin post in the show notes that the more she leaned on ai tools in her work the more she found her writing slipping into prose that sounded robotic and repetitive the better she got it prompting the harder it became to kind of color outside the lines when writing on her own so her response to this kind of feeling of starting to lose her voice a bit her own unique voice was she actually picked up and started writing poetry again which is kind of a pursuit she had had for a while and ai has not necessarily she said freed her up to be more human but given her the
1:15:49contrast to show her what her voice is versus what ai's is and she laid out a few practices for using ai tools without losing yourself that we found super helpful to share so she said that poetry's imperfection helped her deconstruct the structure her writing had taken on and find her voice again um she has taken more time to spend you know time in more in-person communities and events so the friction of being with actual people in person is what pushes us beyond our comfort zones and then she said
1:16:20discernment and intention are key prompting and accepting whatever comes back from ai makes us consumers of our output when it's up to us to be the authors and she extends that last point to companies saying that the company that uses whatever the ai model says starts to sound like everyone else so her bottom line here is that ai has made her faster poetry has made her slower more thoughtful and more creative and the two are not necessarily mutually exclusive so paul this is a really cool read from claire and our team definitely ties into some of the stuff we've talked about this year on the pod yeah i love
1:16:52that she put it out there i mean she and i have had conversations along these lines and so i was really happy to see her you know put her voice to this stuff um the one excerpt i'd highlight was she said we can use these tools without losing ourselves ai has made me faster poetry has made me slower more thoughtful more creative and it turns out the two are not mutually exclusive so just background i mean claire's the incredibly talented producer of this podcast she's very creative she's also very in tune with the impact ai has on creators friends photographers videographers uh any of our ai academy members may
1:17:25recognize claire from her gen ai app review contributions where she often features creative tools and talks about them and the impact but for the context here the most important thing is she thinks deeply about the impact that this stuff has on creative people and she asks challenging questions which i love at our annual meeting this year she actually toward the end of it like our two days together she asked a question that sort of sat with me for a while afterwards about you know what we were doing as a company and our role in you know advancing ai conversations and making sure that we
1:18:00stay human-centered in our approach and that we live that ourselves and so claire along with some of the other people on the team are always pushing me to do more from that human-centered approach and you know for us it comes back to i don't know when i created the tagline i think it was for macon 2019 the first ai conference we ran more intelligent more human was our tagline and that was my belief about the future in essence that everything was going to become more intelligent but in the process it could make us more human and the question about how we bring that to life every day
1:18:35is whether it's through our personal stuff like writing more poetry or in our case with macon like how we create more human experiences where we have artists on site who are doing paintings we have musicians we have time and space for in-person interactions but i mean our event team literally each year sits down and says what are the more intelligent experiences what are the more human experiences and so for me personally like you know again i love just having claire put this out in the world because it causes me to think again more deeply about what we're doing and it's always been
1:19:05about creating more time for me so more time for family and friends more time for personal health and wellness more time to slow down and enjoy and be present in the moments we all experience um but the thing i think the key here is and as claire was illuminating on a personal level at a business level at a leadership level we have to be intentional one we have to be aware that you know we can lose the humanness and all this if we if we let the ai take too much control but employers have to be willing to give some of that time back um because if the expectation from the employer is do more do
1:19:41more do more we're giving you these tools i want you to do more all the time then you're gonna just always feel like all it's doing is just creating more work and that to me ruins the whole potential of ai to to give us abundance which doesn't have to mean wealth and resources abundance can mean time it can mean creative expression it can mean a lot of things and so i think employers have to be intentional about allowing for abundance to be created and personal to people of what is that what
1:20:15does that mean for me what do i get out of all our work with ai so yeah just awesome to you know put a spotlight on claire she does incredible work and i always love when people are willing to sort of take a bit of a risk and like put personal thoughts out there especially on these topics it's really cool yeah and i loved her point about this like almost authorship of like taking control here because that's like what and like determining your approach and your perspective on ai because i just keep coming back to this idea that like the biggest personal imperative is formulating a strong intentional
1:20:49and well-reasoned approach however whatever part of the spectrum you're on whether you like a lot of ai a little ai if you don't decide this someone will decide it for you and that's not a great place to be all right so next up we have our ai use case spotlight where every week we give you a quick look under the hood at some real ai use cases we're exploring here at smarterx so paul i'm going to share one real quick and then hear what you've been working on this week so this past week i uh we released our ai transformation series the first episode of which went live this past week so check
1:21:24that out if you have not already but i was kind of faced with a question here of you know when we record one of these interviews how can that one conversation turn into a much larger body of useful content so i sat down with some uh gpt soul 5.6 uh some extra high thinking in codex to help design a repeatable editorial system for these posts so my goal was kind of to create a little content machine we could run for every interview not just you know spin up random content per episode so basically i gave the system
1:21:57three very different transformation stories that we've already recorded and then i kind of stress tested whether we could find like the same editorial structure through these distinct stories so i could kind of come at this and say hey every time we publish one of these episodes we're going to publish three different types of editorial pieces tackling this from different angles regardless of which direction kind of the interview goes in so so far this seems like it's worked pretty well we're rolling this out right now so we're doing one post that's basically an adoption playbook that explains
1:22:28specifically how a company moved from early experimentation to sustained ai adoption we're going to do a transformation in practice post that isolates one workflow or journey that customers take with ai and shows how the company step by step did it and then do one piece on scaling transformation which is a little more thought leadership around the roles behaviors knowledge sharing operating changes required to make that transformation stick so basically turned each format into its own reusable ai
1:22:58skills so ai each skill can then read the transcript propose angles extract relevant examples and metrics check evidence and draft a first draft in our smarter x voice that gives me clean html to paste into google docs where i do a full human writing and review of it and then once it's ready it converts it into html that pastes neatly into hubspot which takes a lot of time and hassle off our plate so we've just been starting to test this but it's cool to be able to spin up a pretty repeatable system uh pretty
1:23:30quickly which is really fun all right so i was gonna do one this week but instead i want to unpack yours mike because yes sure people who don't know like this is what mike and i did for a living like i owned an agency for 16 years and we largely developed creative content strategies to build awareness audience leads conversions so we did a lot of work around this kind of stuff back in the day and so mike i'll actually ask you um high level break down for me what you just explained which you did since if i'm not mistaken
1:24:09this went live tuesday morning i was driving somewhere tuesday i listened to the episode i was like that was amazing let's focus on an activation strategy because when we used to do this right back in the day we would always say like 20 of the work is the creation of a content asset 80 is the activation of the content asset it's what you do with it so in this case you have a podcast which is the content asset you're starting with but what do you do with that thing besides putting it out on the on the podcast network so you based on what i'm understanding here since tuesday let's just
1:24:45unpack first the creation of the strategy yeah to do this give me what would have been like three years ago versus what it is today yeah so a few years ago we would have i would have sat down in front of a blank sheet of paper and spent a lot of time reasoning through based on my editorial experience and history and expertise okay take it going back by hand or listening again to this episode through the transcript whatever how would i actually take this and turn it into unique different pieces of content not just like summarizing the episode which is great but more like what are the unique spins of
1:25:20like editorial angles that would actually get attention that would make this super compelling and unique almost like you know i used to do as a magazine writer basically um so same idea here except i sat down with a project in codex that keep in mind has already all the context into prepping for these things it's got now the transcripts of the conversations and then i did the same thing i would have done talking to myself but just talking back and forth to codex and kind of hammering out using my domain expertise like hey but also you had provide some cool examples paul of a financial
1:25:54blog that was doing something similar to this which was super helpful seed material which by the way i found doing research last week that was going to be my use case i was going to share i was doing research for a meeting i had and in the process came across a source that was a great example so continue yes so taking those it was like hey here's roughly kind of an example of what we're going for not mimicking it exactly but they had taken some interesting creative angles on a single podcast interview and so work back and forth with codex to be like and especially now with my domain expertise
1:26:25as well just kind of having a sense of what the audience wants and needs and also like what's most valuable to most practitioners i was like okay here's roughly the three angles and then from there it was like okay now let's build skills for each one run each skill went back and forth editing the output and saying like ah you this part was great but you're missing the mark here i think this needs more story and editorial basically just acting like an editorial consultant back and forth with it and then you just say like hey update the skill and now we're in a place where i just did post two
1:26:55this morning uh it's scheduled for tomorrow i think they're coming out pretty well we're still iterating and figuring it out and again it's like these especially are much more heavily human rewritten than i would say some other stuff we do just because this is super important to like really have the human touch on this story but either way it's like we're trying to focus on different angles that are going to be super valuable to the audience but my god like i'm not saying it couldn't have done this without ai have it would have taken it would not we'd not be having this conversation like five days after it so so just ballpark like how much time did you spend putting
1:27:31the plan in place this time versus what would it have been three years well it's interesting the time itself this took me very conservatively a tenth of the time it would have taken probably faster but i would say most of my time was spent just on the plan up front and really refining that as well as refining the outputs it's like the front the barbell it's like the top 10 percent and the last 10 percent were like all my time and energy instead of the middle 80 which was interesting it's awesome yeah so super practical um doable by you know any content yeah creator um anyone can do this if you are if you have that
1:28:10kind of background you're willing to spend time going back and forth with the tools yeah and again a great example of you have the domain expertise and you know decades of experience doing this stuff and so you can you can go in and get the value out of these tools and i think again this is a great example what the future of work looks like someone who is a content strategist and creator by trade can use these tools to accelerate what they're capable of doing and in this case it's so additive because the reality is otherwise we would have just published the podcast and moved on to the next
1:28:41one yeah but we took the time and said well let's use ai to activate this to create more value for people in the authentic voice of you the interviewer and our guest ty we're just taking what they've already created it's not ai slop in any way it's literally like different packaged versions of a great output so yeah it's just an awesome example all right so as we wrap up here paul we've got a bunch of product and funding updates i'm going to run through real quick and then we'll close out this week so first up amazon completed its 50 billion dollar investment in open ai this past week they
1:29:15finalized the remaining 35 billion tranche of the deal announced in february after open ai hit some undisclosed performance milestones this is under an arrangement that makes aws the exclusive third-party cloud provider for open ai's frontier program and expands infrastructure agreements that could total 100 billion dollars over eight years at the same time open ai published a new research report called how ai is expanding what people do at work it analyzed more than 800 000 messages from us chat gpt users
1:29:46it found that 43.5 percent of occupation specific messages involve tasks associated with an occupation other than the user's own this is a pattern they call task crossover and they kind of read it as ai letting workers take on work or at least attempt to that once required other roles real quick i would say it's worth people scanning this report i think this idea of task crossover is something that you're going to hear a lot more about maybe under different terminology but for anyone thinking about change
1:30:20management in relation to ai adoption and scaling of ai this is a critical thing and it basically means that in any given role like a marketer may start doing the work of the salesperson because the ai lets them do it or the salesperson may do work of the customer success team or the or the ceo may do the work of all of them because he or she is impatient and just wants the work done and like so that's what they're talking about is people who couldn't previously do a function now can use their ai agents to do that
1:30:51function and it creates all kinds of change disruption to like how we define roles and org charts and so that's a really important topic even though it's buried here within the product and funding updates indeed uh one more piece of open ai news they also launched chat gpt for academic researchers an initiative giving a hundred thousand scientists and mathematicians free access to the best chat gpt models google deep mind released gemini robotics 2 a family of three models that brings what it calls whole body intelligence to robots controlling full humanoids from feet to
1:31:26fingertips reasoning through multi-step tasks lasting several minutes and adapting to new robot bodies with fewer than 200 training examples they have partners including apptronic boston dynamics and agile robots uh some other google news that's not so positive uh they launched and then pulled a day later an image generation featuring google earth powered by their nano banana model that let users transform satellite and 3d imagery of real places with text prompts they rolled it back after users
1:31:57started generating imagery that violated its policies in all sorts of ways and said they would work on stronger guard rails a munich court in germany ruled that the ai music company suno broke copyright law by training on and reproducing songs from the rep repertoire of the german music rights society called gema and they're holding suno itself liable rather than its users for using those works and ordering the company to disclose related revenue and pay damages still to be determined linkedin added a quote
1:32:32seems like a high slop button that lets users flag low effort ai generated posts from any post menu one of several moves against machine written content uh we also talked about how sub stack i believe it was last week or the week before it started pairing with the detection service pangram to see what posts there on that platform two quick thoughts seems like a slop is basically 90 of linkedin yes and that button is going to be gone within 30 days like you could just imagine seeing the misuse of that thing and it's just going to get to the point where it's like oh my god it's all i
1:33:05slop like my if you have any like um sizable engagement on linkedin posts like the comment section oh my god and and then like the posts from ai influencers like it's there's a lot of ai slop even resolving the comments might be the better play here for them if they can ever do that all right and then our final news piece today here is coursera co-founder andrew ung launched learn vector a new ai education
1:33:35company backed by a hundred million dollar investment from coursera which aims to turn learning from one to many to one to one with personalized ai learning guides rather than chat bars and they have products expected by early 2027. so uh one final announcement here we mentioned the ai pulse survey at the top of the episode go take this week's at smarterx.ai forward slash pulse and in this week's survey we're going to be asking some questions about if you worry about ai use weakening your own skills like we talked about in claire's post and also asking about how frontier ai
1:34:11should be paced or if it should be paced at all so paul another busy week i thought this one would be a little slower but not really but thanks for breaking it down it was slow as the week went on we only had like 18 topics on thursday and then it just blew up thursday and friday yeah all right man well i will uh i'll see you on the golf course shortly sounds good thanks everyone for joining us have a
Conclusion
1:34:34great week thanks for listening to the artificial intelligence show visit smarterx.ai to continue on your ai learning journey and join more than 100 000 professionals and business leaders who have subscribed to our weekly newsletters downloaded ai blueprints attended virtual and in-person events take in online ai courses and earned professional certificates from our ai academy and engaged in the smarterx slack community until next time stay curious and explore ai
More from The Artificial Intelligence Show

#229: Q3 Trends Briefing - The Pope's AI Encyclical, AI Agents Hack Hugging Face, Fable 5 vs. Washington, and the Battle Over Open Weights
Aug 6, 202651 min

Ep. 227: How Good Karma Brands Got Serious About AI and Made It Stick
Jul 30, 202636 min

#226: OpenAI’s Rogue Model, Kimi K3, Open Weights Letter & Demis Hassabis Calls for AI Regulatory Body
Jul 28, 20261h 44m

#225: GPT-5.6, ChatGPT Work, Enterprise Agents, AI 2040 & Apple Sues OpenAI
Jul 14, 20261h 33m

#224: Fable 5 Is Back, Palantir CEO’s Explosive Interview, the Pillars of Business AI Transformation & OpenAI Offers 5% of Company to US Government
Jul 7, 20261h 26m