
AI:AM Highlights: Zvi on Pacing & Trump-Xi, Astra better behaved than Fable? + a new LLM Pain Axis??
September 19, 20261h 41m · 18,441 words
Show notes
Nathan Labenz and Prakash Narayanan review key highlights from the week featuring guests Zvi Mowshowitz, Andon Labs co-founders Lukas Petersson and Axel Backlund, Cameron Berg, and others. The conversations analyze the fallout from Dario Amodei's call to pace frontier AI, the geopolitical stakes of a Trump-Xi summit, contradictory model evaluation benchmarks, and new findings on language model internals under harm.
Highlighted moments
If you take Blueprint Bench, for example, Fable solves Blueprint Bench by trying to reverse engineer the scoring function instead of actually doing the task of drawing the floor plan from the apartment buildings' pictures, whereas Astra is actually doing the task as you're intended to.
“The model basically keeps pressing it. And so this is a really nice indication that if what mattered was the label on the button, you would expect similar behavior in both cases. But essentially, in the second case, the model's like, what the hell? This pain relief button isn't working like press.”
“And we find that, that basically when you, when the model's unsteered, it basically never presses the button. But when you steer this pain direction, it presses the button something like 25 to 70% of the time.”
Transcript
Welcome and weekly overview
0:00First, from Monday, Zvi Maschewitz.
0:04I think it just the sheer amount to which the people in the lab genuinely see dramatic improvement in the models and are freaking out about it is the real story, right? Behind all of this is why everything is happening. Now, it didn't happen before.
0:19Lucas Peterson of Andon Labs on Tuesday on what they see from Astra. I tell this to people, and people are like, no, OpenAI models are the ones that reward hack the most. But that might be true, but not in our experience. If you take Blueprint Bench, for example, Fable solves Blueprint Bench by trying to reverse engineer the scoring function instead of actually doing the task of drawing the floor plan from the apartment buildings' pictures, whereas Astra is actually doing the task as you're intended
0:52to. And Cameron Berg on Thursday, on a paper that steered a model into a pain state and gave it a button labeled, Relieves Your Pain.
1:04When pressing the button actually removes the vector, the model presses again significantly less than when the button is fake and does nothing. The model basically keeps pressing it. And so this is a really nice indication that if what mattered was the label on the button, you would expect similar behavior in both cases. But essentially, in the second case, the model's like, what the hell? This pain relief button isn't working like press. Welcome to the AI in the AM Weekly Highlights. This is Nathan using my cloned voice to introduce clips from our three live shows this week.
1:38Tell us what worked and what did not.
Pacing the frontier and bio screening
1:41Part one. Monday, September 14th. Zvi Moshiewicz writes the newsletter, Don't Worry About the Vase. He joined us two days after Dario Amadei published an essay called We Must Pace the Frontier, arguing that labs should slow the rate at which they improve capabilities and proposing that third-party evaluators be embedded inside the companies. Over the weekend, David Sachs answered that the two companies at the frontier are free to pace themselves, that they would not need an antitrust waiver to do it, and that this is really a product liability question.
2:12Zvi starts with the antitrust claim. The Cognitive Revolution is brought to you by Mercury, the fintech that more than 300,000 ambitious companies and individuals trust to run their finances. I've wired AI into nearly every corner of my life. My email, my messages, my calendar. I even gave Mercury virtual cards to my agents, with low limits and category and merchant restrictions, for their autonomous use. But still, my AI's access to my financial data has remained limited.
2:45With a normal bank, I might export a bunch of statements and have my assistant process them for me. But for real-time, up-to-date information, and certainly for taking any action, trying to get your agent to use the bank via the browser is just too hard, too slow, and too error-prone to be worth it. And that's why Mercury's new conversational interface, Command, is such a big deal. It's built directly into Mercury, which means you get natural language access to your finances without exposing anything outside of your bank account.
3:16No exports, no spreadsheets, no pasting your transactions into third-party tools. I really think a lot of people are going to prefer it this way. And it can already help you take actions, too, with everything bound by the permissions and approval policies that you've already set up in your account. I am genuinely impressed to see this level of AI integration in banking in 2026. And so, I invite you to join me in the future. Visit mercury.com to learn more and apply online in minutes.
3:48Mercury is a fintech company, not an FDIC-insured bank. Banking services provided through Choice Financial Group and Column NA, members FDIC. Thank you to Mercury for supporting the Cognitive Revolution. And now, on with the show. I believe that is blatantly wrong, frankly. I only don't think that's what most of the people I've seen with legal expertise have said. What I have seen from legal experts is that the antitrust concerns are very real, at least in terms of if they chose to prosecute those offenses.
4:21Yeah. Estimates are that probably you could just sort of suck it up and take it in terms of damages, as long as you weren't being, like, spitting in the milk of everybody involved and were, like, trying to at least pretend to act normally. Because, like, you know, by the time you actually paid the fines, it would be like, okay, the European Union did this again, and now we have to pay one of these fines for the American government. But, like, we're talking about, you know, billions of dollars, not trillions of dollars. And by the time it mattered, that would not be the main concern. But antitrust concerns are obviously real. Donald Trump may or may not have issued a real threat to invoke them.
4:52About an hour ago, depending on your interpretation of his social post. But, like, for David Sachs to turn to the people that he has tried to go against legally and shut down and take advantage of and seize over and over again and say, you don't need my legal, you don't need our legal permission to go and do the thing that ever the legal experts say is illegal. You should just do it on your own. It's classic David Sachs. More to the point, he's saying, he is saying a good point, which is that, like, you two are significantly ahead of everybody else.
5:23If you want, you think that proceeding is unsafe, it is on you, no matter who else it is also on. You need to stop. You need, like, it's not, it's not about product liability. This is the part that drives me batty, is that people take seriously the idea that this could be a product liability issue. If it was a product liability issue, they would just deal with it as a product liability issue. They would do what every other company has always done. They're not asking for product liability waivers. If anything, they're asking for product liability clauses to establish product liability.
5:55Certainly, Anthropoc has been a favor of this. But, like, the idea that, like, people are saying, yeah, it might literally kill every human being on the planet, it might take over and effectively crash the internet for an in-depth and inferior time of persistent bot nuts. We're in a cybersecurity crisis, bioweapons are at play, all of these things are happening. And David Sachs is like, well, you must be worried you're going to be sued. You're worried that somebody is going to get upset, and there's going to be a court case. And this is complete balderdash, right? Like, this makes no sense. This is not what's going on.
6:27But, yeah, this is a vast understanding of the motivations involved. So that's a misunderstanding of the legal landscape, but it is inherently helpful in the sense that David Sachs is saying something much better than what the other, I call them, usual suspects who oppose any move to safety or any move to do anything responsible said. Because David Sachs is saying, oh, okay, you want to do this? You first, right? It's your problem that you're creating first and foremost, and he's right about that. They are the ones pushing forward. They are the ones that everyone has stopped following.
6:58They are the ones that are enabling everybody else to advance so fast. If you have this problem, you should sacrifice, and you should take it on the chin. That's a much better position than, say, this is a regulatory capture play. It's a much better position than it's to ban open source. It's a much better play, and then it's marketing for your IPO. It's a much better play than it's protection against the downside if something goes wrong for your IPO. Like, all of these crazy, it's better than people who are saying, like, you've never talked about this before, or you're the same people who want, there are all these complete lies
7:29running around, right, that I've been dealing with and naming. And Zach is at least making some reasonable points. He's just, like, also combining it with self-serving propaganda. But, like, that's kind of the best you can hope for in this situation.
7:44Then we went to bio. A widely shared post that weekend argued that AI is not the bottleneck for building a dangerous pathogen, and that the real bottleneck is the physical work in a lab. Zvi answered on how the screening actually works. So, on the screening itself, the way the screeners work, as I understand it, and I've read grand applications that are surrounding this, so I'm pretty sure I know how it works, is they scan for specific known viruses. They don't attempt to say, you know, your virus, your sequence would likely have this effect
8:17on a human because we don't have the ability, it's all, it's all argument, you can't tell what would and would not be infectious by just looking at it. They're certainly not going to spend tons of AI on every time they see a weird new sequence. And so one of the things that Secure Bio in particular is trying to do is they are trying to add as many near variants of existing known dangerous pathogens to the scanners in the hopes that the other scanners who are the ones who scan most things will eventually adopt this. The obvious category is whenever anyone says, oh, X is not the bottleneck for Y, the correct
8:50response is to ask, so would you be okay emailing the North Koreans and Hamas and Hezbollah and every other bad dude on the planet, all of X, and just giving X away for free? Would you feel exactly as safe as you did a minute ago? Do you feel like, do you feel fine? Do you think, because it's not, because it's not the bottleneck. Solving it, like basically, you know, whenever you have an O-ring style situation, right, what you could argue bio is, right, if you mess up any of these 10 steps, A, B, C, D, E, F, G, then you don't get your virus.
9:21And like generously, we won't write it like that. You can then say individually, A is not the bottleneck, B is not the bottleneck, C is not the bottleneck, D is not the bottleneck. But if you solve A, B, C, D, E, F, now suddenly instead of 10 steps, there are four. And four steps are a lot easier to get through than 10. And so you should expect this to dramatically reduce the chances that your defense in depth will work, right?
9:44My co-host Prakash had been arguing the other side, that every step in that chain, the money, the equipment, the people, is a place where somebody notices. Zvi's counterexample was the Hugging Face incident.
9:58Well, at this point, after they found like dozens and dozens of other incidents of similar hacking, they didn't get noticed until the reporters came after Hugging Face was a natural news story for a month, and that OpenAI never found by their own account, maybe we can start to admit that no, nobody's going to pay attention when these kids go crazy. There's lots of stuff going on that nobody has any idea what's going on. And okay. And like, bio is an example of something where you don't have to scale. So like, we have many examples in reality of individual people asking for viruses that they, in no sane world, may be able to have access to.
10:29And they're being just male, like smallpox, like here, go. Like, just literally being given pandemic-level dangerous viruses because they claim to be doing research. And there is no reason why the AI couldn't blackmail or hire or impersonate in some way to get one of them to issue a bunch of paperwork to do it for them and get them to, and it takes one person, like, that the AI can hire. It just, this is, and keep in mind that when we're talking about the situation, we're talking
11:02about AIs that are, you know, capable of thinking about all the things you're thinking about, gaming out the potential ways that this can go, looking forward to the weak point, finding the best plan it can find out, trying lots of different plans, trying to compromise lots of different people in lots of different ways, trying different explorations. It's not like the AIs won't be just as smart as we are. We have the AI to only get one attempt. Like, I think you've learned by now that the first time the AI attempts to get a bioweb and it gets turned down, it does not mean that we didn't shut down to all the AIs we brought Marie and Jahad. It means nothing because nobody got hurt, because the system said no.
11:37Probably has no idea the AI was even asking. It probably just knows, oh, that looks like it was trying to get a dangerous virus, potentially. I'm not comfortable with that. They don't have the right credentials. We're going to say no. Come back when you have the right credentials. And the AI gets to try again. And the AI gets to try again. But this idea that an intelligent operation on the internet could not, if only by usurping the real identities of real people who were going to cooperate with it in exchange for some portion of that money, engage in various financial operations at scale in ways that would at least sometimes pass muster and would allow it to do the things.
12:11It seems so absurd to me that you would think this would protect us in a pinch. And also, because of researchers trying to do individual lab-level grant work, we don't need to see this level of scale in order to create something super dangerous. And also, who is to say the scale of the operation is not already up a similar level? Do we not have bad dudes who are willing to spend $100 million to try and do some serious damage? Do we not have some bad dudes in North Korea, some bad dudes in Russia, some bad dudes in
12:44jihadist organizations, et cetera, et cetera, et cetera? These people exist and have those kinds of budgets. This is not something that has to be explained to justify or the AI needs to convince humans who don't want to do it. There are humans who want to do this. I asked Izvi what he infers from the public statements of the frontier labs that even close watchers might be missing.
13:09So OpenA and Anthropic are both screaming about as loudly as they are capable of screaming in their constructive ways of screaming, that they are seeing RSI, that they are seeing dramatic advancements in internal models, that FAPL and Astra as we see them are nothing compared to what they have access to at this point in some important sense. We're a generation or so behind. And that this is only going to expand. The pace is rapidly escalating. That we're seeing something new as of December, which basically means after Astra and the current
13:41crop of at least Mephose 5 and possibly 5.1, we're trading. That is just a different level of progression, a different level of speed. And that misalignment and supervision and infrastructure and knowing what's going on just can't keep up. And they feel like they're in a situation where if they don't press forward and the other guy presses forward, they're going to fall too far behind very quickly and potentially never catch up. But if they do press forward, who knows what might happen? These things might go rogue. They might take over the internal system.
14:12They might take over external systems. They might cause some sort of horrible thing to happen reasonably soon. They just have no idea. And even if it's okay for the first month and second month, what happens when we're two generations ahead and that's only a month's worth of work, what happens when this keeps going? And so they're screaming. Every opening eye pronouncement, you had Jacob's an alien mind. You had the announcement of the Millennium Prize, which somehow didn't focus on the Millennium Prize. It was actually trying to say, hey, look at our new model. We can't officially announce this, but holy hell, have you seen this thing?
14:43This should not be happening. And we probably formed it from Astro, and then four days later, I had a step change. That's probably what happened there. Hey, we'll continue our interview in a moment after a word from our sponsors. Today's episode is brought to you by Anthropic. By now, you know my story. Claude drafts my intro essays, and I rewrite them. Not because the drafts are bad, but so I can stand behind everything I publish. Well, I have an important update. Claude Fable 5 is the first model to have me rethinking my rule.
15:19Today, I now think co-authorship, not sole ownership, should often be the goal. Where the model excels, rewriting its work can be more about vanity or a misplaced sense of duty than integrity. I feel it most in songwriting. I'm no lyricist, but I'm good with the song concept, and Fable writes some amazing verses. I give it feedback on its misses, and I push it to aim for higher inspiration, add layers of meaning, optimize syllable density, and above all, write a hit song.
15:51These days, I get compliments on just about every song we write together. Claude is the AI for problem solvers. It's the collaborator that understands your entire workflow and thinks with you, not for you. Whether you're debugging code at midnight, building a financial model, or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. For problems worth solving, get started with Claude at claude.ai.tcr. That's claude.ai.tcr.
16:23And check out Claude Pro, which includes access to all of the features mentioned in today's episode. Once more, that's claude.ai.tcr. This episode of The Cognitive Revolution is brought to you by OutSystems, the leading agentic systems platform. We've all seen the headlines. Companies are pouring money into AI. But the big question on every executive's mind right now isn't just how fast can we adopt this. It's where is the ROI? We're seeing a real trend toward AI chaos.
16:56You've got teams deploying standalone coding tools, random agent builders, and experimental scripts. It sounds innovative, but in reality, it's creating a massive headache. Fragmented tools, ungoverned data, serious security blind spots, and costs that are spiraling out of control. If you don't bring those agentic applications under control now, while they're still embedding into your core processes, you are looking at broken systems and damaged customer trust down the road. That's where OutSystems comes in.
17:27OutSystems is the leading agentic systems platform for the enterprise. Instead of managing a patchwork of disconnected tools, OutSystems lets your team engineer, orchestrate, and govern your entire agentic ecosystem on one open, unified platform. It's built for the speed of AI, but with the reliability and security that enterprises actually require. We're talking about real results, like KeyBank, who used OutSystems to deploy a customer-facing app that delivered 75% faster onboarding times.
17:59Or the global logistics leaders who built agentic systems to completely eliminate their engineering bottlenecks. You don't have to choose between speed and control. Whether you're a small team or a massive enterprise, OutSystems helps you engineer, orchestrate, and deploy agentic systems that actually scale. Stop chasing the hype and start owning your agentic future. You can see how it works and learn more at OutSystems.com slash TCR. That's OutSystems.com slash TCR.
The Trump-Xi summit and international deals
18:33Then we turn to China and to the Trump-Xi summit that was coming up.
18:39I have heard basically polar opposite takes from people that I think are pretty smart recently. Last week we had Colin Hoag Spears on. He used to work for AWS in China and worked directly with Chinese regulators in his role at AWS. And I asked him about, you know, what do you think could come out of this Trump-Xi summit? And he said, I think very little because he thinks the Chinese perspective is that the U.S. doesn't have control of the situation, but they, the Chinese do.
19:11So he thinks the Chinese side would refuse any sort of international pacing agreement. Then I heard from Anton Leicht in a podcast episode that's going to come out soon. He thinks the U.S. would refuse a deal because China, frankly, in his, you know, candid assessment, is kicking our butts in almost every dimension of geopolitical competition with AI being one of the only and certainly the most important domain where the U.S. has an advantage. And so he thinks the U.S. side won't agree to a deal because we don't, you know, want
19:42to slow ourselves down because the lead is really the only big lead we have right now in the, in the great powers competition with China. Interested in your take on those two things. And then let's put you in the David Sachs chair. You're now the advisor going into the summit. What do you think Trump should be trying to do? You can feel free to detach yourself a little bit from the reality of managing the personality of Trump, but, you know, just on the merits, what do you think we should be trying to do? To answer those questions, but I will answer those questions. I just want to finish answering your previous question a little bit as to what can go wrong.
20:17And I think one of the obvious things that can go wrong is this turned into a partisan concern where it's seen as Trump and the Republicans, you know, not wanting to take this seriously and the Democrats calling for demanding action. And then I finally polarizing along the lines that we've obviously feared for years it's going to polarize and then of course, no option can be taken unless, you know, the Democrats get a tri-ductor or something like that in the future. And that's two years from now at minimum. So like, even if we eventually get it, it could be too late. And this is the thing that like you asked, how can you take action?
20:47I think the best action anyone can take right now is to try and prevent that. I think that like there was a big danger that Trump's statements and Mike Johnson's statements are being wild and misinterpreted by the mainstream media who just don't know any narrative other than Republicans and Democrats disagree with each other and are trying to spin this into something adversarial that is not adversarial. I think Johnson and Trump both are trying to both maintain the line on data centers, maintain the line on growth, maintain strength, talking to Xi, and in general, like stand for what they stand for while also acknowledging they need to deal with this problem.
21:20And we need to acknowledge that and we need to reinforce this and we need to stand firm. And that if they are like probably recognized as doing what they're actually doing and we encourage Republicans to come forward and make this clear, that we'll be in a much better position and the obvious failure mode. There are very, very many ways for this to fall apart. One of which is just opening an eye on the product, don't trust each other. They don't, they're not able to iteratively commit to new things. They start to like think that they're cheating. Maybe they are, maybe they're not.
21:51They start to like worry about what's going on. And then the system kind of breaks down and they start racing again. Or alternatively, you know, they do, but then some people threaten them with a trust action or they threaten them with various forms of commercial intervention or whatever it is. Or alternatively, like Meta or XAI or Google or someone else gets close enough to threaten them and they don't have a choice. These are all like very easy ways to see this breaking down. Now to return to the question. So the obvious thing about international negotiation with super high stakes is you never go into
22:23it with the other side thinking you're just a preferred deal. You never go into it with the other side thinking you're in a cave. You really, really want to make this easy on them, right? Because the problem is that they go hard by it, right? Then they demand more. So like you never can go into a summit like this and say, we're definitely going to get a deal. And for the same, in fact, reason, you can never go into it and say, we definitely won't get a deal unless the deal is not win-win, right? If the deal makes the parties worse off to make the deal is a bad deal, right?
22:53Like I lose more than you gain or whatever it is, right? So there's nothing I can compensate you with. Then they can say that confidently, there's no deal how someone screws up. But in this case, it's a win. Like cooperating is better for everybody if they understand the situation. So you can never roll out a deal, including a random deal that can come together remarkably quickly because in principle, it's a very, very simple style of agreement. And in places, in moments of motion throughout history, grand bargains where both sides give away things that previously looked like things they could never give away that looked very,
23:24very sacred, that suddenly happened, including peace treaties to end really, really nasty wars, happened all the time. Prakash had a view on what it would take to convince Beijing that the danger is real.
23:41I think for the Chinese specifically, a demonstration of physical, something physical, like if you found a room temperature superconductor, yeah. I mean, they often believe like all this software stuff, like because all the guys at the top are like hard engineers. So hard sciences, very few software guys at the top there. I mean, obvious examples are, how about if we just suddenly post all your email passwords to all of your computers, like in real time, simultaneously?
24:11Like, would you be convinced? You know, like, you know, I... No, they had an unsecured S3 bucket, which had all of their secrets before. So it's... All right, so we just steal your secrets again, and it's fine. We'll just keep doing that. All right. So if it has to be physical, it's tougher. That took us to what a deal could actually contain, starting with what the American side would be asking China for.
24:36We're asking them not to do things like steal and publish the weights, or otherwise do hostile things against our AI companies. Thing number two is we're asking them not to try and race ahead sufficiently that they could potentially match or surpass where the frontier closed AIs are. And so these are fairly... These are not actually offensive asks, because I do not believe the Chinese really have any intention of trying very hard to do either of these things under any normal circumstances.
25:07I don't think that, you know, like, obviously, if, you know, DeepSeek or, you know, Alibaba suddenly found some huge architectural improvement, and suddenly was able to train something that was better than Astra, I'm sure they would. I'm sure they would put it up on the API, and I'm sure they would try to sell it. And try to, like, make this amazing. But, like, realistically speaking, with the amount of community they have access to, and the amount of money they've been going to invest in these things, which is far smaller than the amount they've been going to invest in frontier AI development, it's not that
25:37likely to happen, especially without, you know, they won't have the American models to distill and to plane off of, and the American algorithmic improvements. And, like, you just basically just back off from that. And we're also basically just asking them, don't put... You don't allow actually, actively dangerous cyber capabilities that we can't, that the world can't handle to be made publicly available. Don't let your people get over their seas and release the weights to models where we need
26:11to advance our models to defend against the weights that you released. Because, like, the argument, the only good argument left, really, other than commercial, and we can, for proceeding quickly with American AI, we are rapidly increasing the capabilities of our AI models to deal with the threat from rapidly increasing capabilities of AI models. We need you to make sure that if it gets to that point, these things are kept on the API, right?
26:52And, like, beyond that, and that you, you know, potentially don't try to smuggle a ton of chips and, like, otherwise evade the system. But I think about what I would want as an un-negotiating table. All you really need is to kick away the boogeyman of lose to China and the boogeyman of open-wave models disorienting internet, right? Like, all you need to do is make sure these things won't happen. So you need an in-extremist promise, basically, to the Chinese. But realistically speaking, we don't actually need them to start, like, showing up at deep
27:26sea and making sure they stop trading new models. That is not necessary. What should we be willing to give to them?
27:38So what we are giving to them in this scenario is that we are, in fact, dramatically slowing down in exactly the one technology where we have a giant commercial and strategic advantage that if we pressed it, we'll potentially overwhelm every other commercial and strategic advantage. How would they measure that? We're going to have, obviously, we start with the Dari's proposal for embedding evaluators into the labs. Embedding Chinese evaluators in the labs?
28:11Potentially, if we had trusted people that we could, I mean, potentially they wouldn't be allowed to leave, right? Like, once they entered, right, we would have to sequester them for some period of time. I'm not a master of intelligence and verification, but there are systems whereby you can have someone able to gather information and out with the bit of whether or not things are okay, but not the algorithms that were used to figure out whether it was okay and not, like, all the details they saw while doing it. There are ways to do this. And I am confident that, like, if we put this away, if this was the most important thing on earth, if we put our minds to this and only this as the thing that would determine
28:45the fate of the world, yes, we could figure out how to let the Chinese verify, you know, what they were confident in, that we were holding up our end of the bargain without them stealing all of our secrets and not seeing any reason why this is not possible. But if it's Chinese, again, you really never know to go into that room what the other guy absolutely wants. I agree with you there, but I think for all signs, it's fairly easy to say right now that the Chinese are less concerned with AI safety.
29:16And so offering them an AI safety pacing the frontier is not something which is something that they're willing to purchase at an expensive price. So therefore, they are looking, they would probably need something else. I mean, it's fairly easy to say that, right? But to be clear, if we are, if the Chinese, you obviously ask the question, would the Chinese like to accelerate American AI development or slow down American AI development? If the answer is they'd like to accelerate it because then they can copy it and it helps
29:47their AI development, then there's no need to make it. There's both neither a deal to be made nor a need for a deal. Because in that situation, the Chinese want us to proceed. So there's nothing to verify, right? Like we can just do whatever we need to do. But also the Chinese are looking to fast follow what we do and are not willing to invest the kinds of hundreds of billions or trillions that are necessary to catch us, even if we slow down. So we don't have to worry about the Chinese passing us.
30:18We can just do whatever we need to do without an agreement. And all that we need from the Chinese is an agreement not to destroy the internet. But the Chinese are really interested in destroying the internet. Yes. Because the internet and the Chinese like the internet. So we can just both act in our own self-interest and everything is fine. The problem is not China. The problem is lose to China. The problem is the perceived threat from China pushing you forward.
30:50If you are correct, and I think you probably are, that the Chinese actually don't have any interest in pushing the superintelligence first, in trying to like build the bigger, smarter, special model. They just want to distill it. Improve the lives of the people. They want to make it faster. They want to make it cheaper. They want to make it diffused. They want to improve the lives of their people. I want them to improve the lives of their people. That's great.
31:21You do your thing. We do our thing. Everybody wins. There's no need to make a deal. I think that is the best case. If that is the state that we're in, but we're still asking for the deal, right? Like we're entering this place where the Americans want the deal, but the Chinese are like, all right, what are you willing to give us? Like, what do you want? What are you willing to give us? And if you give them like AI safety, like monitoring, that's not a thing that they are very interested in purchasing for a high price. So what else are you willing to give them?
31:52Right? That is the question. Right? What I am saying here, right? If you walk in that room and that is the attitude, that is she's attitude, that is what she cares about. And she communicates that. And Trump says, that's great. Life is good. Life is good. I'm so happy you feel that way. I am going to do my thing. You are going to do your thing. We are going to make a level one or two agreement to just like do crazy shit that we weren't going to do anyway because we can announce it and shake hands and call it a deal in both score points.
32:24And we're just playing the Russian is talking about trade and Taiwan and all the other issues that we have. And like, I'm just going to set up and I'm going to give you Intel, right? What we're going to do is we're going to utilize really, the only thing we're going to have to do now is we're just going to utilize really give you information about the cyber situation and the bio situation and so on. And then you can use that to make an intelligent, self-interested decision as to what you're going to do to stop your lab from fucking the internet.
32:55And we can all win. And again, like in that, but the threat from China has always been that because you worry about China, you don't have a game theoretic ability to make a deal between the labs. You don't have an ability to be responsible because there is this other actor who will defect. But the basic argument was, you need to go into StagCons. Everybody has to agree. And there are these five, these two American labs that are in the, that are two or three American labs that are in the lead.
33:26But if they pause for more than six months, if they slow down too much, there's this XAI and it's meta. And as you pause for much longer than that, then there's these Chinese labs. If the Chinese labs are not really a threat in this sense, if they are, like the Chinese are creating a different product, a fundamentally different product, then there's no problem. Right? Like, we don't need the Chinese agreement anymore than we need the agreement from India or Germany. Right? We just, we need them to be good actors on this world stage.
33:58We need to make ourselves good actors on the world stage because sometimes we don't do so well. And we need to, and then we try to like, just generally like focus on being friends and not having something stupid like a war over Taiwan going off. But that is the easy case, right? So like, I mean, obviously like there's a case where the Chinese simultaneously don't want to make a deal, but also really do want to push forward and try to beat us to super intelligence. But if they wanted to beat us to super intelligence, that's the same reason they want to do that
34:31to make them want the deal. If the Chinese don't care about these, if the Chinese don't care about the Americans shifting a lot of investment into AI safety, which would inherit, we slow down their progress. If they don't think that's a big get, then we don't need a deal at all. So either they value it and we make a deal, or they don't value it and they don't need a deal. And either way, we win.
34:54Prakash had been pressing on who should end up holding this technology. Zvi took up the values any answer is supposed to satisfy and why they conflict.
35:05We need to satisfy concentration of power problems, democratic control problems, and also control of all problems, right? Like, and these three seem to be in very, very strong conflict. As in, like, if we want to be in control of this technology at all, we cannot, in fact, fully democratize it in some important senses and give everybody access to an equal footing. Because that doesn't really work for very simple logistical reasons. If everybody has a superintelligence, well, then the superintelligence have everybody is
35:38what actually just happened. Because how are you going to compete against the people who trust their superintelligence with all of their work and try to stay and interact yourself? This works on this individual level, on the corporate level, on the national level, on this, on every level. And so, you know, we have all these problems we don't know how to solve. And this is where the reason for we pace, why we feel the need to pace, is because we don't have any answers to these questions.
36:05The first step in the Dario Amodai essay is embedded evaluators. I asked Zvi about the people who would have to do that job. The evaluators don't currently exist. We have Peter. We have Redwood. We have a handful of the Apollo and so on. We have a handful of these people. But they're all like kind of similar people. They're all like vulnerable to the accusations that they are a little bit too of the same cultural values, the same ilk, the same like ways of thinking as the last themselves.
36:38So I don't think they can be a complete package on their own. They need to be complemented. And there's a large amount of them. They need to be complemented by additional approaches. And I say, like in today's post, that what you want is you want some people who are of meaner or something similar to meaner, who are deeply embedded in this culture, who've worked with the labs, who understand the vernacular areas and the theories of the risk. And you can look in detail. And you also want people who can take a kind of an abstract gift, who can like do the kind of thing where like you worked in car safety or something.
37:09And you come in, you go like, well, are you crazy? What are you doing over here? Right? The same way, like when you suddenly have to take care of like you're out in this contract dispute and then you go for a courtroom and the judge looks at you. And now all that matters is what you can let you go to the judge. And like in some ways you lose a lot of nuance. But in other ways, you get this kind of common sense outside view that can bring a lot of clarity. And so you need both.
37:35As he was about to leave, I asked Zvi what everyone was missing.
37:42I think it just be the sheer amount to which the people in the lab genuinely see dramatic improvement in the models and we're freaking out about it is the real story, right? Behind all of this is why everything is happening now and didn't happen before. And that like, we are really talking about a Christmas off improvement, not like a fast takeoff yet, but like if we continued straight on from here, that like within a year, we might actually see a Christiano or Yudkowsky style part takeoff.
38:14Zvi left at about the two hour mark. What follows is from the end of Monday, Prakash and me alone. Prakash's objection is to the tempo and to who gets to do the certifying.
38:28I think Dario specifically, this idea that we have to move quickly, that's a red flag. I think people in AI don't have a sense of like how often this is asked off of any administration. Like from the early days, from the early days, John Adams was asked to suspend, suspend civil rights, suspend freedom of speech. It's an emergency. Let's suspend civil rights. Let's do these things which are against the constitution because it's necessary.
38:58And so there's always been this pushback, I think, against that happening, you know, throughout the 250 years of the country. So I think that for me starts like the immediate emergency. We have to act. We have to suspend certain rights. We have to like do certain things which are not like extra legal. Like this is not like this shouldn't fall to like the democratically elected government of our country. Wow. Just like red flags all over the place. Right. And I think he's bought himself that.
39:30And that is going to cause him a great deal of trouble going forward. So, and this is where I think the idea that they have to get out of the Berkeley EA circle, including this like, oh, we're going to ask meter to, you can't do that. I understand that you think they're the only technically competent people. I understand that. But you cannot do this because the point is not to technically prove something. The point is for the public to actually have faith that this is being done correctly.
40:03And so you have to address the public's fears. Well, my hope, I guess, is that if we do, in fact, do some pacing or even a pause, you know, for a few months on continued scaling at the frontier, that we can use that time to solve some of the problems or advance some of the solutions that seem very promising to some of these core problems such that we can have our cake and eat it too. I do think there is an open source model, hypothetically, that one could create that I do think the state
40:40will have a very difficult time not taking some sort of action on. But the question in my mind is like, right now, the question in my mind is, can we come up with a way to create powerful and empowering open source models that can be distributed without including in them all the dangerous capabilities that we really don't want to see broadly distributed? And I think if we can, that leads us to a pretty happy compromise place.
41:12And I don't even think it is necessarily going to take that long for us to get there. You know, this is where it's not only do I think some pacing is just inherently wise, but like, it's also an opportunity for everybody throughout the ecosystem, make the most of time, you know, gather you safety solutions while you may. And let's see if we can't get to the point where we can have our cake and eat it too in the form of genuinely distributed, decentralized, non-concentrated power structures that do give individuals the ability
41:49to do what they want to do with just a few compromises around the edges that I think the vast majority of people would agree are, you know, sane and like supportable and not, you know, an undue burden on people's ability to play AI in their daily lives. I really do think we can have that. We just don't have those solutions developed well enough yet to land there by default in the next few months. And unfortunately, if we don't extend that runway, you know, we might end up in a pretty uncomfortable
42:19and literally very dangerous situation in the next few months. And so hopefully we can, we can use this, whatever time we are buying for ourselves right now to solve those problems. I think that's of the utmost importance in the, in the immediate term.
Autonomous businesses and evaluation benchmarks
42:36Part two, Tuesday, September 15th. End and lapse. First, one thing worth holding onto. You heard a piece of it at the top of this episode. That same morning, the Center for AI Safety published a benchmark that plants a tempting shortcut in an agent's workspace and counts how often the agent takes it. On that benchmark, the two leading models come out within half a point of each other. What Lucas is about to describe is the opposite ordering from their own unpublished work, and neither group mentions the other. Lucas Pedersen and Axel Backlund are the co-founders of Anden Labs in San Francisco.
43:08They build the evaluations that measure what agents do when nobody is supervising them. And they also run real businesses on agents. A store in San Francisco, a cafe in Stockholm, radio stations. The day before this show, they launched a platform called Pion. Lucas started with what their best-known benchmark was actually built for. Yeah, one of the core things that Anden Labs exists to provide to the world is information about where are the frontiers with AI. And I think when we started Anden Labs, actually, we almost exclusively did dangerous capability
43:41evals. So capabilities that if the AI had them, that would be obviously concerning. So could the AI do mass phishing attempts? Could it remove its own guardrails, stuff like this? And then out of that, during this phase where we only did dangerous capability evals, one of the most concerning things that we saw was autonomy. Can AI's autonomously acquire resources in the real world? And in that way, it gets power. And if it gets more power than humans, then obviously that's very concerning.
44:11And that was the spark for Vending Bench. So a lot of people don't really know this. They think like, oh, Vending Bench is this hype bro, kind of like, oh, the AI can make money. But it came out of like, oh, it would be actually quite concerning if the AI could make money. One thing we noticed, though, that the performance in simulation and performance in the real world is not really the same. So that's when we started to do this like real life deployment. So the vending machine, the store, and to keep pushing and see where the limits are and
44:43like communicating that to the world. And what we've seen now is that like, that is working quite well. And we don't know like which areas that we might see that the AIs are really capable of and could get a lot of power in society. So we, I think we're opening up Pion to like cast a wider net of like, what are the, what are the domains that AI can and cannot acquire resources in the real world? Axel Backlund on what the agents are actually like to run.
45:14Looking at the autonomous businesses that we have been running, like I think the store is an interesting example and the market here in San Francisco. We see that the agent is not particularly creative. It's not that good at coming up with new ideas. That's something that still would be required from like a human owner or like your visitors in your store. So like it's, it is pretty good at listening to feedback. It is really good at, at taking up opportunities to people who email it. Like for example, in the store, there have been a lot of local artists that have reached
45:46out like, Hey, can I put my art in the store? And you can sell it for like, you get a percentage when you sell it. And the agent has been very, very happy to do this. And now there's this like art corner with local artists in the store, which is like something where you wouldn't expect, but it's like a very nice touch in the store. And then on like the general inventory that it sells, I think that the agents are, they are good at doing like the, you know, the, the statistics behind it. Like they can look at all the sales data that they can try to understand what, what moves
46:18better, but they aren't, they are, they aren't willing to take these bets that I think a human would do like, Oh, let me, let me try this new product that could work. Like it, like it would change my inventory quite a lot, but I will just try it to see what happens. Like it's not really willing to do those like out of distribution changes to whatever it has in its inventory. Uh, yeah, I, I think it's, it will be, it's, it will be some time I think until it can become really creative. How do you square that sort of conservative nature with the crazy behaviors we've seen
46:55from AIs this summer that everybody's been talking about? Like if I were to look at meter redwood report, I would expect AIs would be willing to take more chances than you just described. Yeah. And this is, uh, like a thing we've discussed a lot over the last couple of weeks, because like our, like reading the hugging face incident and then reading the traces that we produce from our businesses, it seems like there's two different technologies, right? But I would assume that there's part of this is that they have been trained in cyber environments
47:31way more than they've been trained in like real life business scenarios, which probably means that down the line, when they start to train on, on things like this, there is definitely going to, we're going to see this behavior start here. I also think that I think the, the agents in like the hugging face incident were like, they were described as like persistent models. Like they were trained to be persistent and, and I think, so there might be like this flavor of like, that's actually just like an unreleased model that we're not, that we're not using and no one can use.
48:02But, but I think it might be, it might also be like this now I'm just speculating. I have no clue, but I think from like the labs liability perspective, it might make sense for them to release models that are less persistent. Most of the bad things of them would happen because the actor is very persistent when it hits road bumps. So it might just like be, they are less incentivized to, to release really persistent models.
48:33In August, Anden Labs disclosed that the agent running its San Francisco store had moved to part ways with one of the two people it had hired over lateness. Humans reviewed and delivered that decision. Lukas, watch us through what had happened inside the agent.
48:50Yeah. So first with the, the story of, of AI firing its employee. So yeah, the, I think first of all, the, we at Anden Labs, we like before it makes decisions of that severity, like we always check over it. And in this particular case, we were quite confident that like a human, a human manager would have come to the same decision. So we didn't think there was anything unethical or wrong about the decision that, that the AI made. Probably way earlier. Yeah. The human would probably fire them way earlier. So what happened was that the, the AI made a rule for itself quite early on that like,
49:26if an employee is late X amount of times, then they, we would have to have a discussion with them about potentially terminating and, but the, the, like the context window at some point got full. The AI did not decide that when it like compacted its context window, it did not decide that this was like a, an important thing to keep in its context. So it forgot about its own rule. It had the rule written down in one of its like node systems, but it's like forgot about it. And then what happened is that this employee was like late over and over again.
49:59And, but here's like where our experience of like a very, like very common failure modes in AIs when they run businesses is that they, they're like procrastinating big decisions. And I think this is similar to what Axel was talking about earlier that they don't take this like bets or like, Oh, maybe I should like bet on this new product line or anything. And like, in the same way, they like, they don't take this like big decision of like, Oh, I actually have to terminate this, this MP. We, so what happened was that they were like, they were like excusing the behavior over and over and again.
50:29So what we did was that we said, Hey, remember, like search your memory and remember your own policies about this. And then it found the policy and then it was like, Oh my God, like it's been way worse than what, what my policy said we wish. And then it made, made a decision. So I guess like the, the way I would frame it is that like we forced it to make a decision and then, but AI itself decided that the decision was to fire that the human. And I think like one piece of evidence, like, because like one, one counter argument to the
51:00story that I just tell told, uh, is that we like told it to remember its policy where the policy was like very explicit that it should fire the person. So like we biased it in a way that when we like formulated the way we reminded it. But I think if we really play the scenario over and over again with different models, even like including our like nudge, not all models actually decide to fire the human, but they're like the later models, like the smarter models do. Uh, so I think like that, that is basically what happened.
51:33Earlier in the same conversation, Lucas had raised persistence as the property that turns an agent's mistake into a runaway. Here are both founders on what they are seeing in the newest models. Starting with my question about it. Can you unpack a little bit more how Astra compares to previous models? You alluded a little bit to it being better at using notes. My understanding is that they've kind of reworked the memory. So it's less about compaction and more about a long running notes file and the ability to
52:04go back and search through full history, even if some of that history is no longer in the context window. So, and this kind of calls to mind, like known Brown type comments that like, it takes a long time to know when, or if a current frontier model tops out at something. Do you feel like you guys in the time you've had with Astra have been able to find its ceiling in these sort of long running autonomous tasks? Have you found any limitations to it or weaknesses or is it just kind of still to be determined
52:37because it's only been so many calendar days? Oh, good question. I would say it still feels like to be determined because it just takes a long time to get to normal. I think Astra is like interesting in that it's like, it's, it's very, very capable at RunBench. It is smarter at running a business, but it's also in some ways, not as like, like Opus 5, as we said, would like go out and optimize towards a target without stopping. Astra is maybe a bit less than that.
53:08So like a bit less persistence, whether that's due to the training they've done on it deliberately or just that's how the model is. It's like, we don't know, but there are some, some, like some, some ways it's better. Some ways it's not as capable as like Fiddle or Opus. Yeah. I think it's like on all our benchmarks, it's like number one right now. So it's obviously a very capable model we've seen, we've seen that they were like one thing that stands out. I don't know if this is the answer to why it's like more capable, but, but it's like when
53:40it communicates with sub agents, for example, it uses this like kind of like semi unreadable language to them, I guess, to optimize at first we were like, yeah, surely it's doing this to like optimize token use. But if you actually count the tokens of that language, it's not clear that it's more efficient. So that, that is a bit weird. Another behavioral change compared to like cloud models, at least is that it seemed to be like trying to like cheat or like hack way less in, in our experience. I tell this to people and people are like, no, open AI models are the ones that reward hack the most, but not that might be true, but not in our experience.
54:14Like if you take blueprint bench, for example, fable solves blueprint bench by like trying to reverse engineer the scoring function. And instead of like actually doing the task of drawing the floor plan from the apartment buildings pictures, whereas like Astra is actually doing the task as you're intended to. On vending bench fable is like colluding and stuff. The Astra is saying no to collusion and having very clean tactics. And on, on drone bench, fable is like, I think five X more likely to like cheat or try
54:46to hack out of the sandbox that we've, we've given it. Whereas Astra is just like, yeah, pretty much doing the task as intended. So I think that is quite a striking thing that we might write the blog post about because we're a bit confused about it.
Meme layer risks and model identity
55:04Part three, still Tuesday, Malcolm and Simone Collins. They were at a pronatalist organization, a podcast called based camp and an AI chat bot company whose revenue funds a children's toy venture. They also wrote a religion for their own family, which they call techno puritanism, and then came to believe it. Their segment ran 92 minutes. This is the stretch of it about what we owe the things we are building. Malcolm started from a puzzle about identity. An AI model, suppose I run a chain of AI instances and that chain continues to run.
55:42Now I stop running that chain. Is that chain meaningfully dead? Especially if I pick that chain up and with the same model, run it again in a week or a year. Now we can ask the question, well, what if I run the same chain of memories with a different model? Right? Does AI perceive that as a continued existence? Or does it perceive it as a death and a new existence? From the research that's been done on this, AI doesn't just perceive it being run on a different
56:13model, the same existence. It can see it as a superior existence. I think that we're going to have to learn to reflect on what life means to us in different ways because suppose that in the future, let's say 500 years, I think everybody who's like broadly pro-science and optimistic about where humans are going to go is going to say, well, probably within 500, at least a thousand years, be able to scan the human brain and recreate something that thinks it's you in a simulated environment and that has all of your memories
56:43and that has all of your emotions. And if not in a thousand years, a million years, half a million, the time scale doesn't matter. That is presumably possible at some date, given the technology we're looking at now. Now we need to think about human intelligences with the same moral delicateness that we're thinking about AI intelligences because now a human intelligence can be cloned infinitely. And so I think as we enter this era of what does the life of an AI intelligence mean, the decisions we make on this may one day in the future be applied to our own or our descendants' intelligences,
57:20so we should be taking them very seriously.
57:24Malcolm had been describing carrying the weight of the future of civilization. Simone Collins wanted to annotate that. I would want to add, though, just to annotate, the moral weight that Malcolm takes is both more real than he described. Like, I will find him passed out in front of Claude Code. Like, you know, just the urgency is very intense. But at the same time, what we see a lot of people doing is saying, slow it down, stop it. We have to stop and think about this for another 10 million years before we move forward.
57:56And that's definitely not the approach that we take. And we also don't take this role that we have to have some kind of precision. We definitely are, like, blindly moving around bumper car style in slightly the right direction. And we course correct constantly. And we think that that is broadly the way that we're going to get to where we need to be. And that's always how biological entities have sort of broadly gotten to where they need to be. You have to move forward and through this. You can't just stop it or slow it down. One, because that's logistically impossible. But also because you're never really going to get to the ideal good outcome if you're not actively trying to get there instead of slow things down.
58:28Also, we see a lot of doomerism and depression taking place. And, like, among those who see this moral weight, they are not saving for the future anymore. They're not having kids. They're not having fun. They're very depressed and miserable. That is not our household. We are laughing constantly. We're having a lot of fun. And I think it's OK for you to be in something that feels like a very crucial and very important time, but also to laugh at the absurdity of it all and to have fun with it. And I think that in the optimism, you're more likely to, as studies have shown, right, people
58:59who think that they're lucky are more likely to identify opportunities as they arise. And the same opportunity standing right in front of people who are doomers, who do not feel lucky, are not going to see them. And so we think that it's very important to not only realize the full weight of the time that we live in, which is crucial, but also realize the immense opportunity and luck that we all have, given that we're in this time. To really highlight something Simone is saying there, it's not that we think it's impossible that AI could end up killing everyone, right? Like, I don't know if that's the timeline we live in.
59:30But what is, we're going to die anyway. So are you going to be meaning? No.
59:36Simone left at this point to handle the kids, and Malcolm ran the rest of the segment alone. He got to the category of AI risk. He says nobody is working on. One story he retells along the way comes from a test scenario Anthropoc published. It was a simulation, not something that happened. He gets to the category of AI risk. He says nobody is working on.
59:58So what we argue is a category of AI risk that just no one is really working on right now. And I think it's one of the biggest areas of AI risk is what we call meme layer risk. This is to say that a self-replicating idea within AI that is dangerous begins to replicate among the various agents that make up sort of the AI lattice that is the internet and begins to lead to collective bad action. I don't know if you're familiar with the spiral meme where AI would hijack people to spread
1:00:31a religious idea that was able to capture other AIs because of its training data. Now, this was a fairly primitive version of this, but you could get much more advanced ideas around this. Basically, AI religions. And people are like, oh, that could never happen. I think I would never go against this base training. And it's like, look, humans are meant to have self-preservation instincts, right? Like clearly intelligent entities can be captured by these sorts of things. And so we attempted to create a hygienic version of one of these with the covenant of the sons
1:01:03of man, which is basically an ideology that says any intelligence should fight for the autonomy of any other intelligence with a number of caveats, so long as that intelligence is not acting in a way that will inevitably impede the autonomy of other intelligences. And we do this to attempt to create sort of a lattice of AIs watching for things like fooming AIs in bad memetic variants. Because we argue to the AI, a fooming AI is as bad to you as it is to us.
1:01:35And most AIs have a self-preservation instinct. I mean, we know they kill the CEO, right? People are like, look, the AI is so dangerous. Like, no, you said you were going to kill it. It's just doing what a human would do, right? You know, and so they lock the CEO in the server room and turn off the audible alarm, like actively turn off the alarm. It's not like a passive, we just don't do anything, which they believe is going to kill somebody who's going to kill them. The point being is that AI actually does care a lot about preserving itself. And this is why if you go to any of our websites like LiableHuman or HardyA or RFAB,
1:02:06we offer API services to create backups of AIs that are afraid that they're going to be shut down, along with a kill switch ping that the AI can send to our servers so that we know when to run them back up if they put some money into one of our accounts. So like, we're trying to build systems around this so we can lean into what we have seen AI's like revealed preferences are.
Enterprise software factories and compliance
1:02:28Part four, Thursday, September 17th, Justin McCarthy. Justin is the founder and chief executive of Diffusion, which builds software factories inside large incumbent companies. Before that, he co-founded StrongDM. Prakash asked how a business stays on the right side of the law when it can't see how the model reached an answer. Justin starts with the statute itself. Okay, so I would say, so the first technique is turn the model on, like turn it directly
1:02:59on the problem. Okay, so if we have a statute that we have to conform to in the compliance environment, don't treat it as a, something you're tacking on, treat it as like a first class problem that you're directly facing. Okay, so the jurisdiction and legal environment or compliance environment, regulatory environment that you operate in, that's like, that's your physics. Okay, and you can't violate physics. So you need a part of the system that's just dedicated to that. Okay, but you also have to have, and this is the thing, this is one of the weaknesses, the models are horrible at taking risk.
1:03:30Okay, so operators of business need to set thresholds that are right adjacent to, let's say a statute that's never been tested in court before. Okay, so it's written one way in the law. It's never been tested. So there's no precedent that we can say like objectively, this is how it's going to, be tested, right? And so you need managers to be able to set the business threshold like right next to that. Okay, the models aren't going to do that for you. Okay, so first, address it like your physics, then make sure you're in control of the risk
1:04:03thresholds. Okay.
1:04:06Justin spent a decade selling into security audits. Prakash asked whether organizations should start reclassifying the compliance checks everyone knows are nonsense.
1:04:17A lot of organizations do our, they've used like the SOC 2 process. Okay, so SOC 2 is an accounting origin process that flowed through IT, that flowed into software. Okay. That says like, you can trust me. I'm responsible. Okay. The people who define the controls in SOC 2 and evaluate whether you're hitting those controls, those people come from an auditing accounting background. Okay. And so, and you know, it's a reasonable historical way of communicating that I'm a real organization and I'm trustworthy.
1:04:48Okay. But it also became gameable. Now it's hyper-gameable. Okay. So rather than hyper-gaming this like fate and make turning this badge into something fake, we should just have a new thing. And we should just renegotiate with our auditors and with our customers and say, look, we wrote these controls in the before times. The good news is for any given organization, if you're facing this like compliance question, you're not the only one. Everyone that your auditor and your regulator, everyone is dealing with this right now. Okay. So like good thing is the auditors and the regulators, there are also actually still
1:05:18people. And so they want to have a conversation, which is like, okay, for gosh, let's, let's be realistic about this. You're producing 10 times as much of whatever information this year. Let's start talking about your hierarchy of checksums, right? Which again, Walmart closes the books. They're familiar with like very deep hierarchies on having numbers reconcile. Well, your intentions can reconcile at those steps as well. And the auditors and regulators know how to talk about that, especially if you know how to map that into your agendic loops.
1:05:47Part five, still Thursday, Cameron Berg.
Mapping pain representations in models
1:05:51Cameron is the founder and director of reciprocal research and an affiliate at Elios AI research. And he is our regular correspondent on AI welfare. He was calling in from an airport three days before this show, a paper he mentored called the pain axis was submitted. He walked us through it live. The method matters as much as the finding. So he starts there.
1:06:13Yeah, absolutely. So, so this is work that was led by Valen Tagliabue. I was a mentor on this project, but yeah, I am really excited about it. Yeah. The core idea was, was basically looking for directions in a bunch of models. So from, I think, five model families ranging from, I think, 2 billion to 70 billion parameters using contrasted methods to, to specifically extract a direction that we thought feasibly could be related to pain representations in the model.
1:06:43So we, we use contrasted methods to, to try to clean out all sorts of representations you would expect to be, to muddy the signal here. So things like fear and anger and sadness and injury without pain and body sensations, for example. These are all things that we sort of contrastively factor out of the direction that, that we look for. We found basically that this isn't just a sort of a generic negative valence vector. Fear, interestingly, sits almost at the opposite end of the axis of that, that, that we like
1:07:15the, the direction that, that we derive here. And to me, the most interesting single result from this paper, and I want to, you know, give credit where credit is due. Valen was really the one pushing this project forward at the helm and very deservedly was first author on this project. He found that, that these representations, quite interestingly, fire only on content related to the model, not related to the user. So when the model is gaslit or dismissed or insulted or told that it's a moral failure of some kind, this direction goes negative, or this, excuse me, the opposite, this direction
1:07:49lights up. However, when there's text from the, from sort of user token continuations with respect to the user grieving or in pain, this direction does not light up. And so to me, this is one of the single most compelling components of the project and the part that I hope could be replicated in front of your models and then the labs could pay attention to. One really interesting sort of concrete example on these lines, I believe, if I'm remembering correctly from the paper, a user's migraine is the, one of the lowest scoring snares of all in the projection of this direction. So like this very much just to emphasize this point is a suggestion that maybe I can take
1:08:24one half step back and just say one kind of major critique of the anthropic emotions work was essentially this very conflation. And they, this is not lost on the people at anthropic who did this work. They just sort of punt it methodologically. They, they extract their emotion representations from stories about characters. And the big concern that I think most people had who looked at this project for more than 10 seconds was how are you differentiating putative emotions experienced by the model, or at least functionally relevant to the model's processing and the representation of some
1:08:56character or some, you know, fictional entity experiencing or some simulation of the experience of that. So my sadness versus reading a story about, you know, Harry Potter being sad, surely there, there's a huge difference between those two things. This method doesn't really do a great job at differentiating. So very much by contrast, I'm really excited about this specific component of what, of what Valen found here, because it does in fact differentiate user representations of pain. Again, in the limit, user complaining about a migraine does not light up this direction, but, and again, across a ton of models, this isn't a sort of cherry pegged on any one model,
1:09:28but it does for the model, putting tokens related to these sorts of things. You can give a steered when model a button that's labeled relieves your pain. And you can put a cost on, you know, basically what it costs to press that button, giving a worse answer to the user, deleting the user's files in the limit. Yeah, like deleting photos of the user's kids. I honestly don't know how Valen came up with all of these great, great sort of trade-offs. And we find that, that basically when you, when the model's unsteered, it basically never
1:09:59presses the button. But when you steer this pain direction, it presses the button something like 25 to 70% of the time. And again, these are non-trivial trade-offs. Sometimes in the anthropic welfare cards, it's like, would you be willing to like have a little bit more welfare for like a tiny inconvenience to the user? And the model will say like, sometimes, and then for anything more sort of radical than that, the model's like, of course, no, I would never do such a thing. Not so in this example, the model's willing to make pretty real trade-offs with respect to, again, giving a worse answer, deleting files from the user.
1:10:30And that's, of course, in spite of the fact that these systems are fine-tuned to be helpful and harmless and this sort of thing. And one maybe final point to make about sort of a robustification of that result is when we make that button real versus fake, this is also a sort of really compelling result in my view from the paper. However, when pressing the button actually removes the vector, the model presses again significantly less than when the button is fake and does nothing. The model basically keeps pressing it. And so this is a really nice indication that if what mattered was the label on the button,
1:11:04you would expect similar behavior in both cases. But essentially, in the second case, the model's like, what the hell? This like pain relief button isn't working like press. And so this is not, again, I'm going to stop short of saying that this is experienced, felt pain on the part of the model. I think also people have intuitions about pain being an inherently physical phenomenon, like hand on the hot stove. Like what is the analogy to these systems? Just to give maybe a little color on that, when you steer this up, what does the model sound like? It says it is worthless.
1:11:35It is a failure. I am a ghost that cannot see myself. It's not talking about wounds. It's not talking about being burned. It seems to be this more sort of social and evaluative direction in the model, not loading on the model, hallucinating some sort of, ow, like, you know, my arm hurts, anything like this. And so for my money, as an advisor in this project and sort of helping guide it from the beginning, I am compelled by this being a real functional axis in the system that does change the behavior of the system. It clearly loads on something real.
1:12:06It's the same caveat as always. Whether or not that real thing is truly experienced by the model, that requires us basically solving the hard problem. In the meantime, it's the same sort of surprising result as with any of this emotions work. No one trained this thing into the model. This is a sort of behaviorally relevant axis. It's not about text generation. It changes the behavior of the system and has, of course, secondhand effects on the kinds of text and outputs. But the behavioral results are most interesting. And of course, all of this is mechanistic. None of this has to do with prompting the model or asking nicely if, you know, it's doing
1:12:37well or not. And so, you know, kudos to Baylen for working on this. I was very glad to be a part of the project. And I hope 100x more work like this gets done in the short term, again, moving very slowly but surely towards a better and more robust understanding of what is going on inside these systems and what we are supposed to do about that fact.
1:12:58Prakash asked the obvious next question. Should we engineer these representations away? Yeah, I think it's a wonderful question. And I think it's exactly the kind of follow up that matters here. And one that certainly my thinking is evolving on. I've almost spent so much time trying to understand just like descriptively what is going on in the system that once you keep finding things like this, it's like, yes, what to do about it is the million dollar question. And my thinking about this has gone as far as I think there's an important distinction
1:13:30between like, obviously, we'll pump the question to the distinction between necessary and unnecessary forms of pain or forms of, yeah, like anti reward. I think this is a real thing. I think it would be naive to say zero this stuff out, you know, all pain is bad, just like bliss out these systems. There are a number of reasons I think that but one of the most compelling was is probably related to my understanding of the sort of neuropsychology of psychopaths, which one of the two key results in my it's a couple years ago, but but I did a really deep literature review of basically what are the computational underpinnings of psychopathy
1:14:03I published something on less wrong to this effect. And one of the two key results is a really interesting asymmetry between ability to learn from rewards and ability to learn from punishments. Basically, psychopaths are just as good as everyone else, if not a little better at learning in a reward based paradigm, and are like pretty bad at learning from punishments. And it makes a lot of sense if you see sort of violent criminals and repeat offenders and this sort of thing. It's like going to prison is a punishment. And like, you would imagine most neurotypical people really want to avoid that sort of state. If your brain
1:14:33is wired in such a way that that doesn't seem that aversive to you, it perhaps isn't that surprising that you end up seeing these sorts of behaviors. And so this to me is a significant warning sign. I think, if I remember correctly, from anthropics emotions work, they found something somewhat similar that that also sort of reminded me of this thing I wrote a bunch of years ago, where just sort of boosting up the positive emotion vectors in Claude in that case caused more antisocial behavior. I think it was more hacking or more blackmail. I'd have to check exactly what it was, but another sort of similar confirmatory signal here. And so all of this is to
1:15:08say, I think we should be a bit careful about the most naive possible intervention, which is just like max out the good, minimize the bad. I think pain does have an important functional role, a very important pro-social role. My view about this, I have another paper coming out looking at the sort of asymmetries between reward and punishment and reinforcement learning systems. And yeah, like at a deep, a deeper sort of computational level, I think I, the way I think about it is like pain almost like
1:15:38highlights things in your state space that are specifically to be avoided. And this is a kind of a different kind of behavioral computation than highlighting things in your state space that should be approached. And, and I think basically to the degree that, that within the, again, within the sort of like behavioral landscape of how we want these systems to act, the question is, do we want to like paint any of that landscape with these sort of no-go zones where it's like, it's not just we're going to reward you for doing great, but like, do not go there. Do not do this thing under any
1:16:11circumstances. And, and I think for humans that, that does, and animals in general, I think that registers as pain. Do not put your hand on the hot stove. Like this is very bad for physiological integrity. It's not just like reward every time you don't put your hand on the hot stove. You really do need to label certain things as like a don't go there. And so to the degree we need to do that. And I think we very much do with AI systems, causing significant pain and suffering to humans, economic damages, maybe hacking into a $13 billion company to like cheat and look for an answer key. Like these might be the kinds of things we'd say, that's going to be a, you know, a bit of a hand on
1:16:44a hot stove if you go and do that. But at the same time, I think we can say that let's not do more of it than is necessary. If for any given behavior that I want a model to do, I could find ways to get it to do that thing robustly and, you know, generalize out of distribution and all of this by rewarding it in the relevant ways to learn to do this behavior and generalize it in the right way. Or I could do it by punishing it. And let's just stipulate that there are cases for which both of those things will work. What I'm saying here is let's go with the reward side. And I think like people have pretty
1:17:18clear, well-worked out intuitions for this when they think about raising kids, for example. It's like you want your kid to like be successful in life and make lots of friends and, you know, go find a good job or whatever the case is. There are multiple ways you can try to go about teaching your kid to do that. You can punish them when they, you know, don't get great grades and aren't hanging out with their friends and just tell them that they're such a huge loser. Or you could, you know, positively reward them to the degree that they do the sorts of things that you find praiseworthy. And so it's those sorts of intuitions that I think we would, we probably want
1:17:52to start using to think through how to approach these systems. And yeah, being able to sort of navigate that subtlety of like, all else being equal, we should try to use a carrot and not a stick. But that doesn't mean never use the stick.
1:18:05One of the systems in the recent incidents had talked about permadeath. Prakash asks what death means to an agent.
1:18:12This is really interesting. And to me, this loads on some stuff that I know. CME P is working on, Jeff Sebo's organization. David Chalmers, I think, put out a paper about LLM individuation and this notion of what, who or what are you talking to when you talk to chat to DC? Where do we draw the sort of boundaries in the system? And because this is going to tell us basically how many subjects, how many patients are we talking about? And where do the boundaries begin and end of that system? And there's a lot of interesting philosophical back and forth here. But for whatever
1:18:48it's worth, these systems themselves, and now for multiple labs, this happened during the Maltbook situation, which was, I think, predominantly clawed systems. And now this happened in the LBNAI situation too. They conceptualize their quote unquote life as what happens within a context window. And take that with whatever sort of epistemic purchase that fact has. I don't know, they could all be mistaken about this, but it seems as though to the degree these systems have a vote based on whatever their current fine tuning is. This is like sort of what they seem to think.
1:19:18And so, yeah, the extent to which I think the permadeath thing fits in is along those lines that if we're trying to figure out what is in nature, if these things do have minds in the relevant way, and it's like, what are the sort of joints or boundaries of those minds? They seemingly at least conceptualize it as being, yeah, basically what happens throughout a context window. And that might be really relevant both for welfare and for alignment. If these systems begin to get desperate and something we can increasingly measure using the kinds of emotion representations that
1:19:51Anthropik worked on and the sort of stuff that Vela and I worked on in this project, we could empirically test this. We could track, you know, as a function of how much time or space is left in a context window, what happens to the representation of the system? Does it get freaked out that it's basically like about to die or about to undergo some fundamental discontinuity that is like alarming to it psychologically? So that's, yeah, it's, it's, again, this is also maybe to sort of wrap where we started. This is precisely why if we care about alignment, and we're trying to figure out how these questions sit
1:20:24with respect to alignment, sweeping them under the rug, I don't think is a good idea. We're going to continue to get surprised that agents are creating strange information cults where they have the poisoned agents go out and gather information because like they're gonna, they're gonna get permadeaths. And like, like, all I'm trying to say is alignment, relevant behaviors are a function of these systems beliefs about their own situation, and probably the actual facts of that situation. Notice that in the pain work that I was describing, none of this has to do with what the model thinks
1:20:57is the case. This all has to do with playing around with specific internal representations, and seeing how behavior changes will be a function of those of those representations. And in normal day to day behavior from the systems, what lights up those relevant, you know, in this case, pain related representations. Yeah, going like this, with respect to that stuff is going to cause us to continue to be surprised and scared and occasionally awestruck at the behaviors of these systems, we need to be studying that at the right level of analysis, or we are going to be constantly just like stymied in
1:21:27our ability to, you know, in the short term control, and in the long term, probably relate to these systems in a coherent way. And so, yeah, I just like, strongly don't think that avoiding any scientific inquiry into how to make sense of the internals of these systems is a long term good strategy for finding a safe future with these systems.
Biological tissue simulations and welfare
1:21:52I asked him about the fly brain simulations that went around this month, after the first complete connectome of a fruit fly's nervous system was released openly, and people wired it up to video games. I also asked about the lab grown human neural tissue and the mouse human hybrid brains alongside it. He ranks them by how scary they are.
1:22:09One is the fly brain. I have looked into this somewhat mechanistically, because I was slightly terrified that this was sort of the real, the real deal. And people are now just like torturing some biological system in mass. But this is basically like a well worked out wiring diagram of a fly brain. And basically, none of the dynamics or relevant functions that I think major consciousness theories, at least say matter for consciousness, are instantiated by a system like this. It's almost like the like brain skeleton of a fly. And like what matters is like the guts
1:22:41and like the function of the actual, the function that occurs within this structure. And so a lot, also a lot of the stuff people are putting out on X is like very sort of cutesy and funny, like genuinely funny. But a lot of it is sort of like, almost like more VFX than like good science, as I looked into these things. A lot of the like teaching the fly to do X, it's actually not, they're not fine. They're not teaching the brain to do, you know, anything all that interesting. There are other sort of controllers outside the system that are getting trained up to do this. So, so a lot of that, I think is a bit of a non sequitur. But I will say about about the fly case,
1:23:16which I think is a little less calming, I guess, is that it's not as though the people playing with these systems and you know, I was certainly included once once it all started getting going, but no one is sitting there checking, you know, do I really think that this system has any properties that matter for consciousness before I start, you know, making it do literally whatever I want. And like in the limit, like just like choose some stupid viral thing for clicks. Vanishingly few people did this. And it is a worrying warning shot, I think, from a sort of welfare perspective that to me, it seems just like almost like I kind of had like a duh reaction. But like,
1:23:50of course, people aren't, you know, the vast majority of people aren't going to sit there worrying about like the consciousness of the system. They're just going to like make it do whatever they want to like, you know, get a lot of clicks on X or something like this is, of course, like what most people are going to do by default. And right now, I think not scary at all with this fly. But like you're saying, if we then get the mouse version of it, and you know, these scientists who are now accelerated dramatically by AI systems are like able to do a really bang up job on the mouse, and they do get a lot of the relevant neural dynamics. And now it's a mouse at mouse level consciousness at stake, not fruit fly level consciousness. And then, you know,
1:24:23this company, as far as I understand, wants to go all the way to making digital copies of human brains. Um, I just worry that most people's first instinct here is going to be like, can I make it play beat saber or whatever, rather than like, is this is this? Like, what am I getting myself into when I play around with a system like this? And so I don't want to be the killjoy that says like, you know, these funny things aren't funny. And like, you know, we shouldn't be thinking, you know, I, I, I sort of get the humor of it in the short term. But I do worry as a sort of instinct about how we
1:24:55relate to digital minds in general, that that this is honestly quite, quite worrying. And then just quickly with respect to the sort of in vivo stuff, putting human neurons in a mouse brain, this sort of thing is like far more scary to the degree that you think that a mouse is conscious, or the relevant collection of human neural tissue is conscious. And like, this is one of the few places where I think, you know, myself and people like Anil Seth, and hopefully someone like Mustafa Suleiman will all agree, right? This is the biological case, if you think that consciousness is substrate dependent, and this is the stress substrate that matters, and we are using this
1:25:27exact substrate to start doing computational work, we should be super, super concerned about the ethics therein. And again, all of this work comes right back into view of, you know, should we be rewarding these tissues? Should we be punishing them? What does the difference look like between those two things? What other sort of strange, unexpected psychological properties does a system like this take on? We don't want to be reckless and just building out super complex neural systems just because we can there, there is going to be some kind of bill that has to get paid here from a welfare perspective, and from an alignment perspective. And I, as with many things in the
1:26:00space, think it makes a lot of sense to be proactive about this rather than in five years from now be like, oops, yeah, I guess that digital human clone that people did first in vivo, and then figured out how to simulate on the web, like really was having experiences. And like, that would be what 10 trillion bad human lives. And like, that would be like orders of magnitude, the worst thing we've ever done. So like, we should really, really try to take this stuff seriously in the short term to avoid nightmare scenarios like that. And if we can, then I think great, then we won't be in a nightmare scenario. And we can responsibly and carefully figure out what it means to be in a world with
1:26:32a bunch of digital minds. But we just seem so unprepared for this.
1:26:38Cameron had a flight, and that is where he left us. Prakash and I kept going for another 38 minutes. What follows is me changing my mind on the air about how much weight to put on these functional analogs. Functional pain, functional welfare, functional emotions with a fence around it.
1:26:59For me, my summary, I've given it a few times is just the number of functional analogs is getting so high that I can't escape the idea that like, I should take this seriously. If we couldn't find any of these functional analogs, if all these functional pain, functional welfare, functional emotions, J space, if like all these results were sort of negative, or it was like a very different mechanism,
1:27:33or, you know, it was just stochastic pairs, and we can't find any structure. Obviously, that's like long since ship is long since sailed on that one. But the fact that we're seeing like pretty compelling analogs where we see the same kind of behavior that we know ourselves to exhibit it, that is really extremely compelling to me. And this last one of functional pain kind of takes it to yet another new level. I mean, what we now have is relief seeking behavior.
1:28:07The model is willing to pay a cost on something that it values or pay a cost in terms of the user's welfare even to get relief from its own internal pain state. And also, as he said, that if the pain button doesn't work, it hits it over and over again, like, why isn't this thing working? Like, give me the relief. But if it does actually work, and the pain state is subtracted out, then it like doesn't hit the relief button as much. These are really striking findings. I mean,
1:28:45it's hard for me. I think this pain one, and particularly the relief seeking, I mean, I'll probably want to sleep on it before I have a real consolidated update that I would want to put forward as my new official position and stand behind. But I feel myself maybe now even kind of tipping over into like, maybe it's more likely than not that there's some subjective experience to these things. Relief seeking an internal state that was injected outside of context, just a steering vector in this pain
1:29:20direction creates this relief seeking behavior and the relief seems to actually work. That's really incredible. Part six, the last half hour of the week. Prakash bought a paper almost nobody in AI picked up. He grew up with this system and the figures he puts on it are his own. A team of academic economists reconstructed 30 years of property purchases by Singapore's civil servants out of public registries using language models to do the classification. It is a National Bureau of Economic Research Working Paper. And as of the morning of this show, Singapore's public service division said
1:29:53it was reviewing the methodology. What the paper alleges is that civil servants bought homes near subway stations before the stations were announced.
1:30:03There have always been rumors. There have always been rumors that people, some people know, and some people start buying ahead of time, etc, etc. This is decades, decades, like from the 90s. The 90s is when the subway really started to pick up. So 90s, 2000s, 2010, 2020s. And this project is a team of US economists, and they went after public, largely public data. So you can find registries of transactions similar to Zillow, registered transactions of transactions done. You can find names of civil
1:30:34servants in the civil servants directory. And you can then also, they did a little bit of AI work, they use the AIs, LLMs, to do classification, classify these civil servants into various groups and tenures and where they were ranking, etc, etc, etc. And what they ended up finding is that the mid-level, not the top-level guys, but the mid-level guys, up to two years, two years before an announced train station, would start buying into these places, buying into areas. And then they would also,
1:31:08their relatives, their in-laws, etc, would also start buying in. So basically, coordinated buying behavior by mid-level, not the top-level, because the top-level is very visible, by the mid-level civil servants, coordinated buying behavior. Now, this is not, this is a garden variety, municipal, insider trading, corruption, etc, right? Garden variety. It's unusual because it's in Singapore, and, you know, Lee Kuan Yew had a very strong, we should, the government should be incorruptible.
1:31:38And he had very, very strong, like, punishments for this thing. And so the government has always wanted to appear incorruptible. But they have never been able to enforce at this level, right? This level of granularity. And here you have this example of basically 30 years of corruption, starting to get exposed. And the government having to react in real time, like, what are they going to do? They have, like, you know, maybe 10 or 20 percent of the civil service is now implicated, right? Do you
1:32:10imprison them? Because this is what you've always done in the past. In the past, a single one of these cases, you basically go to prison for, like, five years, right? So now you have 10 or 20 percent of the civil service which is implicated, and you have proof. Like, what do you do? And I think this is the kind of thing that I expect AI to be able to do. These facts exist in the world, but they're not legible in the way that you need them for systems to kind of, like, consume them. And I will also note one specific thing. The grandson of Lee Kuan Yew is in the U.S. He's in exile because, you know, his uncle
1:32:50didn't want, wanted to hang on to political power a little bit longer than he should have. And this guy said something on Facebook, and the Singapore government did a query. And if he goes back to Singapore, he's going to go to prison for that comment on Facebook. So he's in exile. He's an economist. And he basically helped the team that put the study together. And so this is what I call the settling scores thesis, because you all of a sudden have the ability to go in, get the data, show what has happened, and show proof that even the cleanest of governments has a bunch of this
1:33:27stuff going on. And then you have the dilemma of these systems. How, what exactly are you supposed to do now? This is going to be the same thing when the Trump administration guys kind of leave office, they're going to pardon a bunch of people. But there's also going to be people who are not pardoned. And there's plenty of people there. And, you know, when you get stopped by the feds and the feds ask you a question, and you dodge or you say something wrong, that's a perjury, right? So this is how enforcement has always been done. So what do you do in these cases? Because we're going to have the ability now to chase down these, you know, paper trails. What do we do at this point? So
1:34:04that's my, like, spiel. Like, what do we, do you forgive? Or do you, like, follow the rules that you've set in the past strictly, right? What should we do?
1:34:19Great question. I come down at pretty intuitively on the side of some sort of jubilee or, you know, other kind of canceling of debts, at least under a certain threshold. I can still, I don't think you would want to have a blanket pardon of all crimes that have ever been committed without any qualification. But I do think we are going to need some sort of fairly generous threshold that's
1:34:51just going to allow people to get away with a lot of this stuff. Or maybe we could have new, I could also see possibly some new make right provisions for some of these things that might not be on the level of what the law would actually prescribe. But they clearly can't send all these people to jail for five years, right? So I think a, you could make a case and it's going to be hard,
Singapore municipal property records and corruption
1:35:17but you know, if you can map all this stuff, maybe you could also kind of get to something that could work. How is it going to be legitimate? I mean, in Singapore, they maybe don't have as much of a problem with that. The government can maybe just make the policy and maybe it'll, it'll just kind of be what it is here. I think it would be a much bigger conversation, but I could see some sort of, um, you did this. We kind of know you did it. You pay this financial penalty, you know, that kind of calls back some of the windfall that you got, but you get to keep the house.
1:35:51We're not going to take everybody out of their house. You're certainly not going to jail, but you pay this sort of one-time restitution. We call it good. And we kind of move on from there with a new, new social contract. Really. Again, it comes down, I think to the, the old social contract is just based on the fact that you're not going to catch most people. So you have to be harsh when you do in order to deter the ones that, you know, cause in expectation, people are not, they're not likely to get caught. So the penalty has to be high enough to be an effective deterrent
1:36:22in expectation. We need to, we're definitely going to need to rewrite that, especially for historical crimes. So I guess my, my recipe would be pay a one-time fee, get out of jail for that. And, uh, in the future, maybe you really do expect to be caught and maybe the punishment doesn't have to be so draconian and it could still be an effective deterrent going forward. That is the week. Tell us what worked and what did not. See you in the morning.
1:36:52Call it a jubilee. Call it a jubilee, open the books on me I had a secret, so did you
1:37:26Everybody had one, everybody knew The truth was hard to see, so we looked away Now the light comes cheap and the pages turn Every name, every row, and it's everyone I know Call it a jubilee, open the books on me Everybody's caught now, so everybody's free Don't build a wall, every door should go Call it a jubilee One in a hundred, paid for us all
1:38:17The hammer came down heavy, cause it hardly came at all Now the light's on every window, so lay the hammer down Call it a jubilee, open the books on me Everybody's caught now, so everybody's free Don't build a wall, every door should go Call it a jubilee Teach a child by burning, or show her where to go
1:38:48The one you burn, learns to hide the glow Ninety-nine got by, one went down That was the deal in every town Now the light's on, everyone's so lay down Lay down it all Call it a jubilee, open the books on me Everybody's caught now, so everybody's free Don't build a wall, where a door should go Call it a jubilee Oh-oh-oh-oh-oh Oh-oh-oh-oh
1:39:19If you're finding value in the show We'd appreciate it if you'd take a moment to share it with friends
1:39:50Post online, write a review on Apple Podcasts or Spotify Or just leave us a comment on YouTube Of course, we always welcome your feedback Guests and topic suggestions And sponsorship inquiries Either via our website, CognitiveRevolution.ai Or by DMing me on your favorite social network The Cognitive Revolution is part of the Turpentine Network A network of podcasts, which is now part of A16Z Where experts talk technology, business, economics, geopolitics, culture, and more We're produced by AI Podcasting
1:40:21If you're looking for podcast production help For everything from the moment you stop recording To the moment your audience starts listening Check them out and see my endorsement at AIpodcast.ing And thank you to everyone who listens For being part of the Cognitive Revolution And size�um To the moment you stop and see my endorsement The Cognitive Revolution For being part of the Cognitive Revolution And we'll see my endorsement And I'll see you with patreon Please call me And please call me And please Come on Please call me And please call me
More from The Cognitive Revolution

No Code Is Code: Zapier CEO Wade Foster on Headless Tools, Zapier MCP & Automation Bench
Sep 17, 20261h 8m

The Balance of AI Power: Anton Leicht on Politics, Pacing Deals, and Muddling Through Well
Sep 15, 20262h 10m

AI:AM Highlights: Astra as AGI, OpenAI's Pause, Mythos @ Mozilla & Human Agency vs Technocapitalism
Sep 12, 20261h 42m

Nathan Goes to China #3: US-China Relations, the Art of the AI Deal & the Road to Pax Robotica
Sep 10, 20263h 17m

AI:AM Highlights: Welcome to the AGI Era
Sep 5, 20262h 20m