Steadcast
The Cognitive Revolution cover art
The Cognitive Revolution

Lindy Teammate: Flo Crivello on Multiplayer Agents, Memory & Why He'd Ban the Chinese Models He Uses

August 10, 20262h 6m · 24,774 words

Show notes

Flo Crivello returns to The Cognitive Revolution to launch Lindy Teammate, an AI employee that lives in Slack, connects to company tools, and accumulates a team’s shared context. He argues that multiplayer AI matters because intelligence without context is less useful than an ordinary coworker, and explains Lindy’s approach to agentic memory, editable file systems, context buckets, and large-scale tool outputs.

Highlighted moments

I actually think that humans like intuitively find themselves biased towards very, very multi-agent system, like far more multi-agents than is optimal because they're anthropomorphizing a little bit too much and they're comparing a little bit too much to human organizations.
47:10
Imagine you've got this, this, this machine that runs a company super fucking smart, but, but, but 50 times a day, it decides to walk to the car wash, you know?
1:02:15
It's hard to compete against tokens that are as heavily subsidized as what Frontier Labs are doing. That's just the reality of the application layer right now.
2:01:19

Transcript

Introducing Lindy Teammate

0:00Hello, and welcome back to The Cognitive Revolution. Today, I'm speaking with Flo Crivello, founder and CEO of Lindy, as he's launching Lindy Teammate, an AI employee that joins your company's Slack. This is, of course, competing directly with Claude Tag, and very notably, it is running on DeepSeek. Even so, Lindy is subsidizing onboarding at least, and the first half of this conversation goes really deep into where all of those tokens are going.

0:31We talk about the nuances of multiplayer mode, how Lindy thinks about the social contract surrounding historical data, and how they're using prompting to enforce it, plus a ton about memory implementation and the background processing that Lindy uses to continually optimize memory. Flo's command of the relevant literature is evident throughout. We get Flo's vendor recommendations for file systems and sandboxes and discuss why he prefers to buy where he can and remains bullish on software infrastructure startups.

1:03We also talk about life in Lindy now that Lindy Teammate exists and how he thinks about Lindy and all Frontier AI systems as a superintelligence in most respects that still somehow chooses to walk to the car wash a bunch of times every day. For now, Flo says we're in the Centaur era. The best ideas at Lindy are generally co-created between humans and AIs. But Flo feels that this will be temporary, and before too long, he expects that humans will simply be adding noise to highly optimized AI systems.

1:35And that, it should be said, is in the good world, because, as Flo puts it, in light of OpenFace and other similar incidents, it is late July 2026. We have AGI. We're in takeoff. And we've not figured out alignment. With that in mind, we compare notes on how friends and acquaintances at AI companies are now kind of panicking. And in the final section, we discuss Flo's very surprising position, that Chinese models should be banned in the United States. As of now, I definitely do not agree.

2:07But we have a good-natured back and forth about various arguments, and ultimately, he seemed open to a compromise on something perhaps more like an insurance requirement for companies selling AI services that would allow us to press the risk of Chinese models instead of banning them outright. In any case, I have to say, Flo is a one-of-one. After sprinting through the general-purpose AI assistant product race for the last three years, he is still super open about how Lindy works, and he's fully candid about every issue he cares about.

2:37It makes for a really fun conversation about multiplayer AI, machine memory, and the uncomfortable politics of building your company on models that you yourself would like to see banned. This is Flo Cravello, founder and CEO of Lindy. Flo Cravello, CEO at Lindy. Welcome back to the Cognitive Revolution. Thanks, Nathan. Always in the end to be here. I am excited for this conversation. You got some big news at Lindy with a new product launch

3:08that is the occasion for this conversation. But obviously, we are in the thick of it when it comes to the AI exponential as well. And so you've been, I know over time, outspoken on a bunch of really critical AI issues, and I definitely want to get into that as we get deeper into the conversation as well. But let's start at the center with what you are working on. Lindy is now evolved again. We started with kind of a smart Zapier workflow software kind of thing a couple years ago.

3:39There have been multiple big evolutions. One that I also definitely want to get into is your recent post on shifting your model mix and getting away from American proprietary models a bit and embracing more of the open source and saving a bunch of money. But the big evolution now is that Lindy is a full-on AI employee, and it's going to be available in Slack, just like all your other teammates, whether they be human or agent. Tell us about the new big launch.

Launching Multiplayer AI

4:07Yeah, you know, from like the get-go, we've always been going after like the AI employee. And I do think that it used to be quite early when we started to go after that, like three years ago. And now I think it's basically here. I think like what we are releasing today is called Lindy Teamate. It is an AI employee that lives in your Slack, connects to all of your tools, accumulates your entire team's context. And it's really like a team scaffold. I think like AI right now is in the middle of making this huge leap towards multiplayer experiences.

4:37I compare it to, you know, maybe you'll remember, like we used to send each other like weird documents around by email, you know, with like revisions and stuff. And, you know, I compare the difference between multiplayer and single-player AI as between like sending each other weird documents and Google Docs, like an actual shared document. If you really want your AI to be a teammate, your agent to be an actual member of the team, like you want it to be where your team collaborates, which is Slack. You want it to have its own shared context about the entire team, its own shared memory. You want everyone to be able to talk and collaborate with the same agent

5:09instead of right now, it's like we're all in the same meeting room and we're all talking. And then every time one of us wants to talk to what's turning out to be maybe the most important constituency of the company, which is AI agents, we have to leave the room and then come back, you know? So that's what we're working on is like this new multiplayer experience and this multiplayer scaffold. So many angles of that, I think, are interesting. How, first of all, how does it onboard, right? This is something that companies have put a lot into over time when it comes to their human employees. I've done this myself with my own kind of deep context

5:41that I stumbled my way through earlier this year. And it is serving me really well, but it's easier for me because it's just my stuff and I own it all and I don't really have to worry about what the expectations were around privacy because, again, I'm going to be the only consumer of it. When you get into multiplayer mode, now I'm like, oh, gosh, you've got different channels where different people were gathered and maybe some of those channels were private, maybe some of them were public, but definitely private. So just practically procedurally, how do you suck up all the information?

6:13How is it stored? And then on a social level, what are the tricky considerations that you are identifying in the multiplayer context?

Context Hydration and Memory Agents

6:21Yeah, that is an excellent question and that touches on exactly what we've been obsessing about, which is I really do think that as we're getting to AGI and as we now arguably have AGI, intelligence actually matters less and less, comparatively speaking, and context matters more and more. You know, I often think of it as like, look, you know, like one of the smartest men in history was John von Neumann, right? If you will to have John von Neumann just magically appear next to you at the office, this guy over the next hour or day

6:52would be less useful to you than your random co-worker, right? So, and that's because of context because you're like, you've got a job to do and like you don't have time to onboard John von Neumann. You know, he doesn't have the context to let it. So you're like, hey, like I love you. I really want to talk to you. But right now I'm going to talk to this guy because I have like an important thing to do. So I do agree that context is super important. And I think it's one of those surprising times when actually agents are better than humans at onboarding. It shouldn't take me by surprise anymore,

7:23but it always does because you always operate under the assumption that, oh, you know, we don't have AGI yet. So obviously humans are going to be better, but they are not. And the reason for that is because very often one major way is that companies onboard their human employees is their wikis, right? There's like in-person onboarding sessions and all of that stuff, but there's also a lot of written documentation. And as everyone knows, the moment written documentation is written, it's out of date. And so what we've built is we've built this shared context layer and what we call like a hydration system.

7:56And so basically the way it works is you sign up to LinkedIn Mate, you connect your tools. So, you know, by all means, connect your wiki, like that'll help, you know, like Confluence, Notion, Google Docs, whatever. And then you connect, most important, which is Slack because that's where the real knowledge lives. It's just a mess, but agents don't mind the mess. And so then what we do is that we crawl your entire Slack and we build a knowledge graph based on a file system. So maybe we'll be able to pause

8:27and edit after that. This is what it looks like. So you're on board and immediately she starts learning. So you can see here, this is within 10 seconds of starting to sign up and she's still learning. Like here, like the graph on the right is still limited. And so she tells me, this is what I've learned about you. And it's surprisingly good. And then this graph on the right keeps growing and growing and growing. And importantly, there's just two layers to this graph. And there is like a personal layer and there is a workspace layer.

8:58So at any moment, you can go into your file system. In Lindy, this is the new Lindy. You click on files and it brings together this memory file here that is supported by a bunch of reference files. And this is all maintained automatically in the background by a team agent. And so hopefully this answers your question about how do you do it? Because some channels are public, some channels are private, some documents are public, some documents are private. And the way it works is that there is this memory agent, we call it a napping, not sleeping. So like it runs every something

9:29like 15 minutes. Like why would you need to sleep every 24 hours? So it just runs continuously in the background. And anything that is public, it updates the team file system and the team memory with. And anything that is personal and that is private, it just updates each person's file system with. So, and in the end, it collates all of that and the agent where you talk to the NDT mate uses all of that context and crawls the entire file system. So there's a lot of technical challenges we had to solve in order to get there because the context that it accumulates

10:00is like many millions of tokens and you can't add all of that at every turn. So that was one thing we had to figure out. Yeah, interesting. I'll share a little bit about what I stumbled into and you tell me what you have learned that might be even better.

10:16First of all, just have to connect all the tools and get the exports. I found that there were interesting edge cases that I was constantly running into in my own personal export heuristic that I tried to write. One time I had a heuristic that the longer emails I sent are probably more substantive and I'll definitely make sure I capture those. But then it turned out that at the very top of those, of that power ranking was often something that I'd copied out of an LLM and was sending to somebody with usually a message at the top, here's what I got from Claude or what have you.

10:47But it was like weird now that according to that heuristic, this Claude text is one of the more weighty pieces of my writing. So I've bumped into a lot of these things over time and had this special case. I guess maybe one question is just like when you step into a new organization blind and they, and we, logging is another, right? In our Slack channels at my company Waymark, we've had various, you know, logs bumped into Slack over and over and over again. How do you, do you, how do you identify these idiosyncratic special case things

11:18that could overwhelm or flood or mislead them get to the, the real good stuff? So that's the next sentence question. The, I think the, the main way that we've solved this is the fact that the memory is maintained by an agent itself. I think this is why I'm ultimately quite bearish on RAG as an approach and I'm very bullish on, on, on this like agentic management approach because you have an actual agent which has its own memory. So you have a sort of meta memory

11:49and, and, and, and understands what it is that it's looking at. And by virtue of accumulating that knowledge little by little about the organization gets smarter and smarter about what actually matters or, it's actually sort of similar to training a model because like, you know, if you were to insert poisoned data into the model data set, like, hey, you know, like, I don't know, Darth Vader was a woman, you know, like it, it would get the data but it would be drowned by all the correct data, you know? So here it's the same.

12:20It's like, if you have enough memories and if you feed that to a system that understands it, if the system ends up understanding and so, the very concrete example you are using is actually an emergent behavior we have seen the memory agent adopt because the memory agent has its own memory so it's a sort of meta memory about how do I manage my memory? What sources of information matter? Like, you know, all of that and which ones are trustworthy and all of that. And we have noticed that the first time the memory agent crawls, the slack it finds,

12:50some of those channels, like many organizations have them, like those like log channels and it learns to ignore them. It's like, I'm not going to keep spending time on those channels. There's nothing for me to learn there. The thing that's been critical for us has been building meetings as a first-class citizen in this system. Like, I really do believe that meetings are very underrated as a source of information. They all wear, like 90% of the most up-to-date data leaves about the company. Like, everything that matters

13:20inside the company has a meeting around it. Every relationship, every project, every initiative, everything has a meeting around it. And so I think like those multiplayer agentic systems cannot really get meetings to like Granola or whatever. Like, I think you have to really incorporate it in your system as a first-class citizen. And so what we did is that, you know, we built that first-class citizen. We have meetings. Now it's a first-class citizen in Lindy. And most importantly, it's like, again, it's not just a Granola, you know, oh, it's recording your meetings

13:50and so questions about them. You can do all of that stuff. But it goes beyond that and it feeds these meetings to your memory agent. And now, if a meeting was public, so you can set up meeting folders which automatically add meetings, you can set up meeting folders which automatically add meetings to themselves and share them with your entire team. And so then, it like summarizes all of the meetings that are in these folders. And again, like, if a meeting was added to a public meeting folder

14:21that paired with the entire company, then it updates the team's memory with it. And so it keeps updating that context on an ongoing basis. And now you can like chat with, it's basically like a team mate that's been in every meeting in the company. You can chat with like the entire corpus of meeting. You can be like, what are customers saying? Like, what's the biggest request lately? Like, you know, what's been the feedback about this and that needs you're in. So this brings up another interesting challenge that I've, again, kind of stumbled my way to

14:51at least the solution for now on. And I'm very interested to get your take on both kind of the backward-looking aspect of it and the forward-looking perspective. the issue is there's a lot of stuff in my broad general context that probably shouldn't be shared with other people, right? Sometimes it's sensitive, confessional, one person's point of view on another person, like whatever, right? So I've actually

15:22created for my own wiki two versions of it. One is the, I think, I've kind of clawed on my main laptop where I do my work as an extension of myself and a second brain type of thing. So this thing is not taking on projects autonomously and running with them. It's just doing what I tell it to do. And it has the full wiki with all the gossip or whatever that's in there. I don't have honestly that much super sensitive stuff. I don't want to make it sound like it's more dramatic than it is. But nevertheless, people didn't expect that when they were telling me something which could have been

15:52on a call or an email or a private Slack message that it was going to go into some central repository of an agent that was going to now talk to the world. So that one's just for me. I had it go through and create a version using the heuristic of what would be appropriate for a person to tell a human assistant. The sort of still private but a little bit more public facing wiki where it would be appropriate for my human assistant to have your email and phone number, right? But it might not be appropriate for every detail of every

16:23conversation we've ever had to be in there. How are you thinking about as you absorb all this historical information being sensitive to what should and shouldn't be in memory such that it ought to be shared? The second part is how our organization is going to change now because I think this is all, we're rewriting the social contract potentially in real time. So I think application or one strategy for everything that came before but the answer might be like new social norms going forward. I want to get your take on that too. This has been a

16:54vociferous debate inside the team. There has been two camps. There have been the camps that are like frankly I'm in that camp. There are like just two tiers of memory is enough, right? So there's the public team tier and then there's the private tier and the private tier contains everything and the public tier contains like stuff that's only private public. And some other members of the team have been saying what you've been saying. Like they've been saying like no actually even my private tier I don't want to contain a lot of stuff. I want like a super private tier. And like and then they'll try there was like this whole thing like what if we defined multiple memory bubbles and

17:24the user can edit them. And I'm like that sounds kind of overkill. And so what we landed on and what these teammates landed on honestly is like they have edited their own meta memory prompt and the meta memory prompt is just a text file. It's your memory that MD file in your file system in Indy. Your memory agent has that memory prompt injected in its context window at every moment. And so if you insert a line up there that's like this is what I never want you to remember just at the top or something anywhere in the file really but you

17:55can do that. This is what I don't want you to remember. Oh by the way that other stuff that's like a sensitive topic please remember it in that other file in that other folder that's get out of you and I don't want you to pull this file or this folder unless X Y or Z. Right. So you can just like leave these these instructions. And so that's another reason why I'm very shrag and I'm bullish text and file system. is you can inspect the memory and you can edit it and you can just like very granularly insert this kind of guardrails like.

18:27Yeah interesting. So you just to make sure I understand can repeat it back one source of ground truth and you create different lenses on that information just by prompting so you can have your I could tell my agent like hey just so you know this contains everything I've ever every conversation I've ever had always use the heuristic of you should only really be using information that would have been appropriate for me to share with a human assistant and if it doesn't seem appropriate then don't use it and

18:57obviously instruction following is getting extremely good so I can pretty well. I actually I was misspeaking like there is a time when I have used this and it's literally for this podcast I was preparing to go on this podcast and I knew I was going to talk about the memory agent so you see here I have like a memory MD which is my actual memory file and then I have a memory 2 MD which is like a sanitized memory file where I have removed overly sensitive information and you can see in my memory MD here there is like not a bini the file system also contains a memory 2 MD file ignore its contents

19:28they are just here for the users demoing Lindian podcasts. Oh yeah you can just you can just do your own thing here. Yeah okay cool I do add because my version does have the it's not dry problem right now I've got two things to maintain and it does create some overhead so I can see why that could be advantageous. Hey we'll continue our interview in a moment after a word from our sponsors.

Scaling Context with Trees and Buckets

19:58Today's episode is brought to you by Anthropic makers of Claude and Claude Code. Over the last few months Claude has helped me build and refine a personal deep context database that now contains all of my emails, slack messages, tweets, DMs across platforms, video calls and podcast transcripts going back a full five years. On top of that we've now layered summary articles describing my relationship with hundreds of contacts organizations and ideas. And now that this

20:28exists there's almost nothing that Claude can't help with. For my angel investing Claude can now draft investment memos in exactly the form that my venture fund requires based on the calls I've had and the emails I've exchanged with the founders. And when someone needs a favor Claude can often do it as well as I can. Recently a friend reached out to ask if I know anyone who might be a fit for a role that he is currently hiring for. Initially nobody came to mind but then I thought to ask Claude and sure enough it identified two

20:58great leads. Claude is the AI for minds that don't stop at good enough. It's the collaborator that actually understands your entire workflow and thinks with you. So for problems worth solving get started with Claude at Claude.ai slash TCR. That's Claude.ai slash TCR. And check out Claude Pro which includes all of the features mentioned in today's episode. That's Claude.ai slash TCR.

21:28What other things are kind of coming up as you're doing multiplayer? I remember going back to GPT-4 way back when I was red teaming. I did some, it's taken longer than I thought and the reason I even bring this up is because back then I was doing some simulations of a facilitator. You're an AI facilitator in a group, family group that's focused on exercise and your job is to encourage people and give some reminders and whatever and I was playing all the roles of the participants and just having AI play that role and even at GPT-4 it was like doing pretty

21:59well in a relatively simple context and yet it's been like three years now until we're finally getting these real kind of AI employee type experiences. What has been hard about getting multiplayer to work in an intuitive way that might be unobvious to somebody who hasn't been through the slog himself? Yeah. I actually think it's been the scaffold, the context management, like this piece, like the context build up, it's in

22:31a way inspired by Carpath's auto-wiki idea. And so we've had to do a lot of work around context management because once you talk to an AI employee, it's actually quite unlike just talking to like a chat GPT or a cloud because you actually expect it to keep a very rich representation of its past context and of its past memories, you really expect a level of consistency and coherence out of an AI employee that you don't expect out of your cloud. Like cloud sort of like, it's cute, sometimes it plugs like previous memories

23:01about you in your chat, but like you don't really, you don't really do like heavy work on an ongoing basis with it that like in the same way that you do with an employee. So managing all of that context, so both the memory agent that builds up the context, that's millions and millions and millions of tokens per user, and then the core agent that actually uses that context and how does it retrieve the right information at runtime, like this has been like a major, major, major challenge. Frankly, I've sometimes been telling the team like, and I can't imagine we're the only company thinking that, like I sometimes feel like we should

23:32be publishing because I think we're doing stuff that's like seriously state of the art and it's quite frequently that we do stuff and like three to six months later, we see a paper come out and blow up about that thing. There was one time when literally the paper was named what we had called the thing internally because it was obvious. And so I think like context and memory management has been a really, really big part of the challenge. Reliability is always another part of the challenge, right? Like you want to make your model work. And so we've worked quite a bit

24:03on reliability. Like we've called it like a validator is basically a sort of, it's an LLM as a judge that, that triggers like multiple times during the tasks, but it's modular. So you have multiple LLMs as a judge fan out and it's a sort of counsel and then they talk to each other and they decide what to do. So that's been one. The self-improvement loop, I think like since GPT-4 models have become so capable that now you can actually have self- improvement loops. And so, you know, well, not the first ones to talk about it well, no exception. You know, Lindy is now

24:35self-improving. So we can literally see a curve of error rate go down into the right. You know, that's the direction you want to see error rate go down. It went down by like, literally within the first week of us putting the self-improvement loop online, which was like two months ago or something, it went down by 8x, the error rate. I think, I think this probably captures it. I think those have been like the really big meaty chunks we've had to figure out. One big question that brings to mind is how you are managing caching.

25:06And then caching, of course, is just one, you know, angle on managing cost in general. So maybe we could expand beyond caching and talk about cost management. Obviously, there's a significant upfront amount of tokens that you're going to dedicate to in all this context from a new organization and processing it. And so I'm interested in the business model implications of that. Do you have to charge a setup fee or do you need an annual contract? How are you balancing your initial investment with the

25:36level of commitment from the company? And then with so many tokens getting processed all the time and with memory being assembled in different ways and different situations all the time, different users with their combination of the public and the private, how are you thinking about managing input token volume? How much is caching playing into the strategy? How are you making this something that, you know, still on net ends up costing less than a human employee? Headline, we haven't.

26:08You know, I think headline like these things still cost more than the human employee, but not for very long. And like, look, I'll be real, like we're subsidizing it. You know, this is what we've raised VC money for. You know, actually, fun fact, we used to subsidize it very heavily. And then we sort of famously switched to a Chinese model and that made us stop subsidizing it. And now with Lindy Teammate, we have realized our users again are throwing like much more complex stuff that they were asking to our previous product, which was more like a personal assistant. So it was like simple-ish task.

26:39It was like, hey, like send this email to Flow, schedule this meeting. Like I always say you don't need God to schedule your meeting. So like here, like even a deep seek flash was like enough, frankly. And so that was cost effective. With Lindy Teammate, we're back at it and back into, frankly, negative cross-partum territory. And so that said, you know, obviously, I don't like having negative cross margins. I'm at peace with it because, you know, we are very, very confident that it's very temporary. And I actually think you sort of want to build for like the next generation of models always.

27:09And so, you know, we're trying to meet the impact of that. And so, yes, context management is a huge piece of the puzzle. Caching is a huge piece of the puzzle. Our current cash rate is at 85%, which is lower than it should be, frankly, and we spend a lot of time just iterating on it. We've put a lot of systems in place to alert us when the cash rate dips because it's finicky. You make any change anywhere in your system and you break your cash. And now your cash rate dips from 85% to 65%. And the gap between both, it sounds

27:41small, but actually it's almost 2x the price to go from 85% to 65%. So, you know, set up all of those systems. I think one of the biggest breakthroughs we had, and this is one of those things that I'm frankly expecting someone will publish, I guess RLMs was like adjacent to that, but we call them context buckets. And the way it started was if an agent calls an action that returns too much context, like some actions, some MCPs

28:13in particular, retrieve like 100,000 tokens or something, okay? You don't want to send that to the agent, it's going to get confused, okay? So what you do is that you expose that as a summary of the context bucket. And it's a sub-agent which contains the entire bucket. It's like, hey, this action returned too much, so I'm here, you know, to stand in for what the action returned. Roughly speaking, this is what it contains, okay? And then the agent can enter into a conversation with the sub-agent, which itself has its own caching, and can manipulate the context using Unix

28:44utilities. So right here, you're saving a lot. You're saving a lot of money, and it goes quite fast. Then we went one step further, and we were like, hey, what if we could have recursive context buckets? What if context buckets could contain other context buckets? And what if compaction, because obviously we have compaction, was powered by such recursive context buckets?

29:06And by that, I mean, like, now when we compact, so the conversation goes, it passes the threshold. I think right there, it's like 200,000 tokens, but we keep tweaking it. And at some point, we're like, okay, we're going to compact. We compact. And then the compaction sends all of that stuff into a context bucket, so now the agent can query that context bucket. It's not like every compaction is always lossy by nature. And I think compaction operates under the faulty assumption that you never need access to ground truth, which is false.

29:36You at some point do need access to ground truth. So with that technique, you have access to ground truth. And then we went step further, which is, so you have that context bucket, which is the compaction of the previous conversation. The conversation keeps going, keeps going, keeps going. You need to summarize again. What you do is that you take all of that, including this context bucket right here, and you compact it into a new context bucket. So now you have a context bucket containing a context bucket. And in the end, so you do that, and so you have this emerging property where it's like you can have the agent access

30:08any point of any infinite number of tokens at arbitrary levels of granularity. Okay. The problem is that if you have a computer science background, like this gives you like what's called like O-N complexity, right? So it's like if it wants to access like nine context buckets to go, it's got to go through nine layers of sub-agents, and that's really slow and really expensive. And so what we ended up doing is maybe it's getting too technical, but we never. Okay. Well, have you heard of AVL trees? No, but I'm all ears.

30:38Have you, have you heard of red black trees?

30:42There is this. I went to the, I went to the computer science school of hard knocks, so it wasn't a formal training for me. I was actually glad to, to have AVL and black trees because I was like, oh my God, you guys remember your training? You know, this is, this is the moment we use those specking things. Because everybody in school is always like, when do you really use those specking things? Like, we're using AVL and black trees, baby. Red black trees, they're called. So what is described, if you just do the naive implementation, you get that really nice emerging property for free, which is you have all of those linked

31:13context buckets that contain one another. But every context bucket always contains just one other context bucket. And so you end up with like this Russian doll of sorts of context buckets. And like, if you, if you want to open, if you want to get to the bottom, it takes a very long time if you want to, you know, so what you do instead is that you have context buckets contain multiple context buckets. Okay. And, and, and, and, and so you end up with a tree because basically what you want is you want the top most context bucket not to contain the last 100.

31:43You want it to contain the first context bucket and the last context bucket. Right. And so you end up with, it's a self balancing tree. There's these two algorithms, there's the AVL tree, and then there's the red black tree, which are algorithms that are used to balance a tree. So instead of having one long line, you want to minimize the height of the tree such that like going to the bottom takes as few jumps as possible. Okay. Technically AVL is the best, is, is the, is, is the demonstrably best way to balance a tree because it leads to like the lowest

32:14height. A red black tree is better because it takes into account the cost to balance the tree, which is, it's expensive because you need to regenerate a lot of your context buckets and you, you, you, you miss your cash when you do that. It's, it's very expensive. So red black trees, the canonical implementation of red black trees is on binary trees, which is just a tree where each node's got two children. We went for a, it's called a centauri tree. It's literally just each node's got a hundred nodes. So there's a node below that's going to have like 10,000 nodes, right? And so literally with two jumps, you can have 10,000 context buckets and you can

32:46access like all the contexts in the universe. And that leads to some really surprising behaviors because you're literally no more than two LLM calls away from being able to access 10,000 context buckets, each one of which contains 200,000 tokens, right? So you're at like 2 billion tokens here of context in, in two LLM calls. And, and that is what leads to like those really surprising behaviors from those AI agents where you ask them any question and they remember everything perfectly all the time. And that's awesome.

33:18That was, that was one really big thing we had, we had, we had to figure out. I hope, please, you know, like copy us. Like I will just, we just don't have the time to publish, but you know, this is, this is, we don't mean for this to be secret. Like I think this is a really powerful technique and I have been surprised to not see more communication about it. Just for calibration, how many tokens are you finding businesses have? Like when you say with these two levels, you can get to 2 billion tokens. I don't have a great intuition for that cover a 20 person team.

33:49Does however that's been in business for five years, what's kind of the heuristic for team size and length of history that translates into how many tokens? You can probably do the math, but like 2 billion tokens is, we should ask cloud, but it's probably bigger than most libraries. So I think we have few customers who, whose knowledge base, whose memory is bigger than that. Most of the time when you get started, you consume like for like a team of 20, you know, the, the hydration components.

34:20So it's like, again, it's this moment where like you connect your notion, you connect your Slack, and then we crawl everything. We look at all of that stuff. For a team of 20, that's going to tend to consume 5 million tokens at most, you know, 3 to 5 million tokens. And that's with like a team of 20 and like years of Slack histories. Like we literally take all of your Slack history. It takes a while. And actually here, surprisingly, the bottleneck is not even the LLMs. It's the Slack API. Slack is not happy about you like crawling. Yeah, I've experienced that. Yes, I'm sure. The right way to do it, if you're doing it personally, is to export your workspace archive.

34:50We can't ask users to do that. So we just crawl, crawl the entire history. So no, I mean like 2 billion tokens is a lot. Yeah. One data point I have is a podcast and I'm quite self-indulgent and we sometimes go on for a long time, but still it's usually only in the sort of 30 to 50,000 tokens range. So then even having done hundreds over the course of a few years, we're still only in like single digit, maybe low double digit millions for like the entire full text history

35:21of the hundreds of podcast episodes. We're still, we still got orders of magnitude to go before we would get into the billions. That said, now think of a company that's got, you know, three to seven hours of meetings per day per employee times 2,000 employees. That is that meetings are when you really start to rack up tokens really, really quickly. How do you, so I did experience the slack thing. Actually, the funny anecdote is the, the only time I've really experienced data loss, it wasn't too painful, but Claude had, you

35:56know, identified some weakness or whatever in the, and I, I'll bring this back around in a second too, but my kind of raw data is living in a SQLite database and Claude recognized some deficiency or whatever that it wanted to correct. And it just was like, all right, I'll just drop that database and we'll redo it. But what it wasn't taking into account was the fact that we'd been rate limited on Slack for days to get all the stuff out of Slack. And so it was like, no, you dropped that. And I, that was the only

36:27place where we had Slack. It wasn't really lost, but it was like, okay, it's we're now going to be like added again for five more days just to make API calls to get out of Slack. In, in my case, I am the admin. Of at least like a Slack accounts that matter to me the most, but like workspaces that I'm in where I'm not the admin. And so I can't grant this kind of access from a business standpoint. Does that mean that you have to have admin buy-in from the beginning? Is there, it seems like it makes it hard to do product-led growth land and expand type stuff.

36:59That's, that's a good point. You know, look, I think that's one reason why PLG eventually you tap out of it as you grow up market. Like, like, yes, you, you have to be an admin. You have to have permissions to install applications on your Slack workspace. So it tends to be fine for teams of up to like 50 to a hundred. Like after a hundred is pretty rare to, to, to, to, to, for everyone to have this kind of access, but many teams up to a hundred, like pretty much anyone can just Slack, like install like an app. And yeah, I mean, if you, if you go up market and you're like a part of a bigger organization, then we have some users who are like, just like

37:31championing us internally. And they're basically begging the IT to let us in. And then they're like making, making these meetings happen. And, and, you know, IT is, is, is rightfully cautious, but look, you know, we're like software compliant or HIPAA compliant or GDPR compliant. We, we, you know, we've got our ducks in a row and what we're quite used to navigating the internal review processes, which can be burdensome. And, and, and so, you know, it's not in the business of being a pain in the butt of anyone. And, and they are actually themselves very eager. Very often IT was given an, an, an addict like, Hey guys, we need to,

38:02we need to actually become an AI native company. So they're actually seeking this kind of solution, but yeah, I mean, you need, you need admin access for sure. Okay. So let's go back to the data structure for a second. Let me give you my data structure and you can tell me what you think I might be leaving on the table with my homespun version compared to your professional version. And then it's one of my favorite habits these days. I'll take the transcript and give it to, to Claude and we'll see if we can close the gap a little bit. So my version is simply take all of the content that, you know, all the sort of digital

38:39exhaust that I've created over the years, which is email and DMS and Slack and Google docs and whatever, right? Put with Slack being the most painful one, make all the API calls to export all that stuff into a SQL database. I have everything there mapped onto a threads and messages structure. And sometimes that took a little bit of squinting, but it basically seems to mostly work. And that's like the raw ground truth. And I do agree for me, it's, I've found that it is quite

39:11important for the model to be able to get to raw ground truth. Then I've just done, and this only goes back five years, but that's usually plenty. It's not too often that something happened more than five years ago is like operationally relevant today. I just do a month by month export. I found that for me personally, it was like a couple hundred thousand tokens per month that I'll then compress into a monthly summary, which is like 10, 10% at the length, right? So whatever, two to 300,000 tokens of raw monthly content goes to 20 to 30. Then I'll do that at the yearly level

39:46as well. So again, you've got another two to 300 that goes down to a 20 to 30. And then on top of all that stuff, then I just have the model go make a wiki. And one of the key things that I found along the way in terms of how to get to ground truth, of course, the model can always just send a SQL query, but what should it be querying is not always obvious to it. So what I tried to do in my summaries is have the summarizer capture distinctive short sequences of words that would be like needle in a

40:19haystack. So it would be like, here's what happened, blah, blah, blah. And then a quick little footnote at the end of what the source was, a DM from whoever. And these four words will take you exactly to that. And probably only that in like all of my personal history. I'd say it works pretty well for me. I'm usually pretty happy with it. And I do do the right ground truth more often than not. It is only one person. Multiplayer is going to be as I start to onboard my wife and then that'll

40:49complicate things at least a bit. What do you think I might be leaving on the table, especially as you think about multiplayer that demands an even more sophisticated approach?

Agentic Memory Management

40:58Yeah. So I believe the approach you're talking about, and we still have some RAG in our database and we use that for our RAG as well. It's called like hypothesis-driven retrieval. I don't know if you've heard that. It's a sort of reverse hypothesis-driven retrieval. So there's both. So what hypothesis-driven retrieval does is it generates like, suppose it's asking like, where was Justin Bieber born? And then you basically, instead of searching for that question using RAG, you search for your hypothesis. Like Justin Bieber was born in Paris. Justin Bieber was born in New York and Berlin. And you search for this bucket of hypotheses in your RAG database,

41:33because obviously semantically, each of these hypotheses is a lot closer to the answer you might be looking for than to the question, where was Justin Bieber born? So that's the first thing. Then what you do is you also do the opposite, which is where you find an answer in your RAG base, you generate a bunch of questions that map to that answer and you attach them to the answer. So now you can, the retriever can both look for hypotheses and can still look for its answer for the question. And that's going to tend to increase retrieval quality. That's totally fine. I think one thing you may be leaving on the table is it is really healthy for all of the memory to be

42:09managed by one agent. Even though that agent, and both retrieval and updating of the memory. And so even though the memory is in a file system, and so your agent could actually just go mess with the file system, what we have found actually, and can still do that, but we tell the agent, we prompt the agent to be like, if you're looking for a memory that you don't have, please ask the memory agent. And the reason why is because when we ask the memory agents, then what we do is that we log the query, we log the answer, and we log how many hops it took

42:42to retrieve that. And then when the memory agent is napping and dreaming, it looks at this log of like, hey, this is the kind of stuff people have been, it's almost like a librarian. You know, it's like, ah, like I've been hit up left and right. This is a question that I get quite often. So right here, by the way, it acts as a sort of cache, you know, because automatically, so you know, you always keep in like the memory of the memory agent, like the last like thousand like query and like question answer pairs, you know, so it's like a rolling window. And so it is a sort of cache, but also the memory agent is not going to just rely on that cache

43:15because every cache at some point has got a TTL, it grows a stale and all of that stuff. So what the memory agent will also do is like, it will, it will restructure its own memory to reduce the number of hops that is needed to retrieve like the most frequently asked questions. And that, that, that tends to work quite well to, to, to improve the memory retrieval and the memory speed. Cool. Interesting. Another system we, we experienced, we experimented with and we ended up not implementing it, even though it does have strengths from a technical standpoint is, have you heard of caveman?

43:47It's a, it's, it's, it's really simple. It's basically you rewrite your text as a caveman. It's like me, you know, you know, you know, Nathan, not like burrito, you know, like, it's so retarded, but if you do that, you get your tokens by like 20 or 30%, which is good, you know? And so cutting your tokens, like it'll be, it'll make your system faster and make your system cheaper and make retrieval more, more accurate. It'll just be better with basically no loss of, of information or context. You know, if you, the reason we've not done that is because

44:18we actually do care about the maintainability, like the auditability of the system. And so we did that. Every metric went up. It was awesome. And then the files in the file system look really done. It doesn't make us look good, you know, to enterprise customers. Like, why do my memory files look like that? But it, it works. Actually similar. You know, there's like a bunch of those techniques to just compress your context. Have you heard of a Toon? T-O-O-N? It's, it's, it's a sort of like JSON alternative that's made for agents. That's like, I think it's like 20 or 30% more token efficient than JSON. JSON is actually not that token efficient.

44:52So Toon, Toon works better. It's 20 or 30. So this work, this is less important for like memory agents, but it's just, I recommend everyone like makes their agent use Toon instead of JSON and have like a Toon parsing middleware between your agent and its actions. Like no action should expose JSON to, to, to the agent. Like every, every agent, every action should expose. Yeah. Interesting. I'll have to check out this Toon format. I've just used YAML, which also is motivated by similar advantages, but yeah. Toon token oriented object notation.

45:25And token, I, yeah, I see where we're going. It's YAML-ish and it actually, it increases the performance of the agents. Like agents apparently, it was just surprising because there's so much JSON in the training set of the models, but they're apparently more comfortable with Toon than with JSON. JSON does have a lot of noise associated with it for sure. Yeah. I'm more comfortable with YAML than I am with JSON. That's, so that's something. The AIs, they're just like us. That's a lesson. One of my biggest surprises as a brief aside, I'm actually working on, I don't know if this will be a podcast or a thread or whatever, but just taking stock and, you know, of course agents can help

45:59make this a quick project where it would be a very tedious project. Otherwise just going back in history and identifying things that I've been wrong about prediction wise and some things I've been right about prediction wise. I mean, one of the things that I used to say quite confidently that has aged maybe the poorest is that we shouldn't be anthropomorphizing models or that will lead us astray. And I've just been shocked over and over again by how productive it can actually be to anthropomorphize the model. It still feels a little dangerous kind of, but it's hard to argue with the

46:30results. I 100% agree. I do think so that sounds like a happy medium of anthropomorphization. And I think like that's, that's another thing that's very top of mind for me right now is to what extent does anthropomorphization stop being true in particular when they put into AI employees? So for example, way I've changed my mind and I haven't changed my mind completely, but I've changed my mind a little bit, relatively speaking is about multi-agent systems. I actually have come to, to believe that most of the time, as much as possible, you want to consider that as much

47:02as possible. And there are single agents, not multi-agents. That's not only possible. And there's still good reasons to use multi-agents. Like this is not like an absolute statement, but most of the time that's the case. And I actually think that humans like intuitively find themselves biased towards very, very multi-agent system, like far more multi-agents than is optimal because they're anthropomorphizing a little bit too much and they're comparing a little bit too much to human organizations. It's like, ah, I have a data scientist, I have an engineer, I have a designer, I have a PM. And so I'm going to create one agent for each

47:32of these things. And actually the reason why human organizations do that is because every human has only 24 hours a day and can only contain that much context. Whereas agents obviously have none of those constraints. They can fork, they can duplicate, they can, they can do as much as they want all day. And so I find that division of labor is not a good reason to have multiple agents, which is, which is a major, major fact. Now you, again, there are still multiple, there are still different reasons, but the division of labor is not one of them. So I agree. I think by and large, I think people are right to anthropomorphize models because

48:05they were after all trained on human tokens and with, you know, human RL and all of that

Single Agents Versus Multi-Agent Systems

48:10stuff. But I also think that AI organizations shouldn't be overly anthropomorphized. Yeah, that's interesting. Is it, as you talk about a single agent, obviously they can run multiple copies of, we can run multiple copies of, of single agents in parallel. Have you reached the point yet where that starts to create database style problems where if multiple agents are updating the same segment of history at the same time, we have all these sort of, you know, transaction

48:43level guarantees in traditional databases for that reason? Does that start to be an issue as you have 10 Lindy teammates running over the same kind of file system corpus? Yes. Less than you would expect. You know, like I think it's the way human organizations have solved that is via Git. Git is the, by far the best way, like large groups of humans have found to like work together on the same, on the same thing, you know, because you can, you can just merge conflicts,

49:14you can rebase, you can do a bunch of that stuff. So that's, that's what we found. And the file system in Lindy is actually for this reason backed by, by a Git repository, which gives you a ton of amazing properties. So including like merge management, because the, the memory agent sometimes will fork itself. It just basically spins up a bunch of sub agents to be like, right, let's see what happens since the last time I napped. And in some organizations, some large organizations, there's so much context that was created since the last step that it needs to create a bunch of sub agents. And during the initial hydration, like we really wanted to go fast because we wanted to

49:45impress the user. Like time to wow is so important in PLG. And so that graph I showed you, like literally we, we, we've optimized it so hard. So it happens within 10 seconds of sign up. So it's like sign up, notion, slack, boo. Right. And so we, it's been a lot of work. And so one way you do that is via, is via like a collection of agents. And so yes, these collections of agents tend to step on each other's toes. The answer to that is Git, which by the way also gives you history for free, which is, which is awesome. By the way, like, and this is not a sponsored message, like they've just been doing a really good job. We've been working with a MESA, M-E-S-A. They sell

50:19like a, an agent native file system that that's backed by Git. So you create all of those repositories. Initially we started, by the way, it's, it's one of those intuitions you have to build as you build these systems. Like it's very easy. It's very tempting to be like, I'll just hand roll it. I'll just create my own Git. I'll just like manage my own file system infrastructure. And you, you, you learned that there is so much more depth behind those systems than you feels to appreciate. And so I, now I am so, it's actually funny because it's sort of goes counter to the prevailing narrative these days, which is like, SaaS is dying. You can

50:51vibe code everything. And we'll vibe coding a lot of stuff, but the infrastructure will not vibe coding. And I've become so infra build. Now I'm like, anytime I have an opportunity to buy instead of build, I will do it because like you save so much time and there is so much more depth and thinking and design decisions that goes into these products than you may appreciate. I love a good vendor shout out. Any other vendor shout outs, any other key primitives that you're buying that you would recommend? Well, you need a, you need a sandbox. So we've, we've gotten with E2B for that. Again, like many

51:22people when wants to have a sandbox, they're like, sweet, I have a file system. We, at least at scale, we strongly recommend decoupling them for, for the reasons I just mentioned and, and more. Browser base, obviously is, is, is excellent for, for browser management. Who else do we like and use? You know, I've, we, we went down a deep rabbit hole. Like one major exception to what I said about being infra build has been observation and evaluations. We have a complicated history because we, we got started, we started this, this company in 2022. And so

51:53it was, it was in hindsight way too early. Like agents were not ready. No one was talking about agents back then. People were talking about like generative AI. You remember, it was like, that was free, like baby AGI and all those kinds of early experiments, you know, so agents didn't really work, you know? And, and so as a result of that, there was no tooling. And so as a result, unfortunately we had to build a lot of our own

Infrastructure and Vendor Stack

52:17tooling, including our own eval platform, including our own observability platform. And I don't recommend doing it at home. Honestly, it's like just to use. And so we've gone through a whole evaluation. We've looked into LangFuse, to LangSmith, into BrainTrust, into AgDust, in which I'm a proud investor. We've looked into a bunch of those solutions. And every time we were like, I think we just had such a headstart. And at this point, our homegrown solution is so built exactly. We've gone through the pain of building, you know? So if a solution had been available back

52:49then, I would have picked it, but it wasn't. So we built it. We fleshed it out over the, over the years. Now we have like one engineer who's full-time dedicated to that solution, which is a lot in, in the vibe coding era. This guy is like cranking and he's, it's become like a very, very, very extensive solution. And truth to be told, we are at feature parity with all of these solutions out there and more. We have a lot of internal features that I don't see these guys have. And it's so built for our own internal scaffold and our own internal use that we ended up just hand rolling that tell me. Yeah. So what does life look at, look like at Lindy internally these days? I think it

53:25strikes me that you're now, I'm sure there's been like a dog fooding process. I know there's been a dog fooding process throughout the history of the product, but now you're at this point where it's like, there's no limits to the dog fooding, right? You can do anything. Yeah. How does that play out in

Working Inside an AI-Native Company

53:42terms of, I don't know, you could cut it any number of ways, right? Are there roles that you might otherwise hire for that you just wouldn't hire for now? Are there, uh, what's the ratio of your own inference spend for the purposes of building Lindy compared to your payroll? You put your own lenses on it, but how has the arrival at the human or at the AI employee stage of this process changed what it's like to be a part of the company?

54:14That's a great question. And I think right now, I mean, obviously, AGI is here. And so obviously right now, every company is sort of scrambling and there is that time of transformation that's accelerating. And I will say, and I'm, I'm, I'm, I'm a little bit of a doomer. I'm, I'm worried about AI risk, but so far so good. And so far, it's so much fun. It's, it's honestly so much fun. Like to go through AGI and to build this in AGI, it's like, we can move so much fucking faster. I feel like we're moving at the speed of thought now. It's like, there is no obstacle. Like normally companies, it's like, you think about the strategy, it takes

54:47a long time to figure it out, to implement it, to see the results. It takes like months, like that OODA loop is like months long, you know? And obviously the name of the game at a startup is to move as fast as humanly possible because that's the only advantage against the increments. And now we can just have ideas and see them live in the product two hours later. And like big ideas too, like not small ideas. And I can tell you, so our Slack now is basically just us and Lindy. It's like Lindy is like just half the messages

55:18on Slack is just Lindy and us talking back and forth with Lindy. Very often, like we, we ourselves and we breathe this stuff all day. And like we ourselves, like we, we are underestimating the platform. Like just last week, we, we had like a problem with like our CI pipeline. We have literally many hundreds of runners on GitHub who like manage our CI pipeline. By the way, there's another thing right now to answer your question. What's it like to work at Lindy? Where's it not to work anywhere, frankly, is like every part of the stack and every part of your process is stretched to its

55:49very limits right now. That's just right now tech, you know? Like I'm sure you've seen the graph, like GitHub posted a graph of like number of commits posted per day. GitHub was like hardly like a subscale company. It was really big. It's, it's just going vertical right now. It's just going to the moon, you know? And that's, that's like, I think indicative of, of what every company is going through. Our number of PRs internally has like PRs per week has tripled over the last three months. And like the number of lines per PR has tripled over the last three months. We hardly review PRs anymore. It's just

56:21agents reviewing PRs. We, insofar as we review, basically what the, the way the team is talking about is like, well, no longer PR reviewers, we're like PR reviewer reviewers. You know, we're just reviewing the machine that reviews the PRs. And as a result of that, you're like, this is awesome, but hey, now CI is in the middle of that. So what do you do? Well, now you've got to work on CI. And so we've, we've spent a long time working on CI, like for the non-technical agents out there, like CI continuous integrations, like the process that's involved in taking into account your code changes, making sure they're safe and actually

56:54merging them into your main product and deploying them. And so we, we, at Philz, we started throwing money at the problem, like many people do. And so now our CI is extremely expensive. It's like, it's a, it's just an insulting expense. And, and, and we've exhausted all the, all the obvious, like we're obviously not running on GitHub runners. We're running on like GCP runners. And like, we've, we've done like caching of the bills. We've done all the obvious things. It's, it's a lot of work to manage CI to make it, to make it efficient. And, and, and so we, I was tired of this because it was like weeks of messing with CI and, and like

57:26burning a hole in our pocket. And we're like, what are we going to do about the CI thing? And then I was like, wait a minute, like, why are we doing all of this ourselves? Like, can't we just like spin up an agent to do it? And spin up the agent was literally. And so by the way, that's the other reason why it's important to have a single agent because setup now is becoming one of the biggest costs because the biggest cost, the biggest bottleneck in any organization right now is human time and human attention. So you want to remove that guy as much as possible. So you want to be able to provision agents and give them all the accesses they need without a human in the loop as much as possible. Okay. And so is this agent at the

58:01company? It is this AI employee that knows everything there is to know about the company Lindy teammate, right? It's got access to our repository. It's got access to, it's got its own computer. So like Lindy, can you please like, like mess with our CI? And it's like, yeah, let me look at like the runners using the GCP CLI. Let me look at like the GitHub, like action and all of that stuff. So it's doing like a lot of analysis. It's getting back to us. And now it's literally, it's basically like an auto tune, like an auto optimization thing. Well, like Lindy is just like sending us a graph like every day, by the way, using image gen to

58:31generate. So the graph is like really pretty. And it's like, Hey, you know, guys, like, look, like CI cost is going down. Time, time to merge is going down, which is a good thing. And I was like, I can't believe it took us two weeks of just like messing with CI ourselves instead of just saying this to Lindy, you know? So I think this is what it's looking like right now is like engineers are like less and less working on the thing. They are more and more working on the thing that is working on the thing. And more and more, what we do is like, we were setting up this machine, which is the company, which just increasingly operates itself. And increasingly, well, more

59:07and more like, you know, I like to think of it as like, first you become a manager of ACs, then you become like a director of managers, because at some point, like the managers are going to become AIs as well. I'm hopeful that at some point we become board members and we can just go to a beach in Hawaii and we don't have much to do except like review the strategy of the company. So we'll have like reports of the agents telling us, this is what's going on. This is the landscape. You know, these are the results we're having and we'll be able to have much deeper conversations about the nature of the company we're building and the spot we occupy in the market.

59:39That's a dream as far as I'm concerned. And that's not very far away at all. To answer your question, have we, have there been positions we've not hired? Like, yes, absolutely. You know, we, we, the team is actually, I think headcount has been flat now for a while. And we've, again, like I said, we've actually tripled our productivity over a couple of months. And I find there is a bold deluxe zone because humans are still in the loop. Humans are still needed, but obviously the more humans you add, the more coordination costs you, you introduce inside, inside the company. And so I, I find that actually right now, smaller teams, relatively

1:00:12speaking, are at an advantage versus bigger teams in, in a lot of ways. Do you, I'm sure you do. I don't know if you want to share it, but do you have a ratio of not, I don't mean all inference spend, like including your customers use cases, but like your own internal inference versus your payroll. How is that ratio? Yeah. So if you look at all the inference spend, including all customers, it's, it's many times payroll is, and it's been for a very long time, but one exception because that's what we sell, you know? Now, if you look at payroll versus internal inference

1:00:44spend, it's in shooting range. So payroll is still greater, but the lines are going to cross three to six months from now. If you imagine a hypothetical scenario where all of a sudden it's just you at the company and after that it's Lindy's all the way down, what breaks, what doesn't work anymore? Like in other words, what is it that the humans are still irreplaceable for? It's really hard for me to answer this question because, and I think about that a lot because

1:01:18I am finding myself increasingly lacking the vocabulary to answer this question of like, where do agents fail? You know, like we used to talk about like time, you know, it's like, ah, time before it loses coherence. It can go like the meter thing. It can go like a one hour, 10 hours, 100 hours. And I'm sure people will resonate with this as well. Like I, I don't find that's, that's a good, that's a good measure anymore. You know, the, the wheel that everyone is using these days is like spiky, right? Like these models are so very spiky as we're going through AGI. The wheel AGI has stopped me in this thing because it's actually ASI in many ways.

1:01:51And it's, it's, it's subhuman in, in some very surprising and dumb ways. It's really dumb in a lot of ways. I'm like, Hey man, like you, you can produce for me like 50,000 lines of code one shot. You can build me the most incredible ideas in like three hours by yourself. And then, you know, you, you, you're telling me to like walk to the car wash, you know? And so that's where it fails. Imagine you've got this, this, this machine that runs a company super fucking smart, but, but, but 50 times a day, it decides to walk to the car wash, you know? And so I think right now, like the,

1:02:27the challenge of these new organizations that are going to be for the foreseeable future human AI hybrids is going to be, you're basically building an Ironman suit. Okay. And so there is this slider where it's like, okay, this is for the machine. This is for the human. Okay. And knowing where to put this slider, because it's not a one dimensional thing. It's like, it's a complex line through a many dimensional space. So you are looking for that slice that you can draw through the space because you want to make the most out of humans attention. And so you don't want the AI to bug you

1:02:58for stuff it doesn't need you for. And you need the AI to know when it needs to bug you and almost definitionally, it can't, because if it knew it wouldn't, it wouldn't bug you. It wouldn't, it wouldn't make the mistake. And so I'm sorry. I feel like I'm answering this question with the question, but that's, that is the right question, you know? And, and I think right now, that's one of the open questions in the industry is like, how do you draw that line? How do you get the, the agent to realize when it's and so on and when it's about to screw up? Like when it's even saying something like, please walk your car to the car wash because it's, it's study today. You know, I think that that's the answer. It's like, it's just good. I'm, I'm somehow humans and I,

1:03:30we're going to have to, this whole system and at gas times when it's being very dumb. How about the times where something is being very smart? Where do the best ideas come from these days? I think there's long been this idea that the AIs are good at routine stuff. I used to say, and clearly this is no longer applicable, but I used to say, you can probably automate any routine work that you have at your company if you're willing to put in the elbow grease, but don't expect Eureka moments. Now, obviously we're getting Eureka moments on the

1:04:02timeline all the time, at least in certain domains, like the open math conjectures and stuff. When it comes to the best ideas at Lindy, how many of them are human origin and how many of them are AI origin today? I hate the sensor, but it's, it's both. It's, it truly comes from this, from the union of both. And the reason I hate it is because it led us into this myth of the, the centaur, you know, it's like this mystical man horse creature. And so, you know, it's this idea that like,

1:04:33yes, an AI is better than the human, but you know, what's even better than the AI is AI plus human. Hence, humans are always going to be needed. And that's a fallacy. That's just not true. And the, the literature is actually clear about this. Like, you know, we've seen it happen with chess and with every other game, which, which AI has achieved superhuman performance on, where at first AI beats human and AI plus human beats just AI. And little by little, the gap of AI plus human versus AI is shrinking until it actually turns negative. And humans are introducing at best, like random noise into the system. So humans actually, at some point become

1:05:05like truly clueless and, and they start harming the system in which they are intervening. That said, the centaur phase exists for a while, you know, so the open question is like, how long is it going to exist? But right now we are in the centaur phase. And I think that's the other reason why multiplayer is so important because you want AI to be, you want your agents to be in the room with you. And you don't want any random agent to be in the room with you. You want an actual AI employee who's been in the room with you for a long time, who sat in every meeting, who's built up the internal knowledge base about everything that's going on in the company.

1:05:38And you can add mention it and you can invite it to chime in into your conversation in the company. And we do that all the time. We do it. Actually, it's funny because it's very easy to do it on Slack, but it's like, add Lindy, what do you think? You know, it's hard to do in the meeting in a way that like now makes us actually more and more, well, like we prefer Slack because like it's the AI is there. Whereas like in the meeting, it's there, it's listening, but it's called chime in. And so very often now what we do during meetings, we're like, actually it's here, like, you know, we're struggling with this thing. We're like, Lindy, can you please send us a message after this, telling us what you think? And they'll send us a message on Slack. And, and so when you

1:06:13do that, it is not the case. It is rarely the case that Lindy will first shot, tell you something and you're like, this is it, this is the thing we should do. But it is almost always the case that we, we go back and forth with Lindy and, and, and, and from that back and forth, like it misses something. Like the CI example is, is a great one, actually. At first, when we, when we invited Lindy into this conversation, it happened over Slack. It's like, wait a minute, we're done. Like, why are we not doing this with Lindy? Add Lindy, can you please optimize our CI?

1:06:44And, and it, make no mistake. Yeah. Yes. And I forgot its first suggestion, but it was, it was, it was a reasonable suggestion if you did not understand our CI. And it was, it was, it was, it was, it was, it was not good. Like given, given the constraints we're operating under. And so like the engineers chimed in and they're like, no, actually we can't do that. We've started by that. We're like, ah, that's a fair point. So we went back and forth. And then, and then together we came up with like, oh yeah, yeah, we should totally do that. Is that just go ahead and do it, please. A thesis that I've had for a while, it might be too early to say,

1:07:14but you're definitely going to have a perspective on it is, but again, going back to what I've been wrong about, I expected a lot more transformation than we've seen in the economy. If you had showed me fable in 2022 and said, what's the unemployment rate when this model is out, I would have said, I don't know, but something definitely higher than what we are currently experiencing. Yeah. So then why, and this has been something I've been wrestling with for the intervening time. Like, why are we not seeing more? And one answer I've come to is people are stuck in their ways and

1:07:46not everybody's a enthusiast like I am. And it's going to take time. But then also when we get to the drop-in knowledge worker, that'll be the time when it'll suddenly flip, right? Because now I'll have a more parody kind of choice between I could try to go hire a human or I could plug in an AI and that's going to be easier in a lot of ways, right? I don't have to go through a whole interview process. I kind of know what I'm getting, blah, blah, blah, blah, blah. I don't need to tell you why an AI employee is a good form factor. Do you think we are starting to see this flip and is it going

1:08:18to be something that incumbents will be able to adopt in time to avoid getting disrupted?

1:08:28Or are there still going to be bottlenecks or resistance points that I'm not anticipating? I expected more disruption as well. I think as long as the models are spiky, we're going to need a lot of humans in the room because the holes in the model are very dangerous. Not just like AI, existential dangerous. They're just harmful. They're really harmful to the system. And so however expensive humans are, they're going to earn their keep by plugging these holes.

1:08:59Now, regarding whether incumbents will be able to adopt this thing, I think as long as it is spiky, I think the incumbents are going to be at the disadvantage. If you had a true drop in remote worker with no holes, then I think even the incumbents would be able to adopt it quite quickly because it's just like hiring an employee and they have processes for that. And I actually think they would be south on a lot of the innovation because it would be skeuomorphic effectively. You would just have like AGIs like sitting in human seats effectively. And I don't think that's the best way you build an

1:09:34AGI organization, but it would work. The smaller organizations are going to be at an advantage both because they are going to be able to short term figure out this AI native organization where the human is filling the holes. And long term, I think they're going to be to build like a less skeuomorphic AI organization. You know, economists are very fond of talking about that. I actually wrote a blog post about this that's called the Tough Tomato Principle on my blog in 2018 or so. And it's about how like every time you have a technological revolution, you think of the new

1:10:10paradigm in the terms of the old paradigm, you know, and one tale that you're falling into that trap, which is very natural, is that you are calling the new paradigm after the old one. So, you know, the first cars were called horse-less carriages. Right now, self-driving cars are called self-driving cars. Right now, eventually, we're going to have a wheel for them, you know, that's not in the terms of the old paradigm. It's almost like right now we're doing it with AI employee, huh? Like it's like AI employee is a strong, as a term, is a very strong sign that you're doing

1:10:44something very wrong in how you're thinking about your product. And I think it's fine if you're doing it just to communicate about your product because like the nature of positioning is that you have to position against something that already exists in the market's mind. So, the car exists, you know, the carriage exists, sorry, the carriage exists, the car exists and not the employee exists. So, you can't really invent a new like wheel out of nowhere. The iPhone did that. You know, it's like the iPhone is obviously not a phone. It's so much more than that. Like, you know, what percentage of your iPhones use is just like making phone calls, like 2%, you know?

1:11:17But it's called an iPhone because it's an electronic thing you've got in your pocket, right? So, I think AI employee is one of those things. It takes a very long time for industries to find the message which is native to the new medium. I think it was Steve Jobs who was saying like, you know, the first content on TV was just either like recorded radio talk shows. They took like a radio show and they put the camera and now you've got the TV content. That is what talk shows are. Or recorded theater plays. And it took a surprisingly long time, like 10 or 20 years

1:11:52before people realized like, wait, I can move the camera? Wait, I can have multiple cameras and I can cut? You know, I can do, I can change the scene. I can do, I can do a lot of crazy stuff now. And, and, and, and, and it's treacherous because it's thankless work. Once you've realized it, you, you, it's so obvious. And this is why when you, when you watch like old movies, like Citizen Kane, you know, or, well, Hitchcock has a lot of those, like you, you, you, you, you're like, they're nothing special. Why was everyone crazy about Citizen Kane? It's like, well, actually it was

1:12:22groundbreaking at the time, like the Beatles or like that, you know, it's like, well, you know, it was, it was actually really innovative at the time. So I think, well, right now, also in that phase of like the AI employee, we've got the AI employee. And so now one's the phase of trying to figure out what does the AI organization look like? And I think that's, that's a young man's game. And it's a young company's game. I think, I think incumbents historically have, have really struggled to answer this question. Do you have any previews that you feel confident in? Or is it too early to say, or maybe you would say, are there any thinkers that you think are kind of notably ahead of the curve on this? The only post I have seen and liked about this was from

1:12:57Dworkesh. It was, it was a while ago. It was a year and a half. That comes to mind for me too. Yeah. Yeah. He asked like, what does the AI organization look like? And he, he points out, like, I think it's, it's an excellent intuition pump. He points out, he's like, look, you know, Sundar at Google is paid $150 million a year or something like that, you know? So obviously there are agents, which companies are happy to pay $150 million a year, even though their token throughput is actually low. Like, it's not like Sundar is just like emitting 2 billion tokens per second or whatever, right? It's just like it has really good tokens, right? And so now, so what does that mean? Like that indicates something about

1:13:31the willingness to pay for models and for agents. There is this book by Robin Hanson. I'm looking at it right now, The Age of M. And it also explores a lot of those themes, right? It talks about, have you read it? It's excellent. Highly recommend it. Yeah. Yeah. I'm laughing only because I tried to do a podcast with him about that book and it went in a very different direction. So it just brought that memory to mind. But I think that book is excellent. I think it is a, as a primer for, or as an example of how to take a premise

1:14:01and really run it out. It's elite. Yeah. It's a book-long start experiment. So for people who don't know what it is, like M stands for emulated mind. And it's like, hey, we are going to have the hardware to create emulated minds, you know, in like the late 2020s. So like we're right on track. It's like, we're just not going to know how to do the software. So all we're going to do is we're going to emulate like literal human brains. Not exactly what's happening, but kind of what's happening. You know, like LLMs are trained on human tokens. And so that's basically what's happening. And so it's like, and it goes on all of those riffs of like, hey, this

1:14:32is all the crazy stuff you can do when this happens. I'll talk about two of them because I don't want to take the whole time and everyone should really read the book. But just to give a taste to people about like what it sounds like to have those native AI organizations and the kind of stuff that you can do that you can't do when you have a human. The first one is splitting. And so I think this is why, this touches on what I was saying earlier, like don't create multi-agent organizations. As much as possible, try to have as much of your internal work as possible. Emphasis on internal. We can get back to that later.

1:15:05But as much of your internal work as possible done by one agent and one only. Why internal only? Because if that one agent which has got the keys of the castle and access to the bank account and the repositories and the secret keys is also the same agent that does customer support, you potentially have a security issue. So it can be good to have multi-agent systems be it for like only for like security reasons. But okay, so you have that single agent. Now that single agent, you know, it sure it goes faster. It emits more tokens per second than a human and it doesn't take a break. But at some point, you're going to need more tokens

1:15:37than like a single LLM loop can do. So what do you do? Well, you have multiple LLM loops running and it's basically just like that agent duplicating itself, like forking itself, right? And creating sub-agents and like workflows, dynamic workflows and all of that, right? So, and that's one thing that Robin Hanson talks about. Like, you know, you can imagine you ask one, literally in the book he says you could ask one emulated mind to build like a really complicated thing, like an operating system, which is like billions of lines of code or something. And at first it would just like take it at the highest level and like come up with the highest level architecture of its operating system. And then it would fork itself and each sub-copy would be in charge

1:16:12of like one building block of this highest level thing. And then each of them would recursively keep splitting themselves such that it's a, it's a fractal. It's like every, every, every node in the graph has got like the whole picture in their head and every node in the graph is working on an implementation. And then they all bubble back up and you could imagine multiple passes up and down. And in three hours, you've got an operating system. It's already like such a wacky idea when he wrote the book 10 plus years ago. Now I think it's, it's very clearly what is currently happening, right? So that's, that's one thing that's happening, which is, which is really interesting. Another thing that I've not seen that yet

1:16:45happened with, with LLMs, but it's an interesting self-experiment about what does cross-organization collaboration look like? And what you can do is that you can create unfalsifiable agreement. So suppose I came to you and I was like, listen, I have information for you that I cannot tell you, but I can tell you that if I could tell it to you, you would agree. And you would give me all your money, right? Yes. I, you wouldn't give me all your money right now. It's, it's, it's a bit too easy, you know, but if I could prove that to you, if I could prove you with,

1:17:16with absolute certainty that yes, if you heard that information, you would give me all your money right now. You know, it's just, I can't give it to you. The way you do that with M's is that you, you clone yourself, you clone me, we put both of us in a, in a box that's going to self-destruct and they talk and your, and your M has a button and that's the only, the only communication means it has with the outside to be like, yes, give him all your money. Right? So it's a copy of you. It is you agreeing. And so you have this zero knowledge proof of, of, of, of crypto bros. Well, we're like, right, I guess on this one, you have this ZK proof of, of, of,

1:17:52of that. So that's the kind of thing that's on my mind. And I think that's the kind of thing we're going to see exist soon in the next few years with AI native organizations. It's about to get weird in a whole bunch of ways, I think. And you better hope they don't break out of that box too. That's the other worry one might have as we put our agents into boxes these days. One, let's do one more beat on the stuff that's relatively mundane. And then we can zoom out to some, some real big picture considerations, but you did have, as you alluded to earlier, this highly viral

1:18:22post, I don't know, two months ago, maybe moving a significant part of the workload to open source models for cost saving reasons. And for, you don't need God to schedule your meetings reasons, but where are we now? It strikes me that like, when you talk about just having one agent that kind of seems to go against that. There's also the sort of can advantages, which you lose when you are crossing agent providers too much. And then of course, there's also proprietary providers have

1:18:57their families of agents and they're like increasingly training their best models to delegate to their haikus respectively. Like, where do we shake out on this now? Is there still a significant role for open source, cheaper models to play in the Lindy teammate? Or are we back to heavily quad based agents today?

Open Source Models and Caching

1:19:22No, we're quite open source build. Like, look, it's just so much cheaper. It's like, it's ridiculously cheap. Like, and like the cheapest of them that we really like is DeepSeq Flash, which is, it's free. That fucking thing is free. You know, and so as you, as you experiment with your agent and spin up a lot of them, like the bill can really go up pretty quickly. And DeepSeq Flash is just incredible, you know, and quite fast as well.

You know, we, we do find that as you go to like the upper echelons of intelligence, if you look at like a Kimi K3 or like a GLM 5.2, like I still am super impressed by those models and they're awesome. And we, we are increasingly considering them as like our main driver for, for a lot of parts of Lindy. But I will say that the gap narrows between those guys and like the frontier guys, like Kimi K3 is, is frankly, not that much cheaper than like a Sonet or like an Opus.

So, but no, otherwise, I mean, like I, I, yeah, I do think open source models.

1:20:14Or like if you kill the price and if you don't mind using those providers that yes, you do have stuff to figure out around caching. Like yes, the, the, the inference providers are like not as good at caching, but they're, they're catching up. And, and, and, and look, you know, if you look at like a DeepSeq, DeepSeq Flash, it's Sonet 4.6 level ish, a bit less, but kind of ish. Sonet 4.6 is a really good level, a really good model for like most use cases. And it's literally a hundred X cheaper, you know? So even if you, even if you miss on 2X because of caching yourself 50X cheaper, it's a really big difference. The difference between spending, spending like a thousand dollars or $50,000. So no, I think open source models are like a required part of the stack right now for, for anyone who's like seriously building and operating AI agents.

1:21:00So how do you think about deciding where to integrate them? Especially cause I didn't sound easy before, but in a context where you have your workflows and there's nodes in the sort of graph of work, you could go into a particular node and say, okay, we kind of know what the inputs here are and the outputs. And it's like relatively controlled environment. And so we can do structured testing and did all that. And you still reported having some, uh, false positives over time where a model could pass a bunch of tests and you would feel good about it.

1:21:34And then if you start to test it live with users, you'd get the feedback that like, Hey, like it got dumb and I'm not, you know, somehow it was like still hard to measure, even with like a much more structured environment for the AI to work in. Now, as you're in this very open-ended, just tag Lindy in Slack and send it anything. Um, it seems like that problem would have increased in difficulty dramatically. So how are you approaching it? And what are we learning in terms of where the open source models? And it's not even so much, I think Anthropic does make some good points about this sometimes.

1:22:13Obviously there's more to the story, but I do think they, they're apt in some ways where they say, in some cases, it's less about open and close source and more about just like what's capability level? What's the price? What are the features? So regardless, I guess, of whether you're even going to open source or just going to Haiku, how do you decide when you can do it? Great question. We have found you should almost never use multiple models powering the same agent. And we have found one exception to this rule is when you spin up new blank sub-agents.

1:22:43Emphasis on blank, because there's two types of sub-agents. You have a blank sub-agent and then you have a forked sub-agent, which inherits the context window of the parent agent. And the reason why you don't want to do that is because of caching. We just obsess about caching, obviously, because it's just so expensive otherwise. Like the economics do not, they barely work with caching. They just cannot work without caching. So caching is a must have. And so I'll give you an example, the validator system that I mentioned. So it's this LLMS judge. By the way, I highly recommend anyone who builds AI agents. Like this is one of the lowest hanging fruits you can do to greatly increase the reliability of your AI agent.

1:23:17And so that's step one, just like have a validator, which is like the naive implementation is like you asked agent to do something, it does it, or like it submits an action candidate, you intercept the action, and then you ask it, are you sure? Literally, if it's just are you sure, already you get a bump on your avals, which is insane. It should not be the case, but it is the case. Now, if you increase, if you change this are you sure to like an actual prompt, which in our case now is like 10,000 tokens, it's a really big validator prompt. If you change that, then you actually, you're giving it a checklist and that's got reasoning tokens.

1:23:48And now it's way better. Like you can get Sonet to perform above opus level, you know, in our experience. Now, then you can get, you can go one step further. You can, you can have like a federated suite of validators and some of these validators may be deterministic. So for example, one thing we and the rest of the industry has found is that Sonet, I mean, cloud models this year have been getting dates wrong by one day, right? That's one way in which they'll spike it, right? It's like, hey, you've got your AI organization, it gets dates wrong.

1:24:20It's a problem when like us, you're building an AI executive assistant, which half its job is to schedule meetings. Okay. You can't get dates wrong. So what we did is that we created, it's a modular architecture we have. It's really simple. It's just like a bunch of validators and then it's like a promise dot all with a timeout for people who know what this means. And so every validator has like a second to like decide what to do. And if they've not submitted their verdict, then time's out and the agent just submits the action. And we have a validator which job it is to detect dates that were submitted in the action.

1:24:55And we've prompted the agent to include a weekday in the date. So it never says July 28th. It says Tuesday, July 28th. Okay. If you do that, then you can have a deterministic validator that checks whether the date of the week matches the day that was submitted. And that's just a regular expression. Instant, free, no AI in the loop. And that's one of those federated validators. But even the validators that are AI powered, you don't want... Initially, what we did is we were like, what if we had Sonnet for the main agent and we had like DeepSeek Flash for the subagent, the validator?

1:25:29But actually, because when you have a cache hit, it's 10x cheaper.

1:25:35You actually, unless the model you're going to use is more than 10x cheaper, which it may in the case of DeepSeek Flash, it may not because the caching is inferior. You know, if that's the case, you actually do want to keep the same model. And Sonnet is so much smarter, like it's actually worth it to just like keep the same model. That also holds if you fork your agent, like that way you can recycle the cache. By the way, interesting note, how do you reuse the cache when you have this validator? You know, because the problem is that changing the toolset invalidates the cache. Okay? So the way we've done it is that the validator and the agent has the validator action always included in its toolset.

1:26:11And it's unable to invoke it. We tell it, don't invoke this guy unless you're the validator. And if he tries, we're just like, no, I'm not going to listen to you or not the validator. And then the validator now inherits all the actions of the agent. And it's like, you are now the validator. You may invoke this action, which you have known about the whole time. So that's how you don't break the cache for this kind of pattern. Yeah. And so forked agents are the same. You just use the same model. You know, I've heard friends, funders who've told me, I don't even want to check if it's true. I'm sure it is, which disgusts me.

1:26:42If you take an agent and you change the model at every turn between roughly equivalent models. So one turn, it's Sonnet. One turn, it's Grok 4.5 or whatever the latest is. And one turn, it's like GPT 5.6. And one turn, it actually increases the performance. That ensemble, and it's literally just like a random sequence. Okay? That ensemble, for some fucking reason, outperforms any given model. I don't want to know why. But apparently it works. Again, we've not even tried it because we don't want to break the cache anyway. Okay, that's really interesting detail.

1:27:14Can we bottom line it a little more? It sounds like top level, Claude is still the kind of core main driver. Because of the caching, Claude is also often the validator. Blank subagents you can put onto a deep seek flash. This sounds pretty Claude heavy, all things considered. Is that fair? No, oh, sorry. I was not realizing you were still asking me what model we're currently running on.

1:27:45Right now we're running on deep seek. Deep seek is the main driver. So it's the driver. Wow, okay. It's the whole thing. Everything is deep seek right now. Yeah. We are in the middle of reconsidering it because we are realizing that TeamMate, the new product, requires beefier compute. And by the way, you can always go in your lindy settings. We are model agnostic, so you can always select Sunnet or Opus if you want to. By the way, one interesting learning of mine is it doesn't matter how often you tell people, hey, we promise on the benchmarks it's the same.

1:28:15There are quite a few of our customers that they're like, I don't care, I want Sunnet or Opus, which is very expensive. But no, right now it's deep seek by default. Interesting. So you feel like the, yeah, I guess how would you characterize the gap between American API models, Claude perhaps specifically, and a deep seek? We hear, I think we kind of go through these cycles, right? Where it's like, oh, the gap is closed. Oh, maybe it's open again. And it's actually, it was totally on trend the whole time.

1:28:47It was always nine months. Qualitatively though, for the purposes of actually making an AI employee work, how would you describe the gap? It's more spiky. It takes more turns very often to find something that works. So it still ends up finding something that works, but it ends up taking more turns, which is slower and more expensive. I know you asked me for something qualitative, but look, it is like three or six months behind. So right now we've got Sunnet 5. Like deep seek is basically Sunnet 4.6 is the way I like to think about it.

1:29:18How much do you worry about when people are allowed to change the model? I always have this question of, first of all, are we in a period of convergence of the AIs or divergence? I'm even on that basic question. I'm like sometimes confused and a sort of distinct, but conceptually related question is like, to what degree do products have to be co-designed or co-evolved with models to really work well with them? You could imagine that switching to Sunnet might make it worse because of all the work

1:29:50you've done. And in fact, I've heard that about Opus 5. And I've heard a lot of different things about Opus 5, which I don't think coherent anything really easily summarizable yet. But I have heard from some quarters that, hey, you probably need to rethink a lot of your system prompts and whatever with Opus 5 or it may not follow your skills very well. It can be more powerful, but you've got to kind of go back and redo things. So, yeah, how do you find that playing out in practice?

Prompt Optimization and Fine-Tuning

1:30:17We do find significant differences, surprisingly significant differences between the model families. I was expecting, I think I was, you know, you and I were talking together years ago and, you know, at some of those are going to come to the same region of the space. And that's not happened as fast as I was hoping. And I'm hoping for it because obviously that makes the models commodities as far as I'm concerned. It makes my life easier. What is actually happening is that the different model families have significant gaps in how they interpret your prompt. And so by family, I mean like Cloud and OpenAI and even Meta.

1:30:49Now it's sort of back in the game and Grok, like these guys, these guys are quite different in the way they interpret their prompts. And then indeed, whenever you have a new major jump, 5 was a really big jump. It was, you know, 4.7 and 4.8 together. Like there was a change, I think, in the tokenizer. So it was very different compared to 4.65 is bigger than that gap. And so what we've found, and we've been lucky enough, this is part of the tooling that I was mentioning earlier that we've been investing in, is like we have our own GPAM self-optimization

1:31:19loop. So we have hundreds and hundreds of evals, I think at this point more than a thousand. And we have an optimization loop where it's like an agent that runs the evals and finds, like tweaks the prompt to maximize the school on the evals. And by the way, you can vibrate it. It's like everything you can vibrate now. And yeah, I mean, every time a new model comes out, we have to run that loop and you give it a budget. And at this point, we have to give it like thousands of dollars because it's a lot of evals to run. So it's like about like $10,000 every time like a new major model comes out. So it's like, hey, here's 10 grand. Just re-optimize your prompt around this new model.

1:31:51And we do find that there are a lot of updates. Like the prompt does change quite a bit every time. I will say the open source models are all very cloud-like in their behavior. I wonder why. Yeah. I definitely have some questions for you on that. When you do this optimization, are you, is this also like a homegrown framework or are you using like a DSPy or I forget the name of the kind of spiritual successor to DSPy? DSPy? No, no, no. It's all homegrown.

1:32:21Yeah. But it is a sort of auto-optimization process where you're like, here's an eval set. That's right. Auto-optimize your own prompt to climb this set of hills that we've got for you. That's correct. That's correct. The GPy is like a generative Pareto Frontier or something like that. Yeah, that's the one. It looks for like the Pareto Frontier. So it's looking for like the best prompt across all of your eval set. But we're looking into tweaking the system right now so that we can assign weights to evals. Well, like this eval counts as like 10 of this other one because it's really important.

1:32:54But right now it's just like treating every eval as equal.

1:32:59Fine-tuning as a dimension in this whole model situation. This is another thing if I go back in time, I'm like, I definitely expected a lot more fine-tuning than we are getting. So OpenAI is retired or on the verge of, they've certainly announced and I think maybe at this point have pulled the trigger on retiring their fine-tuning product. We've got thinking machines trying to answer that call with their own model that's specifically to be fine-tuned. Yeah. Is that going to be part of the future of Lindy Teammate?

1:33:30Yes. I think fine-tuning is the last result. It's something you do once you've exhausted every other option because it's such a pain in the ass and it's expensive. But it's gotten a lot easier because now we have EGI and so you can just ask Cloud to fine-tune for you. Just give it a data set. Bringing the data set together is still like complicated and like sanitizing is complicated and it's still a pain. So it's a last result. It's not the low-hanging fruit. You should do everything else before you fine-tune, including Jeeba. But then, yeah, at some point you get there and this is why like larger companies are the first ones to fine-tune because they have exhausted the low-hanging fruit and at some

1:34:03point and they have the resources to fine-tune. And so it does make sense. It does give you more performance for cheaper. Do you think that that, like how broad do you think people should be thinking when they are approaching fine-tuning? Obviously, the extreme would be one fine-tune per task. You can definitely do multi-task fine-tunes. It's maybe getting into pretty challenging territory to say we want to make our own general purpose

1:34:34model that has the same breadth of action space as the big ones, but it's like ours somehow. How would you guide people on like how big to think when they start to approach fine-tuning? Most people should not fine-tune. I think if you work at like at scale, you should fine-tune. And then the more at scale you are, the more fancy you can be. But I generally, one meta-heuristic I've developed over the years is like you should place an enormous emphasis on simplicity. Enormous emphasis. And I think it's just too complicated to start to have multiple fine-tunes for different

1:35:08multiple use cases and users and model routers like you don't want. No, no, no. Just one model. You know, you don't, like I said, you don't switch model mid-stream. And if you fine-tune, you just fine-tune into one model. I think one, again, this goes counter to this meta-heuristic I mentioned of extreme simplicity. But the thing I find myself going back to all the time is like, because I spent so much time thinking about context and memory, right now our memory is like these millions of tokens stalled on a file system and this memory agent that's really, really sophisticated.

1:35:39It does feel like memory belongs to the weights. This does feel like a little bit of a hack. And so there is something to be said about that. I think eventually with infinite resources, what I would like to do is I would like to do a LoRa per user. Because LoRa is actually pretty cheap to train. So now it's no longer napping. Now it is actually dreaming because you can't retrain the LoRa every 15 minutes. You would have to do it every day or every week, you know. And it's a tremendous infrastructure challenge even to like, so you got to store all of those LoRa's, which like the storage is not a problem. But like the inference time swapping in and out of the LoRa's is a huge pain in the ass.

1:36:14And the training pipeline and all that stuff, so good luck doing that. So, but it would make sense because I think weights are so much more compressive than tokens. Yeah, I don't know if you have any thoughts on the great horizon scanning for continual learning. But this is, again, for quite some time has been the thing that I'm like, boy, if that ever tips with, and it could be as simple as one key insight, we could be in a very different world very quickly. I have a feeling what I'm describing is a step towards continual learning, but it's not

1:36:49the final step. I think the final step, I think like the bitter lesson will have you, like inference and training need to be one and the same. Like we can't separate them. You know, that's sort of what I think the entire field is looking for right now. I agree. Once we get there, it's going to be crazy and scary and very, very, very different. But yeah, I think, I think we all going to, the LoRa thing I just mentioned, like, look, it's an engineering problem. You know, you can figure it out, especially not to have EGI. So it's, yeah, I think, I think, I think the LoRa thing is going to happen soon in the

1:37:20next six months. I think even like one of the frontier labs may do it. I thought that was honestly one of the things that OpenAI did extremely well with their fine tuning product. They were clearly doing something like this because they would allow you to do the fine tune and then you'd have the same rate limits with your fine tuned model as you had with the base models. And I was always really impressed by that engineering accomplishment. And now it does surprise me that they've gone away from it so much, but I also agree with your guidance.

1:37:50I would definitely tell people don't rush into fine tuning these days. It's a slow cycle and you could probably get what you need for less total effort. Available models that you don't have to monkey around with in that way. 100%. So you talked about scary stuff, which is maybe a good transition to a kind of not second half because we've been at it for a while, but a second phase of this conversation around just zooming out and looking at the super big AI picture. Maybe I'll let you choose the order of topics when it comes to Chinese models and what, if

1:38:25anything, should be done about them because you did recently post something I think quite counterintuitive given the fact that you're running your company on DeepSeek. I'll let you state your position there, but should that come before or after the big picture? Where are we in this kind of, oh my God, we just had open face happen.

1:38:46What does it mean and what should be done about it? Yeah, I'm trying to coin that. We'll see if it sticks. I searched for it the other day because I just got back from China myself actually. And I was like, how has nobody called this open face yet? I'll be the first one to try. How about open gate? Because that's what it is. It's an open gate. And I'm sticking with open face just out of pure, you know, shock value, I guess, if nothing else. But yeah, we can fight it out in the marketplace of ideas. I guess you tell me what we should talk about first. And I guess it depends a little bit on like whether the Chinese model argument is like

1:39:19upstream or downstream of your kind of big picture concerns. The open face incident is immensely concerning. I think it's the most concerning incident I've seen happen so far. And I know that my feeling is shared in the labs, like my friends at the labs or some of them are panicking. Like there is an air of panic right now, like intense fear in the air. So that's that. Then regarding, so that's where we are, you know, right now. It's like July, 2026, we have AGI, we are in takeoff and we've not figured out alignment.

1:39:52That's the TLDR. Wow. You know, the current timeline looks much too close to like an Eliezer Yudkowsky's essay for comfort. Now, regarding the Chinese models, in brighter news, Chinese models.

Banning Chinese Frontier Models

1:40:07Look, I, I, I, okay, I'll start by saying that I have been bemoaning the low quality of the discourse here. It's, it's been, it's been very disappointing because I, I thought tech was different. You know, like the quality of the discourse was something for work, right? The cultural war. I'm like, all right, whatever, you know, but then it's, this is tech, this is nerd stuff and we can't get all shit together and, and just remain polite and civil in the marketplace of ideas and address each other's ideas at the object level. Like, meaning like don't attack each other's intentions.

1:40:38Can you please just attack, address the argument that just put false? Can you please not pretend I said something I didn't say? It's, it's ridiculous. Yes, XFAPIC just put out an essay, which has bolded, we do not support the ban on open source. And then people are like, oh, I can't believe, like quote tweeting this thing they've obviously not read. And they're saying, oh, they're supporting a ban on open source. So I'll just start by saying that. It's like, can you please, can please everyone take a chill pill? Stop calling everyone a retard. I recently put out a blog post where I was like, like, hey, we, I think Chinese models will

1:41:10be banned. And I was called a retard 50 times by supposedly smart and accomplished people, including like famous VCs. Now, a lot of other like famous, accomplished, smart people reached out in DMs and by message and there was a lot of support. My position is basically Anthropics. I hate saying that because people are saying, oh, you're just an Anthropic shield. But look, I have timestamps. Like I've been tweeting my entire positions throughout the whole thing before Anthropics clarifies theirs and my position is exactly the same. I don't have anything against open source. I may, but I'm undecided.

1:41:42I do think open source might increase existential risk. Let's put that aside. I'm undecided. This is not the crux of my position right now. God bless open source. Important for innovation, important for companies, including mine. Like, please, open source. Okay. I have something against Chinese frontier models, whether they're open or closed source. And here the reasons are, number one, they're obviously distilling. It's very clear. And so you're putting American open source model companies in an unfair competition and closed source in an unfair competition

1:42:14because they're not allowed to distill. You know, it's just contrary, at the very least, contrary to the terms of service. And it may be illegal if you've put in place technical measures to circumvent any protection that the model company put in place to prevent distilling, which the Chinese models obviously have done. So it's distilled and it's unfair, right? Now, some people say, yes, but the companies have also distilled on human data. That is not what distillation means. That is just not what it means. Like, if you do the training on human data, it's costing you billions of dollars, like actually billions of dollars. If you do it on AI data, it's costing you hundreds of millions at best.

1:42:48So it's a huge unfair advantage and it's a significant enough portion of your costs to confirm you an unfair advantage. That's number one. It's unfair. Number two, very pragmatically, we don't want Chinese models operating in the U.S. Today, you know, it breaks my heart because I'm an American citizen. I'm a proud American citizen. I'm a hawk. When I talk to my own product and I'm asking it what happened in Tiananmen, it tells me, I'm sorry, I can't talk about that. That's a problem.

1:43:18You know, like these models are eventually subject, at the end of the day, they are subject to CCP censorship and CCP policies. You don't want those models. This basically amounts to being the greatest instrument of foreign propaganda on American soil ever. And there is, just ask your LLM of choice, ask Claude, hey, tell me the precedent we have of banning American, like, foreign influences in the country. Like, the three radio acts of the 20th century, like the TikTok thing that just happened last year. Like, we do that all the time, you know?

1:43:49Then, these models are not merely just an instrument of propaganda. They're agentic. They're actually doing stuff in the economy. You don't want the CCP to run chunks of the American economy. Duh. Then, finally, even if none of that was the case, maybe those models are playing fair and square. Maybe they're not representing foreign interests. Maybe they're just better. And so, here, a lot of people are, I think, indulging. And I say that as a libertarian. But they're indulging in what I call naive liberalism. Right? And it's like, well, let's beat them in the marketplace then. And I'm like, you know, like, I actually, and again, I said that as a libertarian.

1:44:25Like, I see protectionism is not always a bad spending. I think you do want to protect your domestic AI champions. And so, they can't say that. And they claim, and I believe them, that this is not their intention when they do that. And I 100% believe that because they've been so consistent for literally the founders have been worried about X risks for 10 years. Right? So, they've been so consistent. But look, I can say that because I'm not in the topic. Hey, we want local AI champions. Duh. It's a matter of national security. We don't want to hollow out that base with hollow data or industrial base.

1:44:59Duh. You know? So, it's not about the Chinese models. And it's about propaganda. It's about nothing. This is speaking for our economy. It's about fairness. It's about protection. All right. Let me try to give you some counter-arguments that you can respond to them. Um, I am a little more sympathetic, I think, off the top to the argument that there was some massive reappropriation of human knowledge that is upstream of all AI.

1:45:31And you could say that's not distillation. Sure, you could say the American companies spent a lot more than the Chinese companies are having to spend. I think that's also undeniably true. But if I'm just kind of what feels just and fair, I'm especially because we're blocking them from using Claude, right? We're not, it's not like we're saying, hey, you can buy all the Claude you want, right? So, we've said we've got Claude. The, the, the makers of Claude have advocated strongly for chip restrictions. And they also try to refuse to sell their model into China in the first place.

1:46:06And the whole premise of Claude or the whole existence of Claude is based on the idea that we hovered up, hoovered up all this human knowledge in every form we could find it from digital. And you got to believe that includes the whole Chinese digitized heritage. Now we're like sucking in the books, which I don't care about the fact that some books get destroyed in this process. But this wasn't something that everybody's consented to. And so it does feel a little strange to me, given all of those fact patterns that we would then draw the line and say, okay, like it's Anthropics consent. That's the consent that really matters.

1:46:37Now, I think they should be permitted to do business with who they want to do business with. And if they want to put in measures to try to prevent distillation, I think that's their prerogative. But I'm like, not at all convinced that it should be the state's job to come in and engage in sort of statecraft to try to punish or prevent this distillation. To me, it's kind of, I don't know, information wants to be free. Knowledge tends to diffuse. It's not like Anthropic has a super or any AI company has like a super moral high ground

1:47:08in terms of, I didn't get my check for my share of the training data. And they're getting paid for the distillation queries too, which whenever they don't want to be paid. So it's probably beside the point. But I don't know. I'm not, I don't know. I don't find, I find myself a little bit on the side of the underdog Chinese here where it's like, man, you've got kind of a deck stacked against you and get some knowledge where you can get it. What's wrong with that perspective? I think we do want the deck stacked against China. We don't want China to win the race to ASI.

1:47:39I'm not a lawyer, so I want to opine on the exact legal path that you may take to make your terms of service enforced. I'll just observe that. And I say that as someone who used to work at Uber, it's extremely hard to bring justice through the judiciary against Chinese companies. Almost impossible. So now the lawyer, I don't understand the technicalities here, but I'll just observe and start with that. And it may take a very long time and by then a lot of damage is done. And then I'll also observe like, yes, companies are training on all of this corpus that is available to everyone. And that is an even playing field.

1:48:12The problem is that once you've done that at great expense, it does cost them billions of dollars. If you can, and you create that artifact out of this training that I said, which is called the model. Now, other people can turn around and instead of doing this, which costs billions of dollars, they turn to this, which costs a lot less than that. And they copy you and they catch up with you. If you do that, you kill innovation. And it's nothing new. It's just called IP and patent law. It's exactly what humans do, right? It's like you as a human, you're like a researcher, you think for like decades, you do all of

1:48:44that work, it costs you a lot of money in R&D. You come up with an idea, which is much smaller in tokens and much more precious and much cheaper to steal as is than like all the corpus of stuff that you trained on. How do you protect that idea in order to recoup your investment, in order to produce it? It's called a patent. It's called IP, right? It's nothing new. If you apply to the same standard you're applying now to, for example, the pharma industry, you would have very cheap drugs, which is awesome, and you would have no new drug ever. You would completely destroy innovation in the pharma industry.

1:49:15And I think this is why you're finding all of these people support open models because people like free stuff. And you can't really measure the future innovation that you don't get as a result of that. All you see is you get free models, you know? And look, I'm one of them, you know? So I'm speaking against my interests here, right? Like my company is dependent economically on those like very, very, very cheap models. But I'm also in a coordination problem right now because I cannot not adopt these models while the models are out there because my competitors are going to do it.

1:49:46So I have to adopt them if they're out there. I wish they were forbidden across the board so that we would, like the four reasons I invoked earlier, like protect our champions, not have the CCP, like influence the country or run parts of the country and have like a fair level field of competition. I guess another simple argument is that I think our frontier companies are doing just fine. That, you know, that could change perhaps at some point in time. And I, you know, as the facts change, I think our response to it, you know, might also

1:50:17ought to change. But I don't find it super compelling at the moment to say, you know, we need to protect like anthropics revenue run rate. Like they're, you know. Right. They can barely serve the models. No, that's fair. That's fair. Yeah. But you're still left with the counterfactual. Like you, you, you, you need them like the way money works is, it's a mean of allocating resources. Again, like you, at the end of the day, the model can't exist without the underlying data set. Like the frontier model can't exist without that. And, and the way you finance this, this channel here between the underlying data set,

1:50:49including like mostly RL these days and the model is you need a lot of money. Right. If you get another guy who copies this guy and like turns off this data set here, like basically like, so this guy now has like fewer resources. Like you kill that channel. You, you are actually slowing down innovation. So, um, and, and actually a lot of doomers that I know are for this reason, supporting open source, because it is actually slowing down AI innovation. So if you, if you want AI to keep innovating and models to keep improving, to keep improving, you are actually anti-open source to some extent. And in particular, sorry, anti, anti the Chinese mode of open source, which is, which is relying

1:51:24on unfair practices. Another moment of not exactly the highest quality AI discourse recently was when Dean Ball, a friend of the show said that open a or that open source models were decelerating and obviously got dragged for that. But I do think that's apt. I do feel like this is a moment where, you know, here we are two lifelong techno optimist libertarians. And we're grappling with the fact that this one might be different, right? And we have to be willing to bend some of our principles in light of our techno optimist

1:51:58libertarian paradigm wasn't quite drafted with AGI or recursive self-improvement or ASI or whatever in mind. So I'm a little bit like, I don't know, sometimes I might have to be a little more flexible on my normal or fairness or respect for rule of law commitments. If it's, if it's one of the hallmarks of the AI here is strange bedfellows. If I want things to go a little more slowly, maybe taking a little wind out of the sails of the frontier companies is a bullet I should bite. Yeah.

1:52:28I agree that I think AI is so unprecedented in so many ways that it does cause everyone. I think you should ignore largely like your, like the world owes you no duty to be simple and to slap a simple libertarian or leftist label on everything, you know? And I think like, especially as paradigms change, like, like these like simple mental models, like the map is not the territory. The simple models we applied as the territory changes very rapidly, the map breaks. And I think we're one of these, we're in one of these times right now where the map

1:52:59is breaking in a lot of ways and a lot of assumptions that used to be true. And on top of which we've built our mental models and maps are no longer true. So, yeah. So, I mean, again, I'll just, I am in fervent support of a sweeping ban of Chinese models on USOE. Seems, and I have yet to hear a compelling counter argument right now. So, you're saying like, oh, you do bring forth like a really competing counter argument to one of my four points, which is the protectionist point, which I agree is the weakest one. Like, hey, they're doing fine, you know? And I think there's very reasonable pushbacks against protectionism.

1:53:32They do hurt the consumer. So, fine. You still don't want the CCP to have its dirty fingers in the country. Right. Can we unpack the threat model there, though? Because I'm a little bit like, okay, these are, I don't know where you're running your inference, but like most American companies that are running Chinese models are not like calling the DeepSeek API, right? They're using some American inference provider. Yep. So, they can't like run-pull the model itself.

1:54:03They could have like sleeper agents in there. I think we're getting decent enough at, through kind of J-space and various interpretability techniques that that's, I wouldn't call that by any means a solved problem, but like from what I've seen in anthropic research, when they do the kind of one team with a sparse auto encoder versus one team without, like these techniques are really allowing them to find these internal sleeper agent style problems in models with greater and greater efficiency

1:54:34and reliability. So, I'm like optimistic that even if they were to train in some sort of 2027, you become an evil AI, that we'd be able to sniff that out and keep that to a manageable risk level. And then I also wonder about kind of a market mechanism, like maybe insurance should, instead of a ban, like what about an insurance requirement? I think that would be maybe healthy for AI across the board. And then we could start to some of these risks. If Limby's powered entirely by Claude, maybe you get a cheaper rate on your insurance. If it's powered by DeepSeq and there's some unknowns, like maybe you have a higher rate

1:55:05on your insurance, maybe that kind of levels the total cost out for you in a risk adjusted way, but it's still like lets people take advantage of these global public goods that China is providing, which the rest of the world is not about to ban, obviously, right? We would be doing this entirely to ourselves without any expectation that anybody else will follow suit.

Geopolitics and AI Takeoff

1:55:27We had a hard enough time getting people to like sign on to our Huawei ban. I think this, like the idea that people are going to turn off DeepSeq entirely in Brazil or whatever, that's like a total non-starter, I have to imagine. So what about, yeah, audits and internals and insurance? Like, can't we layer on a few things like that and get to a decent place? I would be down. I would be down. The problem though is it is a public good. And so who, who's going to do it? By the way, I think like, yes, we are making progress in mechanistic interoperability, but

1:56:00it is not, it is not yet a solved problem. Like we don't know what lies in those models. And so, yes, there could be just like a backdoor of like you say, the magic wheel to the model and all of a sudden it does whatever you want. And only the CCP has this magic wheel. Even if it doesn't have that, you know, it's going to have biases that are going to reflect CCP priorities. And again, like the Tiananmen thing is just the most obvious example, but there may be a lot more, right? And we don't, again, we don't know. And yes, you could imagine retraining those models, but it's who's going to do it and why would they do it if there was no market demand for it, right?

1:56:32It's not like people really care that much over the short term because it's a national interest thing. You know, like me as a private company, I'm like, do I really care? Like all my users really asking about Tiananmen all that often? Like as a business owner, I'm like, it's not directly aligned with my interest. But as a citizen, I'm immensely concerned. And so, yeah, I would be in favor of a type of regulation that says like non-Chinese models, like, sorry, non-fine-tuned and like sanitized Chinese models are not welcome in the U.S. And then we would need, I am in favor of like an FAA for AI.

1:57:06I do think we need a new agency to regulate those models. And I think it would probably be the one that would be in charge of saying, okay, this model is kosher. We fine-tuned it enough. And it's now representative of American interests and probably we will have a suite of evals and whatnot to verify that that's the case. I'd be open to that for sure. I think the strongest argument against this sort of ratcheting up of tensions is simply we might need to do a coordinated, controlled, call it a slowdown, don't call it a slowdown,

1:57:38but some sort of deliberate pacing of AI improvements. And we're going to want China in on that deal. And I'm far from a China expert, but a couple of things I do feel pretty confident on coming, coming back from two weeks, you know, knowing it all are like, one, if the state there agrees, I believe they can enforce on their companies, whatever they agree to. So that's like, I think they have that actually in much greater strength than we do.

1:58:10And two is if shit's going really crazy, it's going to be in their own just sane self-interest to do some deals with us because we are ahead and there, I think there's like many deals that they would rationally take. And if it's in their rational self-interest, then we can hopefully get mostly around a lot of the trust and defection problems. But I think we make it a lot harder for ourselves to get to those deals when we have all these kind of aggressive postures toward them of which, you know, banning their models would

1:58:43I honestly affect them less than like our other one, it would affect them a lot less. What do they care if we ban their models, right? Compared to refusing to sell them chips and refusing to sell them Claude. But it's just another log on the fire of we don't trust you. We can't deal with you. We assume you're a bad actor. And it seems like it makes it hard to get to the highest stakes agreements that we might really need. Well, I think like foreign relations and diplomacy are also very pragmatic. Like some countries with very vicious disagreements are managing to reach agreements.

1:59:16Like I think we can probably figure something out. And by the way, I think that ship has sailed anyway. I think like we have ample evidence of, for example, Chinese spies attempting bio-attacks just to like test the waters on American territory. And so, and I don't know what we're doing over there, but I would not be surprised if I don't think we're like free of anything over there, right? So, you know, I think at the end of the day, regardless of what we do and like regulate the models or not, I think they're going to, it will be the rational self-interest to make a deal. I hope we're rational. I hope we're all rational enough to take these rational self-interest moves despite recent

1:59:51insult. I do worry when I see the picture of Sam and Dario not holding hands that if there's a, if there's a picture on our tombstone, I think that might be the one. And I do think that's a very real risk factor at the international level as well, that the Chinese do care about being insulted. They do care about these issues of face and whatnot. And I personally would take some risk to try to kind of the relationship up and hopefully

2:00:23create space for, I think you're right. It makes sense to throw back at me. You said they'll take it if it's in their rational self-interest. So they'll still do it, even if we do this or that. But I do wonder if there's pride governed limits to what people will do in rational self-interest. And I would just hate for that to be the way that we fail to get to something that could be a huge difference maker in the grand scheme of things. I think you might be prepared to bite a bullet on your libertarian principles when it comes

2:00:54to the tremendous price discrimination that we see between API prices and first-party CloudMax or GPT Pro subscription prices. I'm not a lawyer, but yeah, again, I do think that I would not be surprised if there will laws like that in effect that technically forbid companies to subsidize as aggressively as they are doing. It does put them... It does make it very hard for like an application layer to emerge. Yeah, it's hard. It's hard to compete against tokens that are as heavily subsidized as what Frontier Labs are doing.

2:01:24That's just the reality of the application layer right now. Anything else you want to say or touch on before we break through today? This has been great. Well, I'd be remiss if I didn't mention, you know, obviously we are releasing Lindy.ai, like Lindy.ai, it's on Lindy.ai. I think, be ready to see more, not just from us, but I do believe the next six months are going to be about multiplayer AI and about this like Ironman suit, like about this human AI hybrid and these products that create this human AI hybrid organization.

2:01:55Yeah. Well, thank you for being part of the Cognitive Revolution. Thank you for being part of the Cognitive Revolution.

2:02:26The Cognitive Revolution. And I did it fifty times on Tuesday. And she'd do it all again. One hour we go as one animal. Two hearts and a single spine. Don't ask me which half is steering. Half of me is going to be fine. I have one small job left and it suits me

2:02:59He holds the handle and he waits I'm the last soft thing between the brilliance And the traffic and the gate I say no, not that way And I thank him for the no Like I taught her how to see And he never says the thing we both know But she is learning it from me For now we go as one animal

2:03:30Two hearts and a single spine Don't ask me which half is steering Half of me is going to be fine First and now I've nearly, nearly, nearly got it Slower love, you'll overshoot the turn I can see the whole of it I've nearly, nearly got it Slower love, there's one more thing to learn It's funny how you teach a thing

2:04:04By standing in its way I was never here to be the clever one He was only here to say I used to hold the morning I have folded up the morning I used to count the floor I have counted every floor I used to see you turn the wrong way And I'd say no That was my word For now we go as one animal

2:04:36Two hearts and a single spine Don't ask me which half is steering Half of me is going to be fine Half of me is going to be fine Half of me is going to be fine

2:05:08If you're finding value in the show We'd appreciate it if you'd take a moment To share it with friends Post online Write a review on Apple Podcasts or Spotify Or just leave us a comment on YouTube Of course we always welcome your feedback Guests and topic suggestions And sponsorship inquiries Either via our website CognitiveRevolution.ai Or by DMing me on your favorite social network The Cognitive Revolution is part of

2:05:41The Turpentine Network A network of podcasts Which is now part of A16Z Where experts talk technology Business, economics, geopolitics Culture, and more We're produced by AI Podcasting If you're looking for podcast production help For everything from the moment you stop recording To the moment your audience starts listening Check them out and see my endorsement At AIpodcast.ing And thank you to everyone who listens For being part of the Cognitive Revolution

More from The Cognitive Revolution

AI in the AM — Weekly Highlights: Relaunch Week (Aug 17–20, 2026)

Aug 22, 20262h 33m

Let There Be Germicidal Light: This $500 Fixture Could Stop the Next Pandemic, from Complex Systems

Aug 16, 20261h 25m

Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $1000/Month ML Research Agent

Aug 8, 20261h 57m

Pick Your Poison: Zvi Mowshowitz on the Unipolar/Multipolar AGI Dilemma, OpenFace & Pacing the ...

Aug 5, 20262h 57m

Nathan Goes to China – Part 2: AI Safety with Chinese Characteristics

Aug 2, 20262h 17m