Steadcast
The AI Daily Brief cover art
The AI Daily Brief

Why AI Washing Won’t Work Much Longer

August 4, 202624 min · 4,983 words

Show notes

Corporate AI has spent years rewarding flashy announcements, dubious layoffs, and shallow use cases. But the arrival of powerful open models—and a much more sophisticated conversation about routing, customization, costs, and organizational redesign—may finally make AI washing harder to sustain.

Highlighted moments

caps, which, by the way, most organizations haven't even gotten close to that level yet, are not the same as cuts.
1:03
Every organization in the world is awakening to the risks of handing the creators of language models the keys to their institutions, of letting these models loose within their homes.
1:53
We have people trying to drug addict us to a future they believe they control.
3:02
current AI revenues, quote, don't sustain the capital expenditures we're making so far, creating a, quote, danger we could hit an AI air pocket such that the expenditures happen but the revenues don't show up.
5:04

Transcript

0:00Today on the AI Daily Brief, what a new Chinese open-wait model release has to do with big shifts in enterprise AI thinking, and before that on the headlines, Palantir and the March to AI Sovereignty. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.

0:20All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Airtable, Section, and Blitzy. To get an ad-free version of the show, go to patreon.com slash aidailybrief, or you can subscribe on Apple Podcasts. And to learn more about sponsoring the show, send us a note at sponsors at aidailybrief.ai. This week, tech earnings continue with Palantir, registering another monster quarter. Quarterly revenue came in at $1.94 billion, up 93% from this time last year and beating expectations.

0:52Commercial sales were up 149% year-over-year, up from 133% growth rate in Q1, demonstrating that enterprise AI demand is still booming, despite flashy headlines of token budget cuts. Indeed, it turns out that caps, which, by the way, most organizations haven't even gotten close to that level yet, are not the same as cuts. Palantir also managed to expand profit margins, with net income reaching a billion dollars for the quarter and growing at a 225% annual pace. CEO Alex Karp described the quarter as otherworldly, and used the earnings as a chance to proclaim

1:22the message that he has been getting increasingly loud about on his bully pulpit. He basically painted Palantir's results as an expression of the demand for AI sovereignty. He said, Palantir is the only company that has demonstrated it can transform tokens into actual economic value. Our customers trust us to provide them with maximal control over their operations, data, and decisions. Palantir hiked annual forecasts, sending the stock surging by 10% in after-hours trading, with maximal control over their operations, data, and decisions. This was the message that he echoed in his shareholder letter as well.

1:52Karp wrote, Every organization in the world is awakening to the risks of handing the creators of language models the keys to their institutions, of letting these models loose within their homes. The demand from our partners is clear. It is for control over data, the prompts that the models ingest, and more fundamentally, the organizational and business intelligence, their alpha, that the language labs are not only ready and willing, but structurally designed to capture from their customers. Later in the letter, he continued, We do not get paid for clicks or tokens or chats. The gamification of the most significant development in modern economic history seems to us misplaced.

2:23The usage of a platform may hint at its value, but is by no means dispositive. And many are now finding out that consumption and usage alone have little or nothing to do with the production of results. Just to add a little fire to all of it, he then says, There are Marxist overtones and undertones to our business. Others, including many of those building large language models, intend, knowingly or otherwise, to capture the means of production of their purported partners. The limitations and faults of the token industrial complex, which has threatened to overtake and dominate the world economy, have increasingly been exposed. We have always declined and will continue to decline, entering into a parasitic relationship

2:55with our partners. In a follow-up interview with CNBC, Cart made it clear exactly who he takes issue with. He said, We have people trying to drug addict us to a future they believe they control. Now, I've spent a lot of time with Dario and the Effective Altruism crew. They want to tell you we have to march into a future where we own nothing, where your businesses aren't profitable, where none of us have jobs, and where our adversaries win. Now, of course, these are not the type of comments that everyone is going to take at face value, and they generated a wide range of opinions. Summing up the most optimistic version of the take, Amit is investing rights.

3:26The age of AI is not about valuations, but about empowering workers, enabling agency, and growing GDP. Now, speaking of companies that are taking issue with the frontier labs, Apple right now is of course embroiled in a lawsuit with OpenAI, claiming that via an employee who left Apple to join OpenAI, the LLM lab, stole Apple trade secrets. OpenAI is biting back. In a new letter they wrote, Apple is getting this wrong. Apple, they write, is one of the greatest companies of all time, and built a reputation for obsessing over the smallest details. This careless, aggressive, and oddly personal lawsuit sadly doesn't live up to that reputation.

4:00Now, the big jaw-drop line was this one. Apple had claimed that they contacted OpenAI in February and that we didn't respond. They now admit that their outside lawyers emailed the wrong person after confusing two Asian last names, only after we brought this to their attention. Which honestly, if true, brought up some big questions with the lawsuit for many folks watching this from the outside. Now, at this stage, this is a little bit more psychodrama than we normally get into on this show. It'll be worth paying attention to if it actually shifts the plans that OpenAI can make with hardware. For now, that was just such a wild mistake to make that I had to capture the zeitgeist of

4:33the AIX community by sharing it with you here. Moving on to some other comments that are getting attention, Google DeepMind's chief strategy officer has reframed sky-high CapEx as a down payment on recursive self-improvement. During a panel at UC Berkeley, Jesjeet Sikhan said that RSI was a key part of the investment thesis. Now, if you've been listening closely, this earnings season has been all about how the ROI of AI lines up with endlessly ramping CapEx. Google in particular was punished for their AI income failing to live up with AI spending on the short term as free cash flow flipped negative.

5:04This particular Google executive acknowledged that current AI revenues, quote, don't sustain the capital expenditures we're making so far, creating a, quote, danger we could hit an AI air pocket such that the expenditures happen but the revenues don't show up. However, he argues that the CapEx build-out is not about near-term revenue, but instead the, quote, biggest scientific bet civilization has ever made. Next up, an interesting good news story in cybersecurity as Claude helps researchers uncover a decades-old vulnerability in DNA evidence databases used in criminal prosecution. At labs across the country, forensic scientists keep DNA samples as digital files to create

5:36a searchable database. A group of researchers have found that they could alter the files using code written by Claude in a process that takes around 45 minutes. The core issue is that this database software was created in 1995 and includes none of the tamper-evident protections of modern software. One of the researchers, forensic scientist and New Haven University professor Laura Gadoch-Kom said, Effectively, what we have are data files that are legitimately referred to as the gold standard of forensic science that lack the same level of tamper-evident markings that we require for a paper bag. Now, the ramifications are concerning.

6:07A bad actor with rudimentary knowledge of how DNA testing works and access to the database could corrupt files or even modify evidence to frame an innocent person. None of the labs around the nation have reported any kind of this evidence tampering, but they also haven't been able to figure out a way to detect if tampering has occurred. As part of their disclosure, researchers noted that some of the encryption still being used relies on an encryption key that's been available on the internet for years. The creator of the database software said that they have been working closely with the U.S. Cybersecurity and Infrastructure Agency on mitigations and have pushed a software update to implement digital signatures.

6:37Sarah Chu, the director of policy and reform at the Perlmutter Center for Legal Justice, said that the research highlights how badly behind the forensic sector is. She said, Lessons learned from other industries haven't been imported into forensic science in a serious way. We've been behind the ball for so long. That kind of all rolls downhill into this incident. Now, for the AI folks, this highlights what many have been saying about how critical it is to be very conscientious about guardrails on frontier models. This is an example where defense-focused researchers were able to find serious vulnerabilities in ancient systems still being used in a very high-stakes field.

7:11It is absolutely the case that there are hundreds or even thousands of outdated systems like this running critical infrastructure across society. Cost is a huge part, perhaps the biggest part, of why these systems haven't been overhauled or even properly tested for phone or abilities. While it's obviously not perfect, the ability to do even rudimentary testing using AI on a modest budget could be an absolute game-changer for this sort of software. TLDR, while yes, it is scary that cybercriminals have a new suite of powerful tools, AI is and remains a massive upgrade for cyber defenders as well. Lastly today, one that we will pick up on, I'm sure, tomorrow.

7:43The White House is hosting a set of AI companies to discuss the administration's voluntary review framework for frontier models. I am very interested to see what comes out of that, but let's wait until we have some more solid reporting. For now, that's going to do it for today's headlines. Next up, the main episode. One of the most important AI questions right now isn't who's using AI, it's who's using it well. KPMG and the University of Texas at Austin just analyzed 1.4 million real workplace AI interactions and found something surprising.

8:15The highest impact users aren't better prompt engineers, they treat AI like a reasoning partner. They frame problems, guide thinking, iterate, and push for better answers. And the good news? These behaviors are teachable at scale. If you're trying to move from AI access to real capability, KPMG's research on sophisticated AI collaboration is worth your time. Learn more at kpmg.com slash US slash sophisticated. That's kpmg.com slash US slash sophisticated. This episode of the AI Daily Brief is brought to you by HyperAgent, where you run fleets of

8:48agents your team can manage together. New users get $1,000 in inference. Forget local agents and chat workflows waiting on your laptop to be prompted. HyperAgent deploys always-on agents in the cloud, doing real work across the tools your team already uses. Marketing's agent turns competitor moves into landing pages. Sales's agent enriches leads, drafts emails, and updates the CRM. Ops agent chases the paperwork and tracks the budget. Every agent has access to shared context and follows your rules about scope and approvals. It's time you add agents that feel like teammates. Hire yours at HyperAgent, built by the team at Airtable.

9:19Claim your $1,000 in inference at hyperagent.com slash AI Daily Brief. Here's a harsh truth. Your company is probably spending thousands or millions of dollars on AI tools that are being massively underutilized. Half of companies have AI tools, but only 12% use them for business value. Most employees are still using AI to summarize meeting notes. If you're the one responsible for AI adoption at your company, you need Section. Section is a platform that helps you manage AI transformation across your entire organization. It coaches employees on real use cases, tracks who's using AI for business impact,

9:52and shows you exactly where AI is and isn't creating value. The result? You go from rolling out tools to driving measurable AI value. Your employees move from meeting summaries to solving actual business problems. And you can prove the ROI. Stop guessing if your AI investment is working. Check out Section at SectionAI.com. That's S-E-C-T-I-O-N-A-I dot com. Blitzy deeply understands your code base before it writes code. Here's the first place that pays off. Security in the age of AI.

10:22Vulnerabilities don't live in isolation. They live buried inside millions of lines of interconnected code, where patching one thing quietly breaks three others. That's why surface-level scans fail. Blitzy starts from its knowledge graph of your entire application, identifies and surfaces CVEs across the full estate, proactively recommends patches, and can execute the PR. Each fix is grounded in how your systems connect and validate so nothing new breaks. And the knowledge graph dynamically updates, keeping you ahead of an ever-accelerating threat landscape. One Blitzy customer resolved 21 active CVEs across six core microservices in four days.

10:52Zero compile errors, every validation scan clean, months of planned work fixed in less than a week. Security remediation grounded in real architectural context at the speed of compute. Harden your code base at Blitzy.com. That's B-L-I-T-Z-Y dot com.

11:11Welcome back to the AI Daily Brief. Today we are talking about the latest model release, which is Quen 3.8 Max. But the context we're putting it in is a little bit different than normal. As you know, I am just back from KPMG's Tech and Innovation Symposium last week, and my biggest takeaway was about just how radically the conversation had shifted over the last year. When I was there for their 2025 edition, we were still genuinely talking about things like the percentage of organizations that had one, two, or three AI use cases. This year we were talking about how organizations are handling the complex governance issues that arise from everyone having powerful coding tools,

11:47and how people are dealing with cost provisioning issues in different models across different parts of their organization, and whether they should be looking into fine-tuning open-weights models and have an open-weights models policy. Point being that the level of sophistication around AI, and more importantly, the quality of the questions that enterprises are asking, has increased dramatically. Meaning that my guess is, for the first time, some meaningful number of enterprise buyers, not just early adopter developer types, but actual enterprise IT type folks, are actively paying attention when things like the new Quen 3.8 Max model comes out.

12:19So first let's talk about the model. Like Kimi K3, it is a large model, this one coming in at 2.4 trillion parameters. Indeed, Leighton Space writes that the model would have been the top open model in the world, but for the recent Kimi K3 release. The official Alibaba Quen account calls it a new bar for coding and co-work, and points to examples of its work on autonomous coding, 10 plus days of self-evolving development from empty folder to production without hand-holding, sharing a complete project trace on their GitHub. They also laud production quality deliverables across hundreds of professions,

12:50systems-level autonomous planning with closed-loop adaptive learning, and native multimodal intelligence. Vision they write isn't just input, it's a continuous feedback loop for planning, execution, and self-correction. Now, at least on the self-reported benchmarks, there are some pretty impressive numbers. Alibaba reports a terminal bench score that puts them between Fable 5 and GPT-5-6-SOL, a co-work bench score between SOL and Fable 5, visual reasoning and research reproduction that are totally state-of-the-art, and so on and so forth. They also report state-of-the-art on OS World Verified,

13:20which is agentic computer use, which, if that result is accurate, is one of the more significant when it comes to its applicability as an agentic work tool. Now, interestingly, a lot of the discourse online was not about the model itself, but about the video ad that they released it with. The ad, which is an incredibly simple concept, very well executed, is a laptop in the foreground doing all sorts of different types of work. There's coding work, science work, different types of knowledge work. And in the background, you see the humans presumably paired with that laptop off doing more of the things that make them, well, enjoy their lives.

13:53There's a guy fishing at one point, another man playing tennis, a woman rock climbing, and another woman reading. In other words, without using basically any words, Quen is saying that the value to your life as an individual of these incredibly powerful and advanced models is to give you more time to do the things that you actually want to do. Jason Calacanis from All In tweeted, Not only is China making increasingly competitive models, they're also making world-positive AI marketing. The computers are going to do our jobs for us, and we're all going to the beach. Nansen's Alex Venevic writes,

14:24What is this marketing? No graveyards? No entry-level jobs disappearing? Vittorio also points out that it seems like perhaps they are trolling us a little bit, with using one of the examples being verifying protein sources given these strict guardrails on Fable. And yet, one of the big deals about this announcement is that it represents a return for Alibaba and the Quen series back to OpenWeight's models. Last year, there was a lot of personnel shifting around Quen, and for their largest Max series models, they moved from open to closed. In fact, it brought up the question more broadly

14:54of whether this was something that was just going to be more commonplace for Chinese models as they got more advanced and got closer to the frontier. Well, now they are back, and as Yuchen Jin points out, this marks the first time Quen will open-source the weights of a Quen Max-class model. The full weights will be released next week. Strawberry Labs' Chirag Asarpota writes, This is so huge for open models. Chinese labs are clearly taking advantage of Anthropik's recent PR mess and are going all-in on open weights. The message is loud. They want Chinese open models to dominate globally. Yet, maybe the even bigger deal is the price.

15:26One of the things that I've talked about on this show a lot recently is that many folks, particularly around the policy sector, still have this idea that every Chinese model is cents on the dollar relative to their state-of-the-art American competitors. That hasn't been the case for a while. In fact, to give you one example, while Kimi K3 is cheaper than Opus, it's only cheaper by about 40%. $15 per million output tokens versus $25 per million output tokens. Quen, however, takes it down significantly. Quen is being priced at $2 per million input and $6 per million output tokens,

15:56making it a little more than a third of the price of Kimi and a fifth of the price of Opus. Now, that's still not the pennies on the dollar that some people assume, but it is a much more significant decrease than, for example, Kimi K3. Now, given all of this, it's not surprising that a lot of the first impressions were excited. Alex Volkov from the Thursday AI podcast writes, Alibaba is back? Quen 3.8 Max is about to get open-weighted with 2.4 trillion parameters. This model comes very close to K3 running projects for 16 days, but it is worth being at least a little cautious at this early stage.

16:27When it comes to independent benchmarks, Artificial Analysis did publish Quen 3.8 Max's score of 53, which puts them four points behind Kimi K3 and even a point behind Grok 4.5. Now, adding some intrigue to the whole thing, Artificial Analysis quickly took down those scores without any explanation, so we'll have to see what was going on with that. On some of Ethan Mollick's tests, he writes, Basic impression after a bunch of experiments is that it is a solid model, but not Kimi K3 level in my experience so far. Datum writes, Quen 3.8 Max is unusable, saying that they tested it across coding, design planning, agent orchestration,

16:58and multiple harnesses, and running it against Kimi K3, Grok 4.5, GLM 5.2, GPT Luna, and Opus 5, with Quen coming in last every time. He says it's extremely slow, unstable, and burns through usage like crazy. It often fails on the first attempt, then the second, and sometimes the third. Pavel Hirin also tested it, on his bug bench test, which has two real repos with 105 hidden bugs, and found that it found 19 out of 105 bugs. Kimi K3 and Opus 5 were both ahead, fixing 21, GPT 5.6 Sol led at 42. And while he noted that Quen did fix one bug

17:29that none of the other 10 models he tested found, that, quote, just getting it to run took five attempts. One coding agent used up the five-hour subscription quota in half an hour. The pay-as-you-go key exhausted its free tier, then refused to bill until I opened a different console. All in all, the cost was about $31, leading Pavel to conclude, let me save you some money and time. GPT 5.6 Luna ran the same benchmark for $1.80 and fixed $33. Grok 4.5 judged $16 in 25 minutes. So, first evidence suggests some good reason to be skeptical about the published benchmark numbers. But ultimately,

17:59why it's still an exciting release to people is the fact that it's coming in open weights, which means they're going to be able to get their hands on it in a totally different type of way. Which gets us to my contention from the beginning of the show, that despite this being the type of model release that in the past would only be something that was really paid attention to by developers and early adopters, increasingly, there are going to be folks inside big mainstream corporations and enterprises who are paying attention to this as well. Now, also yesterday, the New York Times published an opinion piece from the former chief information officer for Lululemon, Julie Averill. The piece, which remember is titled

18:30not by the author of the piece, but by the headline writers at the New York Times, was, I helped run Lululemon, companies need to stop kidding themselves about AI. Now, to some extent, this is the platonic archetype right now of the average essay about enterprise AI. The first part is all about how common it is for corporations to want to use AI for PR value rather than actually integrating it deeply into strategy. Julie coins a term AI wishing, which she defines as the belief by company leaders that AI is magic, that you can wave its wand towards a hard problem and skip the work of solving it.

19:00But don't mistake this as a critique of AI. It is not. It's a critique instead of the very real human processes that impact how AI gets integrated. She writes, don't mistake any of this for doubt about AI's potential. It's the most powerful technology I've seen and entire industries are already being remade by it. Pharmaceutical companies are rebuilding drug discovery around it. Banks run fraud and risk detection on it. Governments are treating the chips behind it as a matter of national security. And then in the most important line, she concludes, this kind of work doesn't happen in a quarter and believing that it can is the trap.

19:31And so in this context, not only did we get AI wishing, we got AI washing. As she puts it, the insidious cousin of AI wishing, that is when a company under pressure to immediately show results claims to be doing more with AI than it actually is. Now, the most negative and problematic version of this is the AI layoff. Quote, a company proclaims it needs fewer people because AI made its operations more efficient. Often the efficiency doesn't exist yet. The cut is really about freeing up cash, sometimes to spend more on AI. In May, she continues, U.S. employers announced 97,000 job cuts

20:02and according to one research firm, companies blamed 40% of them on AI. A separate survey found that around one-third of hiring managers who had cut a role because of AI had already rehired for the same or a similar one. Of course they did. They eliminated positions before redesigning the work, so the work shifted onto the people who remained. Then companies quietly hired back some of the capability they claimed AI had replaced. That cycle burns money and loses the talent and experience that was pushed out of the door, not to mention the trust of the people asked to stay. Now, I think at this point, my perspective on this is very well-trodden territory.

20:34I think the organizations that view AI strictly as an efficiency technology rather than as an opportunity technology might be able to eke out a few headline wins in the short term, but are ultimately going to be absolutely pummeled by the companies who understand that this is a redesign moment that opens up massive new opportunities, not a chance to make shareholders excited about cost cuts for Q3. In short, AI is going to take work to do well. It is going to involve organizational redesign, job role redesign, process redesign, and shortcuts are bound to fail.

21:05And so how did these two stories intersect? Quen 3.8 Max on the one hand, and this cautionary tale about AI wishing and washing in the enterprise on the other? Well, what it comes down to is that the change in the conversation that I was mentioning, as exemplified by the KPMG event, and the sophistication that it implicates, shows that more and more enterprise AI leaders are asking the right type of questions. For example, there is increasingly a discourse about whether OpenWeight's models can be part of a complete AI system that involves different types of models

21:36for different types of tasks. Despite headlines which suggest only blunt tools like cost caps, the reality in practice is that basically every enterprise AI leader that I interact with is digging deep to actually understand what their approach to, on the one hand, wanting more AI consumption, and on the other, needing to manage costs, should actually be. One part of that is discourse about OpenWeight models. A year ago, almost no company had any sort of actually expressed policy when it came to these models. Other than, lol, of course we're not going to use Chinese models.

22:07Now that conversation has really changed. Just yesterday, I saw a Forbes guest op-ed about why leaders should learn about OpenWeight models. And of course, regular listeners will know that it's not just op-eds. Increasingly, we have companies that are offering access not only to cheaper models, but to the ability to customize and fine-tune them based on your organization's unique needs and data. Thinking Machines Lab, the spinoff from former OpenAI CTO Mira Mirati, introduced a product to Tinker to do that late last year. And Microsoft, while obviously not using OpenModels, is building a frontier tuning service

22:38on top of their lower-cost MAI models that is explicitly trying to solve for this era of increasing AI complexity. Then, of course, there are the router companies. Every day, it seems, we have some new company introducing their latest router or news that indicates how valuable these companies have come, like reports that Stripe is about to buy OpenRouter for $10 billion. And while you might think that routers are just replacing the last thing as the buzzy corporate word, in my experience, that is just absolutely not the case. Those same enterprise AI leaders that I was talking about before are, yes, exploring routing solutions,

23:09but not in a way where they're just hoping to buy whatever Gardner says is the best vendor and call it a day. As a recent post on DigitalApply.com put it, AI cost optimization is now a discipline, not a hack. Companies are exploring what combination of third-party and internal solutions is going to be the right way to handle this, and whether or not a router is even the right solution. But again, the fact that those are the types of questions they're asking, I think is worthy of a lot of optimism. Now, I know, many listeners who are the AI leaders in their organizations will be shaking their heads, wishing that the situation

23:39I describe among enterprise AI leaders were the situation that they were dealing with inside their company. And I certainly don't want to minimize how steep the hills to climb are for many opportunity AI advocates who are trying to get their organizations to handle this the right way. But for basically the first time since ChatGPT launched, my observation is that the enterprise conventional wisdom around AI is getting more directionally correct. I think our attitudes around the types of things that go into AI wishing and washing are changing radically, and the PR value or board plaudits that people got before

24:10are going to stop, which helpfully will cut off the incentive loop to do AI the wrong way. If I am right, what we'll be left with is a situation where we can actually begin the exciting work of really redesigning around the capabilities of AI with all of the implications for both our personal and our professional lives. For now though, that's going to do it for today's AI Daily Brief. Appreciate you listening or watching as always, and until next time, peace! and I'll see you next time. Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye!

24:40Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye! Bye!

More from The AI Daily Brief

41 Stats That Tell the Story of AI Right Now

Aug 8, 202622 min

The Right Way to Worry About AI

Aug 7, 202628 min

Google’s AI Leadership Shakeup: Disaster or Exactly What It Needs?

Aug 6, 202633 min

Why the Data Center Fight Has Little to Do With AI

Aug 5, 202635 min

What Happens When AI Breakthroughs Outrun Human Understanding

Aug 3, 202628 min