Steadcast
Sharp Tech with Ben Thompson cover art
Sharp Tech with Ben Thompson

(Preview) Nvidia’s Answer to Capital Constraints, Google’s Attrition and Direction, Q&A on AI Writing, Vision Pro, Vibe Coding

August 14, 202626 min · 4,600 words

Show notes

Ben and Andrew begin with Nvidia’s announcements of a new funding model for AI infrastructure, including the differences and similarities with railroad expansion 150 years ago, why LLMs were a gift and curse to Nvidia’s business, the pressure on Nvidia coming from Google and Amazon, and the expanded blast radius as Nvidia works to mobilize third party funding.

Highlighted moments

the crazy thing is we sort of blew through debt in, like, nine months. Like, the amount of debt that was raised in the second half of last year and the first half of this year is in the hundreds of millions, probably soon to be approaching, like, a trillion dollars.
5:40
So sort of the classic example here is the pension fund. The pension fund is that people – you're paying into your pension over time.
6:52
what happens with the doctor is you're in school for a very long time, so you start making money relatively late. But once you make money, you usually make a fairly decent amount of money. And so it's like a catch-up plan where you can put way more money into retirement.
7:31
a big shift that has happened in the last couple generations has been a shift to water cooling. So you have, which requires entirely new kinds of data centers.
12:26

Transcript

Lakers ownership and mailbag shift

0:00Hello, and welcome to a free preview of Sharp Tech.

0:09Hello, and welcome back to another episode of Sharp Tech. I'm Andrew Sharp, and on the other line, Ben Thompson. Ben, how you doing? I'm doing okay, Andrew. I kind of feel a little bit in a funk. There's been some travel going on. It's kind of dreary outside. The brewers are terrible. We're trying to figure out what is causing what, but it's okay. We're here. But here we go. We'll make it happen. That's right. You know what I feel? I feel FOMO, because we were together in Wisconsin last week,

0:41and I feel like we could have put a call in to Mark Walter to see whether he was interested in selling the Lakers. Sounded like an asset he needed to move pretty quickly. Maybe we would have gotten lucky. Could have beaten Kushner to the punch. Alas, here we are, humble podcasters once again. Yeah, not to dive into a totally random aside, but Josh Kushner, not Jared, Josh Kushner is now one of the owners of the Los Angeles Lakers. Thrive Capital is in a lot of, you know, they're kind of on the cutting edge, like the new generation of VC companies are doing very well for themselves.

1:17I do think their largest holding is OpenAI. So maybe the real bubble concern now is, will anything happen to the Los Angeles Lakers if everything goes sideways? We'll have to keep an eye on it. God willing. That would be one benefit of the bubble bursting. So let's see what happens. For now, Ben, we're going to do all mail on this episode, and I'll tell you why. Because the last two episodes we've recorded, we've gotten so deep into various conversations that we've hit hardly any mail. So we'll try to remedy that today.

1:49And we'll start with an article you wrote this week. I need to not monologue so much. That's right. Be on your P's and Q's. Let's hit as many of these questions as we can.

Nvidia financing platform and AI infrastructure

1:59We'll see how it goes.

Nvidia financing platform and AI infrastructure

2:00NVIDIA announced partnerships this week with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR, a super team, to establish an independent financing platform or independent financing platforms designed to mobilize over $500 billion of third-party capital to support the build-out of AI infrastructure over time. That's NVIDIA's announcement. Part of that plan, as I understand it, involves shifting GPU depreciation risk away from traditional lenders in a bit of innovative financial engineering that I hope you can explain for me because I'm still a little confused what the plan is there.

2:41This is American greatness at play. We can invent a very expensive thing to spend money on and invent incredibly… Convoluted ways to pay for it? That's right. Great. God bless America. So, Andrew, in response to the article you wrote about this on Tuesday… Is this Andrew in Washington, D.C.? This is a different Andrew. Although, look, this Andrew also has lots of questions about what this actually entails. Andrew asks, can CUDA really generate earnings growth at a rate that outpaces depreciation of the GPUs?

3:14I'm being very unscientific about this, but it feels to me like there's an order of magnitude difference in there and not in CUDA's favor. The conclusion of your article on Tuesday carries echoes for me of the re-securitization of mortgage instruments into CDOs and credit default swaps that created the conditions for the subprime loans crisis and the global financial crisis. Do you see any parallels? In seeking to expand the breadth of available capital, is Wong creating the preconditions for a subsequent cascading collapse?

3:49Perhaps more interestingly, is there a feasible alternative or is this just the way the bubble expands? So, what do you think, Ben? Take it whatever direction you prefer. Well, let me take it whatever direction I prefer. We may look up an hour later and have not gotten too far through our mailbag.

4:08There's like a macro question about sort of AI infrastructure generally, and then there's like a micro question about NVIDIA specifically. And both are at play, and I think sort of what happened this week. So, at a very high level, and this is where people do reach for the railroad analogy. Like, I sort of reluctantly linked to it. It is such a good analogy this week, just because, like, everyone's sort of talking about it. So, I had to sort of cite, look, even Satya Nadella brought this up on his call.

4:39I'm not anything special here. Everyone's reading the same book. That was just sort of an acknowledgement that this is not – well, not just that. I mean, people have been talking about the railroad thing for a few years now. The book just came out this year, which is sort of fuel on the railroad analogy, fire. But where the railroad point is interesting is the fundamental issue that happened in 1873 is the world ran out of money. Like, which, you know, we've talked about the running out of compute. We've talked about running out of power. But the issue at hand here is what happens when you run out of money.

5:12And that sounds like an incredible thing to say, given how much money there is in the world. But we talked on this podcast even a year ago, like, not that long ago, about, well, you can't really call it a bubble when these companies are paying for this out of their free cash flow, right? Like, what's the spillover that we're worried about? What's the risk they're assuming in that scenario? That's right. It's like, once we start getting into debt, then we need to have a conversation. And the crazy thing is we sort of blew through debt in, like, nine months.

5:48Like, the amount of debt that was raised in the second half of last year and the first half of this year is in the hundreds of millions, probably soon to be approaching, like, a trillion dollars. And it raised by big companies with great balance sheets. And there's just a – or great sort of businesses, I should say. The balance sheets are getting – Money-printing businesses. So they're real businesses. And at some point, just like the – you run out of people willing to give you money.

6:20Totally. I mean, we talked about this a week ago in Madison where we were discussing, like, the lending environment will tighten and hyperscalers are going to have to get creative here, which is what we're seeing. Yeah, good job by us. Yeah. Yeah, good job by us, like, foreshadowing sort of this announcement. And so – but there's still lots of money out there. And there is money that traditionally goes to large, long-running infrastructure projects because that money itself is a long-term liability.

6:52So sort of the classic example here is the pension fund. The pension fund is that people – you're paying into your pension over time. Your employer is paying into your pension over time. I actually know a surprising amount about the mechanics of this because for one-person businesses, actually, like, pensions are the best possible retirement plan. Interesting. Okay. So the – because you can contribute a much greater amount than a traditional retirement plan before taxes and shift sort of your tax liability window, all these things that go into it.

7:26It's actually called the doctor plan because doctors are the most frequent users. Okay, yeah. Right, because what happens with the doctor is you're in school for a very long time, so you start making money relatively late. But once you make money, you usually make a fairly decent amount of money. And so it's like a catch-up plan where you can put way more money into retirement. All the money you weren't saving in your late 20s as you're toiling through school and residency. That's right. Okay. Yeah, it's kind of interesting because it's like a hangover from, like, old-school pension plans that aren't really in favor anymore.

7:57Anyhow, the – but those are – that's money that it has to be there in the long run. But it doesn't have to be paid out for quite a while. And so these are the sort of investments that want to go into, like, a toll road is the classic pension investment where you're putting a lot of money to work, but the predictability and understandability of the long-term payback is very clear. And it's going to pay back over a very long time. You're going to make a lot of money in the long run, but you have to have very patient capital because the – and, like, pensions in theory would have been a good match for, say, railroads.

8:34Because the problem with the railroad is you build it, and you might not really get your money back for 30 years. But – and this is a beautiful symmetry because I think this NVIDIA deal is symmetric with the Google equity issuance in which I wrote about Berkshire Hathaway and their shift from Seize Candy is a very high-margin business using that cash flow to get into BNSF railroads, which is a lower-margin business. But the absolute cash that's thrown off is very high.

9:07Right. And the analogy there is, you know, to what extent is Google making the same shift? And I think that's a very pertinent point to this NVIDIA thing, which we can sort of circle back around to. So you have this long, patient capital that is a very good alignment for long-running investments. And so what you had in this post by Jensen Huang is trying to make the case that actually you're all thinking about AI wrong.

9:37Now, it's not a short-term investment. It's actually a long-term investment. And, you know, if you put NVIDIA GPUs in, they run for a very long time and longer than you think. And we make them better with CUDA over time. And this is sort of building on the hyperscaler's argument, which is, look, the data shells, like the actual buildings, those are 30-year investments. And we're only buying GPUs right when we need them. So they're kind of aligned but a little not aligned in that regard.

10:08And so the case that's being made here is this is a long-run investment that deserves long-run capital. And if you zoom out, it's like, yeah, because all the short-run capital has been used up. And so that's sort of the case being made here. Now, is the case valid is sort of the next question. Ben and Madison wants to email and say, do we buy it?

GPU obsolescence and cooling challenges

10:34Hi, guys.

GPU obsolescence and cooling challenges

10:35Is this case valid? Well, in terms of the invalidity or potential invalidity, one of the concerns is that the GPUs that any of these companies, any of these infrastructure companies are buying from NVIDIA burn out before the patient capital can realize the upside. Or not just that, but NVIDIA comes out with new GPUs that makes your own GPUs obsolete. They're obsoleted. Exactly. Right. And so NVIDIA is trying to guard against that risk, correct, and try to allay some of those concerns? Well, NVIDIA is trying to do a lot of things.

11:05It's most importantly, preserve their competitive position and their margins. So it's kind of an interesting point, a talking point that Jensen Huang raised, and that was repeated on the Corweave earnings call. And I don't think it was an accident that these happened back to back. It's like Jensen Huang comes out and makes this case. Then Corweave comes out in their earnings and says, we have A100 chips that we are contracting out at a higher rate than before. And they're working great. Which I think is, I think that's absolutely believable.

11:36I've, like, it better be they said it in their earnings, right? It's a little more, and it makes sense. Like, compute is in such demand. There's already installed compute. Even if that compute is six years, I think the A100. Yeah, it's a previous generation for anybody who's not clear. But still being utilized. It's a little, right. So on the surface, it's a great case. It's like, look, people are out there saying GPUs only last two to three years. Actually, here's an example of a chip that is six years signing contracts right now.

12:06Those contracts are worth more than what the contracts were previously. They're actually increasing in value. And by the way, these are fully depreciated assets. They're, like, all the cash they're earning is pure profit. This is a long-term asset. And on the surface, it's a pretty good argument. There's just a couple problems. Okay. So problem number one, a big shift that has happened in the last couple generations has been a shift to water cooling. So you have, which requires entirely new kinds of data centers.

12:40So you can't just take in, like, your GB200s or the upcoming Vera Rubin and slot it into the old data center because they actually need water cooling and requires entirely new ways of putting servers together. Like, Facebook had this entire, like, they love the whole open source or open, like, data center thing. They had this concept. They could manufacture data centers very rapidly. And it was, like, this two-story sort of thing. I think it was two-story or whatever. But it all depended on passive cooling.

13:11So one question I have about the A100 case, and I think some H, I think the H generation might also be air-cooled, not water-cooled, or maybe it was half and half. The reason those are staying in place is because there's no replacement for them. You have data centers that are built for a particular assumption around cooling. New GPUs don't fit that assumption. So that data center is actually stranded. They're stuck with the A100s for life because of the way the data center was built.

13:41That's right. So on one hand, in a compute-scarce environment, absolutely they can keep selling them. But that is not, the A100 is not representative of what your expectations should be for GPUs going forward. And not necessarily dispositive as to the question in terms of whether this will still have utility in a market. It's not, this doesn't undo it. The fact of the matter is A100s are being sold for more than they were before because compute is scarce. But that gets to the next question. The available compute today is a function of decisions that were made pre-2024 and before.

14:21Because it takes like two years to bring these online. And of course, back then, the market was, for the record, freaking out about CapEx. And everyone who spent money on CapEx were right. Actually, no, they were wrong. They were wrong because they didn't spend enough on CapEx. They should have spent more in 2024. But so everyone has these signals. Everyone's talking about our demand exceeds our supply. But everyone's getting the same signal at the same time. So a reasonable concern from the market is, okay, if one company was getting this signal, then yes, they can invest appropriately.

15:01Basically, if 10 companies are getting this signal and they invest, do we overshoot? This is how the boom-bust cycle happens, is everyone's getting the same signal. That doesn't mean the signal is a 10X signal. It might be a 5X signal, but 10 companies invest, so you end up with double the capacity that you need. So that's another concern where the supply-demand environment right now is not necessarily representative of the supply-demand environment in two years or three years or in 30 years, however long you want these long-running assets to sort of be considered on.

15:35And that scenario would involve several companies bowing out of some of these infrastructure build-outs and the race to the frontier. Is that right? Well, what it would entail is your A100s are not going to be getting contracts if there's a gazillion GB200s available, right? They are available as a function of there not being compute. If there is an abundant amount of compute, the old compute is going to get retired very quickly. Yeah. Which, again, I'm not saying the argument being put forward is wrong.

16:08There's a lot of weight being hung on these A100 contracts that I'm just saying are not necessarily going to be long-term representative. Now, the pushback is we are so short on compute. We've barely scratched the surface of what these things can do. Actually, it's not just that in two years we're not going to have a surplus. We're still going to be in shortage. And, by the way, that might be true, like the extent to which the possibilities are barely being tapped as far as AI, particularly once we get to purely autonomous sort of functionality, like where you don't need to have the human in the loop, like there is – the bulk case is not insane.

16:52And it's not like a railroad. This is the distinction from the article or the railroad article. Like there's just no way to accelerate the revenue generation potential of a railroad. Like you've got to actually – It's closer to a toll road. That's right. You have to actually build it across a brutal terrain, which takes a very long time. Then you actually have to develop the land that you got for it. The land has to like build up like productive functions such that it starts using – like it's just in the physical world things are slow.

17:23Yeah. That's right. And even then, the number of – say you instantly had total saturation all over the railroad, you can only run so many trains. You have to build trains. So it's not – whereas AI, the scalability capability, if this stuff starts working, you have all the benefits of any digital good, right? What's the idea of software? You write software once, it's instantly, infinitely duplicatable and can be used sort of everywhere.

17:54And there's aspects of that to AI, particularly when you think about the concept of AI improving itself, AI writing its own programs, AI being like set loose on a company and creating agents on its own that figure out all the functions of it. Like again, none of that stuff quite works now, but it's working pretty well and it's accelerating unbelievably rapidly. I think we have an emailer in here saying that I'm a Luddite because I took too long to like vibe code, which I'll push back on in a little bit, but it speaks to the point that I'm sorry, like my six months was too slow for you.

18:25And it's kind of like a valid point, right? Yeah. So like the speed with which this is moving is a very real thing, but it's not a slam dunk case at all. And also, there's a real tension in bringing more supply to market will depress prices. Now, you can argue the demand's so high that the prices will still go up because demand will accelerate more than supply, but they're not going to go up as much as if you did bring more supply to market.

18:56And this is sort of, it's just a math function. Like the price depends on how much supply you have. It also depends how much demand you have. The bet is that demand is going to not just increase faster than supply, but increase even more such that prices, it doesn't matter how much NVIDIA produces, price is going to go up. And maybe that will be the case, but it's not. But the other question is just this timing question. This gets back to the amount of capital in the market. In the long run, you can't be funding stuff with debt forever.

19:29At some point, you need to actually make money and that money gets cycled back into buying new stuff. Yeah. And that I'm sure is going to happen. But like we talk about with stock picking, it's not enough to be right. It's about timing. Yeah. And the big question with this, these capital issues is I believe this stuff will pay for itself. The question is we'll pay for itself in time to avoid like an air pocket where we run out of money, a whole bunch of bag holders.

Nvidia's moat and developer ecosystem

20:01Sure. That's right. Well, and one other question before we move on, there's an element of this that read to me reading your article. As sort of a defensive move from NVIDIA as Google brings all this infrastructure online and you've got two dominant AI players. And as everybody becomes more cost sensitive, there's going to be an increasingly urgent push to get on TPUs as opposed to NVIDIA chips. So NVIDIA wants to facilitate building out with NVIDIA hardware and NVIDIA software.

20:37Does that make sense? Did I read that correctly? Yeah. So that gets the micro question, like the NVIDIA specific question. And this is a question, by the way, we've been talking about for a few years now. I think it was GTC 2024. So it was about 15 months after ChatGPT had come out when NVIDIA is truly like a stock of flame. A stride of the world. Yep. That was the GTC where Jensen Huang was at the San Jose convention or like arena, like the hockey arena.

21:07Yep. And it's like a rock star thing, right? I think he may have also signed someone's boobs at that GTC. That was, no, I think that was in Taiwan that that happened, but I might be wrong. Either way, same era, NVIDIA just owning the universe at that point. And I remember that was kind of a boring keynote in a way that NVIDIA's GTC keynotes were not boring. But because before ChatGPT, they knew they had this incredible computing capability, this sort of like this, you know, highly parallel.

21:41What are the things you can do for it? CUDA lets you program it more easily. And I wrote an update years ago where someone's like, how can NVIDIA announce all this stuff? Why can these keynotes be so cool? Especially because they do a key, like NVIDIA loves doing keynotes. They do keynotes like every six months or actually less. If you could see yes and things like that, they do Jensen's up on stage like every three to four months. That's true. How does he talk about so many new things? And the reason is that it's all the same thing. Everything is just parallel computing using CUDA. And they're just making all these libraries where they're just changing a few things, but they're all the same thing.

22:16And so – but the reason they were doing that is they're trying to – they're throwing everything against the wall. What's every possible application of parallel computing? Let's make a library for it and see if we can find and get the next – a market spinning around this beyond gaming and beyond Bitcoin mining. It's funny you say that because there was – before ChatGPT launched, I remember a GTC that you covered. I don't know whether I was working with Stratechery at that point, but it just seemed like Jensen was throwing all kinds of crazy ideas at the wall to see what sticks.

22:50And it was like, cool. It was like imagining the future. It's great that he's got all these ideas. I don't know how much any of this will actually be real, but he's clearly thinking about where we're going to be and how we're going to be computing 10 years from now. And then ChatGPT blows up maybe nine months later, and it's like, oh, okay, so this is it. And NVIDIA is in the catbird seat. So the weird thing about large language models is they were obviously incredible for NVIDIA. That's why their stock went to the moon.

23:21They have been on and off the most valuable company in the world. It was also very bad for NVIDIA. And the reason it was bad for NVIDIA is that the play with CUDA is to build a developer ecosystem on top of CUDA. But CUDA only works on NVIDIA GPUs. So you get CUDA for free. It's easier to use, and it's a tremendous investment. Like, NVIDIA almost went under trying to build CUDA at a time when no one understood what they were doing or why they were wasting money on it. And that's why, like, Jensen Wang will get bristly, particularly when people, like, question, like, their, like, rent-seeking or profit or whatever.

24:01It's like, no, they earned their spot. He was tasing the risks. Absolutely. Like, and it shouldn't be forgotten. Like, they have earned every dollar they've gotten through 25 years of taking massive risks. And the stock bottomed out several times along the way. It bottomed out in October 2022. Right. I wrote an article, like, three weeks before ChatGPT came out, like, tracing their bottoming out history and their search for what was next. NVIDIA in the Valley. I remember it well.

24:33NVIDIA in the Valley. So there is a, so go back to this GTC. So I wrote an article at the time called NVIDIA Waves and Motes. And what was interesting about that GTC was, number one, it was very boring. Like, all the cool stuff kind of got scrubbed out. Now, Jensen Wang has brought that stuff back. So the last few GTCs, like, he's more talking about other things. Now it comes across as, oh, you're still looking for something to be on the LLM. Because the problem with the LLM is it shifts the developer platform far above where NVIDIA sits.

25:07All the activity is happening on top of LLMs. You're, like, it's not. And so no one who's writing an AI application today is using CUDA. Now, some people are, if you're, like, training your own model and you're doing some, like, low-level things or non-LLM things. But the vast majority of the energy and all the money and the ecosystem is far removed from CUDA. They have no idea and don't need to know or care what chips their application are running on.

25:39They're just on the OpenAI API or they're on the Anthropic API or they're on, like, using Bedrock and Amazon and it's sitting on Tranium and they're using a Chinese open-source model. It's totally abstracted away. And this is why LLMs were bad for NVIDIA. Now, again, all the money they made along the way is worth it. But their moat has been tremendously diminished. Cuda is still a moat if you need to do stuff that requires CUDA.

26:11Right. But the vast majority of stuff and energy doesn't require CUDA.

Free preview sign-off and subscription info

26:14All right. And that is the end of the free preview. If you'd like to hear more from Ben and I, there are links to subscribe in the show notes or you can also go to sharptech.fm. Either option will get you access to a personalized feed that has all the shows we do every week, plus lots more great content from Stratechery and the Stratechery Plus bundle. Check it out. And if you've got feedback, please email us at email at sharptech.fm. Thank you.

More from Sharp Tech with Ben Thompson

(Preview) Astra (and AGI?) Arrives, Meta’s Muse and the Agent Opportunity, Anthropic and the Revival of (P)Doom Angst

Sep 10, 202624 min

(Preview) Fable 5.1 and Anthropic’s Data Retention Pivot, AI Civilizations and Related Matters, Q&A on Meta, Shopify, 3-D Printing

Sep 3, 202625 min

(Preview) Meta’s New Restrictions for Teens, Nvidia’s Open Source Investments, Q&A on Netflix, Druckenmiller, Parameters and Performance

Aug 28, 202622 min

(Preview) The App Store in the Shadow of AI, Offensive and Defensive Cybersecurity, Q&A on Financial Planning, AI Writing, American Sports

Aug 21, 202620 min

(Preview) Microsoft’s Plan for Platform Survival, Meta and the Market’s Permission, A Lack of Situational Awareness

Aug 6, 202628 min