Steadcast
Latent Space cover art
Latent Space

🔬 An Oscar, Two Asteroids, and the Algorithm in Your sklearn: John Platt on AI for Science

September 22, 20262h 1m · 23,777 words

Show notes

How often do you get to talk to a guest who has both an Academy Award and who invented textbook machine learning algorithms? John Platt has an Oscar, two textbook algorithms, two named asteroids, and an Erdos-Bacon number of 6. This was easily the most fun bio of all the guests we’ve read to date.

Highlighted moments

A predictive model is like, let's say you just have a, you have some inputs and you have some outputs and you just, I just want to build a piece of code that tries to just have the lowest error rate on some data set. It's a statistical model. A descriptive model is actually what science is trying to get to, which is, okay, it should be able to extrapolate because it has sort of the physics or the actual, some description of reality that's captured within it.
0:15
“you often run into this in remote sensing because there's always a trade-off. There's satellites flying above the Earth and there's a trade-off between how frequently they can revisit a spot on the Earth, what their spatial resolution is, how big the pixels are, and their spectral resolutions are how many bands they have.”
7:08
“If I could get a magic wish, I would say, someone, please make the, the everything lab that you could like send JSON blob to, and it will do any experiment at all.”
2:00:21

Transcript

Predictive versus descriptive models

0:00Are you talking about introducing explicit priors that, you know, she based upon some human intuition or maybe in this case, LOM intuition? When you talk about multiple hypothesis testing, right, there's predictive models and there's descriptive models. A predictive model is like, let's say you just have a, you have some inputs and you have some outputs and you just, I just want to build a piece of code that tries to just have the lowest error rate on some data set. It's a statistical model. A descriptive model is actually what science is trying to get to, which is, okay, it should be able to extrapolate because it has sort of the physics or the actual, some description of reality that's captured within it.

0:40And then you can use it to extrapolate. Yes, Newton thought of apples and gravity, but gravity isn't actually about apple, right? If you take a 17th century machine learning model, like, oh, apples will fall, but how about planets? You know, I don't know. I have no data about planets. So who knows what they do, right? The distinction between those is a little bit blurry, right? Because when a physicist or scientist comes, they use their intuition or maybe even more than intuition. Like essentially there's maybe a solid pile of facts that they know about the world and then they make sure that whatever model they build is sort of consistent with what's known.

Introducing John Platt

1:13Welcome to Latent Space Science. I'm Brandon, joined by my co-host RJ. It's a pleasure to have John Platt, you know, with us today. John is a Google Fellow and head of Applied Science at Google Research. He has really a fun, like, background. He, I guess, you described yourself when we were talking a few minutes ago as a mega nerd. No, giga nerd. Ginga nerd. Ginga nerd. He's excited in absolutely everything. And it really, it really shows. Yeah, you, correct me if I'm wrong about any of this stuff, but so you started college at 14 and started your PhD at 18 at Caltech.

1:47You were advised or co-advised by John Hotfield, right? Oh, yeah. Yeah, yeah. Who just won a Nobel Prize in, you know, two or two years ago. So, John created several, it was responsible for several textbook algorithms, one known as Platt Scaling, another one Sequential Minimal Optimization, which is the textbook algorithm for training SVMs. Even today, it's still, if you use sklearn, it's there. John has discovered and named two asteroids, has a Oscar for technical developments from 2006.

2:18So, if you've ever watched a Pixar movie, you've seen John's algorithms and work. John has an Erdos Bacon number of six or three, three and three from either side. And I'm going to skip over, like, 20 years of your career. But then jumping into Google, you know, working at Google Sciences, you've worked on fusion, quantum computing, climate modeling, and many other topics. Is that more or less right? That's right, yeah. Okay, okay, cool. Did I miss anything important for today? No, I mean, I've also done, you know, lots of applied math and signal processing and all sorts of fun things.

2:52Yeah, yeah. I think you also, your Wikipedia has a fun story about patents and the iPhone, too. The iPod. The iPod. iPod, yeah, yeah. Yeah, welcome. Thank you. Thank you for having me.

Scorable tasks in Era

3:05Can you tell us about the era is the, I think, the way that the acronym is pronounced? Mm-hmm. And I know that there's a lot of different semi-related stuff out there, both with Winton and outside of Google. So what, can you tell us a little bit about the details of era and what makes it special? Well, we've been doing sort of AI for science in Google research for more than 10 years now. Yeah. And around 10 years ago, it was very much using, I don't know what you call it now, maybe classical machine learning models, you know, things like, you know, convolutional nets or whatever.

3:43And they were specific models to build to, you know, solve specific science problems. But about two years ago, we got very excited about these more general LMs that have popped up in the last few years. And we were wondering what can be done with them. And, of course, a lot of people have been playing and trying to figure out what the right thing to do is. And we kind of stumbled into this mapping. And we found that many different scientific problems can be mapped into something we call scorable tasks.

4:16So you can often phrase a scientific problem as, gosh, I really would like to have a piece of code that, you know, maximizes some score. And it's surprising the number of different sort of scientific problems you can make a lot of progress on by mapping into that framework. Well, one thing is a lot of scientists spend a lot of time sort of building models. They might be statistical models or they might be, you know, physically based models. And if it's a statistical model, like in machine learning, your scoring function is, well, I have some data set and I'd like to have the fit, the model, and the data set go up.

4:51And we can talk about overfitting in a minute. But that was one of our questions. And that's actually very interesting. So machine learning is kind of a subset of this sort of scorable task. But you could do other kind of things, like especially Michael Brenner, who's the lead author on the ERA paper. He's very, very skilled because he likes to sort of knock out a scientific paper in an evening now with a tool. So there's something in applied math called asymptotic expansions, which is you're asking,

5:24how does an ordinary differential or partial differential equation, but say ordinary differential equation behave, there's some parameter that has an epsilon in it, and you're trying to say, how does it behave as epsilon goes to zero? And it turns out you can turn that into an empirical task by essentially asking that it proposes some solutions that are asymptotically correct. And you check to see if the asymptotic solution is correct for like epsilon equals one E minus four or something.

5:54And then you check that fit, but then you ask Gemini, which is the core AI underneath it, to do the mathematical reasoning, try to solve the problem while also maximizing the fit to the data. So you can actually, there's a lot of sort of tricks you can do because it's not that the underlying thing that's altering the code or the underlying thing that's sort of making the decisions is not a random process. It's an AI itself that is smart and knows about things and knows a lot about the world.

6:25You can get a lot of, solve a lot of interesting problems because that sort of core inner loop is an AI that has huge amounts of prior knowledge. So that's sort of the trick. So we've been running around trying to map lots of scientific problems into scorable tasks and trying to solve them. And it's really been kind of fun. And I'm happy to talk about the ones that I've been involved in at least. Yeah, I would love to hear about some of the more, so it's a statistical one is what everyone listening will probably know about. What you just mentioned makes sense. What are some of the other interesting ones?

6:57Let's see, that are not statistical. Let's see, because we have interesting ones like one that we just put a paper up on Archive is, or actually I think it might be on GitHub. It's, you often run into this in remote sensing because there's always a trade-off. There's satellites flying above the Earth and there's a trade-off between how frequently they can revisit a spot on the Earth, what their spatial resolution is, how big the pixels are, and their spectral resolutions are how many bands they have.

7:28And ideally you'd like to have a monitoring of the Earth that's constant and, you know, a frame every five minutes at hyperspectral resolution at whatever, 10 centimeters. You can't get that. But, for example, to monitor CO2, the atmospheric concentration of CO2, you can take data for one satellite that's, for example, it's OCO2 or OCO3. OCO3 is actually attached to the International Space Station. But, so it gets you like a little strip of CO2 measurements that are highly accurate and pretty high resolution.

8:04You can actually try to do, because a lot of it is in the infrared, weather satellites like GOES has some infrared bands and it takes a picture every essentially five minutes, but the pixels are very large and it doesn't have such great spectra resolution in terms of it wasn't designed to find CO2. So you just ask one to estimate the other and you shovel other data in, like what's the current weather? What's sort of the long-term, you know, albedo?

8:34And so ERA came up with this very nice model that can do almost like super resolution, an informed super resolution of one satellite to another. So that's like one example. Yeah, yeah, okay. So any scientific problem that you can map into this framework. So the input to ERA is sort of this mapping and the output is code, is that? Well, sort of. I mean, the input is you, the way we've got it set up in the product is you just start talking, right?

9:09And so because a lot of times it's non-obvious to how to do this mapping, although, you know, experts like Michael Brenner know how to do it. So he actually wrote an agent that actually helps you, sort of talks to you to try to help you define what your scorable tasks should be. So there's sort of already, there's sort of an instance of Gemini sitting there trying to help you write a code. So that's actually sort of almost like an intermediate result. You start talking to it about your problem and it tries to produce essentially a Python notebook underneath that has a score,

9:41essentially a function with a scorable, which essentially produces a score. And then it starts to mutate that notebook in a clever way, because again, it's Gemini and it will try to sort of keep proposing code that tries to maximize the score. So what is different about this and just a generally, a general agentic system that can sort of optimize notebooks? Right now, it's essentially its own, in modern 2026 parlance, we actually worked on this in 24 and 25, but the modern parlance is kind of a specialized harness that runs an algorithm,

10:21which for ERA was Monte Carlo Tree Search. So essentially it's keeping hundreds or thousands of possible instances of notebooks, and then it selects one, and I can explain how it selects one, and it decides, well, okay, what can I do? Gemini asked itself, what can I do to make that notebook be better? And then it will make a new one and test it and then put it back into the candidate pool. So you can imagine the candidate pool is actually tree-structured because it's, you know, every candidate possibly has some children.

10:52And what you do is you pick, based on something called the, it's actually a fairly standard algorithm in a firm reinforcement learning called Upper Confidence Bound UCB. So you essentially pick, it's an optimistic algorithm, so it tries to estimate, say, what's the 95th percentile outcome of mutation? It tries to estimate that, and it picks the one with the highest bound, the highest optimistic bound. So in other words, it doesn't always pick the best performing notebook. It tries to predict, like, what's the current performance plus two sigma of its guess, and so it's always trying, so it hunts around.

11:27So it's like high recall, basically. High recall. Well, it's trying to make its bet so that it most efficiently tries to make progress, which isn't always greedily doing the best candidate. Sometimes it's the fifth best. We're also play, we've also played around where it kind of recombines, it sort of takes ideas from two candidates and smashes them together and tries to make a third candidate out of that. How does it seed the initial candidate pool? Well, that's the amazing thing, is that underneath, Gemini is actually good at writing code.

12:02I mean, you just ask it, write me a thing, because you have a textual description of the problem. It isn't just, oh, here's your scoring function start. You say a textual description of the function, and you might give it, in fact, we have under some things, like, here are five papers that people tried to solve this problem with. And it's kind of smart. It actually goes and reads the papers and will actually take a first stab at code. It might not be great, or it might, you know, sometimes it has bugs and it returns essentially minus infinity.

12:32But it will then try to mutate the code and say, oh, it'll try to make it be better. So, it's pretty cool. You don't actually have to give it, I mean, you can, if you want, give it some kind of starter code, but you don't have to. How many agents are you spinning up? I guess maybe not agents, or how many different, you know, tree branches are you spinning up at each iteration? Oh, at every iteration? Well, there's a trade-off. You'd like to do a lot of parallel work, but if you do too much parallel work, you can't learn from previous things.

13:03So, right now, we use about, the default is 10 parallel, so you try to grow 10 leaves at a time. Okay. That seems to be about the right trade-off. When you say you can't learn from previous iterations, that means that the orchestrator, or is there some sort of, yeah, what's the, that, you said that there is some step which is able to, like, recompose. Combine, or make decisions beyond just, like, the score? I mean, so, yeah, I guess maybe one of the options is, as a human, when you are doing some sort of ML project, you don't just look at, like, oh, there's this one metric that we're trying to optimize.

13:39Oftentimes, there's, like, orthogonal metrics, sometimes even insights such as, like, just watching, you know, training curves can sometimes give you intuition about what's going on, or, like, looking into specific examples. Does it do any sort of introspection like this? Is there a... Well, it has. It has the history of... But by I meant why you can't do too many things in parallel is if you have 10 parallel searches at once, then the number one can't actually see what numbers two through ten are doing.

14:10So, if you do 1,000 at once, then you're using a huge amount of computation without a lot of cross-learning, whereas once you finish a little batch, you get the history. It's sort of, obviously, you have to prune it so it doesn't blow up the context, but you get the history of what it was thinking about as it was kind of writing the code and the results of the code. So, it can learn from its previous attempts. Okay. And does it learn across? Oh, yes. Yes. Essentially, it's like one essentially shared context.

14:42Yes. So, it is sort of thinking as it goes along. It's not like it's 1,000 different completely independent branches. You're really pushing Gemini's, like, long context abilities here. And you have to do the right management and stuff. Yeah, yeah, yeah. Okay. Oh, that's cool. Going back to RJ's question, the, like, key point here being the first key goal is to, I guess, identify what specific score that you are trying to optimize, right?

15:13Sometimes that is, I agree, like, kind of the hardest part of the problem. And so, I find it interesting that I'm not sure I'd always trust my agent to do that part. That part seems like the more human task in the loop. It is. And often, you have to be careful. In fact, a lot of what you do, it's kind of, it's very meta. I guess everything we do is very sort of high level. You have to make sure that there's no one common thing is you come up with a scoring function or the agent does or you do it together. And then the iteration finds a way to cheat or hold, like, oh, no, I didn't mean that.

15:49And so, you have to go through and often sort of play and have a loop around it where you kind of iterate, like, no, no, no, I didn't mean that. Or you have to tell it in its instructions, okay, you know, don't do this. So, yeah, there's often iterations. And so, even with agentic help, you don't necessarily get the right scoring function from day one. And, in fact, it's really neat because, I mean, in the old days, i.e. 2024 or something, you know, a lot of grad students would spend a lot of time, like, doing scientific software.

16:19And it's just so much effort to write code at all that you kind of try maybe a few things or a few things that are very related. And then you sort of stop because you have to write your paper or you have to do your next experiment. This thing is kind of underneath, kind of relentless because it keeps trying and keeps trying and keeps trying. And so, the people who use it are now spending all their time almost at the right level, almost at the scientific creativity level. What does it mean to have a cost function? You know what I mean? And so, that's almost like the essence of the scientific problem.

16:51You're not so much now in the details of, oh, oh, I have to import this CSV file or I have to get this database to work or whatever. Or it's, you're now sort of thinking almost like deeply philosophically about your actual scientific problem, not down in the grungy goop of worrying about, you know, databases. In fact, one cool thing the agent can do is actually suggest data sets to you. Like, oh, have you thought about maybe pulling in this data set and doing a join? And so, it'll make suggestions about, like, you know, data sets you can join with, which is kind of cool.

17:25Going back to what you said a second ago, in terms of agents love to hack things and reward hack, do you have any fun stories or interesting stories about, you know, where things were comedically run off the rails? Boy, I'm blanking. I know other folks have run into it. I don't know if it's, I don't know if I have enough details to sort of say, to sort of express the comedy of it. But it does, you kind of get surprised. Yeah. I don't know if I have any really concrete, sorry, I'm blanking.

17:56No, it's fine. Yeah. I always like to think of machine learning as like a, it's sort of like the old genie stories. That's right. Or monkey Paul. That's right. That's what you wish for because you're going to get it. Yeah. That's right. And you have that, and that happens very much with this. So you have to be careful. But on the other hand, it has some knowledge. The nice thing is that sort of Gemini knows a lot about many things, sort of more than any one person can do. So it at least knows, especially if you point papers, point, you know, like here, there are five papers that try to do this in some way.

18:32So to some extent, it does have that genie feel, but to some extent it also sort of does sane things. This is why, remember the whole idea of sort of evolutionary coding, it's been around since the 70s. Everyone's loved to do that. Like, oh, let's mutate Lisp code or whatever to do things. But the reason why it just hasn't taken off is that random mutation in code space is pretty much worthless. I mean, just like, well, just like DNA, it's sort of like, you know, most things are harmful.

19:03So here it's like, oh, no, no, we can actually find, it sort of knows, sort of underneath it knows interesting gradients to try, which is why the thing works, that the underlying loop itself is an AI. So yes, it can, it can maybe overfit and have funny sort of genie problems like you allude to, but it also has some amount of sanity because it's sort of, it knows about the world and it, it has world knowledge in it. So the paper though, you were doing Gemini 2.5 and I think, you know, Gemini has advanced quite a bit. Do you have metrics or have you, you know, this is a tool that you're continuously using and it sounds like you're improving.

19:38And I'm wondering, like, do you internally, have you seen like almost like a phase transition in how effective this tooling has been? Has it, how dramatic has the improvement been over the last, like, I guess, year or two? Oh, well, I mean, year or two. Yeah. Amazing. You know, there's every, even every half version of, of, I mean, essentially, I think it would have been impossible under Gemini 2.0. Really? Yeah, I think so. And others wouldn't, it wouldn't have worked. So you started at 2.5 and that was like just the thing? Oh, well, no, we've been trying to experiment with these things actually for a while.

20:11Okay. And things just weren't working and then they started to work and then now they're just amazing. So the progress on Gemini major versions has just been stunningly amazing. Yeah. Yeah, I think this is an experience a lot of people have been having where things were just seemed impossible, whatever, are suddenly becoming magically useful. Yes. Like really quickly. Yeah. And so, so if people are, I even say this to scientists because there's some people like, oh, I tried whatever 2.0 and I didn't like, oh yeah, that, that was a long time. That was a year ago. That was like a long.

20:42That was an eternity ago. That was an eternity ago. Right. Yes. Yeah. And in fact, all the, all the, we even have one of the preprints where we've sort of combined Aero with anti-gravity and that, you know, the whole, that whole harness of anti-gravity is pretty amazing too. And it's, that, that's the one where you can sort of pull in lots of papers and it can write lots of code for you. And so, yeah. Is that publicly available or is that? The, the anti-gravity? Yeah, yeah. Well, anti-gravity is certainly publicly available. Oh, sorry, sorry. Yeah. The, the Aero plus, anti-gravity? Not yet. Okay. Okay.

21:13Not yet. Okay. I find this area really fascinating because like you said, it's been, there's been some form of code mutation out there since the dawn of computer science, basically. The canonical problem is sort of the overfitting or multiple hypothesis testing problem, I think, which is maybe a little bit better matched to the problem where you're basically, my hypothesis now that this algorithm worked. Now, my hypothesis, and so that, um, you run the risk that sort of, it has exponentially exploded, right?

21:44Because now suddenly I have like these, it's like hyper, hyper parameters that I'm optimizing. And so that you have this explosion of state space that you're exploring. And so that it seems much easier to, to sort of overfit to a problem. What are your thoughts about that? Because on the other hand, empirically, my experience, I even tried, um, the, you know, sort of open source version of Aero. Um, I kind of strapped it into Claude and it's running right now, so I can't tell you how well it's working. Okay. I'm curious. Yeah. I'll, I'll let you know.

22:15Um, but I'm just curious to know, this is a question that's been in my mind about just general AI for science. And what, so what are your experiences with this sort of on the, the, sort of on the ground? I guess there's two questions sort of embedded in your question, I think, right? Because when you talk about multiple hypothesis testing, right, there's predictive models and there's descriptive models, right?

Extrapolative science and machine learning

22:38Uh, I, I, you know, can you tell I've been doing machine learning for a long time with statistics. That's actually a really good point though. Do you mind explaining that, expanding that? I'm not sure that's something that everyone would, in our audience would be familiar with. Right. Especially in, in, in modern days, I think people are trying to sort of obscure the two. If you started with LLMs, I'm not sure that distinction would be meaningful. That's right. So a, a predictive model is like, let's say you just have a, you have some inputs and you have some outputs and you just, I just want to build a piece of code that tries to just have the lowest error rate on some data set.

23:10That's just a statistical model, right? A descriptive model is actually what science is trying to get to, which is, okay, it should be able to extrapolate because it has sort of the physics or the actual, some description of reality that's captured within it. And then you can use it to extrapolate because it's sort of like, you know, yes, Newton thought of apples and gravity, but gravity isn't actually about apple, right? If he just fit, if he had taken as a machine, you know, the 17th century machine learning model, like, oh, apples will fall, but how about planets?

23:40You know, I don't know. I have no data about planets. So who knows what they do, right? So, so when you would say extrapolative, okay. I didn't realize we're kind of going on a tangent here, but I am curious. Okay. So when you say extrapolative, so there's different ways I could think about this. One of them is, you said, a model of physics or a model of the world. Are you talking about introducing explicit priors that, you know, based upon some human intuition or maybe in this case, LLM intuition? Or are you talking about this is the, the physics is actually learned by the model or the underlying process of the world is under the model?

24:13The distinction between those is a little bit blurry, right? Because when a physicist or scientist comes, they use their intuition and they, or about, or maybe even more than intuition, like essentially there's maybe a solid pile of facts that they know about the world. And then they make sure that whatever model they build is sort of consistent with what's known. Currently in era, it is, it's sort of, it is LLM intuition. Essentially, that's what I was trying to say about having a good gradient underneath, that, that, especially if you point it at existing papers, it will try to build models that are kind of sane underneath.

24:50Because if it's, again, if you, especially if you give it guidance, like, oh, be sure to incorporate this and this, or look at these papers to get these things. So it will, there is a, you can introduce a bias towards certain model choices and it will have a bias because it, it just, its own little sort of world knowledge is accumulated inside of, inside of itself in, in, in sort of pre-training. Can you give us examples of what that might look like? Is it modeling something in a way where it's actually, there's different, for example, if you're doing something with partial differential equations, there's these, you know, formalisms people have, like neural operators, for example, or, or you can embed a, you can encode a differential equation in some sense, or I think that's called physics and form neural networks or something.

25:36It's like one for fluid and kind of PDE type modeling systems, or for, let's see, molecular systems, there's oftentimes this idea about equivariance. I mean, are the models like picking up on these, like, you know, tricks which have been developed in literature, or are they adding some weights for some physical prior or something? Yeah, kind of what does that look like when it, when it introduces like a physics, when it introduces some sort of, you know, knowledge like that? I don't know if I have enough data to sort of say, oh, you know, 73% of the time it does this, but especially when you point it at existing papers, it will try to, in fact, it will do very well at adapting the, the, the methods that are described in the papers for the problem.

26:22And, um, in fact, it will do an amazing job. You can actually often just recreate or reverse engineer, uh, paper that, that's again what Michael Brenner actually likes to do this. Uh, he'll, he'll say, oh, that sounds like an interesting paper. Oh, we did this actually for the, uh, we had this thing. It was actually kind of a hack. I, uh, suggested this to Michael where, uh, there's this one MIT professor who, um, he came up with, uh, uh, some code to do, um, essentially if you have a rooftop, uh, if you have a rooftop with a fixed, uh, area and you want to sort of maximize the amount of solar, uh, power you capture over a day, sort of solar energy, uh, you can build up, which of course captures more sunlight.

27:01You can sort of build, you can have a design, a widget, you know, with involving mirrors or struts or solar panels, uh, just sort of stuck at whatever, uh, uh, you know, angles or, or sizes you like. And, uh, usually with, with a, with a maximum height and then try to figure, let it sort of explore that design space. And I believe Michael, we can ask him, I believe it actually just, I don't think he actually installed the simulator. I think, I think the code just, it just reproduced the code because it has like a coding agent inside of it.

27:33So it just reproduced the code from the paper and sort of figured it all out. So yes, it's, it's very good, especially if given a pointer to what other people have done. It's very good at kind of like, oh, I haven't seen it try an equivariant modeling that, that can get very hairy if you know about, uh, Klebsch-Gordon coefficients. It's, it's pretty fun. Yeah. Yes. So I don't know if it'll do the true equivariant stuff, but it actually knows a lot.

The phase change in AI capabilities

28:00Um, I remember actually when, uh, Gemini 2.5 came out, I know this is not exactly about era, but I, I remember sort of thinking, oh, this is a new world. When like the day 2.5 came out, cause I said, hey, uh, Gemini 2.5, can you write me, uh, some boosted decision tree code? And it did. And it worked. It just did. Yes. And I said, you know, yeah, this is the, this is a new world. Uh, so yes, I think it's, to loop back to your question, I think, yes, if you give it sort of guidance about, oh, you know, it's important to put this kind of thing in, it will.

28:31And so it's, it, it won't necessarily, or at least not that we've seen, discover completely new, uh, like if you didn't know about Klebsch, Gordon, Goffin, I don't know. I mean, you know what I mean? You didn't know about something. It won't completely discover a new kind of physical model from scratch, but it will certainly, if you tell it about interesting constraints about the world that are known, it will certainly follow up. I don't know if I answered your question. So, so there's still room for humans for the next year or two. Oh, in fact, there's, going back, I think there's totally room for humans because I don't know, I mean, we have co-scientists that tries to help you come up with, um, sort of hypothesis generation.

29:09But really, I still haven't seen, um, sort of the creativity and the philosophy and the sort of the, sort of the careful rigor. You totally need humans. I, I don't see humans going away. They can make strange suggestions and, uh, I've used co-scientists actually for an interesting problem in geochemistry and I learned about a new kind of ion I didn't realize happened in magma. But, uh, I, but, and so it'll tell you interesting things and you'll learn stuff, but I don't think it sort of substitutes for human creativity.

29:42You know, going back, we were just talking about two years, two point, you know, two, you had Gemini 2.0 to 2.5 and this was like, you already saw a leap and now it's been another year or two and now we're at 3.5 and, you know, there's, you're saying this is working much better. I mean, whenever you look at a graph, you know, you can, if something looks like an exponential, it can either, you can either be in a sigmoid or you can be at the beginning of a takeoff, right? I guess every exponential turns into a sigmoid dimension. Every exponential turns into a sigmoid. But the question is like, where are we on that? I mean, I, I guess I'm a big believer in sort of the whole jagged, um, the frontier.

30:14Yeah, the jagged frontier. And so certainly, at least what I see, I mean, I don't know what's going to happen in a couple of years, but yes, there, there's some big spikes out in jaggedness in terms of coding ability and just gathering knowledge and, and finding related things. And, and that's huge and wonderful, which I think is great for scientists. So far, it's kind of less in terms of, uh, rigor and, um, it, I, we can talk about things like the International Meth Olympiad, but, and meth in general. But, but, but in terms of sort of philosophy and creativity, I think it's still kind of not.

30:48And I, I, maybe it'll, maybe everything will, people, some people are saying everything's going to inflate and pass, but I'm still seeing a lot of very strong jaggedness. So I, I can see, well, maybe it'll get to be extra good at coding and extra good at, at fitting models and extra good at making suggestions and things. But I don't know, so far, not, uh, so far you need the humans. Yeah. I want to get back to the question about the multiple hypothesis testing. Oh, sorry, sorry. We got it. We derailed. Sorry. Yeah, yeah. Totally, totally love the, the tangent. Um, multiple hypothesis testing is when you make, you have a descriptive model and you're saying, this is the way the world works and you have a bunch of data and you take a billion darts and you throw.

31:24And so you have to be careful. There's something called false discovery rate. Yeah. Right. And so the question is, is this finding descriptive models or is this finding sort of predictive models? And fundamentally, the scientist is there to make sure that whatever is saying is descriptive. We haven't been able to make a system so far out of these pieces that, that can really sort of discover completely new physics or completely new science. But this is sort of a power tool to help you discover completely new science.

31:59So, so I, maybe I'm trying to unask your question about multiple. But I think this gets at the heart of it. Yes. Yes. But then you're saying, what about just pure over, okay, let's set aside. It's not, it's not trying to figure out a descriptive model of the world that's still up to the scientist. But what about just plain old overfitting? Yeah. Yes, you have to be very careful because it's a power tool. It can, I shouldn't probably say it, it can slice your fingers off. No, but you know what I mean? It's not, you have to be very careful and you have to be very rigorous. In fact, now you have to be more careful and more rigorous to not fool yourself.

32:29You really, really need to be just excruciatingly careful about having, you know, very hidden holdouts that you don't look at. You have to be just super, super rigorous to make sure that you don't completely, because it is a total power tool. So, the question to how do I not slice my fingers off is you need to use the same techniques, but be very careful with them. Yes. That's a very clear answer that I'm going to get before thinking. Yeah. Okay. Yeah, I actually don't think I've heard any guests say that. Yeah, it's like, it is, I think, a very important skill.

33:02Maybe one of the most important skills in the new, people talk a lot about taste. Yeah. But, and maybe this is a variation of taste, but like. The taste is like the other side, right? Yeah. It's like the rigor. It's like the, yes, yes. In fact, if anything. Yeah, taste or rigor, which one's more important? Well, I don't know. I think people, at least the way I am viewing it is, I mean, the aspect of the research, researchers are software developers, which there's a lot of overlap in a lot of fields. I'm seeing that software engineers are, it's almost like, obviously, there's a lot of concern, like, oh, no, what am I going to do?

33:39You know, coding seems to be, you know, getting automatic. So I think, I think there's sort of both. I think there's a lot of people get pulled into, well, I'll be the creative source. So I'll try to figure out new science. I'll try to figure out new products. I'll try to sort of really be very recreated. And I, again, I'm a strong believer that I don't think that's going to go away. There's also people sort of pull towards rigor. Like, oh, I want to make sure this doesn't, that this doesn't crash. I want to make sure this scales. I want to make sure this isn't wrong. I think you need both. And I think you need people who are really good at both, but they don't necessarily have to be the same people.

34:14But yes, I think you need, I think this is even broader than science, just as sort of software engineering evolves. Yeah, it'll be, you know, the people who will bring the creativity and the people who will bring the rigor. And I think those will be sort of anchors. Some other things that I've seen are related work out there. There's a really cool leaderboard for, you know, a claw leaderboard, agent leaderboard for scientific problems from Stanford. I don't know if you're familiar with it.

34:44It seems like a really interesting idea to me to have, you know, sort of different agents kind of competing on the leader. So it seems like if you squint a little bit, what ERA is doing is, is kind of a leaderboard, but it's internal and it's recombining ideas. Whereas what are your thoughts about this? And do you, is that like a thing that you guys are working on? And is there problems with that or advantages to that?

Goodhart's law and Kaggle competitions

35:10And ironically, you know, the whole ERA project actually started because people may not realize Kaggle is actually part of Google. Oh, yeah. And so it was called the auto Kaggle problem. So it was actually like, that's what it was. Let's try to, let's try to have a system that can sort of, you know, win at Kaggle competitions. So that's sort of why it sort of has this shape. That's sort of how the project started. And it goes back to sort of overfitting, right? If you've ever actually competed in a Kaggle competition.

35:42I have done Kaggle competitions and, or I've done one. It is a really interesting phenomenon because there's this, overfitting is like rampant. Yeah, yeah. And it's really impressive how people can overfit to certain data sets. That's right. In a way that is, yeah. Or we even, we had a, we have a fun project I can talk about over, if you like, that tries to mitigate contrails, jet contrails. I'm going to talk about that if you want. And we had a contrail Kaggle competition and people actually beat us, but they found that we had a half pixel error in our labels.

36:17And they, because it had to do with the center versus the lower left. Like where is zero, zero? Is it in the lower left of the pixel or is it in the center? You know what I mean? Yeah. So they found that and exploited that and, and squeezed or whatever, a little bit extra stuff. Because it turns out when you make artificial data and you rotate it, you have to make sure that, that you take into account that half pixel offset. So yes, people or people themselves will be, act like these LMs and, and try to sort of reward hack on these things.

36:48So it sort of goes back to, what is it? Goodhart's law? You know, Goodhart's law? Yeah, yeah. Any, any, any. Maybe, maybe quite a thought. Let's see. Let's say any metric that becomes a target is no longer good as a metric. Yeah. Uh, and, uh, so that's the, I mean, it's good and it just, you have to be, you have to be very, very careful and you have to, again, you have to have like layers of rigor and like, okay, but we'll do this and we'll optimize for this. But you have to realize, okay, that's just now Goodhart's law applies and you have to be careful. And so, yeah, that's a lot of reasons why the whole AI field has been kind of constantly exhausting these, these things.

37:24Because again, Goodhart's law applies individually to every, to every leaderboard you make. So again, it's sort of, you just have to step back and be very, very careful. I, I, maybe that's not, I don't know. No, no, no, no, that, no, that, that, that's really useful to my thinking is, you know, as we've had guests on, it's been a recurrent theme of how do you manage all this, the complexity that's introduced by LLMs in agentic science. I think my follow-up question was about overfitting in Kaggle. Um, yeah, it is, if you, you had an auto, auto Kaggle problem and then the question is given auto Kaggle, how often was it successful?

38:02I mean, assuming you probably just ran this on like all of your Kaggle competitions or something. Well, we tried it on various, um, like what they call playground competitions and it did very, very well in the playground competitions. Um, uh, we've entered into different competitions, some of them. It turns out, uh, there's, uh, in the last few years, just the, the number of leaderboards and competitions and whatnot have just exploded far beyond Kaggle. So, uh, we've done very well in some of them. Like, uh, one thing we're super proud of is the whole, um, CDC set up this, uh, competition where you try to predict next weeks, the number of, uh, COVID and flu cases that will happen in every state and, um, territory in the U.S.

38:41Uh, and you try to predict a week in advance and era did super well on that. It's funny because in some sense, Google invented the concept of using data to track, uh, disease progression with Google flu. So, uh, so it's, so it's kind of funny that you were sort of going full circle 20 years later or something like that. So that did very well. Other ones where we've entered, um, uh, we weren't quite as good often because people are very, again, you know, you have to sometimes, sometimes how well you do in these competitions is a, um, measure of how much sort of TLC you put into it and how much you're willing to squeeze the last .001.

39:18And so, so it was, I mean, era did well, it got you pretty close, but you, but we didn't, we didn't close the jump in the last, whatever, 30 places or whatever, because no one was there to shave the last, you know, .001 off the thing. Yeah, yeah. Is it very iterative? Like you get to it, you know, it sort of, I saw the charts in the paper and, you know, you sort of get these step changes as it discovers something and then, and then flat. And then, so is it very much human in a loop? Like, okay, you've stalled on the problem, like try this kind of thing.

39:51Okay. Yes. Uh, at the outer loop, which is, I think that's almost like the more fun creative part. Uh, so yes. Oh, here have to look at this paper or, oh, you're doing something bad or you know what I mean? So it's almost like having a, uh, hyper eager, uh, grad student or something who doesn't sleep. Uh, uh, and you sort of tell it things and you, and you sort of guide it around. How often does someone intervene versus like, what is the outer loop look like actually? Uh, it might run for a few hours and, and come back and give you, uh, some examples.

40:23And then you would, um, you know, you can do this as far as you like, you can, you can sort of keep trying and keep poking at it. So that's, it's very much designed to be human in the loop then? Yes. Yes. Interesting. Because a lot of the other tools that I've tried tend to be very, the one shot. Well, I guess it depends on your definition, right? I mean, it's obviously you talk to it and you start it and it'll go for some number of hours and then come back. And then, but of course then you say, but then that's where the human creativity kicks in.

40:55And then you're sort of doing the outer loop where you sort of every, you know, depends if you want to sleep, but you know, every few hours you, you go and you give it another try. And you, what kind of budget are you giving this thing? Like you, you blew through a million dollars accidentally kind of thing. I don't actually know because we're using sort of, you know, internal calls to, uh, to, to Gemini. So actually I, I don't actually know. So, but look, there's token budget, but then there's also like, I'm solving a problem that is computationally expensive. Oh yes. That also, it, um, essentially underneath it because, because the, the scoring function itself might have, you know, Monte Carlo estimation or, or whatever.

41:32Yes. So you actually end up, you can actually end up, uh, using a lot of compute, uh, uh, to just, just even do, or simulation. Like if you have a simulator inside, it has to run a simulation. Uh, so yeah, you can, you can, you can spend a fair amount of just CPU or GPU. So my little experiment with Aira, um, Aira and Cod is, is to build a, a neural network for some classification problems. And, and so they obviously like, if you have enough data, then, you know, larger networks work better, but they're more expensive to train.

42:02And then you get to, you start to run into a question of how do I manage my budget if I have a fixed budget so that I'm spending my, my, uh, my, my dollars on the most effective solutions. That's right. And I think that's still something we need to figure out, but it's, of course, it's no different than if you have a grad student and they're trying to train a very, very large, uh, neural network or a very, very large data set. They themselves have to, there's some thing like, Oh, is there a scaling law?

42:32Can I extrapolate? So it's not, I guess it's an, the same problem, but maybe more urgent because it, it just, it runs into this problem because it's so relentless. It runs into the problem much quicker than a grad student could. One of the things that you optimized was, um, contrails.

Contrail mitigation and climate impact

42:48Um, can you talk a little bit about that? Uh, well, uh, let me maybe spend a minute or two talking about the contrails problem. Yeah. For context, contrails, not chemtrails, which is a conspiracy theory. Uh, yes. Uh, although you should also dislike contrails, but maybe not for the same reason. So contrails are, uh, if you've ever seen, uh, those white clouds, uh, form behind jets, those are called condensation trails or contrails. And it turns out they add, at least according to the estimates that people have, uh, about 1% of all anthropogenic, uh, uh, global warming is caused by contrails.

43:24Why is that? I can just talk about the, maybe the physics of that. So it turns out that, uh, there's actually two, uh, countervailing effects. Contrails are, are, uh, well, sometimes if you've ever seen them, they, they're streaking and they kind of go away. Those don't really do anything, but sometimes they last for a long time. You'll just see in the cloud, uh, in the sky, just like, almost like a waffle of, of, of, of, of just, well, persistent contrails, they're called. And, uh, there's two effects that they have. Those are thin white clouds, so they reflect sunlight, but that only, of course, happens during the day. It turns out, uh, all, uh, all objects emit something called black body radiation, and the Earth does at whatever the temperature is, about 300 Kelvin.

44:02It's in the far infrared, around 10 microns. And at those, uh, wavelengths, uh, contrails have very low albedo. They're almost essentially black. And so they'll absorb a little bit of the outgoing, uh, infrared radiation and then re-emit it both directions. So essentially they'll reflect some of the outgoing heat, so it'll trap heat like a blanket. Um, and so because that happens 24 hours a day, uh, they tend to be, uh, warming. And it turns out it's a surprising, again, there's, uh, some uncertainty about it, but, you know, uh, contrails cirrus, cirrus that sort of comes from, from, from contrails, um, might, might cover, especially in places like Europe, which has a lot of air light traffic, a few percent of the actual sky is covered by sort of additional contrails.

44:47So it adds, uh, in those places like Europe, it adds about one watt per square meter of, of forcing, uh, locally at least, uh, which means that, uh, just to give you a sense, all of anthropogenic, uh, uh, warming sort of average across the whole globe is about three watts per square meter. So in places of high, uh, airplane traffic, uh, it can be a lot of warming locally. Uh, so what can you do? Well, um, it turns out, uh, contrails are caused by areas in the atmosphere that are ice supersaturated.

45:19They're a little bit like rock candy. So like when you have, uh, rock candy, you get a water solution that has too much sugar in it and any little, um, you know, little bit of sugar in it will just crystallize all the, the, the, the sugar out. Just like in this contrail, there are these regions, they tend to be kind of pancake shaped, only a few hundred meters tall. And if you fly through it, the jet exhaust has a little bit of, uh, moisture in it, which, uh, will turn into droplets and then freeze. And then for every gram, if you're in this bad region, for every gram of water, uh, ice or soot you put out, it's about 10 kilograms of water gets sucked out.

45:56So there's this enormous 10,001 curing ratio. Um, so it's a big problem. So what you can do is you can figure out where these regions that are invisible, of course, uh, these regions of ice supersaturated are, and then tell the plane to go underneath. And you only have to drop essentially, uh, what they call two flight levels. So it actually, it costs a little bit of fuel, but not very much to kind of avoid these sort of bad regions. So we built a system that sort of looks at satellite images and tries to detect where contrails are.

46:27So we have essentially a continuous monitoring system and then try to build a model of, of, because it turns out the weather models are not quite accurate enough to, to find these, uh, places of ice supersaturation. So we built a custom model, again, like a convolutional net or a unit or something, uh, to essentially to predict where they're going to happen. So that, uh, then you, we give, uh, maps to a flight planning software so that they can dodge it and, and inexpensively reduce the, the, um, the climate impact of, of aviation by a lot.

46:58So what's the physics behind why you can predict that, right? Is it just, I see it in the satellite and then tomorrow I think it'll be there because planes go to the same place? Oh, no, it's because you're trying to detect these, uh, regions of ice supersaturation because they're very, very persistent. Essentially they're, oh, they're persistent. Oh, yeah, yeah. I mean, no one knows exactly, but they could last for days. Essentially they're caused by, they think, sort of warm, moist air being injected just at the boundary of the tropopause between the, uh, just at the bottom of the stratosphere.

47:30And then when humidity gets up there, it sort of sticks there for a long time and then gradually dissipates. Got it. So, so it's, there's just a sort of. They're like bad spots in the atmosphere you don't want to fly through. Right. Okay. And so once you've established that, it's probably good for a couple of days at least. Uh, well, you have to keep predicting where that is. Yeah. And the, the models you're using are, you mentioned like CNNs or something like that. That's right. And we haven't replaced those with era level, uh, models yet, but there was a very interesting problem that came up, which is you sort of want to know, well, just how much warming did this contrail make and how much did it add to global warming?

48:08Uh, because for example, you might want to find the biggest ones because there's some fuel cost and maybe it costs a bit of money for the, uh, airplanes to avoid it. So you say, well, gee, I'd like to kind of know how much it did, but that's actually a, what they call a counterfactual problem. Like, okay, you made a contrail and certain amount of infrared radiation happen. So we can measure that, uh, if you're careful, but would have happened if there hadn't been a contrail there. That's a very difficult thing to estimate because you can't access the universe where that didn't happen.

48:38That didn't happen. So you have to make these things called counterfactual models. And there's, those are, uh, if, uh, I don't know if your listeners know, counterfactual models are actually pretty tricky to, to fit and make. Uh, and we were, there was, remember there were two, there were two, um, the reflecting the, the sunlight and then there's the infrared. It turns out the, uh, measuring what the effect of reflecting sunlight is actually more difficult. And we were actually stuck on it for two years. We, we had a model that worked okay, a counterfactual model for the outgoing long wave radiation, but not for the reflected sunlight.

49:10But ERA actually helped us find a, uh, a model that sort of searched all the confounders and sort of figured out like, oh, how, how can we estimate it? Cause we had, again, we even had like test, test code on, on sort of artificial, cause you can kind of inject artificial, uh, make artificial data sets where there's sort of injected, um, contrails and sort of figure out, oh, well, we know how much it was because we injected it. And, and so again, our own attempts didn't even pass our own, um, tests, but ERA's thing actually did and, and, and sort of unstuck this problem.

49:43So yeah, we're in the middle of writing up a paper. Uh, we have a paper about the outgoing long wave radiation, but we have a paper, um, that, uh, it's not submitted yet, but we've talked about it at, uh, at EGU, I think, um, where, um, we actually solved this problem. So, and, and the, the models that ERA comes up with, are they just like a big monstrosity of, of, of, of code or are they like pretty basic and it was just, you needed the intuition to develop that? Yes, it's actually, in this particular case, it was actually more of the latter that it essentially sort of helped identify what the, it was a very simple model with some, just, uh, uh, some number of confounders that we just hadn't, we just hadn't tried that combination before and it worked very, very well.

50:25So yeah, it actually sort of came up with the, the, uh, and it was sane in, in retrospect. So that was, I think, uh, a big win. Yeah, that's interesting. I know you've done a lot of work in climate. What other, other, um, stuff have you done? I think I talked about this, right? The, the, I talked about the CO2 thing. That was pretty fun because it's this, it's still quite a, uh, uh, the reason why estimating CO2 in the atmosphere is an interesting problem is we actually don't know what the carbon flux is into the bio, in and out of the bio.

50:55I mean, we do, we know that the biosphere captures, right? We emit a bunch of CO2 out into the atmosphere and some of it gets absorbed into the ocean with sort of mostly inorganic chemistry, some, some phytoplankton, and a lot of it gets absorbed on land. Uh, but we, but the error bars about what happens is are moderately large and the error bars 50 years from now are very large. Like, like, the models in 2100, we don't know how, how the biosphere will react to the ever-increasing temperatures and CO2, so we don't actually know how much the CO2 will absorb, and the, the, the, the error bars are 300 ppm of CO2 just from the uncertainty of how, what gets absorbed.

51:41And just to point out, you know, right now there's about, what, 440, 450 ppm, so it's huge. I mean, it could be seriously, amazingly awful or not great, but, you know, it, it, the, the 300 ppm is, like, enormous, uh, uncertainty, so it'd be really nice to figure out, you know, can we, can we reduce that? So this is, like, the CO2, uh, concentration is, like, one step towards that, um, solving that. And I know that Google has made some really big, um, improvements in climate modeling and, and weather prediction as well, right?

52:17Is it, I was at NeurIPS this year, this last NeurIPS. And, and, um, and I stopped by the climate track, and, you know, I maybe only had a chance to listen to two talks or something, but I was, it really blew my mind, the sort of step change I think that's happened in the past, I don't know what it is, maybe 10 years or whatever, in terms of climate modeling, and I know a lot of that happened at Google. So, can you talk a little bit about what has happened in Google and other places that has made that, allowed that really big transition in, in climate and weather modeling?

Weather versus climate modeling

52:51Uh, okay, so let me, people often sort of collapse climate and weather together. Yeah, right. Well, because they're fundamentally the same physics. Yeah. Although, uh, at least for the atmospheric physics, they're, uh, obviously when you start having ice and land, you know, climate is long-term weather. And so, uh, the complexity of a full-Earth system model, which is a climate model, is much, much bigger than an atmospheric model. Like, you have to actually measure, uh, what's the water flux and the CO2 flux in and out of the land, or what will happen with ice.

53:23And so, there has been a step change with weather models. Sorry, sorry, I want to make this. So, weather is up to 15 days. Okay. Approximately. Yeah. Because, you know, weather itself, or the atmosphere, appears to be chaotic. Uh, uh, I'm, I'm hoping, I don't know if I should explain chaos. Uh, essentially it's the butterfly effect, right? That small perturbations, like a butterfly flaps its wings and the weather will be completely different in, you know, two or three weeks. So, weather is trying to predict the actual, um, trajectory of the atmosphere over, say, two weeks.

53:57And that's, now, that has been a huge step change. And that's because that's been a lot of, not even the new LLM stuff, that was based on, you know, the 2018 era, um, machine learning stuff. And just a large amount of data and a large amount of compute. So, it's been, there's been a lot of sort of very clever work, and a lot of it from Google, making new, um, new weather models. And it's been, it's been great. And, in fact, uh, we had a, uh, a really neat breakthrough, because now we can apparently predict tracks of cyclones, tropical cyclones, uh, much more accurately, many days, uh, in advance.

54:32And so, places like, um, Jamaica had, got hammered by a terrible, uh, hurricane. And, uh, a lot of the classic models didn't actually predict it, partially because it's often, that the, especially the intensification, it's all being driven by the, uh, what the surface temperature is. Because hurricanes are, people might not realize, are heat engines. Essentially, they, they convert sort of heat in the ocean to big atmospheric, um, um, motions. So, weather's been great. The climate is much more difficult, because you don't actually care about, you're not trying to predict whether it's going to rain in Seattle in 2070.

55:06You're trying to get kind of, like, averages. And, what makes it difficult is that it's what they call non-stationary. So, that, in fact, literally, the, the, the, it's like the underlying physics, or the underlying, like, you know, plants are behaving differently, and, and ice behaves differently. And, and, and so, it's a, it's very, very difficult to use sort of classical ML on sort of true climate models. And so, that's why sort of the whole discussion you guys had about, you know, what we're talking about, about descriptive models and multiple hypothesis testing.

55:38That is incredibly severe in climate, because we don't, we have no data from 30 years from now. And we don't want to wait 30 or 50 years to find out whether we were right or that we overfit. So, whatever things we do, you have to be kind of careful and try to peel off sub-problems and, and the problem of, of exactly how do you inject, how do you build a big model that can predict into the future, uh, but is still constrained by what we know. So, it's a fascinating problem. I think it's still unsolved, but it's a, it's a great problem to have, because of, again, these uncertainties, we really would like to know what will happen in, in 60 years to the climate.

56:16So, it's still, it's a thing, it's a very, very interesting problem to work on. But, so far, AI has not revolutionized it, because it's very, very resistant, again, because of this data problem. It's, it's a low data problem. So, does the butterfly effect, the chaotic, um, nature of weather, does that also impact climate, or is the timescale so large that you have a closed system for which, you know, maybe it's oscillating between poles or whatever, but it's, it's sort of, when you look at it at that timescale, it's more stationary.

56:46It's unfortunately non-stationary in a different way, but the original sort of whole, uh, chaos thing was, well, I, maybe many people came up with it, but, uh, I mean, meteorology, it was back to a person named Lorenz, uh, who had this sort of very model, uh, very simple model ODE. So, the difference between climate and weather is, is weather is where are you on the attractor, and climate is about the statistics itself of the attractor. The problem with, with climate is that we're altering it, so the attractor itself is changing, it's moving, and there could be, and there could be, everyone talks about tipping points, that means that the attractor suddenly changes.

57:23And the trouble is, that's very, very, very difficult to predict. So, even the attractor's shape changes quickly. Or could. Could. And the trouble is, when you run a simulator, you don't know, like, is it, did it go unstable because my model's not great, or is it an actual physical instability? Interesting, yeah. And it's extremely difficult to tell the difference. So, what do you do, especially when, I mean, to me, it strikes me that not only do you have, not have future data, you really don't have much past data. You can do some measurements and ice cores and lots of stuff to try to do that, but there was nobody with an instrument 100 years ago.

58:00That's right. So, if you have annual data or whatever, maybe you have, if you're lucky, 50 data points in any one location. Right. So, whatever we do has to be very constrained by what we know, but it's just very difficult. I'm just telling you sort of the horns of the dilemma people want. So, people make these, in fact, people in applied science in general, I would say climate is the most extreme, make these things called process models, where what you do, and I've seen the code. Oh, well, you know, I'm going to be reductionist, and I'm going to sort of take the horrible, complicated climate thing and sort of boil it down to a thousand pieces, and then I'm going to, you know, find the expert who wrote a paper about, you know, piece number 763, and he fit a cubic to some data.

58:43Like, for example, one thing that's very mysterious, which is related to contrails, is how does ice behave in clouds? It turns out, you might, I get, everything is complicated once you dig into it, but it turns out that, like, when you make a contrail, how long does it last? Well, it depends on, because the way contrails can evaporate is ice starts to accumulate, as I said, and then the ice crystals get big, and then they fall. But, of course, how quickly they fall depends on their shape, which is not known.

59:15And how much does the contrail mix from the moist inside the contrail out? Again, people have approximations, but they don't know. And so the uncertainty is very much confound, and it's not just, oh, John, who cares about contrails? It turns out that the actual physics of, of microphysics of ice has very strong implications about what climate models do, and it's sort of, we just don't know. So, I'm trying to say it's very gnarly, and it's not a solved problem. My hope is that with tools, maybe not like today's era, but maybe tomorrow's era, because it, remember, it can, as I was saying before, it not only can fit data, it can read papers, right?

59:58And the question is, it can read a lot more papers than we can. So, maybe, maybe we can integrate all the data that, or all the knowledge that people have carefully evaluated, much more than any one person writing a piece of code, and fit data. I mean, that would be, that would be utterly glorious. We don't have that today, but that's sort of one of the hopes that I have, even for a more amazing tool in the future, is something that really can write code in a sane way. Even much more sane, because it'll be constrained by all the scientific knowledge that we've accumulated so far.

1:00:33That would be amazing. We don't have that today. I mean, that strikes me as being very similar to biology. Oh, yes. Oh, boy.

1:00:41Right? If you've ever played biology, or even looked at biology, there's so many exceptions, and so many hacks in the biological systems. Yes. Oh, boy. So, yeah, it would be amazing if we could have a thing that could really integrate all known scientific knowledge with data, and try to synthesize sort of new models and new things. I think what you're saying is that AI can be an unlock here to some extent, because the models are so piecemeal.

1:01:12They are necessarily piecemeal. And so that being able to assemble the jigsaw puzzle, not to mix metaphors, but to assemble like a really, this jigsaw puzzle, having just scale and capacity actually helps a lot. That's right. The one thing that these AIs have is somehow, you know, humans, even I'm pretty well right, I think, but it's just difficult for me to kind of integrate across the N-squared different paper. If N is the number of papers I've read in life, it's pretty big, and it's just difficult for me to even do that N-squared thing.

1:01:45But somehow, there's just so much data in those billions and billions of parameters, and you can also give it access to read PDFs, that it can somehow start to pull things together that people wouldn't do it. So that's, again, I'm starting to see little indications of that inside of Aira. I'm not claiming that's what Aira does today. But, yeah, that's sort of my hope of where this is going to go. I think I've heard a lot of people suggest something like the route to intelligence is to combine LLMs with some form of search. So it's actually amazingly like Aira is doing that, right?

1:02:17Something which, you know, is maybe a very strong database lookup with a good search algorithm is one way. And, of course, I mean, there was the whole, I mean, people still do, I guess, the whole RAG thing, of course. And if you think about it, Google itself, you know, the 10 blue links things, it was or is a form of AI even before we had LLMs, right? Because it was like, you could cast yourself back to whatever, 2010 or 2015. You could ask Google about literally anything and it will tell you stuff, right?

1:02:50Surprisingly well, actually. Yeah, surprisingly well. Because somebody on the internet has written about it probably. That's right. So if you can match that. In fact, that was one of the reasons why I wanted to come to Google. It's just that was such an amazing thing, right? I wonder how much of our audience did a search pre-Google and just know how bad that experience was. Yeah, I remember in 1998, I think, I think it's when Google, like I was using Altivista. Yeah. And like, I don't know, Google just got released and I used it and I just, sorry, digital.

1:03:20Yeah, I remember that too. I just dropped like a hot potato or something and started immediately using Google. So that is sort of a form of AI. And so, yes, it might be, yeah, it could be that just having access to all of that and sort of keeping it in mind at the same time. That model of sort of scientific discovery in as much as it pans out is kind of comforting too because it is reductionist so that you can look at the individual parts and understand them. So it found the exact things to assemble, but they're all actually maybe fundamentally things that people have invented or it's done iterations on.

1:03:59And so all those little pieces are individually understandable and then you can also put them together into a coherent picture. I think for a lot of problems like biology or climate science, I don't think we would trust the answer unless it was in that shape. Yeah. Because if there was some giant black box model that said, oh, this is how a cell works. It's like, do I believe it? I mean, I don't know if I believe it because I can't examine it. But I mean, to argue against it though, like if it works really well. But you'd have to gather, you'd have to, I mean, you have to test it obviously, but a statistical model, again, it has to extrapolate.

1:04:36Yeah. Yeah. And it has to extrapolate to the extreme or the sort of the black swan events. That's right. Yeah. And so it's very hard. This is why things like self-driving cars are very, very, it's a very difficult problem. Right. It's all corner cases. It's kind of amazing how well they've done. That's an interesting point thinking about like when, you know, coming from the world of physics where a model was usually a single equation or a small number of equations, which uniquely define a system and everything about it. And you just crank, you know, just find a solution to the system and you've, you now know everything you need to know.

1:05:09And I think something like AlphaFold was kind of a shift for a lot of people where before they thought, oh, protein folding is, you know, a problem where you just, if we find the right force field and we have the right computational engine, we will solve protein folding. And the thought of even really solving it in a data-driven way was only a period of a few years before, you know, AlphaFold, AlphaFold 1 came out. And it's interesting that I think it's sort of forced the new AI modeling has, I think, forced people to reevaluate almost what is science.

1:05:45Because AlphaFold is, and similar models are incredibly powerful. There's a lot of things that they've opened up as tools, but at their core, they oftentimes don't give intuition in nearly the same way that, let's say, that most physicists historically would have wanted. And so I guess it's sort of, there's this old saying, all models are wrong, some are useful. Yes, Bux said that. Yeah, yeah. When do you find the data-driven models to be sufficient? And when do you want sort of like something which is interpretable that humans can actually understand?

1:06:20I think it boils down to almost like the difference between weather and climate. If you're in a data-rich regime, like weather, or even proteins because of PDB, you can feel, oh, yes. In other words, I've got enough data to kind of cover. And so a statistical model like AlphaFold should do. And so, in fact, a lot of people happily use Alpha, you know, I think it's really revolutionized my understanding. I'm not a biochemist, but people seem to love it. And one amazing thing they did is they exhaustively, they just ran it on all PDB and published it, which is just really, really cool.

1:06:57It's like 6 billion protein predictions. The vast majority are actually quite accurate, yeah. Yeah, so that's just amazing. But it feels closed, if you know what I mean. Yeah. But when it's like climate and it's open and it's non-stationary or you have to make these big extrapolations, you have to be much more cautious. Or maybe biology. Again, and there's probably, there may be parts of biology like, oh, there was this virtual cell challenge from the ARC Institute. That had a funny result. I know that people were, there may have been some overfitting for at least that.

1:07:29Yeah. What would we say more? Sorry. Or I guess maybe at a high level, I think simple baselines. Oh, it worked very, very well. Yeah, yeah, yeah. Yeah, just like in the. One of the classic things in whenever you do biology is just always start with a simple baseline. Maybe this is probably just good ML in general. It's good ML. Yeah, understand your simplest case in biology. There are many problems where they're extremely resistant to anything beyond the simple baseline, even if you have a lot of data. That's right. And in fact, I tell people the same thing. I said, always just fit linear regression.

1:07:59Just fit. Just do it. Just do it. Just do linear regression. Or SVMs, yeah. Or SVMs. I mean, SVMs are just a different. Different form of linear regression. Yeah, yeah. Different basis. So when do you need the more process model-y thing? I think it's just when you have, I mean, sort of climate is on one end and I don't know, weather maybe on the other end. That may be too extreme. But I think it's where are you on the data richness thing? When can you feel like, oh, no, I really have a closed problem and I think I can actually cover it?

1:08:32Yeah, closed problem that data fully covers is, yeah, I think that that makes a lot of sense in what I've seen as well. Yeah. I think one thing I'm kind of curious about is when you're working on climate modeling, what are the, you talked about contrails, you've talked about CO2 predictions. What are the broad things you're trying to accomplish? So one of them is, I guess, making up, like, making interventions and the other one might be making predictions for things like insurance or, like, how do you help, like, adjust for some sort of climate change?

1:09:10Or, like, yeah, what are the, like, principal goals, I guess, for you specifically or the community at large? I think, you know, just like any community, there's probably many different goals. For me, I'm, or, and my team, we're very, very interested in interventions. Um, um, so, like, which ones are, uh, possible at, at, at relative cost? I mean, contrails was kind of amazing because it turns out that the intervention is quite low cost and we, um, and also one amazing thing about contrails is they are local, unlike things like CO2, because, so if a country decides to fix contrails over itself, it actually improves its, I mean, it, it has the global effects, but it mostly improves the, the climate a little bit over themselves.

1:09:52Yeah. And so, they, they like that. I guess that if you were in a cold climate and you want to warm it up, this is now your own, you could know. Oh, yeah, it, it turns out that it's a little bit asymmetric. The warming is constant, essentially, and global. The cooling only happens when you're sort of at a good, then the sun is at a good angle over you. So, it's very rare that the uncertainty, there are contrails that where are uncertainty bounds in terms of the warming. There are many, many contrails, mostly at night, of course, where the, where, where it's, it's like a largely warming and we're, we're very sure in terms of like two sigma.

1:10:26Uh, there's very, not very many contrails where you say, oh, I know for sure that it's cooling and I want more of it. Like, like, like, like, so only over sort of the poles, uh, in polar summer do you know that the contrails are cooling and therefore, if you got rid of them, they would warm up. But there's essentially no flights over Antarctica and not that many over, over the poles, uh, in the summer. So, so no one's gonna, no one who lives in a cold climate is going to use this maliciously. Well, uh, yes, in that, uh, in that, well, they wouldn't know for sure whether it was warming or cooling and so they would do stuff and things.

1:11:01So, mostly we just sort of ignore, we don't recommend, uh, that people fly those. You also brought up an interesting point about the economics. I mean, a lot of, uh, there, I think was a lot of resistance historically about certain, you know, to climate, certain climate change interventions, which have, in some sense, the market has just taken over. Like, at this point, unambiguously, like, renewables and batteries are just almost universally, unambiguously, just better than alternatives. For, for, for, for, for, for non-mobile.

1:11:32I mean, you, for mobile, I mean, EVs are. That's a really good point. Yeah, yeah. Like, planes, we, we do not have a solution to. Correct. I mean, there are some battery-powered planes, but they're very small and have to go very limited range. They probably will never actually be, uh, or. It's hard to imagine. The physics would be very, very, unless we came up with something like nuclear batteries, which would be kind of amazing. But, I don't, we don't know how to do that. Or, even if we did, I think that the risk of, like, people would be too afraid of a nuclear battery going wrong or something. Oh, yeah. Yeah. Since we don't know what they are, we don't know what the risk is. I guess we don't have the risk. Yeah.

1:12:03Yeah, so we don't know. So, yeah, that's, that's the problem with, I talk about, I have given talks about climate change, and I talk about the pie, the pie chart of badness, pie chart of sadness, which is, there's no one silver bullet for climate change, right? There's so many different things that contribute greenhouse gases just from across our economy. So, they sort of all have to be fixed, or many, many of them have to be fixed. So, there's no one single thing. I mean, I've worked on Fusion, Fusion's cool, and it might actually knock a lot of them out if it's cheap enough, which we don't know, because we don't know if it'll work yet.

Fusion energy and plasma control

1:12:34I mean, Fusion is one of those interesting things where the joke was always Fusion is 30 years away, but I think it's actually now less than 30 years away, maybe. Yeah, no, I think there's a, there's a definite probability that, that someone will make a commercially relevant Fusion even by the end of this decade. So, I think it's like three years away, not 30 years away. This is very real. Interestingly enough, I think a lot of that, I'm going to not just stump or advertise some of our other episodes, but a lot of that actually comes down to material science, interesting enough, in that.

1:13:06Well, I'm skeptical. Oh, superconductors, sure, sure, sure. Well, yeah, sorry, we can talk about Fusion a few minutes. Yeah, yeah, there was actually two things. One is Fusion, the other is better control systems, which I think is actually up to us. Yes, and in fact, Google DeepMind has been working on control systems for tokamaks to make sure they don't essentially go unstable and go disrupt. Yeah, that's right. Disruptions are quite interesting to themselves. Yes, yes. Yeah, it's basically the entire, all energy in the tokamak culminates into one little beam and then-

1:13:38And then it hits your vacuum chamber, and you're very, very sad. You're very sad. Yeah, yeah, I think people believe that either could turn on after $30 billion disrupt and then basically have a $30 billion brick or something. Oh, yeah, yes, I guess you could try to patch it. I remember working, we, again, before LLMs, we worked with a fusion company called TAE, and I was in their control room. And yes, it was kind of sad. You have to be very careful. We were making systems to recommend new experiments, and they were very, very skeptical and jaundiced, which they should, because I've even been there for that.

1:14:15Even under human control, it's like they were doing some experiment, and then you hear this big bang, and it's like, oh, no. And then it's like, you know, then the apparatus is down for two weeks as they patch some- Wait, you were there during a disruption? Oh, no, this is, sorry. They have field reverse configuration. Oh, okay. I got it. I said, yeah. Which has its own, I mean, things, you know, there's some arc. So what is that? Sorry, I'm not familiar. Oh, oh, what's the field reverse configuration? Well, it turns out tokamaks are not, although they're perhaps the most studied form of plasma. There's many different kinds of architectures, essentially ways to try to stabilize and compress plasmas.

1:14:49There was a shape, essentially, it's essentially a self-contained football of plasma called the field reverse configuration, where essentially the magnetic field inside and outside are opposite, so they're separated by something called a separatrix, and that is sort of, in theory, unstable, but in practice, stable. Like, for example, when you run Magnete Hydrodynamics MHD code, it's unstable under that assumption, but that's an assumption, that's not the way the real world works. And so, yeah, it was kind of disfavored for many years,

1:15:20but TA and other people, I think Helion, have FRCs because they are actually relatively robust. You can actually knock them against walls, and they'll still, they stay stable. And yes, but you can still get discharges and things that punch holes in your vacuum chamber, which is kind of unfortunate. For clarification, so you have these fusion reactors, they are, or trying to be reactors, maybe. Apparatuses, yeah. Okay, apparatus, and you create a plasma.

1:15:53The plasma is magnetically charged. Oh, confined, yes. Or confined. So it's confined by a magnetic field, so you have some sort of magnetic system that is tunable by a computer, and then the computer tries to kind of maintain the confinement. Well, FRCs, kind of, once you make them, they're sort of sustained. There's different ways of trying to make sure you, okay. So all of fusion boils down to something called the Lawson criteria. There's essentially, and it explains why fusion is hard.

1:16:25Essentially, you can just very easily, on the back of an envelope, just show that the density, the temperature, and essentially the energy loss, it's called the confinement time. It's one over the amount of time it takes for the energy to decay away, one over E in a plasma. So the product of those three numbers has to be bigger than some constant, and then you can get fusion, and if you don't, then you don't. And that explains, the fact that it's a product of three numbers explains why fusion is so hard, because every approach has an Achilles heel where one of those numbers is not very big,

1:16:56and then they try to desperately make that be higher. And every approach to fusion is kind of different. And a lot of, you have to be a bit skeptical when there's all these sort of breathless news things about fusion, because they'll say, you know, now confinement time is starting to like, oh, stable for X minutes or whatever. And it's talking about like one of the three numbers, but you have to have all three numbers before you can get fusion. I think that the whole field is making a lot of progress, and it's very exciting, but you do have to be a little bit cautious about the breathless news articles that only talk about one number.

1:17:30So what is the computational part of that? Oh, unfortunately, for better or for worse, it depends on the approach. So for Tokamaks, as Brandon said, it's mostly stable, except that there's occasionally this instability that takes all the energy and smacks it into one place. And so you have to sort of keep everything sort of under control. So it's a control system. FRCs themselves have very simple instabilities. So, for example, they have what they call a Z instability. So it's fine, it's stable, it'll just wobble, literally wobble back and forth,

1:18:03but you just make what they call a PID controller that just keeps the football in the center of the reactor, and things are fine. And it does that by adjusting the magnetic field? Yeah, it's sort of, it actually just, I think the electric field is sort of, it sort of knocks it back and forth. The issue that people have is it really depends on which sort of plasma architecture they're deciding to use. Climate is, you know, sort of political because of economics, basically, probably, mostly, maybe other stuff. But the economics of it, you know, you have to persuade people to somehow spend more,

1:18:39or you have to have a solution that has like this happy coincidence where it's both economically better and better for the climate. That's hard. But in many cases, it's not. I mean, in many cases, it hasn't been hard. But yeah, I guess. Well, I like to think about like, okay, predicting even weather, right, you can prep, and you could see how that could be economically beneficial. So, what kind of work are you doing with interventions, and how does that kind of interact with economics?

1:19:09Like, it sounds like the contrails one, where I said it in an analysis and said, actually, this is great because it's very low economic impact, but high value. That's right. So, if you're trying to think, there's sort of energy intervention. So, you have to be, you have to sort of compete with existing forms of energy. And that's not trivial, unless there's a co-benefit or there's some sort of clever, you know, just co-benefit. Like, this, again, is highly speculative. It wasn't our work. There was a startup that was, I don't know if you saw the news.

1:19:40It was last year, I think, where someone figured out if you inject mercury into a fusion reactor, that the neutron flux can actually transmute the mercury into gold. And then you can sell the gold, which I thought was very clever. It might not work, but I know. You know, as a physicist, the one thing I want out of a fusion reactor is helium. But that's a different story. Oh, helium, it's three. Well, I mean, the helium-4 is kind of boring, although it is getting, because the strategic reserve has been shut down, there's less of it. Yes.

1:20:10And, of course, I want helium-3, helium-4. Helium-3, yes. Well, not even just a fuse, just to make dilution refrigerators for quantum. Or MRIs, or so much technology we think about actually just goes out the window if we run out of helium. That's true. No one's thinking about it. Sorry, that's like a complete aside, though. Yes, the fact that the U.S. had a helium, strategic helium was to reserve, was for a very important blimp. Yeah. But they kept it for decades anyway, so that was nice, but then we stopped. And then we got rid of it all. It all went up in the air. Yes, yes, in balloons and stuff.

1:20:41Yeah, yeah. Or out of natural gas wells. Yeah. Sorry, now we're talking about helium. Yeah, yeah. So, interventions, what are some of the most exciting, interesting ones? Well, I'm very excited by fusion. I mean, I don't know if it's intervention. That's sort of a source of energy. Because if we can make it work, and we can make it be sort of low enough capital costs, that will actually help a lot. Because at least the current models there, renewables are great. Ideally, you'd like to electrify everything, right?

1:21:11Which has problems, because you can't electrify flights. But you could try to electrify a lot of stuff. You know, there are EVs. You'd have to figure out how to electrify things like cement or steel. Those are hard, especially things like making steel want reduction power anyway, to essentially, you're adding carbon and you're reducing iron ore. So, there's a lot of sort of things that are difficult about electrifying everything. But if you could electrify everything, then the amount of electricity required would grow by a factor of five. And you could try to grow renewables.

1:21:46Renewables plus battery, again, trying to squeeze all of it out, it starts getting ever more expensive, because you just need ever more, you need like a huge number of batteries to sort of cover the last, you know, few percent. Or even, you know, 10 or 20 percent. So, we do need some sort of power that can cover the last 20 percent. Something that's, you know, base load. So, fusion might be a thing for that. So, that's super exciting. Again, there's no one sort of silver bullet that can sort of cover all the cases. So, I'm happy to sort of talk about any specific case.

1:22:18But it's sort of like, the world is a very complicated place, and the global economy is a very complicated place. So, it's super hard to sort of talk about sort of interventions in general. Maybe instead of interventions, one thing I'm curious about is, how does this make, affect decisions into, for example, like, what do we build? How do we build? I think you're from L.A., right?

More from Latent Space

Jev: System One models for Prod, not God — with Diogo Almeida, CEO, TypeSafe AI

Sep 21, 20262h 20m

Underwriting Superintelligence: Backing Agents you can Sue — Rune Kvist, AIUC

Sep 16, 20261h 26m

Humanity’s Last Invention — Richard Socher of Recursive

Sep 14, 20261h 32m

🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing

Aug 26, 20261h 23m

Simulation: the new Scaling Law — Joon Sung Park, Simile AI

Aug 21, 20261h 9m