Steadcast
Science Friday cover art
Science Friday

Creating 'world models' for robots + An AI math shakeup

August 19, 202617 min · 3,074 words

Show notes

When you ask an LLM like ChatGPT or Claude a question, the model goes through its massive amount of training data and guesses the answer by mathematically predicting the word most likely to appear next in a sentence. This model, experts say, will not work well for technology designed to navigate the physical world.

Highlighted moments

If you are taking things from one box and putting them into another box, the box size never really changes. They're usually located in the same exact spot. And so this is really simple to program robots.
2:25
Turns out you have to really, the footage has to show both hands in the frame, and you've got to kind of really show it very carefully.
4:54
People should not believe the hype about a humanoid robot coming to their house right now. It is not coming right now.
9:11
the announcement came in the form of a 250-page PDF entirely written by AI, so not with any human mathematical input that we can tell, together with some further verification of each of the 10 solutions in the form of something called a lean proof.
12:35

Transcript

Why AI needs physical world models

0:00Hey, I'm Ira, and you're listening to Science Friday. When you ask ChatGBT or Claude a question, the model goes through its massive amount of training data, makes guesses on the answer, mathematically predicting the next most likely word to appear in a sentence, for example. This model, experts say, will not work well in a world that's not based on textual information. Instead, they argue, you need a completely different approach,

0:31a world model that can understand the physical world around us, stuff we walk by or we bang into, and enable humanoid robots to work in warehouses and improve self-driving cars. It's an area of AI development that's raised billions over the past few years, with big players like Amazon and Google's deep mind, showing more and more interest. So what does a world model mean, actually? And how do you train the AI to recognize what the real world looks like?

1:02My next guest has seen the early days of these models up close, even in her home, and is here to tell us how ready for prime time they are. Joanna Stern is a tech journalist who writes newsletters and creates videos for New Things Media. Welcome to Science Friday, Joanna. Hello. Thank you for having me, Ira. You're welcome. Okay. Tell me more about why tech companies want this different approach to AI. Well, if you think about the physical world and the type of AI we are thinking we're going to get

1:34in the physical world, best examples would be self-driving cars and humanoid robots or robots of any kind. You can see why just studying the text of the world and even just studying basic images of the world is not going to be enough for a car to get you from point A to point B, or it's not going to be enough for a humanoid robot to step into your house and start doing the laundry. So a company like Amazon that has to move a lot of stuff around could really benefit from this kind of robot.

2:05Sure. And Amazon has over a million robots already functioning in some of their warehouses and in their partners' warehouses. And so the big difference there, though, in terms of what Amazon's talking about in robotics, is that those robots are largely industrial work-type robots, right? They have a very, very regimented, simple environment that never changes. Think about it. If you are taking things from one box and putting them into another box, the box size never really changes. They're usually located

2:40in the same exact spot. And so this is really simple to program robots. That's very different than, say, a self-driving car where, sure, you're still going from point A to point B, but a bicycle just went in between that road that wasn't there before, or they just set up a road closure sign that wasn't there before. And so these things are constantly changing in our physical world. The home is even worse, right? Ira, I mean, I don't know about your home, but my home is changing every day. Someone moves something in the refrigerator. Somebody moves a plate. My kids

3:12have put their toys all over the ground. And so is video very helpful? Is that what we're talking about here? Well, so there's some debate about how it's going to be best to train these robots. One way is through something called teleoperation, where the robot actually is in the home. There's cameras and sensors on the robot, and somebody is controlling that robot with a virtual reality headset someplace else, right? They're wearing arm controllers, and they are controlling, basically as a puppet, the humanoid

3:45robot that's in the house. And I've seen this demoed a number of times now. So the robot isn't doing it autonomously. And so this idea of teleoperation is that we are training and gathering data of the movements of the robot. We're also gathering video and sensor data from the robot. And that over time can start to train. There's another idea that's becoming pretty popular, and that's we use video data. We take lots of video off YouTube. We take lots of video from humans doing this in the real

4:15world, and we train it. And in fact, last month, I did a story on a company called Micro AGI out of Germany. And they have a new app called Shift. And you can say, I want to earn money by filming myself doing things in my house, doing the dishes, vacuuming, cleaning the counters. Don't pay me to do these things. They will pay you to do these things. The catch is you have to wear an iPhone on your head, some sort of smartphone on your head, or now they've started to make hats with embedded cameras.

4:46So I tried this. I tried it for about three weeks. I will say I was not very successful at it. Turns out you have to really, the footage has to show both hands in the frame, and you've got to kind of really show it very carefully. But look, I did make some money folding my laundry, putting my dishes in, which is more to say that, you know, look, it was, I would say everyone in the house was happy that I was finally contributing. Well, actually, watch your video. I saw your video

5:18doing that. And it looked very interesting. And, and, and of course, as you say, to train the robot, it has to see physically what you're doing with your hands, so that it could tell the robot what to do with its hands. Exactly. The big thing, and I kind of challenged the CEO on this, I said, well, look, I filmed five hours, you're only paying me for three hours. He said, well, yeah, because only three hours is useful to me, right? Only three hours showed my hands both in frame and my hands doing something useful that they think they can then go train the robot on. And in fact,

5:52they sent me the footage where they apply sort of the, the 3D modeling to my hands. And you can see, okay, the robot has to see these weird kind of like almost skeleton-like lines in your hands doing things for it to start to make sense of what is happening. You know, if I were to go backwards, engineer things, if I were a Luddite, I might say, you know, instead of a robot trying to do this, my 15 year old will do this for next to nothing, don't have to teach anything. Why go with a robot on this?

6:24It's true. It's true. But look, I will say there's many times in the week where I have a big pile of laundry that I need to fold. And I would rather do something different. Last summer, I actually had a company come to my house and set up their laundry folding contraption. And that wasn't a humanoid robot. That was really these two robotic arms. You can kind of picture, you know, when you go to an arcade and they have the claw machine that you can put the money in and it will pick up a stuffed animal or some other toy. The robotic arms looked a

6:55little bit like that. And so these two robotic arms sat over the table, had some cameras. It was a big sort of contraption set up. And these two robotic arms just on their own were folding the laundry. And this, it didn't work great, but it was really great to see because I could see, first of all, it improve over time. And it gave me a sense of where we are at right now. I mean, it was only folding t-shirts. The company now says it folds more things. Sometimes it would take five minutes to fold one t-shirt, right? This gives us a good idea of where we are at in this moment. And look, this was

7:30last year and things have improved now, but this is, this is how the progress we need to see and why this data is so important to these companies and why there are companies specifically popping up that just collect the data to sell to the robotics companies. Right. And that brings me to, I'm glad you brought that point up because I'm sure there are privacy experts who are not thrilled about people sending thousands of hours of footage inside their own homes. Totally. And I thought about that, obviously, before embarking on the experiment I did in that

8:01video with that company, Micro AGI, though they gave me a pretty long list of the ways they're protecting privacy. One of them being that they blur out any identifiable information they might see in your footage. So, for instance, they sent me back a video where some writing on a t-shirt was blurred out. And I said, why is this blurred out? They said, well, we blur any text that we see. Okay. It was my kids' like, you know, Pokemon t-shirt, but that's fine. But these companies are taking steps towards

8:32that. But look, you're agreeing to do this. And if you think the trade-off of putting some cameras on your head and showing them your house is worth the $20 they might pay you an hour, well, that's your decision to make. How do you see this panning out? Do you think this model is going to catch on like the large language models did? Or do you think this is just the beginning? I mean, folding a shirt for five minutes has a long way to go. Well, that's one of the reasons I love covering this space right now. I feel like we saw this big breakthrough in large language models. We've seen how that really

9:05disrupted society and technology and all of the things we see playing out with AI right now. And there's a lot of hype about how this is coming right now. People should not believe the hype about a humanoid robot coming to their house right now. It is not coming right now. I don't think it's coming in the next five years. What we're watching play out is this need for more data and this need for better. We haven't talked about it here, but what we need is better and safer robots too. I don't want to allow a hundred pound robot in my house that can't stand upright, that has a risk of falling

9:37down. So there are a lot of things that have to happen in the next number of years. I think it will happen. I can't give you a timeframe, but I can tell you it's not happening this year or next year. All right. That's a good place to end it, Joanna. Good luck in your research. Yeah, I'm going to keep on. I'm keeping on this topic. I'm going to be, you know, with the robots for many, many years. And they'll remember when they're finally working, they'll say, this was the woman who told the world. So when the singularity comes, they'll keep you in mind. Yeah. That's right. Joanna Stern, tech journalist who writes newsletters and creates videos for new things media.

10:13After the break, we're adding up the drama happening in the AI math world. Stay with us.

AI models tackling math problems

10:18Now we're turning to news in the math world. In the last month, AI models have disproved a handful of conjectures. OpenAI announced that it had solved 10 longstanding problems with an advanced unreleased model. Some mathematicians said the results were impressive, while others said they

10:48were overblown and accused the announcement of sloppy attribution. Play nice in the sandbox, folks. So what do mathematicians make of all of this? We thought it was a good time to check in with our mathematical referee, Dr. Emily Real, professor of mathematics at Johns Hopkins University, who's joined us before to walk us through the AI math world. And she's back with us from Baltimore, Maryland. Welcome back, Emily. Thanks for having me. Nice to have you. All right, let's get into this.

11:18When we last had you on in March, you talked about AI also. Are things progressing slower or faster than you thought? I think most mathematicians would say this summer has been very surprising at the pace of improvement in the models. I think the first one that really caught my eye was a disproof of something called the unit distance conjecture, which was a 1946 conjecture of Paul Erdős, a celebrated Hungarian combinatorialist. I would say that there were relatively few areas where AI had made

11:56important contributions, and that is changing. AI is still not contributing to all research areas in mathematics, but in certain areas it is certainly making progress. It seems to do best in areas where the problems are easy to state, if not easy to solve, and maybe is less skilled in the areas where the problems are sort of harder to understand the statements of. Right. You know, like I said, OpenAI announced their new unreleased model also solved 10 longstanding problems. What's been the

12:29reaction from mathematicians, and what do you think of them?

12:35So the announcement came in the form of a 250-page PDF entirely written by AI, so not with any human mathematical input that we can tell, together with some further verification of each of the 10 solutions in the form of something called a lean proof. Lean here refers to a computer proof assistant that can automate some of the refereeing process. Maybe the first thing to say is that when a new mathematical

13:06breakthrough happens, it usually takes the community sometime to respond. And if the new breakthrough involves a delicate mathematical argument that, you know, might take dozens or even hundreds of pages, then it can take sometimes even a year or even more for the community to really verify the correctness of the solution and also understand the impact down the field. I'm trying to think of what a mathematical prompt to AI would look like. I mean, it must be huge, right? You might think that,

13:39but, you know, one of the surprises is that often these prompts are really short and relatively naive. A recent example is the Dinitz-Gorman's conjecture, which is a question in graph theory. And here, the prompting was sort of four extremely naive prompts, telling the model which problem to work on and to keep working until it found a counter example. So, you know, what do we make of this? I mean, you know, one phenomenon is that there are a lot of folks out there who love math, but are not

14:10getting paid to do math full time, who are kind of rediscovering their love of math by pushing the frontier of research knowledge in this way. You could imagine anybody who was aware of that problem could write that prompt and, you know, then generate a solution. So, in one sense, this is expanding the number of people who are contributing to the mathematics research enterprise. And we've seen this with the Erdős problems too. The Erdős problems were sort of famous for attracting

14:43non-specialists who could contribute in a productive way. And those non-specialists are sort of superpowered now with the latest AI models.

Strengths and limits of mathematical AI

14:51Did AI have a particular strength that allowed it to do this? There's been some discussion of that. So, there are some mathematical problems that essentially involve finding a needle in a haystack. And it does seem that AI, for whatever reason, is good at finding unusual mathematical objects with unusual properties. It's not as good as explaining to us how it found them. The sort of interpretation work for some of these counterexamples has been done by humans.

15:24Yeah. You know, my math teacher always said, show your work, right? Speaking of the latest AI models, are you getting a clearer picture on what the relationship between mathematics and AI models will look like in the future? I think there is a lot that is currently in flux. You know, one open question is how much AI can actually help us understand mathematics as opposed to just tell us what is true and false. You know,

15:56right now, it still seems like there's a lot of work for humans to do to sort of map out the mathematical universe. So, you can think of the mathematical universe as some sort of vast building of unbounded expanse that mathematicians have spent millennia both exploring and renovating. This building is so large, there's no architect who has seen the floor plan for everything. There's no contractor who's directing everybody to do the work. Instead, there are sort of specialists in different areas who are interested in different problems who are going off down some dark corridor

16:26and trying to open the doors that are stuck to get inside. So, when there's a door that's stuck, you know, a problem that we don't know immediately how to solve. It's not clear immediately how to get through. Maybe it just needs a shove in the right corner. Maybe it needs somebody to find a key. Maybe it needs somebody to create a new key. And what the AI models are doing right now is helping us get in some of these doors that mathematicians were having trouble breaking through before. But what mathematicians still have to do once they've opened up a new door is see if they can turn on any lights

16:58in there. Maybe renovate a bit to make sure the ceiling doesn't collapse. See whether this door leads to just a closet or maybe a whole entire new wing of the building where there will be new mathematical discoveries around the road. And right now, at least, AI is not contributing to those sort of larger enterprises of expanding the mathematical universe. Well, Emily, it's always fun having you coming on the show. Thank you for taking time to be with us today. Thanks very much. Dr. Emily Real, Professor of Mathematics at Johns Hopkins University in

17:29Baltimore, Maryland. This episode was produced by D. Peter Schmidt. I'm Ira Flato. Thanks for listening.

More from Science Friday

The ancient Roman doctrine guiding modern environmental fights

Aug 25, 202614 min

What are the mysterious ‘little red dots’ in images of space?

Aug 24, 202613 min

Could we live in a dinosaur world?

Aug 22, 20265 min

Do microplastics in the body affect pregnancy?

Aug 21, 202612 min

Making a case for physicists to join climate research

Aug 20, 202612 min