Steadcast
Latent Space cover art
Latent Space

Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI

July 28, 20261h 9m · 14,284 words

Show notes

There are roughly 100x more people who use code than who can write code. As code that “just works” becomes easier to generate, this group may be the biggest prize of all — if you can get the agentic interface right. A key trend we have been tracking over at AINews is the absolute explosion in Codex usage this year, with MAU now up >10x from Jan 2026.

Transcript

0:00Okay, we're here in the studio with Akshay from OpenAI. Welcome. Thank you. And with our trusty co-host, Vibu. So you recently launched ChatGPT Work. You lead core product engineering. You know, it's been a long journey into all this. I find it very interesting that you started with no code or low code with Walrus and Airtable. And to some extent, ChatGPT Work is kind of like the super app of super apps of, well, here is the ultimate no code.

0:30You just write a prompt. Yeah, yeah. It's funny how things come like full circle. I mean, I think for a long time in my career, I mean, I started my career working consumer fintech, but then after that, like, there's this hypothesis that, you know, the things that we were able to do with code, like, as engineers, like, if we could bring that to many more people in a more accessible way, then that would be truly magical. We were working on a startup. It's actually funny, like, before LLMs, before Vision LLMs, on how to do automated testing with AI.

1:01And it was just kind of jank back then, but, you know, doing what we can. And then worked at Airtable for a while on, you know, the same thesis that, like, if we can bring a database or the primitives behind a database to people, that would be really useful to them. But once, I think, LLMs came onto the scene, it became clear that, like, this was, like, the missing piece, like, the missing technology required to, like, bring the magic of code to everyone without them having to know what's going on underneath the hood. And so, like, I think this launch and, you know, a lot of the stuff that we've been up to is, like, the manifestation of that.

1:33How was stuff when you joined? So, you joined OpenAI 2023. Now we've got, you know, so much more stuff. So, ChatGPT, CodexApp, ChatGPT for work, have things changed? I actually think the more interesting thing is how things haven't changed. Like, I guess, like, one, I joined, I remember when I joined, it was, like, 500 people. One thing I was worried about was, like, I was looking for something, you know, more early stage and, like, was it going to feel startup enough? And I joined, I was, like, this feels even more startup-y than I could ever imagine. And, like, that really hasn't changed even until now.

2:04I mean, I think the, like, level of, like, bottoms-up ambition and, like, the ability of anyone to, like, you know, do anything or have an idea and ship it is really cool. But on the, like, sort of mission side, I think what was really compelling to me is this mission of, you know, bringing Frontier intelligence to everyone, like, building AGI and then bringing it to everyone. And I think technology even back then that, like, that vision is going to, you know, not be a linear progression. Like, we're probably going to, like, try different products and have different things that succeed and don't. But the vision has stayed the same and the mission has stayed the same.

2:34And we're starting to see the pieces fall together. And that's really cool. You worked on Enterprise. What, a lot of people never touch ChatGPT for Enterprise, God. What is something that you learned from there that you're bringing into your work now? I think how there's no, like, one-size-fits-all solution in Enterprise. I remember in the early days of ChatGPT Enterprise, like, we would talk to customers and, like, everyone, that was, like, when, I think it was a year after ChatGPT was released.

3:04And everyone was so excited to bring, you know, AI into their enterprise. And, like, there's all these teams that were being stood up as, like, you know, the AI deployment team, these enormous budgets. And if you asked anyone, like, what were they excited about? Like, what were they excited about? It's all I think, like, at first you'd get, like, you know, kind of, like, the baseline answers of, like, we have all this context and data and all this stuff. But then if you ask them, like, you know, what was, like, a discrete use case that, like, they want AI to enable in their workplace? You get such a different, like, variance, like, explosion of different types of answers.

3:35And it's interesting, like, you know, you using, like, these models and these products, you have this box and you can say anything to it, which is the magic. But on the flip side, it also means that, like, you don't know what to do with it. And in enterprise, I think a big part of that is, like, actually meeting the users where they are, like, what use cases are they trying to solve, and then actually teaching them how they can use AI to, like, gain leverage there. Do you meaningfully differentiate that from forward deployed engineering? Or, I think there's, like, the, like, go-to-market side of it. Yeah. And then there's, like, the product side of it. I think you need to see more in the product side.

4:08And I think, like, however good we get at FDE Motion, like, I think at the end of the day, if we have a user who's, like, looking at their computer or looking at their phone, like, it's our job in the product to, like, be enabling them and showing them where to go. So we're really excited about that. Do you think there's been changes, you know, over the past three years of adoption? So there have been, you know, step function changes, you have reasoning models and whatnot. Is there still the same problems of enterprise has black box, don't know what to do with it, or have things changed?

4:39I mean, we're seeing now that, like, there's this huge uptake, right? Everyone's extremely excited about it. It feels like, you know, many people, like, millions, hundreds of millions of people are using ChatGPT. They understand, like, how generally to work with AI. But then, like, every time, like, a new capability gets unlocked. So now, like, we're seeing with agents, like, there is probably a contingent of, like, early adopters still who, you know, truly get it, who are, like, you know, you can do anything. You just have to make sure the right context is there, it's connected to the right tools, and then that you're supervising it, but, like, anything is possible.

5:09But then there's, like, this, like, 10x or 100x bigger market, or, like, they don't yet get that or they don't yet see that. And so I think that's the next stage here. So I guess to answer your question, like, I think the adoption is there and growing fast, but I think the opportunity is, like, far, far bigger than that. That's where we want to play, especially with ChatGPT work. Yeah. Well, let's skip ahead to ChatGPT work. Only, like, a month ago or so announced, what was the sort of decision process that led into it? You know, there was this overall merging

5:40of the super app. Is that what we're officially calling it? You deprecated the browser as well. So just, I guess, summarize your last, like, couple of months of working on this thing. Yeah. It feels like forever now, but I guess it's only been a few months. I think maybe the one, like, impetus that, like, is most salient is when we release codecs or even internally add codecs. Like, it was really surprising to us. I think we recently put out some stats on this, that there was this, like, real inflection of, like, adoption among non-developers

6:11at OpenAI. And I, you know, through this product development process, like, would go to, like, these UXR sessions to talk to people internally. And the thing that stuck out to me is, like, one, like, you know, you go talk to, like, strategic finance or marketing or whatever, and they're all using codecs for, you know, their use cases. That part's cool. But the thing that really stuck out to me is how proud people were that they were using codecs. Like, how, like... It's like, I'm not supposed to be using it, but I am. It was that. It was, like, that they were, you know, early to this, like, new thing. But it was also this thing of, like,

6:41they felt like they had a superpower, right? And what we recognized then is that, like, the power of codecs, the power of agents, like, we already had this massive distribution base of people who have, you know, come to know and love ChatGPT. Like, how do we show that to them? Like, how do we bring it to them? Which is, like, a hard product problem. And it's, like, a tricky thing, right? There's many ways you can go about it. And so that's what we call the merge and the super wrap over time and ultimately launched in ChatGPT work is how do we do that? But it came from that initial realization that, like, the power was not only for developers.

7:12Like, much, much earlier than probably even we thought. Like, it could be extended to everyone. How do you see the products differently? So, like, who is it for, right? So Codex started out, even CLI, then app. Now there's a merge of ChatGPT, Codex, and ChatGPT work. So is it the opening for the average user, for enterprise, for work? How do you position it? I think we want to get it, to position it for if you're doing worky-related things, for lack of a better word, right?

7:42I think productivity is, like, it's actually what, like, the pillar that I support. Like, that's the name of the team. And the reason for that, the reason we call it productivity and not, like, you know, enterprise or, like, work or something like that is because there's also personal productivity, right? And, like, I think ChatGPT work is, I've seen people do things in their personal lives that you wouldn't classify as, like, work technically, but, like, these agents are, you know, super capable for. Like, one recent example that someone posted about on our site is, like, someone has, like, a missed package. Like, they didn't receive it and then they got, like, the picture of it,

8:14you know, from Amazon or wherever the courier was and they, like, asked ChatGPT work to, like, find out where that package is and, like, the agent, you know, is extremely tenacious and, like, took the image and, like, looked at a bunch of, like, listings around their neighborhood and figured out exactly the apartment complex in which the package was that gave them some information. And so, like, I think there's all these things that, like, you, you know, work-y or productivity-related things. I think that's what we want the product to be. You asked about Codex. I think we think Codex is, you know, a durable brand, but we have a principle that, like, the user,

8:44you know, we don't want a user to get stuck in a tab or an experience where they don't get the power of the product. And so, like, basically everything that you can do, you know, in the Codex portion of the product on desktop you can do in ChatGPT work and vice versa. But we made some opinionated product decisions on, like, you know, how much of the git state, if you're in a git repo, do we want to expose to the end user? Or how much do we want to make the experience of seeing the agents thinking, like, diff forward so that you get exposed to the diff side of the back? And then, like, on the safety side, like, how do we want to think about, like, sandboxing and making sure

9:15that we have the right defaults in one state versus the other? So, there's, like, some opinions that go behind that, but we do want, we don't want the user to need to choose which experience they're in. That is a good goal for AGI, right? Like, people don't want, like, to hide, to choose what version of AGI they want. They just want the AGI to decide for them. Can I get an answer, or, like, it's not super clear to me. Is the Codex harness and the ChatGPT work harness the same? Is it just UI affordances, or are there actually

9:45prompt level or even deeper differences? So, the harness is the same. The harness is shared. In both of the products, we made improvements to the harness to make it good for knowledge work, especially as it relates to plugins or computer use or artifacts. You get that power regardless of what your experience you're in. On the UX side, there's opinionated takes that we have when you're in Codex mode how the UX should behave and some stuff around the sandbox like I mentioned, but the underlying harness and capabilities should be the same. I think I'm just kind of curious,

10:17maybe we can, is there a query that we can run that would look different in the two modes? Yeah, I try to create, like ask it to create like a retirement calculator or spreadsheet or something in both modes. And then in Codex mode, you might have to be in a repo for this, but you'll see like the diffs of like the sheet that it's creating and stuff like that and the file edits, but in RRQ you won't be able to see that. I think that's super clear. And then also, the other thing I wanted to dive into was the productivity team.

10:48What else is there? First of all, what are the top level teams other than productivity? Isn't productivity everything? So, you know, we have a team focused on ChatGPT, like the core chat experience for a consumer, which is like, you know, not, I think, all productivity. Like there's, people are using ChatGPT every day for search to, you know, figure out how to write messages to loved ones, to think about how to like learn a new topic, et cetera. And so,

11:18there's so much more inside to create images. There's so much more in Chat that, you know, the hundreds of millions of users are using that, you know, obviously that warrants like a very dedicated effort. And there's teams focused on enterprise and infrastructure and API and stuff like that as well. I will bring it up. Yeah, so I have them both running. Yeah. This is work. There's a codex version here. I picked 5.6 Sol, so this will take a while. I think, I think we'll just keep it in the background and, you know, as, as they finish, we'll look into some of the differences. Yeah, but immediately, I think if you flip back

11:50to the codex version, you'll see that, uh, it assumes, it assumes Git. Exactly. Like Dynamic Island assumes that you're in a Git repo. and you might miss some stuff because some of it is like in the actual chain of thought with those changes and how we display that. Is there an unintuitive, like, is there a thing that you wanted to ship and then you got feedback and you were like, no, let's not do it. Like, what's the thinking behind that? In, uh, try to be your work? Yeah. I think one direction we could have gone with this is like keeping the experiences like completely separate.

12:21So it's like, why? Different apps. Exactly. Like different apps or even in the same app, like different, completely different experiences. Like why merge at all? Like what is, you know, Codex obviously people love. Like why, why bring these products together? And I think the intuition here is that like all of our jobs are like changing dramatically with AI. Like for the, you know, every few months, like I feel like I wake up and I'm like doing a completely different thing than I was doing a few months ago. And my, my hypothesis here is that, or I should say our hypothesis is that like part of what we're, we're building in this, this technology is giving people leverage

12:52like, you know, the things, maybe it's a more mundane parts of your job or, or parts that like if you were able to automate, you'd be able to share more ideas faster or whatever, like you're able to do now. And because of that, like that might actually blur the lines between someone who's like only writing code or creating strategy docs or, you know, planning events or helping with marketing or doing podcasts or whatever, right? And so like these things are going to get blurred over time. And so like trying to draw a hard boundary based on like the who you are is going to be, is going to be tough. And like we should enable users to choose, but we shouldn't box them in.

13:24And so a lot of the work that went in here, like, you know, keeping the primitives the same, like for example, plugins are like unified across this product and ChatGPT in the cloud was because of that. It's this thesis that like eventually things are going to come together and, and we don't want to be like, we want to be prescriptive about when to be in either experience, but we don't want to box anyone in. I wonder if there's users who are very tuned to the old ChatGPT harness that is effectively now replaced by the Codex harness. I can't imagine what that was,

13:54but maybe they're more, the more conversational side. Can you compare and contrast the two harnesses because only you've seen it? Yeah. I mean, I think ChatGPT, the, the existing harness like still exists today. It like exists in this app. The classic, right? You just start a new chat and you don't go under work, right? Yeah. If you start a new chat and go to chat, then you're talking to ChatGPT with the Instant model. I mean, we can technically do another. I guess, you know, on Instant. Yeah. So this one's not going to code or it's going to be inline. It's not in a sandbox.

14:26Actually, we try to push you to go to work if you're creating a spreadsheet. This is a router decision. Sorry? It's a router decision. This is the decision that, you know, the model is making and then like, you know, it sees that you're able to, or you're trying to do something that would be better served in work mode. But I think your question was like, what are the advantages of like the chat, like ChatGPT chat harness? It's more broadly, like I want to basically do an oral history of harness engineering. Right? You know, the ChatGPT harness lasted us from,

14:57let's call it the 01 era until now. And now it's being replaced by the codex harness effectively. And they're overlapping somewhat. But I'm curious what changed if, you know, if there is. My perspective on this is like there's, there's sort of like a constant process of like divergence, convergence, divergence, convergence. And in chat, like many of the use cases I was talking about before, like, you know, search or learning, I think we're really optimizing for latency and optimizing

15:28for personality and like different things that over time, like the product, the reason people love ChatGPT is because we've been optimizing for those things and working on them for so long. Codex, what we learned was that like if you give the agent access to this infinitely flexible environment as a computer, you can do really, really powerful things. And so when we think about like, okay, well, for knowledge work, like what is, which mode should we choose? It was like, it felt more natural to us to bring that to this like

15:59computer environment and, you know, maybe abstract some of the details of this computer away from you. users who might not be used to that, but like give them that same power. But ultimately, I think that we want the power in all places, right? We want to meet people where they are. So I'm sure there'll be work down the road in order to get things to be equivalently capable in all scenarios. But it's just a question of like what we've been focusing on the product on historically and what we're focusing on now. I think alongside that, outside of just harness and when to use Codex, ChatGPT or work,

16:29there's also the new models you've released, right? Any guidance there? So people love to min-max what to use, like only use Terra on high reasoning versus for this, you know, you want to use Sol here, ignore all these. There's 32 options. Yeah, yeah, yeah. But that being said, you know, for people that are expanding, so productivity, trying stuff for work that don't have the breakdown of what all this is, what's the advice, right? Well, I mean, I think before the advice,

17:00like the first thing is like none of this would be possible without these models. Like I think you asked earlier, like, you know, what was like the inspiration for work and like, you know, early on, like I mentioned like what we were seeing with Codex, but that was also because the models were getting infinitely more capable. That's happening again. I think it's like another step function jump now. And to answer the question on advice, like we want this default to be the best possible, like we want to be opinionated about the default. And so we've chosen a default that we think is going to be the best for everyone. And, you know, we have for power users

17:32options under the hood. We could, one could argue that there might be too many right now and we're, you know, working on simplifying it. But you can extend, you know, the reasoning level and you can change between the different model classes if you need to, but the default should be the best for most use cases. So my advice to most people would be to stick to that. And then, you know, if you reach a situation in which you think that you could, you want to try a different configuration if you're not seeing either the efficiency on the cost side or the quality

18:02on the intelligence side, then you can change the defaults and see if you can get something better. But we think the default should be good enough. I have, I'm just going to run something by you since you have way more experience than me. I've recently been doing so light but with goal. With the idea that the goal basically augments the reasoning effort but with more terminations and turns. Is that a good way to think about it? As opposed to so ultra or so, you know, extra high. Yeah. It's hard to say because it's like an interaction effect.

18:33Exactly. It's like there's a preference on, you know, for you as an individual, like how do you like to collaborate with the models? Like how many of those like terminations as you call them do you want where, you know, you can steer or make sure that it's doing the right thing? I think generally people should try whatever works for them. I think that like using ultra or the like multi-agent setups are best for like when you have like tasks that are either incredibly complicated, like open open expirations

19:03or very parallelizable. I think even for tasks using goal, I think is best for tasks that you know that you'll be able to make consistent progress in a way that's verifiable over time. But I think for most tasks they actually don't fall into either of those buckets. And so like at least when they're starting and so that's why I think the best first step is like trying it with the default configuration and then seeing like where you want to go from there. Right. You guys worked on a slider which actually is super helpful for reducing the amount

19:34of panic. It's nice on mobile at least. There's a nice slider. I haven't tried it. So you have the advanced view there but if you click advanced view, yeah. Yeah. Very pretty, very colorful. Yeah, that idea was to reduce it to like one dimension even though there's multiple dimensions right? Try to project it onto a single dimension for the user. You have like you know something from that represents like you know speed and efficiency on one side and then like sort of like quality and thoroughness on the other side. I am just puzzled that it

20:05uses Sol so much like the lower I think the slider if I'm not mistaken is Tera. Oh it is. So they preset Tera to only be the light one. But like I think a lot of people actually more people should use Tera. One because Sol keeps running out of capacity.

20:21I'm the reason you know. Here's 10 minutes of our retirement calculator. Oh that's the Excel thing working for you. Oh my god look at that. This is work and then Codex is still cooking so we'll get back into it. I think it'll be interesting to actually see the thought process, the reasoning and also you know I guess this is eight minutes on work Codex is still cooking. Yeah and by the way do you know Gabriel Chua he's part of the Open Hestime Port team he showed me this and I was like pretty shocked that this looks like Excel.

20:52It edits Excel files. You never pay an Excel license. Right? But somehow this is like kind of workable and it's agentic Excel. Yeah I mean one of the big pushes that we made for this launch was like artifacts. Both on the model side like I think if you compare this with 5.5 and 5.4 before that you'll see that there's been pretty dramatic improvements in the quality of these artifacts and then also on the product side. The UX side is also crazy. Like hosted sites and whatnot no longer needing to host your

21:22own little web page. Oh I have a story about that. I can do a separate thing. I'll need to take the visuals here but we'll cut to that later. Was there co-training I guess because you were making this big move and you launched 5.6 on the same day as ChatGPC work. Was there influence between the model training teams and the harness teams or did the launch days just happen to line up with the same day? I think we collaborate heavily with the research teams and I think that's one of the

21:52most magical parts of the job. The most fun parts of the job. Just using artifacts as an example a lot of what you're seeing underneath the hood there's a lot of work that went into making sure that we had the right infra to be able to train the models to get better at this and on the product side had the right experience for users to be able to collaborate with the model on an artifact like this. It's not necessarily that you wouldn't need an Excel license. This is stage one. This is probably not what you meant when you're making a

22:23retirement calculator. When you're seeing it and if this thing is high fidelity to what your co-workers would see if you were to send this to Sean that I think makes it so easier and makes you trust the product in terms of iteration. When you say co-workers would see do you see a multiplayer multi-team collaboration with artifacts? Any things you guys think about? You can already share it, right? Yeah. It's something that we're actively thinking about. One thing that we've noticed internally without talking

22:53too much about the roadmap is that there's many times when someone will ping me about something and I'll ask the question and then I'll ping them back the answer. The simplest would be the three of us are just all on one hosted. Exactly. I'll think about was I required in this loop? Maybe it was. I'd rephrase what they were asking or pulled from certain context or whatever but when I gave them back the answer, that process was also lossy. I gave them my interpretation of what Chachibiji worked cooked up but

23:24underneath the hood there's so much context in the rollout and stuff that could be interesting. So the answer was preemptively respond to every inbound request? No, it's just literally this is what I do sometimes as my job. I know, you copy paste and you're just a message forwarding service from AI to AI. I think it's interesting, right? It helps people understand the capability of what you can ask and delegate that oftentimes people don't realize until they try or someone shows you and then you're like, oh, okay, okay. I think there's also a light

23:54security issue where basically you're the permissions layer. Yes, I could query everything that you query and I could get an automated response but maybe I'm not supposed to see it. There's no way I would know because I'm not supposed to know what I don't know. Especially as which IWD works for asking you to connect your plugins and it's pulling from your local files and stuff like that. The amount of context that the agent has access to is deeply personal and that's something we need to preserve. So that'll be a challenge. There's Excel, there's PowerPoint,

24:25there's Docs, the grand trio of work. What other formats of work do you think about? Obviously you worked on Airtable. Is there a future where there's

More from Latent Space

The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten

Aug 3, 20261h 41m

Inside the Model Factory — Eiso Kant, Poolside AI

Jul 23, 20261h 54m

🔬Causal Models Need Causal Data - Xaira’s X-Cell model for Drug Discovery (Bo Wang & Ci Chu, Chief Discovery Officer & Chief AI Scientist)

Jul 21, 20261h 29m

🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences

Jul 16, 20261h 41m

Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO

Jul 8, 202657 min