Transcription
2024 is the year of coding. I think 2025 is the year of coding in the enterprise. Like I'm starting to see like the meaning for adoption. Yeah.
>> What's 2026?
>> I don't know man. Like uh here's a paper I just released. I need to deeply pro uh I think something like 30 minutes like to achieve it. The author of the paper was telling us like it was like weeks of work.
A new model like five or 51 come out. What does that like process actually look like and how do you imagine that changing or evolving over time? The days where you could just like hot swap essentially like one API parameter from like you know the one model to the next are basically gone.
>> What do the really good top [music] teams do when like they're experimenting with a new model?
>> That's probably to me the biggest earnings like of the past 2 years of like you know building for like developers businesses is hey Olivia Goodman is the head of products for enterprise at OpenAI. I'm Jacob Efron and today on unsupervised learning I got to ask Olivier all my top questions. We talked about the progress he's seen in AI for science and how open AI models are helping there. We talked about the scaffolding and harnesses that are emerging around models, the patterns he's seen in different spaces, and the extent to that he thinks it will converge, as well as have foundation model providers provide this for startups. We talked about his reaction to Andre Carpathy and his comments about agents still being a decade away. And we also talked about the next frontiers for models and what Olivier thinks will be most important to unlock further value in enterprises. Olivier also gave a lot of really great tips on what the best enterprises using OpenAI do to adopt new models as well as fully get the most out of these products. Just an awesome conversation with a really brilliant mind. I think folks will really enjoy it. Without further ado, here's Olivia.
>> Well, thanks so much for coming on the podcast. Really appreciate it.
>> Thank you for having me.
>> And it's extra fun to get to do it in, you know, opening IH HQ.
>> I know.
>> I'm probably the the least qualified person to ever sit in like the demo day launch seat. And so it's uh it's it's fun.
>> You're in the in the model ship. [laughter]
>> You know, lots of things to discuss today, but I figured maybe one place to start would be with, you know, 5.1. Obviously, you've got a new model. I think you know some really interesting improvements on the personality side on reasoning as a whole. Maybe just talk about how how do these things actually come about these sets of model improvements and what do you you've seen so far that you're excited about and how people are using it.
>> Yeah, totally. So, we shipped a bunch of new models last week, DP 5.1 and DP 5.1 Codex for coding. Um those models basically were um trained based on the feedback of DP5. When DP5 came around, people love the intelligence. People love like the ability to follow instructions, the streamability. People didn't like the speed. Like the model was really good like you know when it was like thinking for a long time but like for more basic queries GP5 was too slow and so that was one of the main sort of design goals of DP5.1 like keep that intelligence but try to you know compress like the thinking tokens like you know as much as we can in order to respond like you know way faster. Um I think with that regard like the goal has been achieved. uh we're seeing like people like you know switch like pretty seamlessly between like you know low thinking effort for like you know fairly basic queries and like larger thinking efforts for you know more complex queries. The codex model has been like gaining quite a bit of adoption as well. Uh I think Codex is having like a moment at the moment.
>> It definitely seems like it's having a moment.
>> like you know developers are loving it. We're starting to see like you know more and more like frequent use like you know we're seeing companies as well adopt Codex uh in
>> Did that surprise you or like had you played around with it enough internally to know this was coming?
>> It's probably the model that we dog food the most internally just by the virtue of like you know every single like software engineer like researcher open AI like use or has used codeex. Um I think the latest stat I think something like you know engineers like on the team like are able to push something like 70% like more um because of codex. So yeah it's became it's become clear like you know and sort of in part of the tool stack at
>> Yeah. Yeah. Um, no, I mean it's been uh it's been impressive to see and I guess any like you know that kind of improvement in latency and know the improvement in the overall codeex models. Any like new use cases you've seen unlocked or like fun ways that you've seen uh folks start to use these models?
>> I think most of the use case were pretty much what we expected but better coding productivity any like knowledge like query like you know retrieving information customer support customer experience same but more essentially I think the one domain which has surprised me quite a bit over the past few months frankly is usage of deep and 5.1 among the scientific community we're seeing more and more reports of scientists and researchers use um LLMs in order to perform like their job. Um that job meaning like you know condensating aggregating like scientific like literature and knowledge just get faster like test hypothesis. It's just been like a really exciting because it's funny like when I join a couple years ago like that was always like one of the goal of the company is like hey how cool would it be if you could like accelerate scientific research. I would say for the first time in the past few months I'm actually feeling that it's happening. Of course it's early like you know and you know we have way to go but for the first time we're seeing like pretty stellar like you know scientists tell us like yes I have done my job faster you know I have done like you know um achieve the discovery that prove like you know faster because of
>> how much faster like do you think this is you know anecotally as you talk to these folks like you know what kind of speed up do you think we are at now and is that like an eval that you that you you know think about hill climbing on as these models get better.
>> Mark Mark Chen our um chief research officer um was working with that uh physicist working on black holes uh and the physicist like gave it like a really hard task which was hey here's a paper I just released so it's not in a training set or anything just released try to essentially reproduce like the math um and it took deeply 5 pro I think something like 30 minutes like to achieve it and you know I'm not a physicist I cannot judge but a physicist like you know the author of the paper was telling us like it was weeks of work like you know for a professional physicist essentially to achieve that. So um I think we're starting to see some clear like glimmers of like you know acceleration in that work.
On the one hand you know coding support science it feels like there's been these really exciting overall progress in in models and the abilities for uh you know them to do all these things and then you know I think there's been this public conversation about you know where are we with these models? There was this big moment when Andre Carpathy went on Doresh's podcast and uh you know said basically he's like look you know the industry is making too big of a jump trying to pretend like some of the stuff's amazing some of it is slop. It'll take I think you said at least a decade until AI can meaningfully automate entire jobs. I'm curious like your rea like you're every day working with enterprises like and seeing what these things can and can't do. What was your reaction to like that overall take from from Andre?
>> I mean it's no secret that like building a really good agent is hard. like we haven't reached a stage where you know we just take like a day essentially to automate like you know any like you know white color job I think we're starting to see like some quite strong automation use case in like a few specific fields coding is the one that comes to mind um I think at that point we reach the point where you know if I were to take away coding tools um AI coding tools from you know software engineers otherwise like there would probably be a riot or something like you know people would like polish jobs essentially so that stuff is happening I would say the model the like the automation is probably not yet at the level of like automating completely the job of software engineer but I think we have like a line of sight essentially to get there we talked about customer experience like you know sales customer support we're starting to see like fairly strong cases of adoption um I've been working a bunch with the folks at T-Mobile the telecom company in the US to essentially provide like better experience like to their customers u and we're starting to achieve like fairly good like results in terms of quality at a meaningful scale so but that stuff is hard like you know on top of the model like having like a really good model that you train like you have to build like really good harness like you know how do you connect the model to tools like how do you p the model you have to build like really good like evaluations framework um and then you have to build like you know some sort of a flywheel like human in the loop like to constantly improve essentially that model harness it's a lot of work um but my sense is you know uh we'll probably surprise in the next like year or two on like the amount of tasks that can be automated reliably
Are there any that like you feel, you know, any industries or end use cases that you feel like are are on the cusp where you're like, you know, I'd be surprised in a year or two if there wasn't a lot more activity happening?
>> Oh, my bet is often on life sensors.
>> Yeah.
>> Uh for my companies. Uh so I've been working a bunch with Amgen and a few others. It's really interesting. Um essentially when you ask Men like, "Hey, why do you exist as a company?" Their goal is to design new drugs. Essentially design and like develop new drugs. And there's roughly like two big chunks of work that goes into designing new drugs. there's like the actual R&D experiments validation like you know scientists and then there's more admin work the admin work is huge like the time it takes from you know once you lock essentially the recipe of a drug to having that drug um on the market month sometime years um and that's about like generating like really complex regulated documents having committees like review it having regulators review it so you know a lot of information sharing transformation validation which you know turns out like the models are pretty good at that. Like they're pretty good at aggregating, consolidating a sort of a like tons of structured unstructured data, spotting like diffs and you know uh different changes like you know uh in the document. Um so we've been working quite a bit with them and so my hunch is you know once of course like regulating industries so you know taking things like you know very carefully but once we figure out like you know a way to properly um version audit um release like you know the model with like the right permission internally we're going to see like you know a bunch more adoption the life sensors and you know the outcome will probably be like more medication like new drugs like essentially for people which is pretty cool.
>> Yeah. No, I'm struck by obviously any industry that's heavily regulated and requires lots of paperwork and you know uh submissions from you know other sets of documents is it's a pretty perfect uh LLM use case.
>> Yeah. And there are some like all over the place. I was meeting with a a pretty big investment firm the other day. Um similarly like you know their job is to aggregate analyze like mountains of data like in real time and try to make sense of it essentially to make judgment calls you know. Um and again like models like superpower is to you know do like in a split second that's sort of a massive analysis and so uh we're starting to see in hedge funds uh investment funds like banks like quite a bit of adoption uh yeah on that front.
>> it's interesting one thing I'm struck by is so much of the application layer today feels like it's companies you know that go deep in some industry and they like forward deploy and they're just like okay what are your problems oh you have a bunch of documents and you're trying to do this workflow and they they build those workflows and you know even for all those examples you you provided T-Mobile uh you know Amgen um you know this this investment firm there's also like startups you know like Sierra might be wanting to serve that T-Mobile thing uh company like Colate might want to serve uh the Amgen use case how do you think about like when it makes sense for you guys to be the ones working super closely with the customers uh and when it makes sense for like you know the broader ecosystem.
Hey everybody, I'm Ann Maris Walton and I'm the new producer of Unsupervised Learning. Today I'm super excited to introduce a new concept to the show where we open up the floor to you. If you have any pressing questions about the AI ecosystem that you'd want to ask Jacob and an upcoming guest, please submit them using the Google form that's listed in the show notes below. We're super curious to hear about what you have to say and very appreciative of your support of the show.
How do you think about like when it makes sense for you guys to be the ones working super closely with the customers uh and when it makes sense for like you know the broader ecosystem?
>> Yeah, it's a good question. Um what I've learned over the past couple of years working on uh AI adoption among businesses and the enterprise is frankly the the the the size and the depth of the of the problem. Like [laughter] once you pick any of them like you know Amgen, T-Mobile, like BNY like you know any company like the amount of complexity and like use cases that you have internally in order to you know operate at that scale with that level of quality absolutely huge and so I think we at OpenAI have no um illusions that we're going to be like you know the only ones like to you know build like really good agents like create products like frankly like part of the mission is like to enable the ecosystem like third parties like you know to work better um with us um we're thinking on ways to achieve that in the product like you know description wise I think we we announced at dev day this year apps in charge which is one of the first I think like big vectors of you know partnership with ecosystem.
>> like the feedback we've heard essentially from enterprises is like hey employees >> love chpt as a sort of a universal interface you know they want to do more in chptt and so and at the moment is pretty capped you know besid is pretty capped on just essentially open eye features and there are like tons of startups you know that want to build that specific feature of that specific market and just you know benefit from like the adoption like you know and the memory like connectors of chip on day one so um that's one you know vector and I think we'll do you know more of those like in the future.
>> is pretty much every enterprise worker in the future like they open chat GBT like right when they start their day is that kind of how you envision?
>> I think so I think so it's it's hard to predict for sure you know I don't think JP is going to replace every tool like you know for sure like if I'm like um a financial analyst, I'm going to spend my day in Excel or like you know some like tool which was purpose designed for my specific use case.
>> But I do expect like ch to become more and more like the first like place that you check in the morning. Yeah,
>> I don't know how much you've been using pulse. Like pulse has been a pretty monumental feature for me. Like it's very on point now, you know, preparing my day. Like, hey, yesterday night that email came in, that meeting is coming up like should prepare by the way that paper came up. Like it's becoming like a really really accurate and like insightful like source of productive information. Um, and so I think we'll see more and more of that. I think we'll see like more and more like people taking actions in CHP like simple actions as well. So yeah, my hunch is that CHP could become like you know the sort of a first like website in the morning and then of course you know some deep workflows like you would double click and like you know do it somewhere else like coding in ID or like doing like you know data manipulation like in a spreadsheet.
One thing that's fascinating about the space is just as we're discovering the capability of these models they also get way better. I think there's a question if we like froze model capabilities today. Is there, you know, we've only been playing with these models, you know, GPT4 and above for for a few years, like does it feel like there's an endless amount of stuff to go build this current set of capabilities or um and we're just discovering more and more or are you still kind of like waiting for the next leap from the research team to to you know unlock even even further things?
>> Oh man, I feel like I always operate on like two or three different time horizons. [laughter] Time horizon number one is just with like the current model capabilities like you know unlock like use cases which haven't been like you know haven't been enabled yet that could be as simple as like you know having like the right harness right use case right data but you know current capabilities then I'm trying to like think like you know a couple of months in advance like like I mentioned for GP 5.1 like for some of those use cases like you know it takes like a little bit of like smart like post training on radar to really make the experience like just better like you know faster more align like more steerable So that would be like the second time horizon and then the third one is like you know fundamental breakthrough like you know in the shape of like 01 essentially.
>> like hey what could you achieve if the model could reliably think for like 30 minutes essentially.
>> that one is you know harder to predict because you know you don't predict like when research is going to you know is going to land but for sure like we think pretty deeply about you know on the products that we release the use cases that we enable with customers like are they going to meaningfully benefit essentially before through and usually when that happens like that's where you know the magic essentially shows up.
>> Yeah. What are the next like model frontiers that you think about or you're like oh it' be amazing when X or Y?
>> oh man there are many of them from a customer like business use case perspective I think once we crack continuous learning.
>> yeah like that probably will be like a very very meaningful exchange. I think of it very simply like I think of an agent as like hey I'm hiring an intern on day one they have like a bunch of academic knowledge they don't have a ton of practical knowledge on the job because you know I have never been documenting like everything I do you know like you have to learn on the job essentially.
>> and so I think the relationship with agent is very much going to be like you know based on like feedback annotations and like hey you you did that or you know you responded that way you know you should do it slightly differently this time and the model you know incorporates that over time at the moment like you know we autoch it through you know prompting and that sort of stuff but once the model is able to actually like updates like its weights based on you know like human feedback like you know uh infant time sign feedback I think that's going to unlock quite a few use cases uh coding customer experience finance yeah it's going to be all over the place.
>> basically your agent shows up as an intern each morning but with better and better instructions and so it would be nice uh nice if it learned.
>> they slept on it and you know they're smarter as a result.
>> yeah I'm curious from my seat it feels you know if you asked me a year ago like the categories that really had product market fit and AI I would have said coding customer support maybe healthcare and legal too as as applications and and if you ask me today I still feel like those four are kind of like the dominant ones there's been some interesting voice uh applications as well across domains.
>> but you see this all the time on the ground with enterprises I guess you'd add life sciences anything else that you know as you categorize you know insane product market fit kind of product market fit still early like anything else that you feel like goes in Insane product market fit bug.
>> Yeah. Sometime I remember like I remind my team like when we talk about like coding, customer experience, like finance, those are gigantic markets.
>> Oh, 100%. [laughter]
>> Like the market for coding, the market for software. I mean, I was chatting with the team like.
>> we have no idea how big it is, right? We've never had such cheap.
>> how big [clears throat] is the time of like software, you know, like what we know for sure is that there's a shortage of software engineers and so, you know, like probably more than like the current pay of like software engineers, probably way more than that. And so frankly the way we think about it I would say like the first like few years like you know post like GP3 were very much like you know spraying and like trying to see what sticks. Now I think we have a much better picture of the industries use cases domains on which we think one you know there is like a massive customer problem and number two like the models are going to keep improving and make it you know more efficient essentially effective like in that market. And so my current philosophy frankly is to try to double down more on those markets. Of course we expand like you know all the time. Um, you know, we talked about science like you know, I don't know if we'll do but we should probably do something you know on science given like you know the feedback that we're receiving but yeah on coding like we can push it much further customer support like you know okay you automated like you know tier one tickets like you know what does it mean like to go much further? what does it mean to like you know turn customer support into an actual like revenue maker for the company like you know have it like be way more personalized um that sort of stuff. So yeah I would say like similar domains but like going deeper and deeper.
>> Yeah. So you think a year from now we'll like kind of have a similar list of of the stuff.
>> I'm sure there'll be new domains like you know if you had asked me like a year ago I'm not sure life licenses like pharma like healthcare was in there. Now I can see it like um and some of it is you know not just like model the capabilities it's like pure software and change management.
>> totally.
>> um those enterprises like you know have like massive systems lots of employees and so you know having like you know AI being adapted adjusted essentially to work really well like for those cases is a lot of work and so to your question like let's say we froze like you know research like you know for like a few years.
>> there would be like you know many years of like enterprise adoption that would still be quite valuable. And it feels like in some of these industries you kind of just reach a tipping point where enough people are are using it that it starts to become like irresponsible not to. Um I feel like even customer support people were like kind of dipping their toes in and then some people had huge success with it never and now it's become one of these things like if you're not trying uh one of these solutions you know what what are you almost doing?
>> Of course of course I mean that's what we see with like every market and frankly there are like a few pioneers u enterprise startups you know there are companies essentially who are willing to take the risks. Um it feels like shaky at first like you know it requires like a lot of scaffolding and like you know the thing like barely stands but at some point you can see essentially the thing working like you harden it and then you know everyone follows essentially. So yeah we see the same motion like across like basically industry.
Does it feel like the scaffolding and the harnesses that people are building are pretty similar across use cases? Like a common set of patterns for scaffolding, you know, in life sciences and support and coding or like how bespoke to the end problem is the scaffolding.
>> I would say it's fairly bespoke at the moment which is a challenge frankly. Um the way I would put it is frankly people are trying to make it work whatever it take.
>> Yeah. one agent, multiple agents, you know, some deterministic like gates like in between like you know, people are trying like many different things. I think there hasn't been I mean we're starting to see it but like there hasn't been like a really like standard like agent architecture or runtime you know which has been adopted across industries that's something that we are actively working on.
>> why don't you think there has been?
>> I mean it's only been three years you see what I mean like [laughter] um took us I mean I've been at for like two years and a half the first year I was just trying to keep up with the growth and be like okay what the heck is going on like you know [laughter] year two was like okay let's sit down like you know okay what can we achieve what is working not working and now I think we're starting to converge again on like okay we have a good idea of like you know more capabilities customer problems what is being done okay what are the pieces like you know in the stack that you know if they were like standardized would like meaningfully accelerate like adoption uh so yeah my you know like uh sort of a my sort of boring answer is like frankly just time.
>> and how are you thinking about that like uh that what that standard set of scaffolding might look like?
>> I think something that we're seeing is uh code and coding is a much more general purpose capability than just software engineering like the models are really really good generating.
>> writing code on an insane progress path too. So it seems like a great thing to bet on.
>> Exactly like writing script executing like that thing in a tool like in a shell like you know having it back. So I think the whole like effort towards like you know basically giving access to agents over computer it's probably going to become I think like a standard uh in the industry one um on the whole like data like API connections I think MCP was like you know a really good standard uh you know I expect the industry to keep like you know uh normalizing around it um I think on the whole like agent to agent communication like there hasn't been yet frankly like you know like a true uh breakthrough or you know a true standard that you know we um we see being actually used um evaluation I think we're getting much better at like generating traces evaluating them trying to infer like you know some improvements on the traces so yeah it's a bit of a it's a game of inches actually.
>> but you think you know it sounds like behind what you're saying is like moving more of this stuff to code feels very helpful because it's the vector which these models are improving at at the most rapid pace and so anything you can any part of your scaffolding that you can move to that is probably quite helpful.
>> Exactly. It's a bit like a human frankly like you know if you ask me to do some my work like without a laptop.
>> I'm not useless but you know [laughter] I'm not far away from me useless versus you give me like a laptop with like you know internet like you know a shell like you know like you know an ID to good stuff like your capabilities are like you know like way way way larger and so yeah to the extent we can replicate essentially that model with um with an agent with a model um I think we are going to see like you know a bunch of POS and to your point like you know every PROS that we do to like the car coding capabilities of the models is going to basically have like you know um orders of magnitude more impact as a result.
I guess it seems like when these models come out everyone's in just like use case exploration mode and like just finding all the things that the models can do.
I'm wondering you know uh does it feel like cost is a limiting factor today on some of these things like where I heard you say that on on BG2 and I was curious because where does it feel like there's you know there's uh interesting use cases at one price point but maybe not at another.
>> If I step back like the the sort of the the reduction of prices of cost that we've seen over the past two or three years I think is basically unheard of frankly in the technology like we reduced like the cost of a DT4 like level queries by like you know one to two orders of magnitude like into three years and you know without like sacrificing the margin or something like it's pure like compressing the M size having like you know better like hardware like you know being better at like networking together like um GPUs so you know every of the stack um what I see like talking you know with teams who are working you know each of the stack is that there are still like a ton of like room for optimization improvement so you know we are not going to stop there and then what I see on the business side is that you know for some very high stake um use cases say you know coding yeah.
>> the economics work you know like it's so high leverage like you know to multiply by two like the productivity of your software engineer that sure paying like you know dozens hundreds of dollars you know every month like could be worth.
>> Um but then there are many like you know other um use cases which are currently being blocked like you know if I were to take like a very like real one like models are pretty good at personalization and like understand intent and like you know adjusting the content. So why is not every like content website in the world like homepage like you know not like you know infused with I think I know the answer like you know like cost and probably latency and so um I very much see it as part of the open eye mission frankly.
>> to keep driving the cost down um what I've seen is that you know I don't even know over the past two years I think we've done like you know dozens and dozens like of cost cuts every time like you know the increase in like the base like the volume is like larger than you know essentially the the the sort of the the revenue like the price effect, you know.
>> And so that tells me that like there is still like a wild untapped demand which is basically um limited by cost.
Yeah, I'm really excited to see it play out on the Sora models too. Like you have the API and it's clear it's going to be a massive cost corruption over time in those models and I think people will find all sorts of.
>> Exactly. Exactly. So um and yeah, you know, if you look at like some of our most successful agents at the moment like you know run for like you know hours like on you know like tens and tens of minutes like that can get pretty expensive like quickly and so if you truly want to move to a mode where essentially like each of us has like 1000 agents running in parallel async like in the background pretty much all the time. Yeah, we have uh we need as industry to basically bring down the cost at you know every layer of the stack and you know I have good uh good conviction we'll get there.
>> Yeah. What about uh RFT? I feel like it's it's very in in in uh the zeitgeist these days. You know I feel like a lot of people seem to think uh there's been a real step change in improvement in efficacy of you know uh of tweaking these models versus maybe the SFT paradigm. Um are you actually you work closely with these enterprises like are you seeing folks kind of start to use this and where are we in the journey of like more and more people actually you know tweaking these models for their use cases?
>> we we start to see it uh it's not widely adopted yet I mean the way I think about the enterprise opportunity at the moment is that most of the market is catching up really catching up to the frontier like you know to our discussion earlier like they they haven't yet fully leveraged like the capabilities of dt 5.1 base so that's like you know the gist of the market and then you have a few innovators who are clearly blocked at the frontier know exactly what they need like to get there you know um couple of example I have um I was working with um an accounting uh software like firm essentially to um make extremely accurate um accounting tax accounting essentially analysis and the model was just not quite good too slow essentially out of the box um and it took like I don't know something like maybe a few dozen samples of like really high quality environments and navigators, you know, to update um improve the model by I don't know 20 or 30% on their gold like evil standard which was essentially like the the diff which allows them essentially to you know get from like not really viable to actually viable and so we're starting to see like you know some like innovators who are um identifying those things trying everything they can outside of tuning like to make it work and not succeeding and therefore you know having a knock with that said it's still a lot of work like you have to build like a really high quality environment months you have to really like you know make sure that your graders is good once you kick off like you know a reinforcement learning uh job like you know it takes like you know can take hours can take days sometime so you know it's more like heavy-handed I think the we released an API the API which I believe was probably the first one on the market in that so we're seeing quite a bit you know of like excitement but it's clearly not not yet like you know a mass market maybe never will frankly um but that's fine you know as soon as as long as we're like pushing like the frontier of like some customers you know I'm happy about it.
Do you think like over time most people end up using it or is it like you know you can kind of just tie yourself to the general model improvements and you know who who needs to push the frontiers 6 months ahead of that or 12 months ahead of that for most enterprises.
>> I suspect that like most enterprises will have some RF use cases, but I suspect those if use cases are not going to be like the gist of [music] their business like they will want to innovate on like some aspect you know of the model but for like the the sort of the the the the massive like automating your operation making like your knowledge base like you know maintaining it like updating it like I expect like the base model will be pretty good like out of the box.
What about on the startup side? So I feel like for the longest time it seemed really silly to like train your own model or like spend too much time on even on fine-tuning, right? It didn't it didn't help that much. Um now it does seem like RFD can push performance like you know for a while I would have said hey most AI startups they need people that know how to use these models well but maybe not people that are super deep in in doing all on them. Like do you think that's changed in this paradigm?
>> It's a good question. I like to remind like startups and people of what we do at OpenAI. We postrain a lot the models like we fine-tune a lot the models. So of course I like to think of us doing a really good job and like you know the model is like post train like you know for most use cases but of course like there are some use cases and like the model will be like you know less like you know well trained like you know it's going to behave like not exactly the right way. the style, the formatting, the tone, the conciseness, you know, will not be quite it. Um, and so I expect like you know startups who are achieving like a certain scale or you know want to get like really to the next level of um capabilities will continue to fine tune. Um it is a lot of work. Maybe at some day like you know some point we'll be able to achieve that continuous learning or you know some sort of you know automated fine tuning. We're not quite there yet. Um uh but yeah, I do expect that you know a fraction of startups like will continue to uh use venting to push envelope.
>> Yeah, as you think about like the choices developers makes in the models they use. I mean obviously there's you know the overall quality of these models. What else do you think will like ultimately drive you know where startups and developers choose to to build on?
>> Yeah, I think I like to think of it as you know at three big buckets. One is like clearly model like capabilities and behaviors and you know we'll decompose that. Number two is like cost latency. Number three is like vibes like trends like Twitter essentially and you know that's becoming more and more important because you know.
>> that is like the eval these days right like Gemini 3 comes out today and I'm like I yeah they all the percents seem good but like let's see what Twitter says. It's really hard to keep up frankly with like you know everything and so um yeah, I'm starting to see like you know more and more like serious startups and like enterprises like you know like rely like way more on like you know specific like you know influencers like media I don't know um but anyway um capabilities behavior for sure like the biggest like you know um bucket um the trend I'm seeing pretty much all the time is you know um people are not prematurely optimizing for like cost and latency like the first goal is to make it work and hopefully to make it work like you know cheaply and like fast enough. Um the trend I'm seeing here is that um at some point like there is so much like sort of a information you can cram into like academic benchmarks. Um if I'm building a customer support like agent sure like sweet bench like MMLU like probably tells me like give me some indication but frankly like you know probably not that much. Um we're starting to see like some interesting like you know industry level benchmarks. Um, Tow Bench is like a good one um in the services uh industry. Um, we're trying like to build more and more those like we released um a few months ago GDP which was essentially meant to uh analyze the model capabilities on uh real world like economic tasks. So um but I would say here like the industry is probably like in a catch-up mode and the startups like do not have the luxury to wait like for us to come up and so I'm seeing a lot of like frankly like testing like you know yeah basically qualitative testing which is pretty interesting. I'm starting to see like you know among these startups like some people who have like such a good taste for nuances of models you know just like people like who are really good at writing or you know like painting like you know they're not necessarily able like to elaborate why like the framework where everything but they have like you know that that that sort of that sense um I'm starting to see like the same happen frankly like you know for models which is pretty cool cost and latency pretty important cost my expectation is going to continue to be reduced like by multiple like you know over the next like year um or so and then you have like yeah like you know Twitter vibes um I would say uh it's funny I feel in a way like we are like reinventing like you know why like Gartner and like other exist which is at some point like you cannot like compare like all the accounting software in the world like you know you have to like trust someone and so yeah I think um we're getting at that stage and uh I think we'll stay there.
>> I like that Twitter's like the centralized guard yeah [laughter] yeah that's a good way to.
>> one thing I thought was really interesting um And and you'll have to forgive me for referring to I think it was in the concept of anthropic models, but cognition was talked about like when you know 45 Sonic came out I think it was and they had to move everything over uh and like actually required like a ton of net new work from them and I'm wondering like you work with all these enterprises and then you have a new model like five or 51 come out.
What does that like process actually look like and how do you imagine that changing or evolving over time?
>> That's part of the model fatigue. uh I think uh which is the days where you could just like hot swap essentially like one API parameter from like you know the one model to the next are basically gone for like you know non-trivial use cases the idiosyncrasies like of the models are becoming especially among different providers are getting like you know more and more distinct in a way and so some models like respond better to certain types of instructions some models like have been pre-trained like post train sorry for like you know specific tool like signature specific like tool names. Some models um handle like more or less like you know um like very long context um recall and so you have to essentially understand like all the quirks essentially of models and then to adjust your pump your harness essentially to adjust to it and it's a lot of work and even among like you know the top like you know most subop startups what we see is that you know um doing it like every time and do it accurately is like hard like takes a lot of work and so they would rather not do it unless like there's a meaningful change and so you can imagine like in the enterprise like the primary job is not like you know to implementation the primary job is like to I don't design drugs or you know sell like you know cell phones um they are very much eager I think to move to a much more.
sort of a regular cadence with like clear chance logs in a way. I feel we are sort of reinventing, like rediscovering, like, you know, how do we like deploy software? Which is [laughter] if I drop every day to you like, you know, a new like binary, fine, you're like, good luck. You're like, cool, okay. But what should I do about it versus, okay, it's the view 1.1, okay, it's a major version, here's a change log, and you can expect more in three months. People just want like, you know, predictability and like, you know, transparency.
>> Do you think that like over time the scaffolding evolves to be more generalizable such that it can like better absorb changes in these models, or is it always going to be like, hey, you're going to have to take that like week sprint to figure out like what's what changed, what what you need to adapt?
>> I think as we standardize that agent architecture, it has to become like more generalized, uh, in a way. I think one of the reasons why at the moment like there is so much like sort of discrepancies across different models is that each like lab is like training like, you know, for like different harness, different on purpose, different use case essentially like the models. Uh, and there's not yet like a single like, you know, common architecture like framework, you know, to follow. Uh, so my hope is that, yeah, at some point like the industry will converge to like a much more like universal like agent like, you know, sort of, yeah, framework. Um, tool definitions, something like MCP essentially, but, you know, across like every single like dimension of the agent, and that will make it, um, easier essentially for customers like to compare different models, to adopt new ones, like to adopt multiple models for different use cases.
Do you have any advice for builders like what do the really good like the top teams do when like they're experimenting with a new model and try and like, you know, basically getting ready to shift everything to that new model?
>> And any lessons for like the rest of the the world?
>> Frankly, they take the time. Like what I've seen is, uh, teams that are way too impatient and just want to, you know, yeah, I call it like the hot swap essentially of like the mall name and like, oh shoot, that mall sucks, like doesn't work as well. So the best teams I would say have strong taste but come with an open mind, take the time like, you know, to battle test like the model, take the time to work with us as well. Frankly, like, you know, we control the posting and so, you know, if some teams, you know, is like really, really like, you know, ganged up on like, you know, having that tool like specific tool like behave like in a specific way, like we can influence that, and so the more like specific feedback we receive with like specific examples, and we can actually tweak like for next like snapshots, um, models for that. So, um, usually that's what happens. Yeah.
>> We haven't hit on like voice yet, and I want to make, you know, it feels like that's one thing that's just changed so much the last year, like there's so many people building with, uh, you know, the real-time API. How do you think about the next frontiers there and like what's, you know, what are where are like where are you seeing a bunch of fit and where are you like, you know, we still need to, you know, do X, Y, or Z to unlock the next set?
>> So it's interesting, um, GPT4O, which came I think in May or June last year, so a little more than a year ago, I think to me was like probably the second like breakthrough in terms of like feeling the AGI after JDPT, or like man, like the world will not be the same essentially, like having a model like being able to express that branch of like tone, emotions, and like to understand as well, like, you know, as a result, um, from the human like such tone and emotion was like put in my brain. With that said, like I think we clearly haven't like crossed like the the Turing test yet for voice. For text, at that point, I'm pretty sure I wouldn't be able frankly like, you know, to distinguish between a human and and a bot like, you know, on text. On voice, like it still feels like the interruptions, like, you know, the cadence, still not quite exactly it. Um, and so I think that's probably the next frontier on voice, which is, hey, like, you know, make the world more intelligent. Cool. Okay, we'll do that. But second, like get to a point of like naturalness and expressiveness where basically like provided the same level intelligence, like you're basically like, you know, you're okay to be served by AI, it's just be okay be served by a human. Um, I think we have a line of sight to get there, um, and so I expect like, you know, that's going to be like through like to really like, you know, cross the chasm. We are starting to see like deployment, um, I come back to customer support because, you know, that's one of the main like, you know, voice like calls essentially, like, you know, use cases in the world.
>> We're starting to see like, you know, meaningful deployments among like, you know, tier one like customer support calls, um, which is good, uh, for a couple of reasons. Number one is that the model is like infinitely patient. So, you know, for people who need more time like, you know, to have their answer, like, you know, the model can take like five minutes at most, you can. The second thing, which I did not quite realize, um, is how critical multilingual capabilities were for customer support because, you know, let's say you're in the US, you're like, you know, a big like retailer in the US, like, you know, most of your customers speak English, others will speak like Spanish, Chinese, like, you know, there's a long tail of languages, and if you have to staff like customer support agents for like every language, like the day, like you're not going to make it essentially. So you have to make like really hard trade-offs and like literally lose customers as a result. And so we've been seeing like, you know, pretty strong results, um, and like reception from customers on those multilingual capabilities.
>> Yeah, we were talking before about how Codex has kind of like bitten the the main character of the ecosystem, uh, for for the past months, um, and obviously just tremendous improvement there.
>> How do you like see that space playing out and you know, what do you, besides just the underlying capabilities of the models, like what do you think will ultimately determine, you know, whether folks reach for Codex or cloud code or one of these other tools going forward?
>> It's a good question. Uh, the Codex team is incredible. Um, to me, they it's like the sort of epitome of like, um, a really small, talented team, singularly focused on a use case and like cranking at it at like whatever it takes, like, you know, model harness integration, like data, like, you know, so, um, they're really, really good. My read on how the software engineering, uh, space is going to evolve, at the moment, the models are really good at generating code and understanding code. They could be better, but, you know, they've made like meaningful progress. But when I think about software engineers, like that's only part of the job of a software engineer, right? The other part of the job is like to be on call, another part of the job is like to communicate with your teammate, to scope some changes, to make some tough like architectural decisions, to duplicate some APIs, you know, like there is much more to it. And so I think that's, you know, on top of like improving the more capabilities on like writing like understanding code, like that like second axis of like collaboration essentially, probably like a major unlock like to, um, to, uh, to spread like the benefits, um, of AI, you know, uh, more broadly. So that's probably one. Um, a second one, frankly, is like maybe trivial, but like just like for that to gain adoption in the enterprise. Um, when I talked to so many enterprises who have not yet essentially like, you know, are still stuck on like the GitHub Copilot V1 essentially, because they've never like went through like all the security and like provisioning process, you know, to really like, you know, make sure that these agents can be used properly like in code bases. Um, so yeah, my
>> Do you think we're like close to enterprises being like, all right, time for these, uh, agentic coding tools, or does it feel like, god, there's, you know, three years of hurdles, you know, on like the security or or like, you know, compliance side?
>> I think I'm starting to see a critical mass of enterprises who are truly leaning in and like provisioning thousands, hundreds, thousands of licenses, like, you know, engineers, like letting like, you know, them like experiment, iterate on like some specific use cases. Yeah, my gut is like 2025 in a way. Like you mentioned like 2024 is the year of coding. I think 2025 is the year of coding in the enterprise. Like I'm starting to see like meaningful adoption. Yeah.
>> Yeah. What's 2026?
>> I don't know, man. Like, uh, PMs, you know, maybe like an APM, like, you know, automated PM, I don't know. Yeah. My my bet would be like, you know, one, the mods are like more reliable essentially, to do like, you know, write and understand better like code, and second, the mods are getting really good at collaboration, and so you're starting like to have like a more of a multi-dimensional like, you know, AI software engineer that you can work with.
>> How much of like the improvements in collaboration is just models getting better versus the harnesses you put around them?
>> The two, I think my assessment is that it's getting harder and harder like to detangle like what is the model versus the harness. But I see is that, you know, some of the best agents out there are trained like, you know, models are being trained for like a specific harness.
>> I think that's why Codex is so good, frankly, um, and so I think of them like more and more as a sort of a simpies or something.
>> I guess it goes back to this question I was asking earlier about whether startups will train their own models. I wonder whether if there's a standard set of harnesses they can leverage your models, or if they have their own harness, you know, they'll ultimately need to train very specifically on that harness.
>> So that's what we try to do as much as we can, which is we open source the Codex harness essentially, we open source like the actual code on GitHub, we open source like the the tool like definition for people to be able like to fully utilize like, you know, uh, Codex abilities. Um, and today, you know, if you wanted to use like Codex in Cursor, like any other IDE, like you could essentially. So I I think that's how the industry is going to evolve, like, you know, AI providers are going to move from being just like like model like inference APIs essentially to providing both the model, the harness, you know, maybe some UI.
>> The reference design for the harness, basically, and then everyone else can.
>> Exactly, like a sort of a standard architecture, like, you know, sort of a standard like blueprint essentially to use like the model to the best of the capabilities. If I step back, that's probably to me the biggest learnings like of the past two years of like, you know, building for like developers, businesses, is hey, you can't just drop new models like in an API. Like it's going to be really hard for people to maximally utilize like the multi capabilities unless you give them like more of a blueprint essentially, more like documentation or like, you know, a specific harness. Um, it's hard like to discover, like models are so beautiful and weird like in some way, like, you know, unless you have like a really good like recipe or like a ton of experience like interacting with models. It's hard like to massively, uh, leverage them.
>> And do you think that like, uh, in the future, you know, your average enterprise will be able to, you know, directly interact with the models and these harnesses themselves, or like, it seems like there's this whole set of applications that are basically like translation layers between like, here's the capability of the models and like, here's your end industry, and we sit in between.
>> My gut is enterprises are mostly going to buy harnesses. Um, they're mostly going to buy solutions. I think there will be some exceptions. Uh, if you're talking about a use case which is like the core business of the enterprise, I think there is some reason frankly to build your own harness.
>> But you know, if you're like a retailer and you have to operate like sales, finance, it, like frankly, just like you buy software like, you know, at some point, why would you, you know, like.
>> I think people usually tend to underestimate like the amount of effort it takes like to do anything frankly, like in a week, like, you know, great level of quality. The same will be true, or even like more true for like agents. And so, yeah, my bet would be like, you know, buy versus build for most use cases.
>> Yeah, it'll be fascinating to see if the harnesses from the labs converge, or like they're actually end up looking quite different, and then as a result, you kind of have to like figure out which harness ecosystem you want to play in.
>> I don't know. It's a, it's a fun game of like, you know, divergent convergence.
>> Yeah.
>> I don't know, maybe I'm being naive, but like I do expect convergence will win like in the long term. Uh, but, you know, science research, like, you know, you have to let a thousand bloom essentially to then like figure out which one's the most promising and like, you know, like gather like behind it.
>> Yeah. I mean, you obviously work with so many different kinds of enterprises. When you go into like a net new enterprise that's maybe a little newer to the game, um, do you have like a cheat sheet of like, here's your like, the the few things you really should do or know that you kind of take from your most complex customers, or like what, you know, if you could distill or or, you know, share some lessons from totally the most sophisticated ones, what would they be?
>> Oh, at that point, I think I worked with, uh, I don't know, 200 plus [snorts] like enterprises.
>> Um, we talk about T-Mobile, MGEN, like, you know, Salesforce, BNY, um, you know, many of them database, and many others. My cheat sheet. There are many, many tips and tricks. Um, the first one is a classic like enterprise software, but like if your data is a mess, you'll be able to achieve nothing. Like, like I could give you the most powerful coding agent. If you're not able like to plug that coding agent like the right code bases, the right like identity permissions, like, um, database, uh, you know, it's going to be really hard for that agent to be useful. Uh, and so frankly, like, you know, a lot of the work is just like explaining like, hey, how do you structure your data, um, if there is like no API, like no services, like how do you stand up like the right services, how to use MCP or not use MCP, how do you authenticate, like, you know, those requests, how do you log those requests, and so there's a lot of, you know, like data engineering, you know, if you will, um, work to do. That would be step one. Um, step two would be to really explain like, you know, to, um, to, uh, teams like, you know, how to evaluate models. Um, I was like, you know, talking earlier about like, you know, people like vibe like evaluating, which, you know, is necessary to some degrees, but at some point, if you have like a production grade like, you know, use case, like you have to evaluate it. And evaluation, like rigorous evaluation, like doesn't come naturally to, you know, most teams unless they've done it before. And so, um, and I always tell them like, you know, we often talk about OpenAI at OpenAI about the teams that train models, and, you know, they're extremely important. We have like crack teams, we evaluate models like full-time, and they're probably equally important, frankly. Uh, and so spending a lot of time to like document like golden sets, document like the procedures, like the SOPs, like the standard operating procedures, and then build like rigorous evaluations in order to make sure that, you know, you are like hill climbing like the right direction. Um, that would be number two.
>> Um, what's the most common mistake you see on the on the eval side?
>> Too little or too much evals. Like too little, like, you know, you you've done only like five like sets, and you know, first customer comes in and like, it's completely out of distribution, you're like, okay, like, you know, what did I learn? I don't know. I mean, one of my biggest findings like working with enterprises is, um, most of the knowledge is in people's brains.
>> Yeah, like you usually come in and you're like, customer support, I'm sure there is like, whatever a J or confence somewhere with like every procedure being written. That never happens. Like, you know, you're lucky if you have like 20, 30% of the proc being written. The rest is like, oh, you know, Sarah and Mark like know really well about it, and you know, you should talk to them.
>> And so as a result, like building really good evals is like not so much about, you know, converting like text like to eval.
>> Yeah, it's finding.
>> But finding the right people, finding the S and the John essentially, well, you know, the ones who know about it. And so that's like a fairly like iterative process, like, you know, you don't do it like on one day, then you're done, like you start to do it, you ship the thing, you're like, okay, that thing is not quite working, let me talk to John, you know, why is that not? And then you build.
>> I know I cut you off. Was there a third one, uh, on on like the the things that are most common on the enterprise?
>> Change management.
>> Yeah.
>> Like that technology is new to all of us. Maybe for me, sometime like, you know, I'm like weirded out by, you know, how powerful it is, and sometime, you know, how like unequivocal it is. And so, yeah, taking the time like to explain to teams, to customers, you know, how does it work, uh, I cannot emphasize enough, you know, how critical it is.
>> Have you seen enterprises do anything cool with Sora API yet? I know it's, uh, it's only been out for a little bit.
>> Yes, actually. Um, we're seeing quite a bit of energy in particular in two industries. Um, one is like ads/content generation, where people are creating some crazy personalized like, you know, content, which is really fun. The second one is like more like the studios, the production companies. Um, that one is really interesting, actually. I I learned a lot about what it takes like to build like really good movies. And now like being able like to just show to someone in like 30 seconds what you have in mind, and, you know, the sort of the picture like imagery that you have in mind is apparently like, you know, being like quite useful like, you know, for teams like to brainstorm like way more. So yeah, that's a fun use case. But yeah, video generation, I think we're still like in the early innings frankly of video generation. It's expensive, it's slow, and, you know, but we can start to see how it's going to totally like, you know, transform like how that sort of use cases like jobs are being performed. And so, yeah, quite excited to see what's happen next.
>> That's awesome. Well, I always like to end my interviews with a standard set of quickfire questions where we get your thoughts on some overly broad questions that we stuff into the end. Uh, and so maybe to start, what's one thing that you think's overhyped and underhyped in the AI world today?
>> Oh, shoot. I did not prepare for that one. [laughter] Uh, overhyped, underhyped, underhyped. I keep coming back to science, dog design, dog discovery. I think we tend to talk less about it because for us who've been working in software like, you know, for a while, doesn't come super naturally as a use case. But at the end of the day, like, you know, if you look at like the arc of history, like that's like the substrate of like progress. Totally. And so if we're able to accelerate, if only by like 5%, like the rate of discovery, like the sort of the the implications on like the rest of the economy and like technology are just like, you know, humongous.
>> And so yes, making models, um, harnesses extremely good for like, you know, scientific.
>> You have to build a lab to do that. I mean, it feels like, uh, to really build a biomodel, you need some sort of feedback loop of experiments, right?
>> Yes, you need a lab, you need data, like you need people like, you know, who are like at the at the intersection of like, you know, LLMs and like, you know, their field. So, you know, it's hard. But probably like, you know, the highest like compounding, uh, like, you know, benefits.
>> What's one thing you've changed your mind on in AI in the last year?
>> The model is everything. A couple of reasons. Number one, just like we discussed, like the harnesses are getting like extremely important as well, and, you know, it's getting harder and harder like to detangle the two. But I do expect like those, you know, extremely powerful agents will have like, you know, extremely goodnesses which are evolving like, you know, even faster than the models are. Yeah.
>> Um, that's important. The third thing is like in the enterprise, like the whole like goal of like the model is like to get more and more data like to perform outputs. And so if you don't get like high quality data, like, you know, the more output will not be that good. And so I think equally important for like AI to be adopted in the enterprise widely is, you know, a set of like standard infrastructure framework to present the right data at the right time to the model. And so I think once we crack that as an industry, yeah, like we're probably going to see enterprise adoption like, you know, scale lock it.
>> Yeah. Do you think that most of the advances in harnesses, uh, will happen within the model companies because they're as close to the models that are posted on these things, or do you think will happen in the startup world?
>> I think we'll see a bunch of adoption in the startup world as well. I think as a result of like open sourcing the harnesses and sort of giving like the recipe essentially to best utilize the models, like people are going to innovate. They're going to see some quirks of the model, like some way essentially to tweak the harness definition to get like more results out of it. And so, yeah, I know I do expect a fair bit of innovation.
>> From the outside, it feels like everything just always goes right at OpenAI. Anything looking back on your last two and a half years that you're like, "Oh, we got we got that thing pretty wrong."
>> Oh man, where do I start? Uh, I think the world has been pretty forgiving to OpenAI because we are first on many things and so I think people like, you know, have like a higher like sort of a tolerance or, you know, flexibility like, you know, to, um, to sort of pioneers. What have we gotten wrong? Uh, I mean, there are plenty of like product features, like tools that we ship that, you know, did not like, you know, find product market fit or like, you know, did not like, you know, utilize like the models as well as we could. Um, we had some like definitions of like what an agent was like in 2023, which was probably too early and like, you know, didn't catch like adoption, and so we had to less set like, you know, some APIs, um, as a result. Um, we invested a bunch in like, you know, like different kinds of like audio technologies that did not really pan out as well, frankly. So yeah, frankly, I think at the end of the day, like history remembers like, you know, the sort of the successes, but there's like a fair amount like a healthy amount of, you know, like failures or experiments that didn't pan out, you know, in the in in the meantime.
>> Yeah. Um, what do you think of Gemini 3?
>> I haven't played with it yet. Um, the benchmarks look really good. Uh, I saw some couple example visual examples that look like really strong. So, it seems like the Google team like really cooked like a great model. But yeah, I cannot wait like to actually test it myself.
>> Yeah. What do you do to test it? Like, what are you going to do?
>> I have like two or three like different kinds of tests I do myself. Um, number one is usually on style, tone, like, you know, personality essentially, when, you know, I do have like some personal questions which I like to ask by dumping like, you know, a lot of context like to models. That's one. Number two is the classic like, you know, generating like an app from scratch, and, you know, look at the front end, like, you know, testing it a bit. Um, number three is I love to test like very long horizon capabilities, like dumping like literally like hundreds of thousands of tokens in there and try to ask it like very hard questions on like the the, you know, the second to last like, you know, token.
>> [laughter]
>> That's awesome. Um, well, this has been a fascinating conversation. I want to make sure to leave the last word to you. Uh, the mic is yours. Where can folks go to learn more, uh, about you, about anything that OpenAI has shipped that you want to point folks to? Uh, I'll leave it to you.
>> We tweet a lot. We tweet. Maybe not enough, but we should tweet more, but yeah, on Twitter, thank you. Twitter OpenAI. I'm also on Twitter. I respond to everyone. Any feedback, any feature request, just tweet at me and, you know, I'll be happy to, uh, connect.
>> Amazing. Well, thanks so much. This a ton of fun.
>> Thanks so much.
>> [music] [music] [music]