Michelle Pokrass · OpenAI API 平台技术负责人

从 API 到 AGI:Structured Outputs 与 OpenAI API 平台(Michelle Pokrass)

2024-09-17 · Latent Space (swyx & Alessio) · 1h12m · 原文链接

→ 在 AI 访谈库中阅读(可切换中英、记录进度)
Pokrass 详解 Structured Outputs 的训练与工程实现、API 平台全景(Assistants、Vision、Voice Mode 等),外加 o1 发布要点速递。技术上最硬的点:OpenAI 如何做到 100% 符合 JSON Schema 的输出——靠受限解码与模型训练结合,而非单纯提示词。

yeah hey everyone welcome to the laden space podcast this is celesio partner and c and residents and deel partners and I'm joined by my co-host swix founder of small AI hey and today we're excited to be in the inperson studio with Michelle welcome thanks thanks for having me very excited to be here this has been a long time coming uh I've been following your work on the API platform for a little bit and uh I'm finally glad that we could make this happen after you you ships structured up how does that feel yeah it feels great uh we've been working on it for quite a while so very excited to have it out there and have people using it we'll tell the story soon uh but I want to give people a little intro to your backgrounds so you've interned and or worked at Google stripe coinbase Clubhouse and obviously open AI what was that Journey like uh you know I the one that has the most appealed to me is Clubhouse because that was a very very hot company for a while you basically you seem to join companies when they're about to scale up really a lot and uh obviously open the eye has been the latest but uh yeah just what what are your learnings and your history going into all these um notable companies yeah totally for a bit of my background uh I'm Canadian I went to the University of watero and there you do like six internships as part of your degree so I started uh actually my first job was really rough I worked at a bank and I learned Visual Basic and I like animated bond yield curves and it was you know not not me too oh really yeah that was a derivative Trader interest rate swaps that kind of stuff yeah yeah so I liked you know having a job but I didn't love that job uh and then my next internship was Google uh and I learned so much there it was tremendous but I had a bunch of friends that were into startups more and you know water has like a big startup culture and one of my friends uh interned at stripe and he said it was super cool so that was kind of my I also was a little bit into crypto at the time and I got into it on Hacker News and uh so coinbase was on my radar and so that was like my first real startup opportunity was coinbase I think I've never learned uh more in my life than in the 4-month period when I was interning coinbase they actually put me on call I worked on like the a rails there and it was it was absolutely crazy you know crypto was a very formative experience and this is 2018 to 2020 kind of like the first that was my fulltime uh but I was there as an intern in 201 16 yeah and so that was the period where I really like learned to become an engineer learned how to use git got on call right away you know managed production databases and stuff so that was super cool after that I went to stripe and kind of got a different flavor of payments on the other side learned a lot uh was really inspired by the cson and then my next internship after that I actually started a company at watero so there's this thing you can do it's an entrepreneurship Co-op uh and I did it with my roommate uh the company's called readwise which still exists but oh yeah yeah rewise what yeah awesome you co-found it rewise yeah premium User it's not even on your your your LinkedIn yeah I mean I only worked on it for about a year and so Tristan and Dan are the real Founders and and I just had an interlude there but uh yeah really loved working on something very startup focused user focused and and hacking with friends it was super fun eventually I decided to go back to coinbase and really like get a lot better as an engineer I didn't feel like I was you know didn't feel equipped to be a CTO of anything at that point and so just learned so much at coinbase and that was a really fun curve but yeah after that I went to clubhouse which was like a really interesting uh time so I wouldn't say that I went there before it blew up I would say I went there as it blew up so not quite the Starling track record that it might seem but it was a super exciting place I joined as like the second or third backend engineer and you know we were down every day basically you know one time Oprah came on and absolutely everything melted down and so we would have a stand up every morning and be like how do we make everything stay up um which is super exciting also one of the first things I worked on there was making our notifications go out more quickly because when you join a clubhouse room you know you need everyone to come in right away so that it's exciting and the person speaking thinks a lot of my audien is here but when I first joined I think it would take like 10 minutes for all the notifications to go out which is insane like you know by the time you want to start talking to the time your audience is there it's like you can totally kill the room so that's one of the first things I worked on is making that a lot faster and you know keeping everything up I mean so already we have an audience of Engineers uh those two things are useful it's keeping things up and notifications out notifications like is it a Kafka topic it was a postgress shop and you had all of the followers in postgress and you needed to like iterate over the followers and like figure out is this a good notification to send and so all of this logic it wasn't like well batched and parallelized and our job queuing infrastructure wasn't right and so there's a lot of like fixing all of these things um eventually there were a lot of database migrations because postrest just wasn't scaling well for us interesting and then think uh keeping things up that was more of a I don't know reliability issue Sr type A lot of it yeah it goes down to like database stuff um everywhere I work databases at coinbase at Clubhouse and at openingi postgress has been a perennial challenge it's like the stuff you learn at one job carries over to all the others because you're always debugging a long running post Quest Creer at 3:00 a.m. for some reason um so those skills have really carried me forward for sure why do you think that not as much of this is prized obviously post crisis an open source project is not aimed as like gigascale but you would think somebody would come around and say hey we're like the yeah I think that's what planet scale is doing kind of it's not on postgress I think it's on my squl but I think that's the vision it's like they have zero time zero down time uh migrations and that's a big pain point I don't know why no one is doing this on postgress but I think it would be pretty cool their connection poers like PG bouncer is like good enough I don't know yeah well even I mean I've run PG bouncer everywhere and it's there's still a lot of problems like your scale is something that not many people see so yeah I mean at some point every company gets the scale every successful company gets to the scale where postgress is not cutting it and then you migrate to some sort of nosql database and that process I've seen happen a bunch of times now mongod DP redis something like that yeah um I mean we're on Azure now and so there's we use Cosmos DV Cosmos DV hey at Clubhouse we I really love Dynamo DV that's probably my favorite database which is like a very nerdy sentence but that's the one I'm using if I need to scale something as far as it goes yeah DB I I um when I learned I worked at AWS briefly and it's kind of like the memory register for the web like yes you know if you treat it just as physical memory you will use it well if you treat it as a real database why you run into problems right you have to totally change your mindset when you're going from postest to Dynamo but I think it's a good mindset shift and kind of makes you design things in a more scalable way yeah I'll recommend a Dynamo DB book for people who need to use Dynamo DB but we're not here to talk about AWS we're here to talk about open ey you join open ey pre chbt I also had the opportunity to join and I didn't what was what was your Insight yeah I think a lot of people who joined open AI join because of a product that really gets them excited and for most people it's chat GPT but for me I was a daily user of co-pilot GitHub co-pilot and I was like so blown away at the quality of this thing I actually remember the first time seeing it on Hacker News and being like wow this is absolutely crazy like this is going to change everything and I started using it every day it just really I even now when like I'm I don't have service and and I'm coding without co-pilot it's just like 10x difference so I was really excited about that product I thought now is maybe the time for AI and I'd done some AI in college and thought some of those skill skills would transfer um and I got introduced to team I liked everyone I talked to so I thought thought it would be cool why didn't you join it was like I was like is Dolly it we were there we were at the dolly like launch thing and and I think you were talking with Lenny and uh Lenny was at open ey at the time and you were like we don't have to go into too much detail but this is one of my biggest regrets of My Life um but but well I was like okay I mean I can I can create images I don't know if like this is the thing to to dedicate but obviously you had a bigger Vision than than I did it was really cool too I remember like first showing my family I was like I'm going to this company and here's like one of the things they do and it like really helped bridge the gap whereas like I still haven't figured out how to explain to my parents what crypto is um my mom for a while thought I worked at Bitcoin so it's like it's pretty different to be able to tell your family what you actually do and they can see it yeah yeah and they can use it too personally so you were there were you immediately on API platform you were there for the chat gbt moment yeah I mean API platform is like a very grandiose term for what it was there was like just a handful of us working on the API yeah it was like a closed beta right not even everyone had access to the G3 a very different access model then um a lot more like tiered roll outs but yeah I I would say the applied team was maybe like 30 or 40 people and yeah probably closer to 30 and there was maybe like fiveish total working on the API at most so yeah we've grown a lot since then it's like 60 70 now right no applied is much bigger than that applied now is bigger than the company when I joined okay yeah we've grown a lot I mean there's so much to build so we need all I'm a little out of date yeah any ched gbt release kind of like all hands on deck stories had had lunch with um Evan morawa a few months ago it sounded like it was a fun time to get build the apis and have all these people trying to use the web thing like how are you prioritizing internally like what was the helping scaling when you're scaling non GPU workloads versus like postgress bouncers and things like that yeah actually surprisingly there were a lot of postgress issues um when chat GPT came out because the accounts for like chat GPT were tied to the accounts in the API and so you're basically creating a developer account to log into chat gbt at the time cuz it's just what we had it was lowkey research preview and so I remember there was just so much work scaling like our authorization system and that would be down a lot yeah also GPU you know I never had worked in a place where you couldn't just scale the thing up it's like everywhere work compute is like free and you just like Auto scale a thing and you like never think about it again but here we're having like tough decisions every day we're like discussing like you know should they go here or here and we have to be principled about it so that's a real mindset shift so you just really structured outputs congrats you also wrote the blog post for it which was really well written and I love all the examples that you put out like it really give the the full story yeah tell us about the whole story from beginning to end yeah I guess the story we should rewind uh quite a bit to Dev Day last year Dev Day last year exactly we shipped Json mode which is our first foray into this area of product so for folks who don't know Json mode is this functionality you can enable in our chat completions and other apis where if you opt in uh we'll kind of constrain the output of the model to match the Json language and so you basically will always get something in a curly brace and this is good this is nice for a lot of people you can like describe your schema what you want in prompt and you know we'll constrain it to Json but it's not getting you exactly where you want because you don't want the model to kind of make up the keys or like match different values than what you want like if you want an enum or a number and you get a string instead it's like pretty frustrating so we've been ideating on this for a while and like people have been asking for basically this every time I talk to customers for maybe the last year so it's really clear that there's developer need and we started working on kind of making it happen and this is a real collab between engineering and research I would say and so it's not enough to just kind of constrain the model I think of that as the engineering side whereas basically you Mass the available tokens that are produced every time to only fit the schema and so you can do this engineering thing and you can force the model to do what you want but you might not get good outputs and sometimes with Json mode developers have seen that our models output like Whit space for a really long time where they don't because it's a legal character right it's legal per Json but it's not really what they want and so that's what happens when you do kind of a very engineering biased approach but the modeling approach is to also train the model to do more of what you want and so we did these together we trained a model which is significantly better than our past models at following formats and we did the entor to serve like this constrained decoding concept at scill so I think marrying these two is is why this feature is pretty cool you just mentioned starts and an with a curly brace and maybe people's minds go to a prefills in the cloud API how should people think about Json mode structured output prefills because some of them are like roughly starts with a curly brace and ask you for Json you should do it and then instructor is like hey here's the rough data scheme I us should use and how do you think about them so I think we kind of designed structured outputs to be the easiest to use so you just like the way you use it in our SDK I think is my favorite thing so you just create like a pantic object or a Zod object and you pass it in and you get back an object and so you don't have to deal with any of the serialization the pars helper yeah you don't have to deal with any of the serialization on the way in or out so I kind of think of this as the feature for the developer who is like I need this to plug into my system I need the function call to be exact I don't want to deal with any parsing so that's where structured outputs is tailored whereas if you want the model to be more creative and use it to come up with Json schema that you don't even know you want then that's kind of where Json mode fits in but I expect most developers are probably going to want to upgrade to structured outputs the thing you just said you just use interchangeable terms for for the same thing which is Tool uh function calling and structured outputs we've had uh disagreements or discussion before on the podcast about are they the same thing semantically they're slightly different they are yes because I think function calling API came out first yes then Json mode and we Ed to abuse function calling for Json mode right do you think we should treat them as synonymous no okay yeah please clarify yeah and by the way there's also tool calling yeah the history here is we started with function calling and function calling you know came from the idea of like let's give the model access to tools and let's see what it does and we basically had these internal prototypes of of what a code interpreter is now and we were like this is super cool let's make it an API but we're not ready to host code interpreter for everybody so you know we're just going to expose The Rock capability and see what people do with it but even now I think there's a really big difference between function calling and structured outputs so you you should use function calling when you actually have functions that you want the model to call right and so like if you have a database that you want the model to be able to query from or if you want the model to send an email or like you know generate Arguments for an actual action and that's the way the model has been like fine tuned on is to like treat function calling for actually calling these tools and getting their outputs the new response format is a way of just getting the model to respond to the user but in a structured way and so this is very different like responding to a user versus like you know I'm going to go send an email a lot of people were hacking function calling to get the response format they needed and so this is why we shipped kind of this new response format so you can get exactly where you want and you get kind of more of the models for boss it's like kind of responding in the way it would speak to a user and so less kind of just programmatic tool calling if that makes sense are you building something into dsdk to actually close the loop with the function calling because right now it Returns the function then you got to run it then you got to like fake another message to then continue the conversation they have that in beta the runs yes we have this in beta in the node SDK so you can basically python it's coming to python as well that's why I didn't know see yeah I'm a node guy so Javas it's already existed it's it's coming everywhere but basically what you do is you write a function and then you add a decorator to it and then you can basically there's this run tools method and it does the whole Loop for you which is pretty cool when I saw that in the node SDK I wasn't sure if that's because it basically runs it in the same machine yeah and maybe you don't want that right to happen yeah I think of it as like if you're prototyping and building something really quickly and just playing around it's so cool to just create a function and give it this decorator but you know you have the flexibility to do it however you like like you don't want it in a critical path of a web request I mean some people definitely will um you know it's just kind of the easiest way to get started but let's say you want to like execute this function on a job Q async then you know it wouldn't make sense to use that prior art instructure outlines Json former what did you study what did you you know credit or learn from these things yeah there's a lot of different approaches to this there's more fill-in theblank style sampling where you uh basically preform kind of the keys and then get the model to sample just the value there's kind of a lot of approaches here we didn't kind of use any of them wholesale but we really loved what we saw from the community and like the developer experiences we saw so that's where we took a lot of uh inspiration there was a question also just about constrained grammar this is something that I I first saw in llama CPP which seems to be the most let's should say academically permissive forevel yeah for those who don't know maybe I don't know if you want to explain it but they use back as nor form which you only learn in like college when you're working on programming languages and compilers I don't know if you like use that under the hood or you explore that yeah we we didn't uh use any kind of other stuff U we kind of built you know our solution from from scratch to meet our specific needs but I think there's a lot of cool stuff out there where you can supply your own grammar right now we only allow Json schema and a dialect of that but I think in the future it could be a really cool extension to let you supply a grammar more broadly and maybe it's more token efficient than Json so lot of opportunity there you mentioned before also training the model to be better function calling what's that discussion like internally for like resour it's like hey we need to get better Json mode and it's like well can't you figure it out on the API platform without touching the model like is there a really tight collaboration between the two teams yeah so I actually work on the API models team I guess we didn't quite get into what I do an API yeah what do you say it is you do here yeah so yeah I'm the I'm the tech lead for the API but also I work on the API models team and this team is really working on making the best models for the API and a lot of common deployment patterns are research makes a model and then you kind of ship it in the API but you know I think there's a lot you miss when you do that you miss a lot of developer feedback and things that are not kind of immediately obvious what we do is we get a lot of feedback from developers and we go and make the models better in certain ways so our team does model training as well we work very closely with our post training team and so for structured outputs it was a collab between a bunch of teams including Safety Systems to make you know a really great model that does uh structured outputs mentioning Safety Systems you have a refusal field yes uh you want to talk about that that seems like a yeah it's a little it's pretty interesting so you can imagine basically if you constrain the model to follow a schema you can imagine there being like a a schema supplied that it wouldn't it would add some risk or be harmful for the model to kind of Follow That schema and we wanted to preserve our model's abilities to refuse uh when something you know doesn't match our policies or is harmful in some way and so we needed to give the model an ability to refuse even when there is this schema but also you know if you are a developer and you have this schema and you get back something that doesn't match it you're like ah the feature's broken so we wanted a really clear way for developers to program against this so if you get something back in the content you know it's valid it's Json Parable but if you get something back in the refusal field it makes for a much better UI for you to kind of display this to your user in a different way and it makes it easier to program against so really there was a few goals but is mainly to allow the model to continue to refuse but also with a really good developer experience yeah why not offer it as like an error code because we have to display error codes anyway yeah we flaff for a long time about API design as we are want to do and there are a few reasons against an error code like you could imagine this being a 4xx error code or something but you know the developer paying for the tokens and that's kind of atypical for like a 4xx error code we pay with errors anyway right 4xx is not that's that's a that's a u error right and it doesn't as a 5xx either because it's not our fault you know the way the API the model is designed and I think the HTTP spec is a little bit limiting for AI in a lot of ways like there are things that are in between your fault and my fault there's kind of like the model's fault and there's no you know error code for that so we really have to kind of invent a lot of the Paradigm here make a 6xx yeah that's one option there's actually some like esoteric error codes we've considered adopting 328 my favorite yeah there's uh yeah there's the teapot one we're still figuring that out but I think there are some things like for example sometimes our model will produce tokens that are invalid based on kind of our language and when that happens it's an error but you know it doesn't 500 is fine which is what we return but it's not as expressive as it could be so yeah just areas where you know Web 2.0 doesn't quite fit with AI yet if you had to put in a spec just change what would be your number one proposal to like rehaul the HTTP committee to reinvent the world yeah that's going I mean I think we just need an error of like a range of model error and we can have many different kinds of model errors like a refusal is a model error 601 model refusal yeah again like so we we've mentioned before that chat completions uses this chat ml format so when the model doesn't follow chat ml that's an error um and we're working on reducing those errors but that's like I don't know 602 I guess a lot of people actually don't no longer know what Chad ml is yeah because that was uh briefly introduced by open the eye and then like kind of deprecated everyone who introd who implements this underhood knows it but maybe the the API users don't know it basically the API started with just one endpoint the completions endpoint and the completions endpoint you just put text in and you get text out and you can prompt in certain ways then we released chat gbt and we decided to put that in the API as well and that became the chat completions API and that API doesn't just take like a string input and produce an output it actually takes in messages and produces messages and so you can get a distinction between like an assistant message and a user message and that allows all kinds of behavior and so the format under the hood for that is called chat ml sometimes you know because the model is so out of distribution based on what you're doing Maybe temperature super high then it can't follow chat ml yeah I didn't know that there could be errors generated there maybe I'm not asking challenging enough questions it's pretty rare and we're working on driving it down but actually this is a side effect of structured outputs now which is that we have removed a class of Errors we didn't really mention this in the blog just because we ran out of space but uh that's what we're here to do yeah the model used to um occasionally pick a recipient that was invalid um and this would cause an error but now we are able to to constrain to chat ml in a more valid way and this reduces a class of errors as well recipient meaning so there's there's like a a few number of defined roles like user assistant system so like recipient as in like picking the right tool um so oh so the model before was able to to hallucinate a tool but now it's uh it can't when you're using structured outputs do you collaborate with other model developers to try and figure out these type of Errors like how do you display them because a lot of people try to work with different models yeah is there any yeah not a ton we're we're kind of just focused on making the best API for developers a lot of research and Engineering I guess comes together with evals you published some evals there I think I think gorilla is one of them what is your assessment of like the state of evals for function calling and structured output right now yeah we've actually collaborated with uh bfcl a little bit which is I think the same function calling leaderboard kudos to the team those Evils are great and we use them internally yeah we've also sent some feedback on some things that are misgraded but and so we're we're collaborating to to make those better in general I feel evals are kind of the hardest part of AI like when we talk to developers it's so hard to get started it's really hard to make a robust Pipeline and you don't want evals that are like 80% successful because you know things are going to improve dramatically and it's really hard to craft the right eval you kind of want to hit everything on the difficulty curve I find that a lot of these evals are mostly saturated like for bfcl all the models are near near the top already and kind of the errors are more I would say like just differences and default behaviors I think most of the models on the leaderboard can kind of get 100% with different prompting but it's more kind of you're just pulling apart different defa defaults at this point so yeah I would say in general we're missing evals you know we work on this a lot internally but it's hard did you other than bfcl would you call out any others just for people explor into space sbench is actually like a very interesting eval if people don't know you basically give the model GitHub issue and like a repo and just see how well it does with the issue which I think is super cool it's kind of like an integration test I would say for models it's a little unfair right what do you mean a little unfair cuz like usually as a human you have more opportunity to like ask questions about what it's supposed to do and you're giving the model like way too little information a hard job to do the job but yeah s bench targets like how well can you follow the diff format and how well can you like search across files and how well can you write code so I'm really excited about eiles like that because the pass rate is low so there's a lot of room to improve yeah and it's just targeting a really cool capability I've seen other evals for function calling where I think might be BFC as well where they they evaluate different kinds of function calling and I think the the top one that people care about for some reason I I don't know personally that this is so important to me but it's parallel function calling right I think you confirmed that you're you have don't support that yet why is that hard just more context about it so yeah we put out parallel function calling Dev Day last year as well and it's kind of the evolution of function calling so function calling V1 you just get one function back function calling V2 you can get multiple back at the same time and save latency we have this in our API all our models support it or all of our newer models support it but we don't support it with structured outputs right now and there's actually a very interesting trade-off here uh so when you basically call our API for structured outputs with a new schema we have to build this artifact for fast sampling later on but when you do parallel function calling the kind of schema we follow is not just directly one of the function schemas it's like this combined schema based on a lot of them if we were kind of do the same thing and build an index every time you pass in a list of functions if you ever change the list you would kind of incur more latency and we thought it would be really unintuitive for developers and like hard to reason about so we decided to kind of wait until we can support a no added latency solution um and not just kind of make it really confusing for developers mentioning latency that is something that people discovered is that there is an increased cost and latency for the first token for the first request yeah first request is that an issue is that going to go down time is is there just an overhead to parsing Json that is just insurmountable it's definitely not insurmountable and I think it will definitely go down over time we just kind of take the approach of of ship early and often um and you know if you if there's nothing in there you you don't want to fix then you probably ship too late um so I think we will get that latency down over time but yeah I think for most developers it's not a big concern CU you're testing out your integration you're you're sending some requests while you're developing it and then it's fast and prod so kind of works for most people the alternative design space that we uh explored was like pre-registering your schema so like a totally different endpoint and then passing in like a schema ID but we thought you know that was a lot of overhead and like another end point to maintain and just kind of more complexity for the developer and we think this latency is going to come down over time so it made sense to keep it kind of in chat completions I mean hypothetically if one were to ship caching at a future point it would basically be the Super set of that maybe I think the in space is a little underexplored like we've seen kind of two versions of it but I think yeah there's ways that maybe put less onus on the developer but you know we haven't committed to anything yet but we're definitely exploring opportunities for making things cheaper over over time is AI in agents just going to be a bunch of structure output and function calling one next to each other like how do you see you know there's like the model does everything where do you draw the line because you don't call these things like an agent API but like if I were a startup trying to raise a c round I would just do function calling and say this is an agent API so how do you think about the difference and like how people build on top of it for like a gentic systems yeah love that question one of the reasons we wanted to build structured outputs is to make agentic applications actually work so right now it's really hard like if something is 95% reliable but you're chaining together a bunch of calls if you magnify that error rate it makes your like application not work so that's a really exciting thing here from going from like 95% to 100% I'm very biased working on the apepi and working on function calling and structured outputs but I think those are the building blocks that we'll be using kind of to distribute this technology very far it's the way you connect like natural language and converting user intent into working with your application and so I think like kind of there's no way to build without it honestly like you need your function calls to work like yeah we wanted to make that a lot easier yeah and do you think the assistance kind of like API thing will be a bigger part as people build agents I think maybe most people just use messages and completion and so I would say the assistance API was kind of a bet in a few areas one bet is hosted tools so we have the file Search tool and code interpreter another bet was kind of statefulness it's our first stateful API it'll store you know threads and you can fetch them later I would say the hosted tools aspect has been really successful like people love our file Search tool and it's like saves a lot of time to not build your own rag pipeline um I think we're still iterating on the shape for the stateful thing to make it as useful as possible right now there's kind of a few end points you need to call before you can get a run going and we want to work to make that you know much more intuitive and easier over time one thing I'm I'm just kind of curious about did you notice any tradeoffs when you add more structured output it gets worse at some other thing that was like kind of you didn't think was related at all yeah it's a good question yeah I mean models are very spiky and RL is hard to predict and so every model kind of improves on some things and maybe is flat or neutral on other things yeah like it's it's like very rare to just add a capability and have no trade-offs and everything else so yeah I don't I have something off the top of my head but I would say yeah every model is a special kind of its own thing this is why we put them in API dated so developers can choose for themselves which one works best for them in general we strive to continue improving on all evals but it's stochastic yeah able to apply the structured output system on backdated models like uh 40 may as well as mini as well as August actually the new response format yeah is only available on two models it's 4 and the new 40 okay so the old 40 doesn't have the new response format okay however for function calling we were able to enable it for all models that support function calling and that's because those models were already trained to follow these schemas we basically just didn't want to add the new response format to models that would do poorly at it because they would just kind of do infinite white space which is you know the most likely token if you have no idea what's going on I just wanted to call out a little bit more in the in the stuff you've T in blog post so in blog post just use cases right I just want people be like yeah we're spelling it out for you use these for extracting structured data from unstructured data by the way it does Vision 2 right so that's cool Dynamic UI generation actually let's talk about Dynamic UI um I think gen UI I think is something that people are very interested in yeah is your first example what did you find about it yeah I just thought it was a super cool capability we have now so the schemas we we support recursive schemas and this allows you to do really cool stuff like you know every UI is a nested tree that has children and so I thought that was super cool you can use one schema and generate like tons of of uis as a back-end engineer who's always struggled with Javas script in front end like for me that's super cool I've now we've now built a system where I can get any front end that I want so that's super cool the extracting structured data like the reality of a lot of AI applications is like you're plugging them into to your Enterprise business and you have something that works but you want to make it a little bit better and so the reliability gains you get here is like you'll never get a like a classification using the wrong enum it's like it's just exactly your your types so really excited about that like maybe hallucinate the actual values right so let's clearly State what the guarantees are the guarantees is that they fits the schema but the schema itself may be too broad because the Json schema type system doesn't say like I only want to range from 1 to 11 you might give me zero give me 12 so yeah Json schema so this is actually a good thing to talk about so Json schema is extremely vast and we weren't able to support every corner of it so we kind of support our own dialect and it's described in the docs and there are a few trade-offs we had to make there so by default if you don't pass in additional properties in a schema by default that's true and so that means you can get other Keys which you know you didn't spell out which is kind of the opposite of what developers want you basically want to supply the keys and values and you want to get those keys and values and so then we had a decision to make it's like do we redefine what additional properties means as the default and that felt really bad it's like there's a scheme of that's predated us like you know it wouldn't be good it would be better to play nice with the community and so we require that you pass it in as false you know one of our design principles is to be very explicit and so developers know you know what to expect and so this is one where we decided you know it's a little harder to discover but we think you should pass this thing in so that we can have like a very clear definition of what you mean and what we mean there's a similar one here with like required by default every key in Json scheme is optional but that's not what developers want right like you would you'd be very surprised if you passed in a bunch of keys and you didn't get some of them back and so that's the trade-off we made is to make everything required and have the developers spell that out is there a require false can people turn it off or they're just getting all so developers can basically what we recommend for that is to make your actual key a union type and so yeah make it Union of int and null and that gets you the same behavior any other of the examples you want to dive into math Chain of Thought yeah you can now specify like a Chain of Thought Field before a final answer this is just like a more structured way of extracting The Final Answer yeah one example we have I think we put up a demo app of this math tutoring example or it's coming out soon I miss it oh okay well basically it's this math tutoring thing and you put in an equation and you can go step by step and insert this is something you can do now with structured out in the past a developer would have to like specify their format and then write a parser and parse out the model's output which would be pretty hard but now you just specify like steps and it's an array of steps and every step you can render and then the user can try it and you can see if it matches and go on that way so I think it just opens up a lot of opportunities like for any kind of UI where you want to treat different parts of the model's responses differently structured outputs is great for that I remembered my my question from earlier I'm basically just using this to ask you all the questions as a user as a daily daily user of the stuff that you put out so one is a tip that people don't know and I confronted it to you on Twitter which is you respect descriptions of Json schemas right and you can basically use that as a prompt for the field totally I assume that's blessed and you know people should do that right one thing that I started to do which I don't it could be hallucination of me is I Chang the the property name to to to prompt the model to what I wanted to do so for example instead of saying topics as a property name I would say like brainstorm a list of topics up to five or something like that as as like a property name I I could stick that in the description as well but is that too much yeah I would say I mean we're so early in AI that people are figuring out the best way to do things and I love when I learn from a developer like a way they found to make something work in general I think there's like three or four places to put instructions yeah you can put instructions in the system message and I would say that's helpful for like when to call a function so it's like you know let's say you're building a customer support thing and you want the model to verify the user's phone number or something you can tell the model in the system message like here's when you should call this function then when you're within a function I would say the descriptions there should be more about how to call a function so really common is someone will have like date as a string but you don't tell the model like do you want year year month month day day or do you want that backwards and that's what a really good spot is for those kind of descriptions is like how do you call this thing and then sometimes there's like really stuff like what you're doing it's like name the The Key by what you want so sometimes people put like do not use and you know if they don't want you know this parameter to be used except only in some some circumstances and really I think that's the fun nature of this it's like you're figuring out the best way to get something out of the model okay so so you don't have official recommendation is what I'm hearing well the official recommendation is you know how to Cola model system instructions exactly exactly that function yeah do you Benchmark these type of things so like same with date it's like description it's like return it and like ISO a or if you call the key date in ISO a6001 I feel like the benchmarks don't go that that deep but then all the AI engineering kind of community like all the work that people do is like oh actually this performs better but then there's way to verify right you know like uh even the I'm going to tip you $100,000 or whatever like some people say it works some people say it doesn't do you pay attention to the stuff as you build this or are you just like the model is just going to get better so why waste my time running evals on these small small things yeah I would say to that I would say we basically pick our battles I mean there's so much surface area of llms that we could dig into and we're just Mo mostly focused on kind of raising the capabilities for everyone I think for customers and we work with a lot of customers really developing their own evals is super high leverage cuz then you can upgrade really quickly when we have a new model you can experiment with these things with confidence so yeah we're we're hoping to make making evals easier I think that's really generally very helpful for Developers for people I would just kind of wrap up the discussion for structured outputs I immediately implemented we use structured outputs for AI news I use instructor and I ripped it out and I think it I saved um 20 lines of code but more importantly it was like we cut it by 55 % of API cost based on what I what I measured because of we saved on the retries nice love to hear that yeah which which people I think don't understand when you you can't just simply like add instructor or add outlines you can do that but it's actually going to cost you a lot of retries to get the the model that you want but you're kind of just kind of building that internally into the model yeah I think this is the kind of feature that works really well when it's integrated with like the llm provider yeah actually I had folks even my my husband's company who works at a small startup they thought we were just retrying um so I had to make clear we are not retrying you know we're doing it in one shot and this is how you save on latency end cost awesome any other behind the scenes stuff just generally unstructured outputs we we're going to move on to the other models yeah I think that's it oh look that's an excellent product and I think everyone will be using it and we have the full story now that people can try out so road map would be parallel function calling anything else that you've called out as like coming soon uh quite soon but you know we're thinking about does it make sense to expose was custom grammars um Beyond Jon schema what would you want to hear from developers to give you information whether it's custom grammars or anything else about structured output like what would do you want to know more of just you know always interested in in feature requests what's not working but I'm I'd be really curious like what specific grammars folks want I know some folks want to match programming languages like python there's some challenges like with the expressivity of our you know implementation and so yeah just kind of the class of grammars folks want I have a very simple one which is a lot of people try to do use GPT as judge right which means they end up doing a rating system and then there's like 10 different kinds of rating systems there a lyer scale is whatever if there was an officially blessed way to do a rating system with structured outputs tot everyone would use it yeah yeah that makes sense I mean we often recommend using log probs with classification tasks so rather than like sampling you know let's say have four options like red yellow blue green rather than sampling you know two tokens for yellow you can just do like ABCD and get the log probs of those you know the inherent randomness of each sampling isn't taken into account and you can just actually look at what is the most likely token I think this is more of like a calibration question like if I ask you to rate things from 1 to 10 a non-calibrated model might always pick seven just like a human would right so like actually have a nice gradation from 1 to 10 would be the the the rough idea yeah and then even for structured outputs I can't just say have a field of rating from 1 to 10 because I I have to validate it and you know it might give me 11 yeah absolutely yeah so what about model selection now you have a lot of models when you first started you had one model endpoint I guess you had like the and then but like most people were using one model endpoint today you have like a lot of competitive models and I think we're nearing the end of the 3.5 run rip how do you advise people to like experiment select both in terms of like task and like cost like what's your playbook in general I think folks should start with 40 mini that's our cheapest model and it's a great work Workhorse uh works for a lot of great use cases if you're not finding the performance you need like you know maybe it's not smart enough then I would suggest going to 40 and if 40 works well for you that's great finally there's some like really Advanced Frontier use cases um and maybe 4 I is not quite cutting it and there I would recommend our fine-tuning API even just like a 100 examples is enough to get started there and you can really get the performance you're looking for we're recording this ahead of it but like you're announcing other some fine tuning stuff that people should pay attention to yeah actually tomorrow we're dropping our GA for gbt 40 fine tuning so 40 mini has been available for a few weeks now and 40 is now going to be generally available and we also have a free training offering for a bit I think until September 23rd you get 1 million of free training tokens a day this is already announced right oh was am I talking about a different so that was for 40 mini and now it's also for 40 so we're really excited to see what people do with it and it's actually a lot easier to get started than a lot of people expect they think they might need tens of thousands of examples but even 100 really high quality ones or a thousand is enough to get going oh well we might get a separate podcast just specifically on that but um you know we haven't confirmed that yet it basically seems like every time I think people's concerns about fine tuning is that they're kind of locked into a model and I think you're Paving the path for migration of models as long as they keep their original data set like they can at least migrate nicely yeah I'm not sure we've said publicly there yet but we definitely want to make it easier for folks to to migrate it's the number one concern you know I'm just you know it's obvious absolutely I also want to point people to you have official model selection docs where it's it's on in the guide we'll put it in the show notes where it says to optimize for accuracy first so prompt engineering rag EV Val fing this was done at Dev Day last year so I'm just repeating things and then optimize for cost and latency second and there's a there's a few sets of steps for optimizing latency so people can read up on that stuff yeah totally yeah we had one episode with um Nicholas Carini from Deep Mind and we actually talked about how some people don't actually get to the boundaries of the model performance you know they just kind of try one model and it's like Oh llms cannot do this and they stop right how should people get over their hurdle it's like how do you know if you hit the model performance or like you hit skill issues you know it's like your prompt is not good or like try another model and whatnot is there an easy way to do that that's tough some people are really good at prompting and they just kind of get it right away and and for others it's it's more of a challenge I think there's a lot we can do to make it easier to prompt our models but for now I think requires a lot of creativity and not giving up right away yeah and a lot of people have experienced now with chat GPT you know before chat GPT the easiest way to play with our models was in the playground but now kind of everyone's played with it with the model of some sort and they have some sort of intuition it's like you know if I tell you my grandma is sick then maybe I'll get the right outut and and we're hoping to kind of remove the need for that but playing around with chat gbt is a really good way to get a feel for you know how to use the API as well will prompt engineering be here forever or is it a dying guard as the models get better I mean it's like The Perennial question of software engineering as well it's like as the models get better at coding you know if we hit 100 on S bench what does that mean I think there will always be Alpha in people who are able to like clearly explain what they're trying to build most of engineering is like figuring out the requirements and stating what you're trying to do and I believe this will be the case with AI as well you're going to have to very clearly explain what you need and some people are better than others at it and people will always be building it's just the tools are going to get far better last two weeks You released two models there's GBC 40224 806 and then there's also chat GBC 4 latest I think people a little bit confused by that and then you you issued a clarification that was one's chat tuned and the other is more function calling tuned can you elaborate just yeah totally so part of the impetus here was to kind of very transparent with what's on chat gbt and in the API so basically we're we're often Trading models and and they're different use cases so you don't really need function calling for userdefined functions in chat gbt and so this gives us kind of the freedom to build the best model for each use case so in chachu BT latest we're releasing kind of this rolling model the weights aren't pinned as we release new models this is literally what we use yeah so it's in what's what's in chat GPT so it's very good for like chat style use cases but for the API broadly you know we really tune our models to be good at things that developers want like function calling and structured outputs and when a developer builds their application they want to know that kind of the weights are stable under them and so we have this offering where it's like if you're tuning to a specific model and you know your function works you know it will never change the weights out from under you and so those are the models we commit to supporting for a long time and we think those are the best for developers but we want to give it up you know we want to leave the choice to developers like do you want the chat gbt model or do you want the API model and you have the freedom to choose what's best for you I think it's for people they they do want to pin model versions so I don't know when they would use CH GPT like the the rolling one unless they're really just kind of cloning chbt which it's like why would they I mean I think there's a lot of interesting stuff that developers can do when unbounded and so we don't we don't want to limit them artificially so it's kind of survival of the fittest like whichever model is better you know that's the one that people should use yeah I talked about it to to my friends as like this isn't that new thing like and basically opening has has never actually shared with you the actual chat gbt model uh and now they do that's well it's not necessarily true actually a lot of the models we have shipped have been the same okay but you know sometimes they diverge and it's not a limitation we want to stick around anything else we should know about the new model I don't think there were there were e there was no evals announced or anything but but people say it's better I mean obviously lmis is like way better above on everything right is like number one in well done yeah we're um we published some release notes they're not as in depth as we want to be at because we're still it's still kind of a science and we're learning what actually changes with each model and how can we better understand the capabilities but we are trying to do more release notes in the future and and keep folks updated but yeah it's it's kind of an art and a science right now you need the best evals team in the world to help you figure this out yeah evals are hard we're hiring if you want to come work on evals hold that thought on hiring we'll come back to the end on what you want what you're looking for because obviously people want to join you and they want to know what uh quality you're looking for so we just talked about API versus chbt what's that guess like the vision for the interface you know the the mission of open is like build a gii that is accessible like a where is it going to come from totally yeah so I believe that uh the API is kind of our broadest vehicle for Distributing AGI um you know we're building some first-party products but they'll never reach every niche in the world and kind of every corner in community and so really love working with developers and seeing the incredible things they come up with I often find that developers kind of see the future before anyone else and we love working with them to make it happen and so really the API is a bet ongo really Broad and we'll go very deep as well in our first-party products but I think just that our impact is absolutely magnified by every developer that we uplift they can do the last M where where you cannot like CHT is one type of product but there's many other kinds uh in fact uh you know I I observed I think in February basically chat gpt's user growth stopped when the API was launched because everyone's like kind of be able to take that and and build other things that has not become true anymore because chbc growth has has continued to grow but then you're not confirming any of this this is me quoting similar web numbers which are have very high variance well the API predates CH the API was actually opening his first product and the first uh idea for commercialization that predates me as well wide release like GA everyone can sign up and use it immediately yeah that's what I'm talking about but but yeah I mean I I I do believe that uh and you know that you also have to expose all of openi models right like U all the multimodal models we'll ask you questions on that but like I think that that API mission is is important it's interesting that hottest new programming language is supposed to be English but it's actually just software engineering right it's it's just you know we're talking about HTTP error codes right like yeah I think you know engineering is still the way you access these models and I think there are companies working on on tools to make engine ing more accessible for everyone but there's still so much Alpha in just writing code and and deploying yeah one might even call it AI engineering exact I know yeah so like there's lots of War Stories from from building this platform we started at the start of your career and then we jump straight to structured outputs there's a whole thing like two years that we skipped in between right what have become your principles what are your favorite stories that that you like to tell we had so much fun working on the assistance API and leading up to Dev day you know things are always pretty chaotic when you have an ex Al like a date that is hard and there's like a stage and there's like a thousand people coming um you can always launch a weight list I mean we're trying hard not to um because you know we love it when people can access the thing on day one and and so yeah the assistant API we had like this really small team and just working as hard as we could to make this come to life but even actually the morning of I don't know if you'll remember this but Sam did this keynote yep and Raman came up and they free credits to everybody so that was live Fully live as were all of the demos that day um but actually maybe like 2 hours before that we had a little outage and everyone was like scrambling to make this thing work again so we're yeah things are are early and Scrappy here and you know we were really glad we were bit on the edge of our seat watching it live what's the plan B in that situation if you can share it's like play video this is classic de right I don't know I mean I I actually don't know what the plan B was no plan B no failure but we just you know we fixed it we got got everything running again and uh the demo went well just higher cracked water loop skill issues as usual sometimes you just got to make it happen I I imagine it's actually very motivating but I I did hear that after Dev day like the whole company got like a few weeks off just to relax a little bit yeah we um we sometimes get like we just had the week of July 4th off and yeah it's hard to take vacation because people are working on such exciting things and it's like you get a lot of fomo on vacation so it helps when the whole company's on vacation mentioning assistance API you actually announced a road map there and things have developed um I think people may not be up to date what's the offering today versus you know one year ago yeah so we've made a bunch of key improvements I would say the biggest one is in the file search product before we only supported I think like 20 files per assistant and the way we used those files was like like less effective basically model would decide based on the file name whether to search a file and there's like not a ton of information in there so our new offering which we shipped a few months ago I think now allows 10K files per assistant which is like dramatically more and also it's a it's a kind of different operation so you can search semantically over all files at once rather than just kind of the model choosing one up front so a lot of customers have seen really good performance we also have exposed more like chunking and reranking options I think the reranking one is coming I think next week or very soon so this kind of gives developers more control and more flexibility there so we're trying to make it the easiest way to kind of do rag at scale yeah I think that visibility into the the rag system was the number one thing missing from De day and then people got their first impressions and then they never looked at it again so that's that's important the ranker is a core feature of let's say some other Foundation model Labs is opening ey going to like offer a reranking service ranker model so we do reranking as part of it I think we're soon going to ship more controls for that okay got it and like if I'm an existing Lang chain llama index whatever how do you compare do you make different choices like where is that exist in the spectrum of choices I think we are just coming at it trying to be the easiest option and so ideally like you don't have to know what a ranker is and you don't have to have a chunking strategy and the thing just kind of works out of the box so I would say that's where we're going and then you know giving controls to the power users to to make the changes they need awesome I'm going to ask about a couple other things just updates on stuff also announced at Dev day and we talked about this before uh determinism something that people really want Dev day would announc the seed pram as well system fingerprint and like objectively I've heard issues yeah I don't know what's going on yeah the seed parameter is not fully deterministic and it's kind of a best effort thing yeah so you'll notice there's more determinism in the first few tokens that's kind of the current implementation we've heard a lot of feedback we're thinking about ways to make it better but it's challenging it's kind of trading off against you know reliability and of time other maybe on rated API only thing loger bias uh that's another thing that kind of seems very useful and that maybe most people are like it's a lot of work I don't want to use it uh do you have any examples of like use cases or like products that are made a lot better through using it so yeah classification is the big one so logic buys your valid classification outputs and you know you're more likely to get something the matches we've seen it people logic bu like punctuation tokens maybe trying to get more succinct writing yeah it's generally a very much a power user feature and so not a ton of folks use it I actually wanted to use it to reduce the incidence of the word delve yeah have people done that probably I don't know is Del one token you're probably you got to do a lot of permutation it's do so much may it is depends on the tokenizer are there non-public tokenizers that I guess you cannot answer or you would ad made it are the 100K and 200k vocabs like the ones that you use like cross models or yeah I think we have docs that publish more information I don't have it off the top but I think we publish which tokenizers for which model okay so those are the only two rate the the tring rate limiting system I don't think there was an official like blog post kind of announcing this but it was kind of mentioned that like you started tying like fine tuning to to tearing uh and like feature rollouts just on the from from your point of view like how do you manage that and like what should people know about the tiering system and rate limiting yeah I think basically the main changes here were to be more transparent and easier to use so before developers didn't know what tier they're in and and now you can see that in in the play in the dashboard I think it's also I think we published like how you move from tier to tier and so this just helps us do kind of gated rollouts for the fine tuning launch I think everyone tier two and up has full access that makes sense that you know would just advise people to just get to tier five as quickly as possible sure like a goar customer you know like I don't know it seems to make sense do we want to maybe WP with future things and kind of like how you think about designning and everything so you just mentioned you want to be the easiest way to basically do everything what's the relationship with other people building in the developer ecosystem like I think at maybe in the early days it's like okay we only have these apis and then everybody helps us but now you're kind of building a whole platform how do you make decisions yeah I think kind of the 8020 principle applies here we'll build things that kind of capture you know 80% of the value and and maybe leave the long tail to other developers so we really prioritized by like how much feedback are we getting how much easier will make this will this make something like an integration for veler so yeah we we we want to do more in this space and not just be an LM as a service but kind of AI development platform as a service oo okay that ties into a thing that I put in the in the notes that we prepped there are other companies trying to be AI development platform so you will compete with them or they just want to know what you won't uh build so that they can build it yeah it's a tough question I I think we haven't you know determined what exactly we will and won't build but you can think of something if it makes it a lot easier for developers to integrate you know it's probably on our radar and we you know stack Rank by impact yeah I so there's like cost tracking and model fallbacks model fallbacks is an interesting one because people do it I don't think it adds a ton of value but like if you don't build it I have to build it because if one API is down or something I I need to fall back to another one yeah I mean the way we're targeting that user need is just by investing a lot in reliability and so we just don't fail I mean we have improved our uptime like pretty dramatically over the last year and it's been you know the result of a lot of hard work from folks so you'll see that on our status page and and our continued commitment going forward is the important thing about owning the platform that gives you the flexibility to put all the kind of messy stuff behind the the scenes or yeah how do you draw the line between what you want to include yeah I just think of it as like how can we on board the next generation of of AI Engineers as you put it right like what's the easiest way to get them building really cool apps and I think it's by building stuff to kind of hide this complexity or just make it really easy to integrate so I think of it a lot as like what is the value ad we can provide Beyond just the models that makes the models really useful okay we'll touch on four more features of the API platform that we prepped batch Vision whisper and then team Enterprise stuff so you wanted to talk about batch yeah so the rough idea is you give a the contract between you and me is that I give you give you the bat job you have 24 hours to run it at it's kind of like spot in for for the API what what what should people know about it so it's half off which is a great savings it also works with like 40 mini so the savings on top of 40 mini is is pretty crazy like the stuff you can do 7.5 cents or something per million yeah I should really have that number top of mine but it's like staggeringly cheap so I think this opens up a lot more use cases like let's say you have a user activation flow and you want to send them an email like maybe every day or like at certain points in their user Journey so now you can do this with the batch API and something that was maybe a lot more expensive and not feasible is now very easy to do so right now we have this 24-hour turnaround time for half off and and curious would love to hear from your community like what kind of turnaround time do they want I would be an ideal user or batch and I cannot use batch because it's 24 hours I need two to four two to four hours okay yeah that's good to know yeah just a lot of folks haven't heard about it it's also really great for like evals running them offline you don't generally don't need them to come back within you know two hours I think you could do a range right 2 to four for me like I need to produce a daily thing and then 24 for like the the average use case and then maybe like a week a month who cares like for people who just have a lot to do yeah absolutely so yeah that's fatch API I think folks should use it more it's pretty cool is there a future in which like 6 months is like free you know like is there like small is there like super small like shards of like GPU run time that like over a long enough timeline you can just run all these things for free yeah it's certainly possible I think we're getting to the point where a lot of these are like almost free that's true already why would they work on something that's like completely free I don't know okay so Vision Vision got G last year people were so wilded by the gpc4 demo and that was primarily Vision um what was it like building the vision API yeah the vision API is super cool we have a great team working there I think the cool thing about vision is that it works across our apis so there's you can use it in the assistance API you can use in the batch API and chat completions it works with structured outputs I I think it just helps a lot of folks with kind of data extraction where you know the spatial relationships between the data it is too complicated and you can't get that over text but yeah there's a lot of really cool use cases I think the the tricky thing for me is understanding how frequent to turn Vision into from like single images into like effectively just always watching and right now I think people just like send a frame every every second will that model ever change would will there just be like I stream you video and then yeah I think it's very possible that we'll have an API where you stream video in and maybe you know to start we'll do the frame uh sampling for you the frame sampling is is the default right right but I I feel like it's hacky yeah I think it's hard for developers to do and so you know we should definitely work on making that easier it's there in the batch API do you have like um time uh guarantees like order guarantees like if I send you a batch request of like a video analysis I need every frame to be done in order for batch you send like a list of requests and each of them stand alone so you'll get all of them finished but they don't kind of chain off each other well if you're doing a video you know if you're doing like analyzing a video I wasn't linking video to batch but that's interesting yeah well a video is like you know if you have a very long video you can just do a batch of all the images and let it process but but Ser sequential true yeah yeah yeah exactly but the whole point of batch is you're just using kind of spare time to run it let's talk about my favorite model whisper all of our I I built yeah yeah build this thing called small podcaster which is a open source tool for podcasters and why does whisper apbi not have diarization when everybody is transcribing people talking that's my main question yeah it's a good question and you've come to the right person I actually worked on the whisper API and shipped that that was one of my first apis I shipped long story short is that like whisper V3 which we open sourced has I think diation feature but there's some like performance trade-offs and so whisper V2 is better at some things than whisper V3 and so it didn't seem that worthwhile to ship whisper V3 compared to like the other things in our priorities I think we still will at some point but yeah it's just you know there's always so many things we could work on it's it's tough to do everything we have a python notebook that does the diarization for the Pod but I would just like you can translate like 50 languages but you cannot tell me who's speaking that was like the funniest thing there's like an XKCD thing about this about hard problems in AI I I forget tell me taken in a park and like that's easy it's like tell me if there's a bird in this picture and it's like give me 10 people in a research team it's like you never know which things are are challenging and diarization is I think you know more challenging than than expected yeah yeah it still breaks a lot with like overlaps obviously sometimes similar voices it struggles with like I need to like double read the thing totally but yeah great model I mean it would take us so long to do transcriptions and I don't know why like small podcast has better transcription than like mostly every commercial tool itats the script and I'm like I'm just using the model I'm literally not doing anything you know it's just a notebook so yeah it just speaks to like sometimes just using the simple openi model is better than like figuring out your own pipeline thing totally I think the top feature request there just would be I mean again you know using you as a feature request dump is uh is like being able to bias the vocab I think there is like in raw whisper you can do that you can pass the prompt in the API as well in but you pass it in the prompts okay yeah there's no more deterministic way to do it so you this is really helpful when you have like acronyms that aren't very familiar to the model and so you can put them in the prompt and you you'll basically get the transcription using those correctly we have the AI engineer solution which is just a dictionary nice we like all the way misspelt it in the pass and then G sub and like replace the if it works it works like that's that's Eng it's like you know llama with like one L like all all these different things or like length chain it like transcribes length chain and like capitalization does a bunch of like three four different ways um you guys should try PR feature I love these like kind of pro tip okay fun question I know we don't know yet but I've been enjoying the advanced voice mode it really streams back and forth uh and it handles interruptions how would your audio endpoint change when when that comes out we're exploring you know new shape of the API to see how it work would work in this kind of speech to speech Paradigm I don't think we're ready to share quite yet but we're definitely working on it I think just the regular request response probably isn't going to be the right solution for those who are listening along I think it's pretty public that open uses life kit for the for the chat gbt app which like seems to be the socket based approach that people should be at least up to speed on like I think a lot of developers only do request response and like that doesn't work for streaming yeah when we do put out this API I think we'll make it really easy for developers to figure out how to use it yeah it's hard to it'll be a paradig change okay and then I think the last one on our list was team Enterprise stuff audit logs service accounts API Keys what should people know the Enterprise offering yeah we recently shipped our admin and audit log apis and so a lot of Enterprise users have been asking for this for a while the ability to kind of manage API Keys programmatically manage your projects get the auto log so we've shipped this and and for folks that need it it's it's out there and happy for your feedback yeah awesome I don't use them so I don't know I imagine it's just like build your own internal Gateway for you know your internal developers to manage your deployment of openi yeah I mean if you work at like a company that needs needs to keep track of all the Epi Keys it was it was pretty hard in the past to do this in in the dashboard we've also improved our SSO offering so that's uh much easier to use now the most important feature of Enterprise company love SSO all right let's go outside of openi what about just you personally so you mentioned waterl maybe let's just do why is everybody a waterl cracked and why are people so good and like why people not replicated it or any other commentary on your experience the first is the co-op program it's obviously really good you know I did six internships learned so much in those I think another reason is that waterl is like you know it's very cold in the winter it's pretty miserable there's like not that much to do apart from study and like Heck on projects and and there's this big like hacker mentality you know there's a hack the north is a very popular hackathon and there's a lot of like startup incubu it's kind of just has this like startup and hacker ethos then that combined with the six internships means that you get people who like graduate with two years of experience and they're very entrepreneurial and you know they're down to grind I do notice a correlation between climate and the crack this of Engineers uh so you know no coincidence that Seattle is the birthplace of Microsoft and and Amazon I think I had this compilation of Denmark where uh people like so it's the birthplace of C++ PHP turbo Pascal standard ml BNF the the thing that we just talked about md5 Crypt Ruby un rails Google Maps and V8 for for Chrome and it's CU according to beond Sr the kto C++ there's nothing else to do yeah well you have Lena stal in Finland right so I mean you hear a lot about this like in relation to SF people say you know New York is way more fun there's nothing to do on SF and maybe it's a little by design that Al alltech is here the climate is too good if we had if we also have fun things to do nature is so nice you can touch grass why why we not touching grass you know restaurants close at like 800 p.m. like that's what people are referring to there's not a lot of like late night dining culture yeah so you you have time to wake up early and and get to work you are a book recommender or book enjoyer uh what underrated books do you recommend most others yeah I think a book I read somewhat recently that was very formative was the making of the Prince of Persia it's a stripe press book that book just made me want to work hard like nothing I've ever read it's just like this journal of of what it takes to like build you know incredible things so I'd recommend that yeah it's funny how video games are for a lot of people at least for me kind of like the some of the moments in technology like when I played the sense of time on PS2 was like my first PlayStation 2 game MH and I was like man this thing is so crazy compared to any Playstation One game and it's like wow my expectations for like the technology I think like open eyes a lot of similar things like the advanced voice it's like you see that thing and then you're like okay what I can expect from everybody else is kind of raised now you know totally another book I like to plug is called misbehaving by Richard Taylor he's a behavioral Economist and talks a lot about how people act irrationally in terms of decision- making and I actually think about that book like once a week probably at least when I'm making a decision and I realize that you know I'm falling into a fallacy or you know it could be a better decision yeah you did a minor iny I did yeah I don't know if I learned that much there but it was interesting did you is there like an example of like a cognitive bias or misbehavior that you just love telling people about yeah people so let's say you won tickets to like a Taylor Swift concert and I don't know how much they're going for but it's probably like $10,000 oh okay or whatever sure and like a lot of people are like oh I have to keep these like I won them it's $10,000 but really it's the same decision you're making if you have $10,000 like would you buy these tickets and so people don't really think about it rationally like would they rather have $10,000 or the tickets for people who want it a lot of the time it's going to be the $10,000 but their bias is because they want it the world organized itself this way and you should keep it for some reason yeah yeah oh okay I'm pretty familiar with this stuff that there's also a loss version loss AV verion version of this where it's like if I take it away from you you respond more if I it to you yes yeah people are like really upset if they like don't get a promotion but if they do get a promotion they're like okay few it's like not even you know excitement it's more like we re react a lot worse to losing something which is why like when you join like a new platform they often give you points and then they'll take it away if you like don't do some some action uh in in like the first few days yeah totally yeah uh the book references people who work like operate very rationally as econs as like a separate group to humans and I often think like you know what would an econ do here in this moment and and try to act that way okay let's let's do this our llms econs I mean they are maximizing probability distributions minimizing loss yeah so I think way more than all of us they are econs whoa okay so they're more rational than us I think their optimization functions are more clear than ours yeah just to wrap you mentioned you need help on a lot of things any specific roles call outs and also people's backgrounds like is there anything that they need to have done before like what people fit well at open yeah we've hired people all kinds of backgrounds people who have PhD and and ml or or folks who just done engineering like me and we're really hiring for a lot of teams we're hiring across the applied org which is where I sit for engineering and for a lot of researchers and there's a really cool Model Behavior role um that we we just dropped so yeah across the board would recommend checking out our careers page and and you don't need a ton of experience in AI specifically to to join I think one thing that I'm trying to get at is like what kind of person does well at open a I think objectively you have done well and I've seen other people not do as well and and basically be managed out I know it's an intense environment I mean the people I enjoy working with the most are kind of low ego do what it takes ready to roll up their sleeves do what needs to be done and and unpretentious about it yeah I also think folks that are are very user focused do well on on kind of API and chat gbt like the YC eths of build something people want is is very true and opening ey as well so I would say low ego user focused driven um cool yeah this was great thank you so much for coming on thanks for having me