从 OpenAI 到 Core Automation:Jerry Tworek 谈推理模型与自动化 AI 实验室
From OpenAI to Core Automation: Jerry Tworek on Reasoning Models and Automated AI Labs

Hi uh my name is Rocky. Uh welcome to the AGI house. Uh so this summit is about one question and uh what happens when AI started doing research itself. So uh our next guest is uh speckling about that that future isn't speculating about that future. He's betting his company on it. He spent seven years at open AAI where he led the team that taught language model to think 01 03 and the reasoning inside GP5. Before that he was a primary contributor to codeex the model behind the co github copilot and before any of that he taught a robot hand to solve a rubric cube.
In January, he walked away from one of the most powerful research posts in AI and April, he launched the core automation with the explicit goal of building the world's most automated AI lab. Please welcome Jerry Tro. >> Thank you very much. Thank you for inviting me and happy to be here. >> Absolutely. So, um we'll just get straight into it. It's like a technical peer conversation. So, it's uh lot to hear a lot of your your stakes on state of art of AI.
So was you spent five years at quant in Amsterdam extracting signals from noisy future markets. What did trading teach you about the most ML researchers never learn? I think that's the first questions. >> That's a great question and and I think there are some skills and there are some some necessary bits that that transfer. What mostly trading has taught me about machine learning research is about being paranoid about your algorithm and finding bugs.
When you are implementing a trading algorithm slightly slightly similar like ML, you are building code that manages a whole bunch of numbers and depending on like where you get good output, good input, you have good output. But if something gets wrong, what usually you just get a wrong number in the end. That number may mean you buy instead of sell or or or do something wrong with large sums of money. So being a good trader means having to understand your algorithm very well and knowing whether it works well, whether it is doing what you intended it to do or or or not because in the end what you see the output is just just P&L.
And the first thing when I started doing AI research, I was very paranoid about everything that that happened inside of a of a network of a learning algorithm of hyperparameters and I was always double and triplecheing everything to make sure I know what is going on and making sure that the algorithm is doing exactly what what I want to be doing because previously my algorithms were were uh trading large sums of money and that's that's a superpower in many ways to uh doing AI research very well.
This is one thing that I would recommend every researchers that are becoming researchers in the current times of AI research being largely automated and technically models can do a lot of your but there's no way to outsource your understanding. So it's it's it's good to good to know what is what is inside of the box. >> I see. Yeah. Coincidentally, we actually observed a lot of the the the great researcher come from the trading background.
If you remember the company DeepS I think the the the whole company was around there and now they are doing that. Uh I remember last time we have this like AJI dinner series. So uh one of the co- leader Google Gemini Oreo Vinet he was here we're talking about uh actually one of our guest asked the question he was also doing trading backgrounds like they have a clear objectives maximize the return does that create this kind of right environment to for for for researchers to do the the experiments do the do the research >> yeah yeah I think trading is a great environment and there are there are obviously there's no perfect environment in the world everyone has every environment has its upside sides and downsides for trading this like relatively small signal to noise which the together with sample inefficiency of RL does create some amount of challenges and another one is that feedback is not often very quick sometimes you have to you have to wait a little bit but beyond that I think I think some people are already very successful at applying reinforcement learning to research this was this was in some way the decision I have made in 2018 do I want to go in the AI lab and try to apply the AI for programming for doing research or do I want to apply reinforcement learning to trading?
I decided to join an AI lab. I probably would have been one of the one of the early people to doing to doing reinforcement learning in trading but it does seem to be a very successful field right now the current state of technology. >> Yeah, absolutely. We've seen that. Uh so once you said you're a mathematician at heart and uh where does a mathematical taste actually show up in the frontal research nowadays? A lot of it is thinking from the from the first principles in some way in some way mathematics is the language of abstraction and a lot of how at least like how human mathematics works because I I think when the models will take over the field and will probably happen pretty soon.
Mathematics will will will very differ and and and the models have the ability to operate much more objects in their head and much more objects in their in their in their attention span. But what human mathematicians do, they essentially build towers of abstractions to be able to talk about complex objects relatively easily for them. And everything has to be as you know in mathematics completely formal. Everything has to be tight and all all arguments have to have to like be checkable and verifiable.
Follow follow the rules of logic. And mathematics is about constructing how do you follow the rules of logic to com to construct more and more abstract objects. In some way it allows you to have a good bird's eyee view on machine learning research. What are we actually computing? How how are we computing it? If you can if you can build a good abstraction around it and start manipulating this abstraction, I I think it's a very powerful field to to start understanding why our animals can do certain things, why they cannot do certain things, what are the actually loops, what what are the inputs, what are their outputs, like all those things are in some way mathematical object that we can that we can keep manipulating and building understanding of.
>> Cool. Yeah, that's goes back to the root mess is still uh very very useful on this. Well, we we're going to um the the the core one of the things that you did at OpenAI. So, you started OpenAI research into training on code and when there's a primary contributor, you are the primary contributor to the codecs. Why was code the the first verticals that you are working on? Was it the data the verifiability or something else that you see at that time?
>> Great question. In some way though, it seems to have been a very good decision back in the 21. But but but like my my take was you have to understand where the where the field was at that moment. It was a moment GPD2 was trained and kind and already presented to the world. But there was this whole narrative of GPT2 being too powerful to release and people being worried about it. Today's open weight models are are like thousands millions of times more powerful than GPD2.
But it was at least worthy worthy discussion to have. But what was GPD2? If you look at that model, what it could it could write a few coherent sentences and that was that was about it what what what the limits of it of it were. Then internally at OpenAI started seeing GPT3 and what it could do and well lo and behold it could do a few short prompting and another technique that at that moment felt felt kind of groundbreaking but at the same time from today's perspective it is it is not not that much and I was in some way there was there was one trend at opening at that moment of excitement of new technological development and models being able to do something that we couldn't do before but I had this inherent belief that those models mod are are very vague.
When they when they generate a few sentences, how do we know if those sentences make sense? If those if the models actually smart where are are the models stocastic pars just regurgitate the training data? They are trained on on everything in the world and and everyone tries to put as much as much training data as possible and then the model writes for you a few sentences and you don't know if you really have hit some interesting like thoughts if if the model came up with something if it just if it just came it through the through the data.
So the [snorts] the way of like write teaching models and evaluating models for writing programs. We we kind of started firstly with measuring even what do we have our first paper is I think I think thing called evaluating large language models for for writing code. It came from perspective of like we have to understand if they can reason. Yeah. And if they can come up with interesting program then I I think verifiability was the part of pinning down the model if they actually understand what they are doing because in some way you cannot write a correct program if you don't understand the concepts like it's it's always easy to talk around in various hyper balls around the topic but it's very hard to to write a program that is that that that is true.
So I think I think that was the main thing and looking at various scaling laws already. Then we started looking at things like test time compute and how does how does that scale how does size of a model scale with understanding we were able to generate a whole set of problems that were kind of out of distribution. Pretty sure no one no one thought of something like that before and measure how those how those things do.
And I think from that perspective, at least to me, a lot of my understanding of how large language models think and solve problems came came from that time and that that era. >> Yeah. So I we still remember here that so that's a company called a cursor was built on on right on that I think it's the end of 223. So we host a hackathon here with some of our uh people who host hack discovered that and uh we at that time we know like a coding is still very toy kind of applications.
Oh, people suspect is it working or not kind of situation. So you guys already doing that three years ahead made that work. Yeah, >> we were all the things that like the world doesn't see the things that open AI is working until until it it comes out. But yeah, some of those things we started. >> Yeah. So uh go back to the the question about the human evolutional correctness benchmark. How much of the last five years of progress is downstream of uh having the verifying like reward?
What do you think? >> I think in in my mind and I will I will be like a little bit spicy but I think the era of evil is done. >> It's dumb. Okay. >> I think the evils don't make sense anymore. But they did they did matter for quite a while. for for for a while they were the main driver of progress partially because they weren't yet gamed super hard and partially the moles were not very good at evil like problems >> like human evil what what literally me and Vixa we we sat in the room and ran wrote like 30 problems each first like I I just wanted to see how the models would do on on on on a problem like this and and it was was part of our experimentation part of our understanding and then then we saw better and better models, better and better reasoning techniques could keep improving those models.
And I think I think there was a golden era of human evolves for for a year or two or maybe even three. Hard for me hard for me to say exactly. But with the they're reasoning models today and especially like with RL already at its height and everyone building environments to to almost whatever whatever they want and the training frameworks that allow you to spend a lot of amount of compute and pretty good base models like almost any task that you want you can put into the model.
Almost every evil gets gained immediately like you you look at the model benchmarks and like sometimes when a benchmark is published it has 50%. that that's that's probably the lowest the benchmark has these days and shoots to 90. We have like two or three model generations. Like we we live in the world that whatever benchmark you have you will you you'll max it like there there's an algorithm there's a loop for it and you you you can just you can just do it.
So I don't think realistically the evils are the right way to look at models anymore. We need to we need to look at it models more from a from a continuous process perspective. Eval eval is like a single point in time. Does does it this model do but like >> the question is a lot about the about the improvement of the models now. How how how does the mall do from from like today to tomorrow to the next month because we know that the models are not doing everything for us yet but it is they are they are kind of growing and this this direction of growth is also like steered by people in the labs who are trying to think what people want the most and how do we direct our attention.
But like no no no no no individual eval like right now kind of kind of like can can tell you a good good story about mostly mostly how much like this particular EVA was in the mind of people training the model. >> Absolutely. I I mean totally agree but but right now like continue learning like you need the model should be improve itself is the next frontier right but then the high though the thing comes back to like how do you know the model is improving you still need some kind of uh verification mechanism eval mechanism right so how do you see that right yeah >> how I see it and it's a great question comes back to trading >> which is how do you how do you evaluate how do you evaluate trading algorithms and there there's always you do you do back test and then then you do um real world test.
Both of those things are are are necessary to really validate algorithm. How does back test work? You choose a date in the past. You train the model on old data from the past and then you test it on like a new fresh data that the model has never has never seen through training. Um that is in many ways the best way to um to measure if your model like has some ability to generalize to to to the new data and new time and this causality and this arrow of time is is very very important in that in that respect and I think I think in in the same way when working with LLMs we need to be able to like have some data from the past and then and then try to figure out how do we do like like how do we how do we adapt to data from the future that there may be some distribution shift there may be some differences when doing when doing trading you have to be very careful to not leak any information about the future and your training data at all every every trader has made this mistake that was thinking oh I have a good algorithm and then you realize you are leaking future information somehow and some of those can be can be very very weird and very very very difficult to to predict and then there comes the real world test where you actually train the model on all your data you deploy it in the real world and then how you see how the model does in the future?
How does a model does tomorrow? So like the the kind of the process evaluation to me is like you actually deploy the model in real world scenarios and then you see like you you could do like live AB tests this model or that model which which which one does better like know tomorrow this this model or that model which one which one climbs faster. The only the only the only test is the real world and the only the only actual measurement is is using those models for automating work >> having the real world feedback.
Yeah. Cool. Uh so next question is like co-pilot was arguably the first mass market LM product. Uh so what did this deployment teach you that benchmark couldn't do in in terms of yeah the real world feedback? I I think I think there there there's a main answer which is which is very real and very very kind of boring is that real world and evolves are not the same. can have can have a lot of the evals but whenever whenever you deploy product you start realizing there are ways to improve the eval that don't improve the product and there are ways to improve the product that don't don't measure in the that don't measure in the eval there are a lot of things that we did that were improving the model we were shipping them to Microsoft and Microsoft like oh our users don't don't care about that so whenever whenever you actually want to provide usefulness to the world you need to you need to like start thinking how do I how do I measure that that usefulness and how do I make sure we are we are building like the models that that that people need which is like you know latency, performance like speed, depth of dep of reasoning the the type of use cases.
It was like very quickly we started doing things like harness engineering because you kind of you kind of realized like training walls is slow and a lot of thing you can build around it and you can do harness to make better GitHub copilot and like those things compound in many ways. >> Absolutely. Yeah, we do agree. Actually, I was hosting Sergey here last month. He we have a very interesting conversation about how Google is training the coding agent.
He asked this question are those like GitHubs the code and GitHubs are the right training data. He was asking I asked you did you train all the Google codebase? He didn't have the answer. >> Worried exactly. So uh we go to the next question is about uh your uh 2021 verifier paper on mass world problems. Now looks like the seed of test time compute. So you've done so many groundbreaking work on those things. Uh walk us through the intellectual leap from um train a verifier to rank samples to run directly on the chain of sort.
So what actually happens inside 01 when when it thinks I think that's something they're very curious. We see the trades but I think you would love to see from the first principle how how you uh have the intellectual things like that make this work. >> Yeah. Yeah. Yeah. Like the the dream of scaling RL was with us from a from a for a very long time at open AI. I I joined open AI because I wanted to work on reinforcement learning and one of the very first research all hands that I came Ilia Sudskver stood on stage and said our research road map is to train large generative model and then do reinforcement learning with it which was like extremely precise.
this is exactly what we are doing today seven years later and and he he he he saw it and he knew that this is this this is the plan but we didn't we didn't we didn't at that moment we didn't know what generative model we are training what's what kind of how do we do reinforcement learning what was what is the algorithm like ppl was all the rage at that moment and we were we're were doing a whole lot of ppl mostly on data and robotics and like it took us a while to get mostly like a lot of that was about the algorithm and and and And some of it was about a bit about conviction.
>> Um like a lot of times we wanted to scale up RL but there's always this problem where do you get GPUs from how how much GPUs do you do you allocate and the the discussions with open AI leadership a lot of times where uh oh like give us some proof that make sure that that we actually it is worth because we could be training like GPD4 GPD5 on those GPUs. Why why would we why would we scale RL like I was RL pills but there there not that many people at open AI were were RL pill there were there naturally some of us but I think I think I think for many years it was was minority and and there there was always this chicken and egg like how do you how do you make a proof without having a GPUs but you need to like catch 22 kind of algorith kind of problem and I I think a lot on on actually the birth of scaling of reinforcement learning was was the day when when when Makoup came to me and say Jerry hey we want to allocate GPUs for scaling up reinforcement learning like take what we have so far and start and start scaling it is what this will I hope I hope this will I hope this will this will work out and I think as some this this became a self-fulfilling prophecy we we had the GPUs so now I started to figure out okay how do we organize we need to scale our so we need we need more data we need to get problems we need to an algorithm that scales what is what what has a chance of scaling we need to we need to set up like babysitting rotation and and then a people structure to make sure we can execute big RL runs which which no one not has done before and I kind of like I assumed success started putting everything in place and like things started like clicking and things started started happening it was a a lot of progress in research as you you assume okay what do what do I need to to to to get there and then and then it's and then it all and it all comes comes together but like everyone everyone already today like roughly knows the formula from a from a high level which is like you know the idea that you need a reward for the environment is not particularly new like rewards in RL have been for for a long time but a lot of the difficulties is like one one one thing this simple idea of you do multiple rods for a sing single prompt to build a baseline is a is a really really powerful one and that that helps a lot there's a lot about numerical stability there is there there is a lot about about like you know some things had to be removed from PO to to make the the algorithm work better and scale better there is always this notion that simplicity is the is what what what makes scaling work.
Your algorithm needs to be needs to be simple. So like you know the today's like set of algorithms for scaling are are actually that simple in multiple ways but this is what I was was able to to get it to scale. >> Great. So uh that when the model backtracks and say hey wait let me reconsider is that emerging from area or do you think it's shaped by the reward design >> completely completely shaped by >> no like the rewards have nothing nothing to do with it and the model just just figures out and it's very very quickly that that you know is what helps it think think longer in some way what is what is happening in reinforcement learning is that the model takes all the behaviors and all the types of activations that we learn through pre-training.
It sees all the texts and all the all the bits where where someone was thinking, someone said, "Hey, let me let me reconsider this." Tries to tries to recall those things and use it almost like like a function call chain of thought to to steer it reasoning and then those that are more successful get reinforced. >> Got it. I see. So you described the 01 as essentially a tech demo that was great at puzzles uh and 03 as a tool you shipped.
What a change in training to make tools tool course part of the reasoning loop. >> Yeah. Yeah. Yeah. It's like 99% about systems and 1% about algorithms. We had a few interesting algorithmic problems but a lot about doing tool use at scale. It's just very hard software engineering problem. Uh every every lab and every company doing it right now knows how painful it is. It's um at some moment you are operating distributed system that is the size of the largest distributed systems ever built.
I think like our open AI training framework had more virtual machines had more data moving than like Netflix, Amazon like imagine imagine whatever like Facebook very likely like imagine imagine any any big website that people were thinking oh this is this is scale this is what we are scaling and our team had to build it in a few months to train to train models and then and then we were we're were pushing a lot of data there were there was a lot of synchronization a lot of uh a lot of trying to make sure that we don't hit bottlenecks of scale and a lot of trying to trying to shard things and all trying to communicate things.
That's um that that that's the main problem that we did. We know that that when we do that it will work because the reasoning already worked. We we we we knew that things like programming do require that and that feedback from real world will be very useful. But the engineering part there was brutal. >> Yeah. I was always thinking is like the toy is more like architect you trying to conduct whatever tools you use for this specific problem.
Yeah. >> Yeah. >> So um the follow question is like what does actual test time computer stop paying uh when does like extra step uh test computer stop paying what what does that scaling curve for syncing time actually look like across like a different task types? >> Yeah. Yeah. Yeah. like we haven't seen any meaningful limits to test and compute. The limits the limits are twofold but I I think the biggest limit cuz just latency doing very long rollouts take a long time.
Not that many people want to wait that long time for very long rollouts that's a that's a pretty big deal. If you for example in your training start having rollouts that take a day of time you know it's it very much limits how many steps you can take what what you can do there are only so many days in the air so uh so that's a that's a pretty pretty big limiting factor of of of how long overall you can you can take [gasps] and like second thing is a bit more architectural is it is obvious but transformers like their generalization with longer and longer context is not perfect.
It is slightly difficult. So both of those things make this make this slightly like hairy and gnarly and gnarly problem. We just don't have very long rollouts because no one wants to wait for them and it's hard to iterate on very long rollouts because they take time. I think like we already see currently the trend of people doing like more parallel rollouts which are called sub agents that that kind of clears the latency but it's also much less useful.
It's not it's not serial computation like the the the the usage of those tokens is much is much worse. Those are more expensive usage uh happens. People complain that oh cloud or or or codex 8 usage because it's on those sub agents and they weren't that useful. It's um it's it's it's a hard problem. It could be architectural. It could be it could be RL problem in many ways. But um it's it's we we we don't see as much progress here as as we would like largely because because those those reasons the limitations but the models they they kind of can generalize very far.
Yeah. It's it's hard to optimize it. >> Got it. So uh for GPD5 they merged the reasoning and non-reasoning into one systems and from a research management perspective what was harder is the science or the shipping a unified model into a billion users while the paradigm still moving >> how would you think it's it was largely an organizational problem and it's some some of it is innovator's dilemma which which hit open AAI very hard around that moment which open AAI had one of the most successful consumer products ever at GPT.
We have a very specific form factor of >> uh we want models that respond very quickly with a pretty reasonable answer to a pretty small and easy problem. Yeah. And like there there's a huge market in that and there's a there's this is a really really good product and a lot of the early discussions was okay we have this model it is a reasoning model it thinks for longer technically it is it is smarter and we have this product called chhat GPT and it was was trying to fit like a triangle peg in a square hole for a while which like you know I think I think in many ways it was a good idea to bring them closer together in a in a seamless fashion but chant GPT like what wasn't wasn't the right thing and we already only see right now it played out.
It took world and took open AI to figure out like cloud code was maybe one of the first goods agentic product actually while openai was trying to shoehorn agentic models into chpp and that didn't that didn't make any any any sense productwise. So uh like right now Chpd is reinventing itself into an agentic product because like it is clear this is a much much bigger thing and that that the like the question and answer bot was was was a good thing but at this moment it becomes like almost a smaller smaller market of the of the of the AI AI landscape.
So, so kind of like it was the question was how do we use resing mods to bring the most the most utility and it didn't have the product form factor fully nailed down yet and it shows only how important is to innovate not not only on on the machine learning aspect on the models aspect but also how how do we like get that usefulness into into into the hands of the makers. >> Yeah, I see. So we switch a little bit out of the reasoning part of things.
So um that's a fun question you are going to so we talk about continue learning is necessary for AGI. So you AGI house. So uh guest always ask this question. Yeah. So today's model becomes hopeless when they got stuck well where the human would learn from the failures and they come back stronger. I think why can't a bigger context window better memory scaffolds like can fake that like do we need like a new architecture I guess that's what we talking about right >> that's what I believe yes like the the story is pretty answer the answer is pretty quick like if it would be enough we would already have it we have context that is like pretty long at least in many in many ways we have >> like we have all the markdown files and we have all the all the search methods and does it does it work I've talked to a lot of people who like have a lot of markdown files and are searching for them a lot and everyone says it's it's brittle doesn't generalize doesn't work but what what it kind of reminds me is people trying to build agents in 2023 you remember autog or or lang chain those were those were people saying let's take GP4 it's a pretty smart model and let's build agent on top of it by building scaffolding and harnesses and function calls and in the end it was a model issue the model was not trained end to end to be a good agent and right now we are we are thinking oh like you know the model like has so much context and we have we have can do so many markdown files why can't we use that for learning and the model is not optimized to end for for for learning I think those are those are machine learning problems not markdown file problems >> got it I guess is it right to say that you believe in this so much that you are okay I want to build a company around it that leads you to that the building of a call automation yeah >> yes yes Yes.
Yes. Whatever the expression is, I'm I'm putting things where they need to be because I decided to start a company around that that thesis and that idea. >> Got it. So, uh I think that's make automating research concrete. So, I guess that's one of the missions of a core automation. Today at a core automation, what does the agent actually do when a research engineer did at OpenAI two years ago like writing experiment code, running like elevations, uh reading results or proposing like a next hypothesis or >> Yeah.
Yeah. We we're trying to build an AI lab automation first which is >> the idea is every work that we are trying to do first the question is how does how do we make agent do it? How do we make AI? You you need to you need to have a every time you start a research program, you need to have a baseline and baseline is try to use existing agents for everything we do at the company and see how good how good they are and then starts the process of what we what we think of and what we what essentially want to do at test time training.
You you build an eval which is which is your lab and the mall does everything and how do we how do we improve that? How do we make it next day a little bit better? How do we we make it next day? next day next day better and and like that's that's that iteration and that's that kind of like optimization on the thing that you actually care about if if your lab operation is what you want to make faster. This is this is what you optimize.
No no improving evals that you that you don't care about that they're not your product but actually the thing that you that you want to do and as it gets better and better I hope we'll become better and better company. >> Yeah, absolutely. We we we're looking forward to that. What do you guys come out of next? So, uh there's some speculation that reports describe a flagship concept called a sir. I don't know whether I pronounce right.
Uh it's a model that keeps learning in production updating weights from real world experience targeting something like 100x less training data. Wow. Uh without asking you to confirm like any code names, anything. How do you make a model that change very every day safe and very evaluable and debugable? I think that's like the next challenges right to to do what we aspire to do. >> Yeah. Yeah. Yeah. Like debugability is hard.
I am I don't think I am very much like bullish about about things like mechanistic interpretability. I don't think you can interpret humans in their thinking. models have have trillion parameters like what we what we think we can we can even try to try to grasp them and try to try to understand them but but we can we can definitely we can definitely see and measure their behaviors that is that is that is kind of like I I I very much treat models behaviorally the models are they behavior and what they what they do in real world and I want to build a lot of like I want to treat models almost like rats in a maze where where we analyze everything they are doing and try to catalog it and try to try to decompose it and try to understand where is it getting sucked, where is it where where it needs to improve.
You need to build a system like almost you think of you have an Olympic swimmer the mall as your Olympic swimmer. You wanted to become the best of research and you have to analyze like every every minute of its life. How is it eating? How is it what is it technique? Where at which moment it is it is like utilizing energy efficiently at which moment it does not and try to give it feedback. try to make it improve to become the best the best version of itself.
>> Great. Uh so I I guess like we come back to the question about architecture. I guess uh you're also exploring architecture that scales better than transformers. So uh what's actually wrong with the transformer for the future that you are describing? Is it attention cost static weight or separation of training from inference or all of the above? >> Yeah. So the main thing which comes comes to the comes to what we what we talked before is we want to do test time training and I don't think transformers are capable of test time training in the way that that I want to I want to do it.
Transformer is a really great model at memorize memorization. It it's like builds all of its understanding of the world from pre-training data. It's all it all happened in the in the lab and then it builds memory in the form of it if it's it's KV cache but a transformer is not meaningfully learning at test time all the learning happened in the in the lab and like it's a it's a good like thought experiment to think how useful transformer would be if openanropic stopped training models today if they said okay this this is it we are we are taking down training clusters you have you have inference we have this model and and and and use it and everyone knows that those models like they are they are limited.
There are there are even those those like ideas called freshness fine tunes and knowledge cutoff right that you need to need to keep retraining the model to make it to make it up to date and have have new skills and new and new capabilities and this is this is what we want to make online what we want to make happen at test times but you don't need to do freshness fine tunes and I don't think like transformer just doesn't have the operation within its model that can can do that this is what we are what we are searching for >> got it do we have any general interactions or >> you have some ideas.
Okay, cool. >> Maybe we'll talk about >> some next time. Yeah. Okay, cool. Yeah, we actually have some our team at AJ House like a very big on like a neuro inspired different kind of uh uh architecture. Even I think I met Ben Spectre back the the founder of Flappy Airplane. He he was also tackling some of the data efficiency issues on the architecture. So uh while we're looking for the public announcement maybe in the f not distant future we continue that conversation.
Yeah exactly. So um I think before we we we ended we have a few other questions maybe around uh let's see the I I I think from the website or co- automation I think one big part of the focus of uh coordination is uh uh around the enterprise application that that part of things right that that how uh can you help elaborate that vision like from from the that perspective? Yeah. >> Yeah. Yeah. Yeah. The vision is very simple which is that everyone knows they want to have a model trained on your own data but like training the model is currently like bad and weird.
The finetuning process is not like very very easy and also the companies are worried about sending their own data out there to other companies. No one no one wants to do that. Sometimes they are held hostage because either rather you get a good model train on your data and send your data or not. What do you what do you do? And I think I think like the test time training is a very nice answer to that because you get you get your own deployment of the model.
No one else sees it. No one else has to touch it. You send your data to that model and to no one else. It can be on your on your own hardware and then it improves on your data. It learns it like it becomes genuinely your model because it learns about about you and about about your your your data. like continual learning is solving the problem of personalization in a in a way how it's how it should be solved and I think like like you know as every every company you need to try to imagine what what product do you have how do you how do you eventually sell your research in some way and this is this is how I think I think it really it really makes sense and and like current like there is a narrative oh we need open weight malls so that we so that we fine-tune them ourselves but I don't think I don't think anyone wants to be fine-tuning the malls in reality you want to have continuous learning malls.
>> I see. Cool. If call automation the wildest success, what do you envision? What's it look like? >> I I imagine the world where we don't have to work like like it does sound dystopian to some people, but I generally think it's it's where where our future should be. Like like work is nice, but at the same time, people still play chess even though computers are very good at playing chess. People still run even though even though the cars are cars are faster than us.
Like it is like like does feel a little bit more like a world without responsibility. But in some way it is the world of superpowers. If the machines really can do for us all the work that we currently need to toil and sweat and like carry heavy things and think very deeply on something. If we can have it on top whenever we want, however we want like it is is a much more like the world of agency in some way. how how I think about whether we can can dream of the things easily and they just they just happen in the real world because AI AI does it for us and like a lot of new opportunities come from that and how how the world could be run how our experiences could be how our relationships could be a lot of the things are very different sometimes I'm I'm jokingly say but it's not not very much joke that the future will look like high school because in many ways high school is a world where no one yet has to work no one has real responsibilities ities and people just just learn about things and spend time socially and try to figure out the world and AI can always be this teacher teaching us about stuff and spending time with with our friends and I don't think it's it's that horrible.
>> Wow. I guess that's very close to what do we call the AGI future. I guess how would you think about the timer? I guess you're being like a very uh >> AGI [clears throat] field at some point. GI it definitely is coming like I am I'm not joking when I say probably we have like few more years left of work >> at all. I don't think I don't think I don't think that the state of like us having to work and like like it's is a thing that will that will last forever a long time.
The question is like how like the question is how quickly like I have a tendency to to be pretty optimistic about those things. Uh but but like the reality is that the technology usually takes a while to deploy. The technology takes a while to to to percolate like the AI progress feels both fast and slow at the at the same time. In some way in seven years we made made a huge progress on top on on what our models can do and in many ways I already feel like living in the future I am mostly vibe coding and mostly self-driving.
>> Yeah. right now. But at the same time, still like our lives look like only a little bit different than that than they used to be in the in the past. But I think I think I think like you know I I definitely think in the in the next five years somewhere in there like probably the work will transform in a way that it's hard to like recognize it from from >> amazing future. Yeah. Look look forward to to that future.
That's what uh we are we are doing here. But so uh I wish I have all the time to continue the conversation. We have another programming for you. So uh we will uh end up the talk for now. I I look forward to welcome back at some other point where we see the model architecture progress. So uh >> I would love to tell you more one day about about what we are doing at core automation. >> Amazing. Yeah. So um so Jerry uh the CEO of uh core automation the man who taught machine to think and now teaching a lab to run itself.
Find the team at callation.com I guess. So, and let's uh work together and see help him to to the build the future that he's he's envisioning and definitely want to be part of that. All right. Thank you so much.