LangChain 创始人 Harrison Chase:构建 AI 智能体的编排层
LangChain's Harrison Chase on Building the Orchestration Layer for AI Agents | Training Data

it's so early on that like it's so early on there's so much to be built yeah like you know GPT 5 is going to come out and it'll probably make some of the things you did not relevant but you're going to learn so much along the way and this is I strongly strongly believe like a transformative technology and so the more that you learn about it the better hi and welcome to training data we have with us today Harrison Chase founder and CEO of Lang chain Harrison is a legend in the agent ecos system as the product Visionary who first connected llms with tools in action and Lang chain is the most popular agent building framework in the AI space today we're excited to ask Harrison about the current state of Agents the future potential and the path ahead Harrison thank you so much for joining us and welcome to the show of course thank you for having me uh so maybe just to set the stage agents are the topic that everybody wants to learn more about and you've been at the the epicenter of agent building pretty much since the llm wave first got going and so maybe first just to set the table what exactly are agents I think defining agents is actually a little bit tricky and people probably have different definitions of them um which I think is pretty fair because it's still pretty early on in the life cycle of everything llms and agent related the way that I think about agents is that it's when an llm is kind of like deciding the control flow of an application um so what I mean by that is if you have a more traditional kind of like rag chain or retrieval augmented Generation chain the steps are generally known ahead of time first you're going to maybe generate a search query then you're going to retrieve some documents then you're going to generate an answer and you're going to return that to a user and it's a very fixed sequence of events um and I think when when I think about things that start to get agentic it's when you put an llm at the center of it and let it decide what exactly it's going to do so maybe sometimes it will look up a search query other times it might not it might just respond directly to the user um maybe it will look up a search query get the results look up another search query look up two more search queries and then respond and so you kind of have the llm deciding the control flow um I I I think there are some other maybe more buzzwordy things that fit into this so like tool usage is often uh associated with agents and I think that makes sense because when you have an LM deciding what to do the main way that it decides what to do is through tool usage um so it's so so I think those kind of go hand in hand um there's some aspect of memory that's commonly associated with agents and I think that also makes sense because when you have an llm deciding what to do it needs to remember what it's done before um and so like tool usage and memory are kind of like loosely associated but to me when I think of an agent it's really having an llm decide the control flow of of your application and Harrison a lot of what I just heard from you is around decision making and I've always thought about agents as as sort of action taking uh do those two things go hand inand is agentic behavior more about one versus the other how do you think about that I think they go hand in hand I think like a lot of what we see agents doing is deciding what actions to take um for for all intents and purposes um and I think the big uh difficulty with action taking is deciding what the right actions to take are um so I do think that solving one kind of leads naturally to the other and after you decide the action as well there's generally the system around the llm that then goes and executes that action and and kind of like feeds it back into the agent um so so I think that in yes so I do think they go kind of hand in hand so Harrison it seems like the main distinction then uh between an agent and something like a chain is that the llm itself is deciding what step to take next what action to take next as opposed to these things being hardcoded is that like a fair way to distinguish agentes yeah I think that's right and there's different gradients as well so as like an extreme example you could have basically a router that decides between which path to go down and so there's maybe just like a classification step in in your chain and so the lm's still deciding like what to do but it's a very simplistic way of deciding what to do um and you know at the Other Extreme you've got these autonomous agent type things and then there's this whole Spectrum in between so I'd say that's largely correct although I'll just note that there's a bunch of nuance in gray area as there is with most things in the llm space these days got it it's like a spectrum from control to like fully autonomous decision making and logic um all those are kind of the spectrum of Agents interesting uh what role do you see Lang chain playing in the agent ecosystem I think um right now we're really focused on making it easy for people to create something in the middle of that Spectrum um and for a bunch of reasons we've seen that that's kind of the best spot to be building agents in at the moment um so we've seen uh some of these more fully autonomous things get a lot of interest and and prototypes out the door and and and there's a lot of benefits to the fully autonomous things they actually quite simple to build um but we see them going off the rails a lot and we see people wanting more constrained things um but a little bit more flexible and powerful than chains um and so a lot of what we're focused on recently is the being this orchestration layer that enables the creation of these agents particularly these things in the middle between chains and autonomous agents um and I can dive into a lot more about what uh exactly we're doing there but at a high level that's that being that piece of orchestration uh uh framework is is kind of where we imagine Lang chain sitting got it so there chains there's autonomous agents there's a spectrum in between and your sweet spot is somewhere in the middle enabling people to build agents yeah and obviously that's changed over time so it's fun to like reflect on the evolution of Lang chain um so you know I think when Lang chain first started it was actually a combination of chains and then we had this one class this agent executor class which was basically this agent thing and we started adding in like a few more controls to that class um and but eventually we realized that people wanted way more flexibility and control than we were giving them with that one class so like recently we've been really heavily invested in Lang graph which is an extension of Lang chain that's really aimed at like customizable agents that sit somewhere in the middle and so kind of like our Focus you know has has evolved over time as as the space has as well fascinating maybe maybe one more final kind of setting the stage question um one of our our core beliefs is that agents are the next big wave in AI um and that we're moving as an industry from from co-pilots to agents I'm curious if you agree with that take um and and why why not yeah I I generally G with that take I think um The Reason Why That's so exciting to me is that a co-pilot still relies on having this human in the loop and so there's a little bit of almost like an upper bound on the amount of work that you can have done by an external kind of like uh by another system um and and so it's a little bit limiting in in that sense I do think there's some really interesting thinking to be done around what is the right ux and human agent interaction patterns um but I do think they'll be more along the lines of an agent doing something and maybe checking in with you as opposed to a co-pilot that's constantly kind of like in the loop I just think it's I just think it's more powerful and and gives you more leverage if the more that they're doing which which is very paradoxical as well because it comes the the more you let it do things by itself there's more risk that it's messing up or going off the rails and so I think striking this right balance is is going to be really really interesting I remember back in I think it was marchish of 2023 uh there were a few of these autonomous agents that really captured everyone's imaginations like baby AI autog GPT a few of these um and I I just remember Twitter was very very excited about it and it seems like that first itation of an Asian architecture hasn't quite met people's expectations um I think why why do you think that is and and where do you think we are in the agent hype cycle now yeah I think um maybe thinking about the agent hype cycle first I think Auto GPT was definitely the start and and and then so I mean it's it's one of the most popular GitHub projects ever so one of one of the peaks of of the hype cycle um I think and and I'd say that started in the spring 2023 to Summer of 2023 is then I personally feel like there is a bit of kind of like uh law slashdown trend from the the late summer to basically the start of the new year um in in 2024 and I think starting in 2024 we've started to see a few more realistic things come online um i' I'd point out some of the work that we've done at linkchain with elastic for example they have kind of like an elastic assistant an elastic agent in production um and so we're seeing that we saw kind of like the Clara customer support bot um kind of like come online and get a lot of hype we've seen Devon we've seen Sierra these other these other companies start to emerge um in in the agent space and so I think uh with that hype cycle in mind talking about why the auto GPT style architecture didn't really work it it was very general and and and very unconstrained um and I think that made it really exciting and captivated people's kind of like imaginations but I think practically for things that people wanted to automate to provide immediate business value there's actually a lot it it's a much more specific thing that they want these agents to do and there's really like a lot more rules that they want the agents to follow or specific ways they want them to do things and so I think in practice what we're seeing with these agents is they're much more kind of like custom cognitive architectures is kind of like what we we call them where there's a certain way of doing things that you generally want an agent to do and there's some flexibility in there for sure otherwise you know you would you would just code it um but it's a very like directed way of thinking about things and that's most of the agents and assistants that we see today and that's just more engineering work and that's just more kind of like trying things out and and and seeing kind of like what works and what doesn't work and it's harder to do so it just takes longer to build and I think that's kind of why you know that that's why that didn't exist a year ago or something like that since you mentioned cognitive architectures I love the way that you think about them maybe can you just explain like what is what is a cognitive architecture and like is there a good mental framework for how we should be thinking about them yeah so the the way that I think about a cognitive architecture is basically what's the system architecture of your llm application um and so what I mean by that is if you're building an LM application there's some steps in there that use llms what are you using these llms to do are you using them to just generate the final answer are you using them to route between two different things are you uh have do you have like a pretty complex one with a lot of different branches and maybe some Cycles repeating um or do you have uh kind of like you know a pretty a loop you basically run this LM in a loop these are all kind of differents of cognitive architectures um and cognitive AR is just a fancy way of saying like from the user input to the user output what's the flow of data of information of llm calls that happens along the way um and what we've seen more and more especially as people are trying to get agents actually into production is that the flow is specific to their application and their domain um so there's maybe some specific checks they want to do right off the bat there's maybe three specific steps that it could take after that and then each one maybe has an option to loop back or has two separate substeps um and so we see these more like if you think about it as a graph that you're drawing out we see more and more basically custom and B spoke graphs as people kind of try to constrain and guide the agent along uh their application um the reason I call it a cognitive architectures is just you know I think a lot of the power of LMS is around reasoning and thinking about what to do um and so you know I would maybe have like a cognitive mental model for how to do a task and I'm basically just encoding that that mental model into some kind of like Software System some some architecture that way and do do you think that's the direction the world is going because I kind of heard two things for me there one was it's very bespoke and second was it's fair barly Brute Force like it's fa hardcoded in a lot of ways do do you think that's where we're headed or do you think that's a stop Gap and at some point more elegant architectures or or a series of default sort of reference architectures will emerge that is a really really good question and one I spend a lot of time thinking about I think so like at at an extreme you could make an argument that if the models get really really good and reliable at planning then the best thing you could possibly have is just this four Loop that runs in a loop calls the llm decides what to do takes the action and Loops again and like all of these constraints on how I want the model to behave I just put that in my prompt and the model follows that kind of like explicitly um I I do think the models will get better at planning um and and reasoning for sure I don't quite think they'll get to the level where that will be the best way to do things for a varet of reasons one I think uh efficiency if you know that you always want to do step a after step B you can just put that in order um and two reliability as well like these are still not a terministic things we're talking about especially in Enterprise settings you probably want a little bit more comfort that if it's always supposed to do step a after step B it's actually always going to do step a over step B um or after step B I think it will get easier to create these things things like I think they'll they'll maybe start to become a little bit less and and less complex um but actually this is maybe a hot take or interesting take that it had you could say like so the the architecture of just running it in a loop um you could think of as like a really simple but General um cognitive architecture and then what we see in production is like custom and complicated kind of like cognitive architectures I think there's a separate access which is like complicated but generic custom or complicated but generic cognitive architectures and so this would be something like a really complicated like planning step and reflection Loop or or like tree of thoughts or something like that and I actually think that quadrant will probably go away over time because I think a lot of that generic planning and generic reflection will get trained into the models themselves but there will still be a bunch of not generic training or not generic planning not generic reflection not generic control loops that are never going to be in the models basically yeah no matter what um and so I think like those two ends of the spectrum I'm pretty I'm pretty bullish on I guess you can almost think about it as like the llm does the kind of like General the very general um agentic reasoning um but then you need domain specific reasoning uh and and that's the sort of stuff that that you can't really build into one General model 100% like I think I think a way of thinking about like the custom C of architectures is you're basically taking you're taking the planning responsibility away from the llm and putting it onto the human and some of that planning you'll you'll move more and more towards the model and and and more and more towards the prompt but I think they'll always be like I think a lot of a lot of tasks are actually quite complicated in some of their planning um and so I think it will be a while before we get things that are just able to do that super super reliably off the shelf it seems like we've simultaneously made a ton of progress on agents in the last six months or so like um I was reading a paper the the Princeton swe paper um where their coding agents can now solve 12.
5% of GitHub issues versus I think 3.8% um when it was just rag um so it feels like we've we've you know we've made a ton of progress in the last six months but 12 and a half% is like not good enough to you know replace even an intern right and so um it feels like we still have a ton of room to go um I'm curious where you think we are both for General agents and also for your customers that are building agents like are they kind of getting to I I assume not 5 9's reliability but are they getting to kind of like the thresholds they need to kind of deploy these agents out to actual customer facing deployments yeah so so theu agent is I would say a relatively Gish agent in that it is expected to work across a bunch of different GitHub repos I I think if you look at something at like v0 by versel um that's probably much more reliable than 12.