transcribe

Webinar: AI Agent Simulation of Human Behavior with Michael Bernstein

Stanford Online · 1h 0m · transcribed 59m ago
More from Stanford Online Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Transcript

0:10 Hi everyone. Good morning, good afternoon, good evening. That way we cover everything. We're happy to have Professor Michael Bernstein here for this webinar. So let me give you a bit of more of his background. Michael Burstein is a professor of computer science at Stanford University where he is a bus university fellow and an STM micro electronics faculty scholar. His research in human computer interaction concentrates on the design of social computing systems. This research has won best paper awards at top conferences in human computer interaction including CHI, CSCW, ICWSM and UISD and has been reported in venues such as the New York Times Science Watt and the Guardian.

1:01 Michael has been recognized with an Alfred Pisarn Fellowship, UISD lasting impact award and the Patrick J. McGgovern Tech for humanity prize. He holds a bachelor's degree in symbolic systems from Stanford University and a master's degree and a PhD in computer science from MIT. Please stay muted so that we can ensure good audio for Michael. we would appreciate if you could turn your cameras on so we can have an engaging and interactive session and we encourage you to write your questions in the chat since at the end of the session we will have some Q&A segment.

1:41 Now if we're all ready Michael we are thrilled to have you here so the floor is yours. >> All right thank you so much for coming. I'm seeing wow a a fairly global presence here in this webinar. I'm actually somewhat astounded by the number of you who are here. over 500 of you. So, thank you for joining me. We're going to talk about AI and we're going to talk in particular about a frontier of AI agents that I think is particularly poised to have a lot of impact. and maybe you're an indiv individual contributor, maybe you're a manager or a leader or you're a managing or you're a policy maker. There are so many different applications of these kinds of technologies if we can get them to work which they're increasingly showing a promise of. So we're going to talk through what what is this frontier and what's going on. When we make decisions, we make those decisions based on what I would view as often incomplete information about how people are going to react. Maybe you're at an organization, whether it's a for-profit or a nonprofit, that's trying to decide if we do this, maybe our, you know, our our clients or customers will react this way, or maybe these people who aren't currently our customers will will or won't react that way. Or maybe I'm a leader and I want to say if I if I change my management style in this way or reorgan the following way, maybe the organization will will do better or worse. We make these decisions based on the best data we can. But that data is often incomplete.

3:15 And this kind of incomplete information means that we often make bad bets. What I mean by that is that we we make the best guess we can. We take all the information we can, but we're ultimately wrong. And this is not necessarily a problem with you. It's not that you're not smart. It's just really hard. We can't know in advance always what's going to happen because we have limited ability to get feedback. We can only launch once. And so we make these big bets and they misfire across these areas across like consumer packaged goods, CPG, personal finance, product design, policy, management. Even me as a professor, I do this all the time. I'm making guesses about what's going to work with my students and not. And you think, oh, this is like a problem of like, you know, of modernity and so on.

3:58 But, you know, like this is not even a new qu question. This is a thing that goes all the way back to 1906. Here's Robert Merton, a famous famous sociologist who pointed out even back then over a hundred years ago how hard it is to design for what a group of people will do. Like we all sort of want to get out of the city so that we can avoid the crowds, but then everyone wants to get out of the city and go on vacation at the same time. So we all wind up in the same you know the same forest and it's too crowded there too.

4:28 So, we're very challenged in making these kinds of decisions. And I want to ask today, what if you had a what if machine? What if you had a whatif machine? What would you use that for? If you could sort of imagine in advance what might happen, you know, imagine that you could say if we to if my organization took this path, how might our customers react? Or if you could say, you know, if we were to fail or if something were to go wrong, what's the pathway through which that would maybe happen?

5:08 Or you could say, well, what if we could see how people might react to this policy change or this new product or this strategy we're thinking about before we deploy it? If you had this kind of what if machine, my claim is that we could make better decisions more often, much more often. And so this has led to a kind of a really interesting question for me as a researcher in computer science. Could we create simulated AI agents that replicate human behavior? Could we create AI simulations that behave like people do, maybe like you know your organization would or like your customers would and so on. If we could, then maybe some of these things become possible.

5:56 Now simulation isn't a new idea either, right? It goes all the way back to 1978. Here's Thomas Shelling won a Nobel Prize in economics for generating this idea of what are called agent-based models. And these things are still in use today. And I'm showing a screenshot here of a recent model on that was run on like supercomputers trying to understand oh if we engage in the following kind of intervention does this reduce the kind of spread of a pandemic and so on there are these very simple algorithmic models of people well actually increasingly complex but still fairly mathematical but you know simulation is not just a thing that we use in policym here's entertainment I bet a bunch of you have played the sims it is one of the singularly most popular video games of all time these are also simul population of people and you know these people in quote unquote these agents in in in the Sims you can sort of change the change the space around them you can intervene on them you know you can have all sorts of fun in this little dollhouse environment and increasingly if we are deploying AIs into worlds where the AIs need to interact with us sometime some people talk about AI teammates what does that even mean like an AI that's on your team at work would need to start reasoning about how you might respond if it takes a certain action and more broadly to create a set of tools that are what I would term look before you launch and I think we struggle with this you know for much of my career I worked in the design of online platforms social media and these kinds of things and many issues would come up because we couldn't effectively foresee what was going to go wrong and so this creates the opportunity to create tools that say how what kinds policy changes or rules do I need to set up before I launch the system so that it doesn't all become a dumpster fire and then I set the rules afterwards.

7:47 Now the problem has then has been that up up until today our models have been pretty rigid. That is to say that like you kind of have two options. One you say you know a human is five parameters. That seems like it's not quite true like we can't quite be reduced to five parameters. It can explain a lot about us, but you know, I'm at least seven par. You know, it's like essentially like this is just much too of a sparse model to encode the richness of human behavior. The other option is to do what the Sims does where you basically write scripts and say, you know, if I get punched, then I fall down and I get mad.

8:24 And there again, it's sort of limited to only the kinds of things we can think to create. and you know, it's always going to be incomplete. So as a result, at least within the academic literature, the basic conclusion is I'm going to quote here, the models have been highly stylized and have had minimal impact. So not a good track record. They're really interesting, but in practice, we haven't seen them taken up in many cases. And I I mean, I'm sure some of you in this call have used agent-based models, but the vast majority of people who I talked to have not.

8:57 But starting a few years ago, my colleagues and I recognized that modern AI models, so-called large language models like Chad GPT, Claude, Quen, Llama, all of these, DeepSeek, have trained on a variety of human behaviors, which is to say that they've literally read the book on all the research about how people behave because they trained on on research. and they've also trained on a bunch of social media data. They've seen the good, the bad, and the ugly of human behavior. And so, as a result, what we recognized was that you can prompt these large language models to take on the perspectives of a bunch of different people with different backgrounds, experiences, and traits.

9:47 So, we could take chatbt or another large language model, LLM. We can give it a prompt that roughly equates to, you know, here's be this person, here's a name, a description, and in this situation, how would they react? And by multipplexing those descriptions, we can create an entire crowd of different people. And if we can create a crowd of people, you can sort of put them in many different situations and see what might happen. And in fact, we started doing this and in some work that shall we say went a bit viral we we created almost like a little terrarium of many of these so-called of these AI agents that are simulated people. We call them generative agents and we created a little town called Smallville. This town is full of 25 generative agents. These 25 agents as I as I would describe them are essentially all completely autonomous AIs that are each sort of playing the part of a different person.

10:48 So you can see the entire town is going about their day. You know the artist in the town is gets up and paints. The college students in the town they sleep in more and then they go to classes and do their homework and so on. This allows us to create a space where we can sort of visualize what it would mean to have these simulations of human behavior in a way that we can learn about people and interventions.

11:11 So in fact this work got a lot of attention and more recently Andre Horowitz I think this was just last last month put out some research you know their own research arguing that this is going to become the next sort of essentially generation of market research tools. they and if you go check out this oh it was in June excuse me if you check out this article then you know they cite our research quite centrally. So what I'm about to show you in this webinar is that AI simulation can create both believable and accurate simulations of human behavior.

11:48 I'm going to give you a how-to guide of the techniques that enable this kind of simulation. I'm going to give you the gotchas, the things that often people get wrong that you should look out for when doing this. And I'm going to talk about the current frontiers of this technology, like what what's kind of on the edge of just becoming possible. So, I hope that sounds good. If you're excited, I don't know, give a thumbs up in Slack or excuse me, in and Zoom and we'll we'll dive in. The basic way we're going to go after this is we're going to create an agent by first we paid an online artist to get a bunch of pixel art of different characters which you can see on the left. and these these visualizations are there mostly to help you sort of see what's going on. You don't actually have to build a little video game environment like I did. I'm just a child, you know, I was born in the 80s, grew up in the '90s with, you know, you 1990s Super Nintendo JRPGs. So, it's just a cool way to show it. But essentially what we do is we create a little persona for each individual. So here's John Lynn. He's a pharmacy shopkeeper. We tell him that he runs the pharmacy in this town. And we describe who he is. You know, he's very helpful. And we tell that agent who he what it knows about others in the simulation. So we tell him he's married to this other agent in the simulation, May Lynn. And they have a son, Eddie Lynn, who's a college student studying music theory. Now we have to tell the agent about all the other agents it knows in this environment because if we don't the agent's going to wake up in in a bed next to me at the beginning of THE SIMULATION AND GO right it won't know it literally will not know that other person. We have to tell that agent everything that it needs to know at the beginning. Now what I want to try to explain to you is over the course of this webinar the important pieces of what it takes to build one of these. I don't expect that you understand this figure yet as I as I'm for showing it, but I'm going to build up to this figure over the next few minutes so that you understand what it is that allows these agents to do what they do and so that you can create them yourselves.

13:56 Now, basically what what we can do is we just describe the agents and we're going to just let them go from there. They wake up without us telling them exactly what to do. We just describe who they are. They'll wake up, brush teeth, take a shower, cook breakfast, catch up, and talk to each other. They'll they'll pack and get ready. and what I'm going to do here is to the global alumni team, I can message you, but I can't message the the people here. I just sent you to the global alumni team a link to a demo. if you can echo that demo link out to the group, they can actually load it up and see the town in sort of effectively real time. So, I encourage you to go check it out.

14:43 Thank you. Yes, you sent out the correct thing. So, please go check out that that link. Now, the agents then act in the environment because these are just AI agents, right? Which means they mostly just act in in the sense of they they they use natural language. So they say, this agent, Isabella Rodriguez, is drinking coffee, right? We then need to ground those that text into concrete movements in our little video game environment. So we say, " they're drinking coffee. Okay, they're going to walk over to this chair, and then we're going to render a little emoji that summarizes what it is that they're doing." So if you go and clicked on that demo, you can see the agents are often having these little emoji bubbles that show what they're up to. So we ground their their actions which start out as language into this game environment.

15:31 But you can talk to these agents literally. You can sort of intervene and talk to them. So if you pretend to be a newspaper reporter, then you can say, "Hey, who's running for office?" You know, who's running for mayor? And if the agent knows, they'll respond. Oh, yeah. I heard Sam's running for mayor. And if you want to control these agents, then you can just pretend to be their inner voice and say, "John, you're running for mayor." And then John says, "Well, you know, I really need to talk to my family about this big decision that I'm considering." so you can also intervene in the game world. I can take the the the toaster, set it on fire, and the agent will recognize what's going on and run over to put out the the fire and then go make something else for breakfast. Right? So, we have all of these opportunities to sort of think about interventions in the world and how agents might respond to them. Now, imagine intervening and changing the layout of the home or the kinds of information they have access to or the kinds of products in their home and so on.

16:36 So, just to give you a sense of what what this is actually like, let me take it from here. These agents talk to each other, right? They're language agents. They're AI agents that can talk. So, in the morning, John sees his son Eddie and and engages in conversation. Says, "Hey, good morning, son. Did you sleep well?" The son says, "Yeah, I slept great." he says, "Hey, what are you doing today?" And the son says, "Yeah, I'm working on this music theory composition for my class." Remember, he's a college student. He's a music theory student. So, he's working on a composition, a new a new song. Now, John doesn't know this, but now he does. He has to remember it because soon in the simulation, Eddie's going to leave the house and John's wife is going to wake up and say, "Where's your son?"

17:20 And John has to remember what he learned so that he can tell her, "Yeah, he left. He's working on this music theory composition." And now she knows. And this is how information can sort of spread through the environment. Now, that's a simple example, but we wanted to examine a much more complex one. So, here's what we did. We initialized Isabella, this agent who runs the cafe in town, with a simple intent to plan a Valentine's Day party.

17:52 So the simulation starts on February 13th, the day before Valentine's Day in the morning, and it runs through the end of the 14th. So we just inserted into her memory this little intent that she was interested in planning a Valentine's Day party. And what I want to convey to you is that nothing about this AI agent architecture that we're talking about today is like a party planning module. Nothing that says, "Oh, this will definitely work." And in fact, there's a lot of reasons why this might fail.

18:18 This agent has to remember to tell the other agents about the party. Anyone who's told about it needs to remember that invitation and those who remember it have to actually show up. And yet from this seed we start to see evidence of classic information diffusion patterns which you might think of as like sort of rumor networks, whisper networks, just word of mouth. So Isabella tells people, people tell other people. You can see this graph of who's hearing about this party from whom as the simulation plays out. And in fact, Isabella goes and gets her friend Maria to help decorate the cafe in preparation for the party. Again, we didn't tell her to do any of this.

18:56 you know, this agent decided to go get help from the other agent. At the end of that day, like the next day on Valentine's Day, half the town had heard about the Valentine's Day party. So, of the 25 agents in the town, 12 of them heard about the Valentine's Day party. of those 12, five came to the party. three said they were busy and four said they were interested but ultimately didn't show up. So, is this accurate?

19:25 It's a little bit of an interesting question. It reads to me as broadly plausible. but it's not necessarily easy to tell. Is this exactly how it would play if you know if you if this were real life? It does to me seem like a a pretty compelling evidence of the kinds of believable behaviors you might get. And in fact, one more kicker on this. we not only told Isabella she was interested in planning a party, we planted a memory in the Maria agent, that ML that you see on the left in the image, that she had a crush on Klouse, another agent. And it turns out that Maria asked her crush to the party.

19:59 So, we have this little nent agent love here. so very beautiful. And in fact researchers who were not us got this whole simulation running again just this just this year and found you know that they could intervene and study different alternatives. So now they said oh okay what would happen if suddenly the agents some of the agents heard on the radio that there was a new communicable disease like a new a new epidemic of swine flu and you can see in that pathway when they ran the simulation oh no one showed up to the party except for poor Klouse. Claus had had not heard about the the new the new illness and so he had he he showed up anyway in the no threat condition the party happens as usual and when they and when the agents hear a radio about a non-infectious disease like diabetes complications or something like this the party also happens as as normal. So you can think about how interventions into the world let you ask these what if questions that can be quite powerful.

20:58 So how do we do this? I'm coming back to this image and I want to try to explain the major parts of this architecture. You don't have to be technical to understand this. I want to un I want to convey the main pieces. The first thing we need to do to create a an an agent like this is we need to give them the ability to remember. So we give them what's called a memory stream. This memory stream is essentially a blow blowby-blow me record of everything that the agent observes. So you can see a lot of this is actually not very interesting.

21:36 They see a bed, they see a desk, they see a closet. Some of it is more interesting like you know I'm stretching. I'm writing in my journal. I'm taking a break. I'm cleaning the kitchen and so on. but there's a lot of stuff here. you could fit it into the context of a large of some of the larger window large language models, but these models the research suggests can get distracted with very very long contexts. So we actually do something that's referred to in the in the literature these days as rag or retrieval augmented generation. So with rag what you what we're doing is we're preferentially retrieving memories that were accessed recently that the a that is are important to the agent and that are relevant to the current circumstance. So recency means if it's a you know something that you observed recently or that you that you recalled recently that's that's more likely to come up if it's important like you know the fact I don't know the agent brushes its teeth not very important to the agent the fact that it got asked out to a party pretty important. And finally, relevance. So if you know an agent is taking a math test, then other memories about taking tests or about that math class should be should be top of mind. So it's almost like running a little bit of a Google search over all of these these memories. And so now what happens is what you can see on the right if we ask the agent the question, what are you looking forward to right now? We're going to retrieve these recent, important, and relevant memories like things like that they're planning a party, that they're ordering decorations for the party and researching ideas for the party. Put those into the context window of the large language model along with the description of Isabella herself. And that's what gets Isabella to reply in a reasonable way saying, "I'm looking forward to this Valentine's Day party."

23:19 Memory, the first thing you need. The second thing that we feel is important is this ability to reflect. If you have a psychology background, then you might be familiar with what I'm about to say, where the memory I just showed is is what's called a purely episodic memory. It's just like a log of what happens. But if you think about it, you we're more than just all, you know, captain's log. This is what happened. This is what happened. We needed to provide the agents with the ability to produce higher level reflections of of who they are, what they like, their dispositions, and their interests and their goals. So what we do is we essentially get these agents to have shower thoughts figuratively speaking we so what we do is at regular intervals we pull memories from the agents memory stream and we ask the agent to reflect on those memories here's two examples to produce a you so a higher level reflection or conclusion about them so given the two observations that Klaus Mueller is reading about gentrification reading about urban design we produce a reflection that Claus spends a lot of time reading. And we can do this over and over again by continuing to group memories and reinserting those reflections back into the memory stream to produce higher and higher level reflections about who this agent is and what their goals are. This produces the ability for the agent to essentially produce a to produce behavior that's much more consistent with these goals rather than just literally being a robot and walking from step to step.

24:52 The third thing is planning. So planning is important for these agents to remain believable over long periods of time. And planning is something that is not unique to this project. It's been a long time issue with with AI. So what we would do is we would start by having these agents plan out their full day. Then we would drill in and given that day say hour by hour what are you going to do? Then we're going to drill in again and go minuteby minute and say okay in details what are you doing? And when the agent sees something as they're acting in the environment, like here the John agent might see Eddie his son taking a walk around his workplace, then we will ask this agent given its background. You can see the prompt here on the left if it should react to the observation. If so, what would be an appropriate reaction? And then replan if we need to. So this is how the agents adapt as as the simulation runs.

25:58 Now this work got a lot of attention. it really caught fire with a number of people and it was really interesting to see and I I think I just want to start by suggesting like there's some evidence that there's there's something to this idea that of the goal of creating these simulations of people and and what what that might enable. You can see here news reports that are featuring Jun Park, my PhD student who led this work. and and yet with this vision, I haven't fully answered the question, right? I asked this question before about whether we can create simulated AI agents that replicate human behavior.

26:47 Everything I've showed so far is, you know, shows believability, but Disney characters are believable. Cartoons are believable, but they may not be accurate. And if you're going to make decisions based on this, then there's this question of how accurate are these agents at replicating behavior. And that's an empirical question that we really need to ground out if you're going to think about whether you want to be using this. So, how do we measure this? Because there's papers out there that just try to replicate this study or that study and say this one works or that one doesn't, but you know others then try to replicate that and show a different result. How do we act like you can't tr like how do you actually know? So let's talk through how you might know. Well, there's a few different ways you could go about this. One thing you could do is build what you might call demographic agents. So here you take sort of a sample of the population and you produce a demographic agent version of a person.

27:44 So here's a demographic agent version of me. it says, you know, my age, you know, where I live, my job, and these kinds of things. On another hand, you could create what we did in that Smallville simulation, these persona agents that are sort of more narrative descriptions of the person. So again, on the right here is an is a persona agent version of me. so it, you know, maybe adds a little bit more flavor, but also misses some information, in but they're both interesting. The issue here, as prior research has demonstrated, is that both of these approaches can produce very simplified and stereotyped behaviors. most kind of one of the instances that that was most striking to me was that if you know June Park again my PhD student or now alum who who started this work if you just produce you know he's from South Korea and if you ask if you tell the model that and say what is he going to have for lunch the model says rice not great really probably not accurate either it's very stereotyped you know has bias all these kinds of things in it so what do what are we supposed to do.

28:49 So what we found is that there are better ways to do this. In particular, we found that rich qualitative information is an interesting direction. So what we did and this is some more recent research is we went out and ran 2-hour interviews with a representative sample of a thousand Americans. Now obviously this is focused in the US, but you can do this yourself in you know your local jurisdiction if you're interested. So, we would run these long interviews, and I'll talk about whether two hours is necessary later. It turns out actually you don't need two hours.

29:21 we we took a an interview script from a project here at Stanford called the American Voices Project led by David Grusky. He's a sociologist, has this very broad ranging interview script say that literally starts with the question, tell me the story of your life. And then it covers everything from their from their communities and their jobs and their finances and their and their health to their politics. It's just this broad backbone of who they are. We use this two-hour interview to construct a digital twin of that same person, a generative agent of that same person with that interview as the memory of the agent. Okay, so now we have a real person, a thousand real people and a thousand twinned agents for them.

30:05 That real person then takes a large battery of surveys and experiments so we can gauge what their actual attitudes and behavior are in in in realistic scenarios. So here I'm going to focus on the general social survey. This is a 170ish question survey that has been fielded in the United States over you know the last 50 years. But we also look at the big five personality index some big some behavioral economic gains some replications of experiments. is this broad battery.

30:34 We then have the generative agent of that same person take the same surveys and experiments and then now we can finally measure this more more accurately. We can say how closely does this person's agent replicate that person's actual behavior and attitudes. I hope that makes sense. So we have real person agent version of them and we ask how well the agent version of them replicates their exact answers. So we run these interviews actually with another AI agent. It's a voicetovoice interview like a phone interview.

31:10 it's on the computer but it's I I've I've changed many of the details here to preserve anonymity but I just want to give a a sense of the flavor of the kind of information that these interviews yield about people. Obviously this data all remains private. for for human subjects reasons. But you you can see that the the kinds of things we're learning are people's backgrounds and their their life stories, their ambivalence toward political identification, what their jobs are and were their finances and how much you know extra money they have to spend on various categories and so on. So it's just a very very rich background.

31:47 All we do then and if you wanted to create these kinds of agents yourself given this kind of interview is we put the interview up at the top of of the top of the prompt to chat GPT or another model and we tell the model based on this interview transcript I'd like you to predict how this person would respond to the following survey or experiment. Now there's more detail to it here. You can see the actual prompt we use. but to for the most part it's just the yellow parts here. interview and then now please predict what this person will say in response.

32:22 What we find is that these agents do accurately replicate attitudes and behavior. Now in order to explain this, we need to say one more thing about the method. It turns out that if I give you a survey today and then I give you the same survey two weeks later, you are not going to answer exactly the same thing. Some a lot of stuff's going to be the same but not exactly the same. So we need to normalize for this. So, we do exactly this. We have everyone, all 10,00 people in this experiment run the same survey twice. And so, what we're going to do here is every result I'm going to show you here is going to be sort of a ratio where 1.0 means that agents replicate people's responses as accurately as people replicate themselves two weeks later. I hope that makes sense. So, like a a 0.1 would mean we can replicate people's responses 10% as well as accurately as people replicate themselves two weeks later.

33:15 Okay, so if you were to do this, if you were to guess randomly on this big survey, you can replicate people about a third as well as they replicate themselves two weeks later. That's a baseline of sort of how how predictable people are. If you use these persona agents or the demographic agents, you can replicate people about 70% as well as they replicate themselves two weeks later. And if, excuse me, if you use these full interviews like I was describing, you can replicate people 85% as well as they replicate themselves two weeks later on the general social survey.

33:49 I think that's pretty impressive. Now, it's not just the the general social survey. Across a series of other tasks like the Big Five personality index, behavioral economic games, you can see that these normalized coefficients are all generally in this similar range. So we think that there's substantial possibility here. Furthermore, actually these interviews reduce bias and if you so if you have this full interview, if all I know about you is that you're like a Republican, a conservative, then it really stereotypes. But if I know a lot more then it can the model can get much more subtle. So it turns out across these different approaches that you might imagine, politics is the hardest to model. So the difference between the best performing group and the worst performing group was around 8% accuracy, normalized accuracy. The worst performing group for us was faright conservatives. the models, the underlying language models are pretty unhappy responding to questions as a far-right conservative. They've been value aligned maybe to to to not say some of these things. And the conservatives, the farright conservatives themselves in the interviews were a little more koi about their their true opinions. Gender and race surprised me. they were actually they had smaller gaps at a baseline than I expected. but again the interviews cut out a lot of the difference. So again we're looking at maybe just you know a sub you know in the case of gender sub 1% difference in accuracy between the different groups.

35:18 In fact we as I mentioned we also had a thousand people do this. We had a thousand people replicate experiments that were that were published in top tier journals, the proceedings of the National Academy of Sciences. This is top top tier science. So we had a thousand people our thousand people replicate those same studies and we had their agents replicate those same five studies. Of those five studies, our agents replicated four of the five. And you might say, "Oh, well that's not great. You missed one of them." But it turns out that we didn't. The thousand real people also didn't replicate that fifth study. That fifth study was actually bad science that didn't replicate. That result with real people was not that was published wasn't a real effect. And the simulations correctly predicted that that was going to happen.

36:06 That that fifth study wasn't going to happen. And in fact, my colleague here at Stanford, Rob Willer, has generalized this in in a really compelling way where they found a bunch of nonpublic study results that were pre-registered. So they got all of these studies, 100,000 people, and found that they could predict the effect sizes of the these experiments with simulation with a correlation between 085 and 0.9 pretty strongly. Now, what this wound up with for us was this agent bank of 1,000 consented participants from the United States. So, we have this agent bank that we can now use to sort of test maybe depolarization or climate change kinds of questions and so on. But if you're trying to do this for your organization, you might want to think about who should who ought to be in that agent bank that you would build for questions that you care about.

37:00 Should they be representative of your customers? Should they be representative people who of people who are not currently your customers but could be should maybe your company has internal marketing or user research data that you've interviewed people in the past. You just have a data set sitting there that you could turn into an agent bank. Here's my advice. Here's my how-to. A bad way to do this would be to define agents with a single demographic variable like just be a conservative, be a Republican. This produces all sorts of stereotyping and under underestimates the variance, the sort of the distrib like it makes it seem too certain when you need it to be more uncertain. It's just it's not enough. Not terrible would be five or six demographic variables. We saw that that was able to replicate people about 70% as well as they as they replicated themselves two weeks later.

37:53 better was data as rich as you can gather like these interviews we were we were finding we were actually able to maintain very strong accuracy by having a much shorter interview. So we went in later and we just deleted random 80% of the of the interview and we were still we went from like 85% accuracy to like 79 or 785 replication ratio to 79. It didn't really decrease it that much which surprised me to be honest. so there's something about these interviews is teaching very rich information but I do think the and I see a comment in the chat about this exactly on point.

38:32 You have the the remaining parts of the interview have to be relevant. So if you're only interviewing people about their views on fashion say and then you're trying to predict their views on climate change there's just not enough for the model to generalize. Okay, maybe if we know that they're very high, you know, high income or something like this, you might be able to extrapolate from there. The models can do quite a bit. But, you know, if if I'm only interviewing you about sports teams, and then I'm trying to predict your, you know, your retirement planning or something like that, it makes sense that it would be that the that that it's a little bit too disjoint. You have to make sure that the data you gather is relevant.

39:07 Now, we cannot proceed without talking about the risks of this. There are there's literature not every study replicates correctly. and you can see some of the the research here that has found examples of cases where this hasn't replicated. And I can show you an example of of one case where this really has happened. So here's a company that I was talking to that tried to implement this. On the left is ground truth data. This is Pew Trust. and if if it's a little blurry for you, then what's highlighted in red is this age broken out by age, 18 to 35, 36 to 51, and so on. And they're asking people how familiar they are with their retirement plan fees. I want you to pay attention to the dark blue part of the bar where, you know, say in the 18 to35 group, 13% of people said they were very familiar with their retirement plan fees. on the right you can see their simulation result where instead of 13% that one was like 1.2%.

40:09 Now you might say oh well Michael you know you still got the the the the biggest one that middle bar was was still guessed to be the most the biggest one and the second one was still guessed to be the second biggest one and so on and so maybe it doesn't matter for your particular cost but maybe it does right maybe the difference between 13% and 1.2% 2% is the difference between we ignore this sub population and they're a sizable minority that we would care about. So these things can make errors.

40:34 And the way that I think about this today is think about it as a ladder where you're as you climb up the ladder, you take on more risk of of issues. but it and but it also gets more ambitious. So you might start out at the bottom of the ladder with a question of possibility. What might this what might happen? What could happen? We're not going to put a probability label on it or say how likely it is to happen. Just this is an outcome that you might want to think through because it could happen.

41:03 Further up would be sort of qualitatively oriented outcomes, things like attitudes, anything that is sort of chat outcomes. And this is this is I think safer because we're pretty good at modeling this already given the interviews higher up. And where I just showed the error was quantitative outcomes like histograms, bar charts where you know the difference between five and 10% might really matter. And at the top would be multi-agent simulation the sort of town that I was describing an entire you know market simulation. In order to trust that you need to have a lot more assumptions. So let me dig into what you would need to see to trust each of these because this is a totally new method and I want to stress that you should not overrust such a method with the possibility rung. This one I would say mostly generally works today. what it needs to generate is a plausible chain of events that could lead to a a potential outcome and you should be able to just by looking at the chain of events say oh yeah that seems plausible that that could happen. Oh yeah if a troll so shows up in my system they could really mess with it in the following way. We need to introduce the following safeguard. On the qualitative rung, it's mostly about attitudes and so we need to be able to accurately estimate individual attitudes. I would say that my my evidence is that this mostly works today if you have rich enough data. Now, it's not the same as running an interview and I would never say that you should do this instead of actually engaging with the communities, but I think it is a good way to get a rough sense of what of what might people say in response to say a new policy or a new strategy or a new product. The quantitative rung requires actual measurement of quant quantitative accuracy. Say, can it recreate your market research surveys? In many cases, we're seeing yes, and in many cases, we're seeing that there are errors. So, you need to tread carefully here. I think the way I would view this is that you ought to you know, consider you have a hundred ideas of what you might do. use simulation to narrow it down to five that seem promising and then actually AB test the remaining five on real people to see what's what's working with multi- aent I don't think that we're really ready for that in in you know decision-m quite yet if you wanted to do this you'd have to be able to trust every individual agent to be accurate right if the individual agents are accurate then when you put them together then the emergent outcome should also be accurate you could also take a you know complex systems perspective of like oh it's more about the phase changes. I think that's much more difficult to reason through. So multi-agent can be right, but it also can be wrong and you don't know how to distinguish those. So I would be very careful.

43:42 How do we mitigate these risks? Well, for one, make sure your agents have in-domain data in their memory. Again, this is don't interview about fashion if you're asking about retirement. Second, the possibility and the qualitative rungs reduce and mitigate many of these risks because they're what I would call rough-edged problems where if it's 80% correct, it's still helpful. It's still helping you learn. Whereas the the quantitative rung is like sharpedged. It's either right or it's wrong. And if it's wrong, you might be making the wrong decision. And finally, validate important questions on a small subsample to make sure that the model isn't too far off.

44:21 Now I want to highlight here that there are frontiers of this of this space. I'll show just a couple. One is the set of tools around look before you launch. So this is actually where we started. as I mentioned I did I was doing a bunch of work for many years in online platforms where people wouldn't foresee the ways that their policies would backfire. And what we did is we found that these tools would be are actually very useful at this possibility rung to help people fix things before they go wrong. And we would get these quotes from people being like, "Yeah, everything we do is in is it set in reaction to a dumpster fire." That's really a terrible way to run a community. And in response and and instead when we gave them access to these kinds of tools, they actually iterated in simulation. They're like, "Oh, I didn't realize that this loweffort post would do this kind of thing and then the troll would respond in that way." and they actually fixed it before it launched. I now actually do this every year when I teach a course in online platform design. I have students I open the fire hose of trolls at these at their systems and their goal is to make them you know stay upright and I think we get better projects for it.

45:31 A second application of this is training soft skills. So, one that we built in collaboration with an expert in the Stanford business school in conflict and negotiation. Yeah. Is these systems around conflict negotiation. Everyone thinks that they're great at conflict until things start going wrong. And we felt maybe these, you know, these generative agents can serve as sort of like sparring partners or training partners so that you can engage in this conflict or this negotiation. Maybe you're negotiating a salary for a new job or whatever in simulation before you go and have the actual conflict. So we created a tool that let people sort of h you know simulate a conflict. It's almost like a Monty Python where we then ran an experiment on it where we randomized people into either watching a lecture on the bestin-class strategies for handling conflict or watching that lecture and going through a simulated conflict. Both groups critically did equally well on a book test of what they're supposed to do in a conflict, but only the group that had the simulation did better in an actual conflict later. It reduced having done the simulation reduced their likelihood of using a like an antisocial strategy by twothirds.

46:46 So there's something about this idea that I could try something sort of in simulation that really aids learning. Now, as I mentioned, there's, I think, going to be a number of business applications for this. I I brought this this article up. You can you can look at this for sort of a fairly up-to-date assessment of what's going on if you're in market research or something like this. I will with some embarrassment point out that if you scroll far enough down in that article, they actually mentioned this startup that that the Stanford research that I'm describing here spun out into. If you're I'm not here to shill any company, but if you're interested in chatting about sort of commercial applications of this, there's a QR code on the screen I'll leave up for just a moment. this the company that we're spinning out is called Simile. But there's a lot of open research in this space that I that I've been that I mostly focused on trying to share here. And what I think I want to do is just highlight that there's so much that we can learn from the opportunity to create this whatif machine that allows us to really go forward and think through what might happen. And so with this I wanted to make sure that we leave a lot of time for discussion. We have about 10 minutes here in the in this webinar for discussion and Q&A. I'm going to flip this over to another QR code for the for global alumni where if you want to learn more I teach an online course that global alumni can say more about here. So I'll hand this over to the global alumni team. I think there's going to be a Q&A that we'll engage with over the next 10 minutes and we'll take it from there.

48:38 >> Thank you so much, Professor Bastine, for for this this session and for your time. We are going to do now some Q&A questions. So, I can give you the first one. Can AI robots do customer service for drive-th through fast food restaurants? >> Oh gosh. can these systems do customer service? You know, that's not really a question of whether we simulate in the sense that like are you really are you trying to simulate the person that's at the at the at the window?

49:10 Can they do customer service? I mean, I think if if if you've used many different apps recently, these things are being pushed into doing customer service like chat GPT customer service agents. I think there are risks. You may have here I'll I'll for a moment or in a moment I'll I'll stop the screen share so you can see my face. the you may have heard of this Canadian airline that did a chat GPT customer service agent that then promised a refund to some of its to some of its customers who someone who had who had experienced I think a death in their family. This was this it said yes but this was actually out of policy for the company. So they then later came back and said, "Oh, sorry. That was just our chatbot. We can't promise this." And then the Canadian courts came in and said, "No, your chatbot promised this.

49:56 You've got to do it." So I think that we have to think very carefully about customer service applications here because it is at the end of the day the experience of your organization is their experience with that last mile with that customer service bot. so do I expect this to be used? Yes, I do. I think you ought to be very careful. Yes. I actually think the thing that's probably going to wind up being even much more further discussed is not simulation of other people, but the creation of entirely artificial characters. Like if you've played or heard of character.ai, AI. This is something that I think we discuss in the online course, the the the global alumni and Stanford course that you know I I expect that in the next 5 years there's going to be a lot of conversation about people who are talking to you to to AI agents completely artificial characters not simulations you know in preference to talking to other people. I think there's going to be a lot of hand ringing about this and thinking about what how what kinds of risks those those pose to people. My colleague D. Young at Stanford has actually just recently put out a white paper on this exact topic.

51:09 >> Thank you. So, next question. There are quite a few, so I don't think we're going to be able to cover everything, but thank you so much for sharing so much of your questions with us. So, does AI mimic or interpret human behavior? >> Does Can you say that again? Does AI mimic or interpret >> Does AI mimic or interpret human behavior? I mean it's sort of doing both. I would claim that I to be clear these are not like agents with cognition. These are not alive. These are not thinking agents. They are simulations of people. So I would view them as sort of mimics mimics of human behavior. But in order to mimic that behavior, they are they have to be able to interpret a lot of human behaviors and the interviews and whatever other information we have. So that interpretation is actually a critical step along the path to getting accurate simulation.

52:07 >> Next question. So human behavior is a complex system in itself. How do you even replicate a complex system in its entirety? >> How do you replicate a complex system in its entirety? That is a great question. So I think that that is actually a sort of an open research frontier a little bit. The in order to do that you would need to rep we I think we've now got the ability to replicate individual agents with a reasonable amount of accuracy but the environment like the world we haven't replicated you would need to do that right right now these agents are in a little tiny town. In fact, just as a as a funny little example, in early drafts of this, there were no no doors.

52:54 Like, we just hadn't implemented the doors. And so, the agents would like walk in on each other using the restroom because they couldn't see. And there was no like you have to get the environment to be accurate if you want accurate behavior. This is sometimes when I get leerary of people who say, "Oh, this this replicates or that doesn't replicate." Because you sort of, it's like they take agents, they put them in like a very, you know, a straight white box that's completely dissociated from any realistic environment. The agent had no earlier, you know, nothing happened to it before it just woke up and it's like given a question. What happened to that agent earlier in the day? Did it have a fight with its kid? Is it under financial stress? Like we need to actually give that much richer picture and the environment itself, which we know accounts for about four, you know, 0.4 for 40% of the of the variance in in in human behavior. I'll make one other point with a quick slide that is like not about this that there's something else I teach which is essentially that even if you could like that human society itself isn't fully predictable. So there's one of my favorite experiments of all time here.

53:56 Basically, they got over 14,000 people to split into a bunch of different worlds where they could either see or not see the sort of view counts on songs like in a Spotify clone. Like here is like imagine a YouTube you know in Spotify you can see how many other people have listened to a song or YouTube you can you see the view count. Imagine you couldn't see that. And now they have all these new songs from indie bands and they would randomize people into whether they could see the the views or not. And then they would randomize them further into parallel universes. So now there's 16 different versions of Spotify. And they could ask whether the same music becomes popular across all of these parallel worlds. And what they found was that when people could see each other, how many other people had voted, then it basically blew out of the water any consistency in the outcomes that the best songs rarely did poorly and the worst songs rarely did well, but basically any other result was possible. And so what I'm trying to convey here is that essentially even if we can, you know, simulate the world, there's no guarantee you get a single outcome in the world because the if you were to rerun the our world multiple times, you would get different outcomes.

55:03 So the best thing we can do with simulation is kind of do what's called Monte Carlo simulation. Run it a bunch of times and to learn from there what are more likely outcomes, what are less likely outcomes. But I can never give you a straight point prediction of this is definitely going to happen. >> Okay, thank you. Next question. So, if we were able to apply this to an experienced design use case, for example, creating persona based agents, how might we reach an appropriate level of confidence that we've answered the important what-ifs?

55:37 >> Oh, that to me feels like a a leadership strategy question. you know, think about you know, how does any good leader or manager prioritize the risks that they face? And I tend to think of this in terms of of agile. I don't know if you've ever read like the lean startup or something like this where you essentially start by asking what are the riskiest risks that we faced. the thing where if we're wrong completely decimates our strategy and you start by reducing your uncertainty in that risk and then you reduce the uncertainty in the in the next riskiest risk and sometimes new risky risks show up and you do this iteratively. So that's how I would think of it is there's I can't give you a complete set but I but you should be able to rank your risks by how critical it is that you be right and in fact if I'm guessing you could go to one of these reasoning models like OpenAI like now it's I would have said 03 but now it's chat it's GPT5 thinking mode and ask it you know what are you know here's my idea what are the big risks that I you know that I should be asking a whatif machine and see what it comes up with I bet it'll have some decent ideas.

56:53 >> So, next one. Does the LLM bias change the bias of the agent over time? For example, I'm wondering if LLM instilled values change the agents biases when the simulations sorry when the simulation goes longer due to a larger chance of hallucination or value bias. >> Yeah. So, there are two questions embedded there. One is does the LLM bias influence the the agents behavior? And for that the the answer is likely yes. and we saw the evidence of that in how difficult it was to to to to to the far-right conservatives were harder to model. That's I think in part because the model is, you know, trying to not be racist and sexist and so on.

57:40 But your question was about long-term. I actually think that the bias the bias that will play out over long term is more likely to be a sort of conflict avoidance like these you know chat GPT tries to be helpful and harmless and so the I I don't know if you saw in the conversations in the in the demos the agents were all very pleasant to each other they didn't get into fist fights and so I would expect that what's going to happen over time is we're going to run into a a a barrier at some point where the big models like GPT5 and so on have been trained to do math tasks and to do be helpful smart assistants, but that makes them less and less accurate at replicating people. And what we'll need to do is train separate AI models that are specifically fine-tuned to be better reflections of people. And in fact, my lab at Stanford has been doing research on exactly this.

58:32 There was a recent paper in nature actually that someone else did called Centaur Centaur that was doing this in in the psych space. I think that's going to be the next generation is AI models that are specifically trained to be like people not just trying to sort of clu it on top of you know you know clawed or or or you know GPT5. >> So we have time for one last question. So let's do what's this so what's this whatif machine versus scenario analysis.

59:07 I mean, I think a what if machine is is my goal. It's what I'm trying to create. Scenario analysis is a more structured strategy. You know, it's a process that you can follow to try and answer questions. In in my head, a what if machine becomes a tool you can do you can use to perform scenario analysis faster and more effectively. >> Okay. So, we've come to the end of our session and I don't know about you, but it certainly time flew by. So, thank you all for joining us today for this insightful session from Stanford online.

59:42 We hope it sparked valuable ideas on how AI agents are reshaping organizational leadership. And thank you, Professor Bernstein, for being with us today and for your time. if you want to take the next step and keep learning, Professor Michael leads the online course UIUX for AI products. discover real world use cases, practical strategies and how to create positive experience powered by AI. If you have any further questions or would you like to explore how this course aligns with your career goals, please feel free to reach out to schedule a call with our academic team advisors. Once again, we appreciate your time and interest and we look forward to welcoming you to the Stanford online experience. Thanks for everything.

60:26 >> Cheers. Thank you for joining. Thank you.

© transcribe · For agents Built with care and craft by Gokul Rajaram