Transcription
Last year, a paper from Stanford put thousands of AI agents in a fully simulated environment and let them live their lives. The agents formed relationships, built up memories, and developed their own personalities. It was stunning, and this paper allowed us to envision what could be possible with simulating entire societies or even the future of video games. Imagine a video game world where NPCs had actual personalities, backstories, and lived their lives in real time—that was the promise of that paper.
But now there's a new paper by the same author, also from Stanford, showing it's possible to get real human personalities into these agents, to then live their lives in these simulated environments. It is mind-blowing. Let me break it all down for you.
This was the previous paper: "Generative Agents: Interactive Simulacra of Human Behavior." What they found is that by putting all of these AI agents, powered by ChatGPT, into this simulated environment with a little backstory, they would develop their own personalities and relationships. They would form friendships; they would develop plans. For example, one agent threw a birthday party and invited all their friends. But not only that, those friends invited some of *their* friends, and they coordinated and showed up to the birthday party. It's really incredible to think about. But imagine this: you could actually put your own personality into these agents.
Now let me show you the new paper. By the same lead author, Jun Sung Park, we have this new paper: "Generative Agent Simulations of 1,000 People." The gist of what they've done is essentially taken that other paper but interviewed a thousand people—two-hour-long interviews of various questions—trying to extract the personality of a real human and then place them in the simulated environment. Listen to this: "We present a novel agent architecture that simulates the attitudes and behaviors of 1,52 real individuals." They not only replicated the personalities of these real individuals, but they were able to test and prove that those agents behaved and had the same thoughts and personalities as their human counterparts. The results: the generative agents replicate participants' responses on the General Social Survey 85% as accurately as participants replicate their own answers two weeks later.
So what does all of this mean? Let me break it all down. They used very common social science tests, such as the General Social Survey, the Big Five personality inventory, well-known behavioral economic games, and other social science experiments to try to extract the essence of what makes up somebody's personality. They then took that two hours worth of interview, converted it into memories for these agents to base their own answers on, retested them against the General Social Survey, the Big Five personality test, and the social science tests, and they found that those agents behaved essentially 85% as accurately as the humans did when asked those same questions two weeks later. So that is extremely accurate.
But why would they do this? What is the point? According to the paper, these simulations could help pilot interventions, develop complex theories capturing nuanced causal and contextual interactions, and, in my opinion, most importantly, expand our understanding of structures like institutions and networks across domains such as economics, sociology, organizations, and political science. Essentially, we can start to predict how people, organizations, and societies will behave without actually implementing something extreme first. So if we have this idea about a completely new tax plan, for example, and we want to see how people will behave based on this new tax plan, rather than actually having to go implement it and then seeing or predicting based on not-as-accurate methods, we could actually set up entire societies of AI agents based on real people and then see how they might react to these massive tax changes.
Today's video is brought to you by Mamut. Mamut AI brings all the best models together in one place for one price: Claude, Llama, GPT-4, Mraw, Gemini Pro, and even GPT-1. Rather than having to pay for each of these AI separately, you pay $10 to Mamut, and they bring it all together in one place. Plus, they have image generation: Midjourney, Flux Pro, Dolly, and Stable Diffusion—again, all for $10. Models are frequently updated as soon as they're released, so be sure to check out Mamut for access to all the best models for one low price: mamut.ai (that's M-A-M-M-O-U-T dot A-I). Thanks again to Mamut. Let me tell you how it works. We present a generative agent architecture that simulates more than 1,000 real individuals using two-hour qualitative interviews. The architecture combines these interviews with a large language model to replicate individuals' attitudes and behaviors. By anchoring on individuals, we can measure accuracy by comparing simulated attitudes and behaviors to the actual attitudes and behaviors. They tested it against the General Social Survey, Big Five personality inventory, five well-known behavioral economic games (like the dictator game and the public goods game), and then the prisoner experiment (which is, funnily enough, out of Stanford)—five social science experiments with control and treatment conditions that were sampled from a recent large-scale replication effort.
So how did they create these agents that essentially replicated people's thoughts and behaviors? Well, they turned to in-depth interviews—not surveys, not simple question-and-answer sessions. They actually gave them real, long-form, and at times dynamic interviews, combining pre-specified questions with adaptive follow-up questions based on respondents' answers—a foundational social science method with several advantages over more structured data collection techniques. They were semi-structured interviews, meaning there's a set of questions, but then they allowed for dynamic and not predetermined follow-up questions. The interesting part: they actually used AI to do all of the interviews. This freedom to answer questions and really dynamic follow-ups give interviewees more freedom to highlight what they find important, ultimately shaping what is measured.
Here's how it works: they have the human participants; they give a two-hour voice-to-voice interview (that is, both sides are using voice). The interview script is then transcribed and those are given to the generative agents to serve as the agents' memory. Now, if we remember back to this previous paper, they basically started the agents with just a very brief description of their background, and each agent, as they interacted with the world, would develop these memories—long-term memory, short-term memory. They had some really cool techniques using RAG (Retrieval Augmented Generation) that essentially allowed them to draw from the memory to determine what actions that agent would take next, whether that action is where to go, what to do, or even how to interact with other agents within the simulation. But now, in the new paper, they took that two hours of interview that essentially gets the essence of the thoughts and behaviors of a real human, and they use that as the memory for the agents. So technically, those agents should behave how their real human counterpart would behave. Then they had the actual participant responses two weeks later, and then they tested what the simulated responses from those agents would be. So basically, they gave the test once, used that to give memory to the agents, then two weeks later they gave those same questions—those interview-style questions—to the humans and then they also gave those interview-style questions to the agents and compared the results, and it turns out they were really accurate.
The interview script explored a wide range of topics of interest to social scientists: from participants' life stories ("Tell me the story of your life from your childhood to education to family and relationships, and to any major life events you may have had") to their views on current societal issues ("How have you responded to the increased focus on race and/or racism and policing"). Those are just some examples. Then the AI interviewer dynamically generated follow-up questions tailored to each participant's responses. Then they took all those responses and gave it to the agent as memory. So when an agent is queried in the simulated environment, that entire interview transcript is injected into the model prompt, instructing the model to imitate the person based on their interview data. For experiments requiring multiple decision-making steps, agents were given memory of previous stimuli and the responses to those stimuli through short text descriptions. The resulting agents can respond to any textual stimulus, including forced-choice prompts, surveys, and multi-stage interactional settings.
Let's look at the actual results. For the GSS (the General Social Survey), the generative agents predicted participants' responses with an average normalized accuracy of 85%—stunning. These interview-based agents significantly outperformed both demographic-based and persona-based agents. So what does that mean? They basically took a bunch of information or a bunch of knowledge about what a demographic might respond with, rather than interviewing individuals, and then they used that as the memory. And what they found is when they used that more generic knowledge rather than interviewing individuals, they didn't perform nearly as well, and that actually shows bias in that data, in that generic data. For the Big Five questions, the generative agents achieved a normalized correlation of 0.80—again, stunning—and once again outperforming the demographic-based and persona-based versions of those agents. On the five well-known economic games designed to elicit participants' behavior in decision-making contexts with real stakes (they have the dictator game, the trust game, the public goods game, the prisoner's dilemma), the generative agents achieved a normalized correlation of 0.66. And just to make sure that the agents didn't just answer the questions right or accurately just based on their existing knowledge before all of this additional knowledge was given to them, they randomly removed 80% of the interview transcript (basically 96 minutes of the 120-minute interview), and the interview-based agents still outperformed the composite agent, achieving an average normalized accuracy of 79% on the GSS, with similar results observed for the Big Five.
To investigate whether the predictive power of the interview stems from linguistic cues or the richness of the knowledge gained, they created "interview summary generative agents" by prompting GPT-4 to convert interview transcripts into bullet-pointed summaries of key response pairs, capturing the factual content while removing the original linguistic features. They are just reaffirming that the knowledge gained from these interviews is real and not just from different language cues. And once again, it outperformed. These findings suggest that when informing language models about human behavior, interviews are much more effective and efficient than survey-based methods.
They also talk a lot about bias in artificial intelligence, which I know a lot of you have asked me to comment on and make videos about. I've done it a little bit in the past, so we're going to talk about that a little bit right now. There is concern about AI systems underperforming or misrepresenting underrepresented populations. Basically, if there's not enough training data for underrepresented populations, how will the model know how to respond based on what those populations would have responded with? To address this concern, they conducted a subgroup analysis focusing on political ideology, race, and gender—dimensions of particular interest in relevant literature. So what did they find? Interview-based agents consistently reduced biases across tasks compared to demographic-based agents. So for political ideology, the bias dropped from 12.35 for demographic-based generative agents to 7.85 for interview-based generative agents, and similar results across the other benchmarks that we talked about.
One thing that we've talked about on this channel: essentially, all public data has been used to train models. Now imagine if AI agents were able to interview humans across the world—different questions, interview-style with follow-up questions dynamically—and all of that data was then used to train different models for different purposes. That is a huge amount of additional data. And then if all of those agents were responding how their human counterparts would, or at least highly accurately compared to their human counterparts, then imagine we can create all of this synthetic data by interviewing the actual AI counterparts. It's really cool to think about. But that's the gist of this paper. I hope you enjoyed it. I thought it was fascinating. If you enjoyed this video, please consider giving a like and subscribe, and I'll see you in the next one.