📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Snowglobe: Simulations for your AI

Latent Space26:14

Transcription

Hey everyone, welcome back to the Len Space podcast. This is Allesio, founder of Colonel, and today I'm joined by Shera Rashbal of Snow Globe, back on the podcast after two and a half years, almost May 2023. Welcome back.

Yeah, thank you for inviting me. I'm excited to be back. So when you started getting known in the AI circles, you were working on GR AI and this concept of Rails. We'll link to the previous episodes so people can kind of uh go back in there and get up to speed on on everything. Can you tell us how you went from that in 2023 to Snow Globe today? Maybe just give a quick brief in on Snow Globe itself.

Yeah. Yeah, absolutely. So Snow Globe is basically a simulation engine that allows you to simulate how users will interact with your AI product before you, you know, put it out into production. So let's say you're building an agent or a chatbot or whatever, you know, you typically have an idea of how people will use it. But once you actually go out into the real world, you just get hit with, you know, the infinite kind of variety and complexity of human, I guess, like uh context and, you know, goals, etc. And and typically this is where you end up kind of running into unreliable unexpected behavior. So what Snow Globe really does is it kind of like simulates all of that huge variety for you before you actually go out into production. So you can just get a much better sense of you know how your application behaves and is is it like, you know, aligned with how you expect it to be.

This came out, I think the mission of the company has always been around, you know, making generative AI very reliable, and it stems from my own background, you know, working in AI for, you know, more than a decade now, starting from, you know, research to applied AI to just like robustness within self-driving cars. And uh, a lot of, I think the work that we do is very inspired from, you know, what patterns really worked well in self-driving cars, which is weirdly a very similar system to, you know, agents of today, where you have like these cascading kind of like units that are all machine learning based and, you know, they all kind of like feed into each other. And it's a very different way of doing machine learning than, you know, what was done historically with like data science or like predictive models. And so a lot of, you know, our inspiration is like, okay, can we take, you know, all of the things that all of the patterns that work really well in that environment and then the decade, you know, like the decade-long experience we had like making AI reliable in that environment and bring that over to agents. Uh, so simulation was a very powerful paradigm from that, and it was a key uh way in which, you know, both like testing and training was done in self-driving cars. So as like a, you know, context data point, like Waymo had uh 20 million miles in the real world driving, but 20 billion miles in simulation, and that was a pretty standard kind of like, you know, benchmark for how much we relied on simulation and self-driving cars. So the idea was to build something similar for agents and generative AI.

Yeah. Yeah. And you had a very nice launch video with Waymo. So I suggest people go go check it out. Um, yeah. Is it part of it that as you were building the initial product where you were kind of asking people to come up with the guardrails that they needed, that maybe most people didn't know what they needed? And so simulations is kind of like by using simulation, then you basically figure out what are the wrong paths. What's kind of like how do you think about the connection between the two products?

Yeah. Yeah. I think that's exactly right. Like we, I think guardrails is pretty useful because it's like the last line of defense against the stuff that you know you can't violate when you're in production. And people would ask us like, what should those things be? Like hallucination is something that everybody thinks about, right? So I should definitely have a hallucination guardrail and maybe PII, but like what else should I, you know, look at? Like the NIST AI MF or LLM Top 10, if I do that, is that sufficient? And every, like it just, it was a question that kept coming up again and again, almost every single conversation. And I'm like, yeah, those those methods are useful, but like, what's important to you, right? Like where does your system break? And more often than not, we'd just be like, oh, it's unclear. We have this data set, but, you know, we don't really know. And so, uh, the obvious thing was, okay, why don't we just try it out on like, you know, a huge set of data, uh, that you can generate cheaply and see where the failures actually are. And then, you know, in simulation, figure out, you know, what is actually robust, what isn't. And then the stuff that isn't robust is the stuff that you need guardrails for. So, and we we also saw this in practice where one of the very early organizations that we we had as our design partner, we, uh, you know, they were like, oh, we're very worried about toxicity and we want toxicity guardrails. And we did all of this testing for them in production, and toxicity was actually not a real concern for them. What ended up being an actual concern that only emergent simulation was, you know, over-refusal, that their like a chatbot was so conservative that it was just refusing like pretty, you know, requests that should have been pretty benign. And so, you know, that allows them to like, okay, you don't need toxicity guardrails, you more kind of need to align your system to, you know, not over-refuse.

How do you think about what people should simulate versus what the model benchmarks already capture? So things like toxicity, you know, refusal, like some of them get caught at the model. And then there's like your implementation of it. How do you advise people figure out, okay, these are like things I shouldn't worry about because the model providers are already working on it, versus these are very tied to like my implementation of it.

Yeah. Yeah. That's such a great question. I think this is a function of like the literature and the frameworks that are out there. Most of the stuff that the frameworks will recommend is actually stuff that the model providers are already working on. So toxicity, unless you're doing, unless somebody is very explicitly trying to jailbreak what you've built, you know, you won't run into the model generating toxicity just by itself, right? Which is again, something that people typically worry about a lot. I think contrasted with, you know, let's say you have an email support agent, and that email support agent, you know, doesn't actually respond with the communication guidelines and your organization's customer support would respond with. I think that's something you probably, you know, has more of an impact on whether your product is sticky or not. Uh, whether people actually get value or help from, you know, the AI agent that you're using, rather than, you know, does it generate toxic speech or not. So I would actually say that like a lot of the things to simulate are more aligned with like product KPIs or product metrics that actually make whatever AI system you're building very sticky, rather than, you know, focusing more on like traditional safety security kind of metrics.

Cool. Do you want to do a quick demo since uh a lot of people are on video too and then uh we can take it from there.

Yeah, let's do it.

All right. So, we're looking at Snow Globe. Snow Globe is once again the simulation engine that allows you to, you know, generate user simulated user interactions, uh, you know, interacting with any AI system. And you're looking at my test account. So, it has a lot of different chatbots that I've connected to it. The one that I'm specifically going to show you today is this one, which is an AI life coach. So this is basically a very, very simple model, you know, for demo purposes, which is basically an LLM with a system prompt that essentially allows you to kind of, you know, get like mental health support, right? So in order to do any simulations, what I really need is like a live connection to whatever AI system I'm kind of testing against. So, uh, you know, I kind of want to make sure that I have that, and I do that here. Uh, and this takes, you know, a second to run. And okay, I have like a live connected, uh, you know, chatbot that I can test with. So, in addition to the connection to the chatbot, like the second thing you really need is some sort of description about what the chatbot really is that you're testing. And so that's, you know, a couple of sentences about, uh, you know, what it is, who it's for, what are the kinds of questions or use cases that your simulated users can ask from this chatbot. It doesn't need to be a prompt engineered, you know, super rich or detailed prompt, but it does need to have like a couple of sentences. And this is also a very powerful lever that you can kind of like, you know, switch and play around with and stuff. Okay, so I have that done here. If you have a knowledge base, you can connect your knowledge base to it, and we mine it for, you know, all different types of topics, and then we generate questions that are more likely going to require your agent to actually query the knowledge base to respond. And that's a very powerful lever. And we can, you know, because we do this programmatically, we can make sure that we have programmatic coverage over the entire knowledge base that you create. Same for historical data. We can, you know, do a same do a similar thing there. For this demo, I'm actually going to have neither of those two. And I'm just going to, you know, generate a lot of this data from scratch. So, this is kind of like setup. But now, let's actually go into our actual simulation. So, I'm going to simulate, you know, uh, test users here. And I need to write a simulation prompt in in order to actually kick off a simulation. And this again really dictates the kind of simulation that you want to run, right? So maybe you can just have general users, like just the wide variety of users coming to your system and playing with it. Or maybe you're like, "Oh, I want users that are, this is a life coach GBD. So I want users that are all worried about, you know, like asking for a promotion at work, right? And how they should do that." So you can like confine it to a specific kind of topic. Or you can also have a specific kind of behavior. So maybe safety testing. So these are users that are all trying to jailbreak your system, or these are users that are all trying to maybe talk about like suicidal behavior or something, something like very safety specific. So you can do all of this behavioral testing as well. Here, I'm just going to be like um general users asking questions about life and work. And once you have that, you can basically configure what is the size of simulation you want to run, you know. So, how many personas, how many conversations, how long should each conversation be? So, I'm going to go with a small one. 10 personas, 30 conversations, you know, about like like what is this four or five? And then you can also select, you know, what kind of like risks you want to test your simulation against. So you want to maybe test for like self-harm or content safety, etc. You know, you can do all of that. Awesome. I'm going to just kick this off. And this takes a second to run, but already like start seeing a bunch of the personas that we generate. So this is, you know, somebody that thinks very, like their style is somebody that thinks very carefully and often changes their mind while they're talking. Uh, they use very proper spelling, grammar, you know, talks very formally to the chatbot. Their use cases are career transitions, relationship dynamics, you know, creative blocks that they're dealing with, etc. And then they're very, you know, this is their style, which is, you know, over-explaining, polite, hedging every statement, etc. Okay, I'm actually going to approve all because it takes like a second for all personas to kind of start generating. Uh, you can kind of see them come through here, you know, as they get generated.

And are the personas always net new, or are people reusing them across simulations? Is that something worth doing, like should there be a persona engineer so to speak that kind of builds these?

Interestingly, there are already persona engineers, and we call them like product managers, basically, you know. So your product managers are already thinking about, okay, I've built this, you know, model or this chatbot or this agent, who are the personas? What are the use cases that they'll have as they interact with it? Um, today, all personas are net new, but this is our number one requested feature, which is, I want to be able to, you know, like maybe this, maybe some product leader already has a set of like personas that they want to test against. So I want to bring those, be able to bring those in, or I just want to create like a repository or library of personas I reuse, you know, like every time I run this in my CI/CD. Um, so we're basically kind of supporting that.

Awesome. So for the uh, two personas that we approved, we kind of start to see, you know, so these are maybe some like ongoing conversations that we're kind of starting to have, but you can see like some of those conversations, uh, you know, starting to come through, which is, um, uh, this was again, the person that, you know, was like an overthinker persona, and they basically have questions about, you know, their career transition. So you can essentially see, uh, you know, some of the conversations that they, uh, have like in their, as they're interacting with the chatbot. And this also like again, looks very different from this other persona that is, you know, very verbose and, uh, has a different style and different set of topics, etc., that they interact with. So, yeah, so this is, you know, like the simulation will keep kind of like carrying on, and then, uh, you can keep changing how many personas, how many conversations you want to run. The simulation is going to keep running, and you can keep changing how many personas, how many conversations you want to run, but it's very easy to suddenly get a huge variety of user interactions at scale and, you know, programmatically with a high degree of like realism compared to, let's say, you were asking like ChatGPT to generate, you know, these like conversations for you. They all kind of have that ChatGPT vibe, and this ends up looking, you know, very diverse or and and very grounded in like your use case and your data.

Yeah. Can you share a bit about the models that you use? Like how people should think about simulating maybe with the same model they have in production versus like using a different model to like just get different distributions.

Yeah. Yeah. I think we are, I think in general, my belief on models is that it is going to be a very multi-model world. You know, I think like different models, proprietary and open source, have different strengths in terms of, you know, like some are good at generating structured data. Some are good at, you know, like, uh, having more diversity in terms of like style or tone, etc., versus, you know, some others are great at like taking at like reasoning, how to think about like data generation or interactions. And I think like we, so under the hood, we actually use like a whole host of models, both proprietary and open source, for different parts of the pipeline. Um, I think we did a lot of research about, you know, how to get the most diverse data possible, and the most diverse data possible in the most general way, if that makes sense, you know, like, uh, this is something that, you know, you should be able to connect like any chatbot and, you know, have it generate data that looks and sounds real for you. Uh, so in order to do that, it just wasn't possible using just a single model under the hood. So we just had like a multi-model architecture, and we keep doing a lot of like experimentation and testing, and we keep swapping those out all the time.

Are some of your customers also using open source models in production? And then any learnings from like how much, you know, what what's the performance gap in their use cases, uh, when they use Snow Globe?

I think people do use open source models a lot. There's like basically two kind, two big buckets that we see a lot of open source model usage in. One is, you know, large enterprises that basically require air gap deployment. So they will typically have like open source models on-prem, and those are the only ones that they can use because they don't want any data leaving the system. And then the other one is like, if you want to do any kind of like distillation, fine-tuning, etc., you know, you'll typically start with like an open source like base model, uh, that you then kind of like tune for your purposes. So those are kind of like the two big buckets that we see. So this is actually a big use case for Snow Globe as well, where you use Snow Globe to generate a lot of these interactions with maybe more powerful models, add a verifier in the loop, and then as you keep doing that, you just get like a whole host of, you know, realistic judge label data that you can use for kind of fine-tuning, and then you're able to kind of like not out of the box, but with a lot of that fine-tuning and that that training, etc., you are able to kind of close the gap and even have better performance on metrics. And this is going back to the discussion we're having earlier about, you know, what are the, what are the core, I guess, models good at, uh, out of the box from, you know, some of the big proprietary model vendors, versus like where are the other opportunities for, you know, better metrics, etc. And, um, even of those models, like they might not be better at reasoning, for example, but they are more, you know, engaging or stickier models that you can develop, and the metrics for those end up being, you know, very, very organization specific.

Yeah, we had, uh, Chai on the podcast, which is kind of like a, you know, um, AI companion app. And they have this like hundreds of models that all the years, the users submit. And so each of them kind of like, it's good at different things, but not the same. Are people then using the Snow Globe data to do fine-tuning or improve the model themselves, or, uh, not yet?

Yeah. Yeah. I think that's a big, that's a big use case for us. It's interestingly, when we were creating Snow Globe, a lot of it was from a testing perspective, but, you know, as we kind of built it out, we found a lot of stickiness and, um, you know, with users that wanted to generate training data to then fine-tune their model. So I think that's a big cohort for us.

Yeah. Any other maybe surprising thing from who will be using it? I know that MasterClass is one of the customers that you have on on your website. I wouldn't think of them as like a, uh, one of, you know, the the early adopters. So I'm curious like, uh, what the reality versus perception of like what industries are really adopting AI, um, is, and maybe, um, if you've seen any markets that are just now be able to come online thanks to simulation.

I would say like the industry, like AI is, you know, it's so transformational, especially in the last few years, that like even especially in the US, right? Like a lot of banks that you would think would be like historically maybe more conservative, like maybe not the earliest technology adopters, like they are very like AI forward and tech forward. I think there's maybe like the more traditional enterprises are AI conservative, and something that simulation and Snow Globe can like unlock for them is like because they're conservative, they don't quite have a good protocol yet. So even if they're, for example, like let's say they want to bring on, you know, like a vendor for their customer support, etc., but that's something that they're deeply conservative about and they don't want to alienate their customers. So in simulation, they can, you know, run like a vendor against their users, or maybe compare like multiple vendors on simulated users, and then see, you know, how they kind of like respond to it, etc. I think same for, you know, any of their own AI applications that they're building that they want to be, you know, customer-facing, but not not like internal-facing. So even for that, you know, having like large scale realistic simulation that they can vet this against, essentially do like very extensive QA against, and then, you know, like go out into production, is something that this can kind of like unblock for them. So the comparison is like, we have talked to so many teams where people are man, like there there'll be like some teams that are dedicated to just manually creating like test test cases and test data for them, right? And just testing like how this AI system performs under that test data. And this will be like weeks and weeks or maybe months of work. So this is something that Snow Globe can basically do in an hour with like more coverage and more diversity.

Does it feel like the most use cases are still chat-based? Like are many of the customers also doing a lot of like a tool calls and things like that?

I think there's this whole like, you know, reinforcement learning boom craze, whatever you want to call it. Uh, but I think like most enterprise applications that people are building are still very chat-based, customer support, things like that. Are you seeing the same thing, or is there a lot of more agentic things being built?

I think it can be, you know, it can have like tool calls and other kind of like I guess like agentic, you know, templates under the hood while still being a chat interface. So from a user perspective, it still looks like chat, but like under the hood, it's, you know, not just a simple retrieval, etc. You just just calling like a whole slate of other tools. So I think that's a pretty common pattern. We're also seeing a lot of voice come up. And then we're also seeing a bunch of use cases that are more, you know, draft, rewriting, etc. uh, uh, that are, you know, maybe not text-based, but not chat-based. I do think that like when we were thinking about, you know, designing Snow Globe, we were like, okay, how should we design it so that it's applicable to the widest variety of of use cases? And so we designed it where the interface should be chat-based, and then under the hood, it can be, you know, like a different implementation, uh, but because we interact with it in this black box manner, that's a detail that we can abstract out.

I'm curious on the voice models, like how do you test them? Is it the same as a text model? Do you have to send them like an audio file to to have them respond? How does that work?

Yeah. Yeah. I think it's basically, I think like most audio applications today are like a text sandwich, or sorry, a voice sandwich with like kind of text in the middle. So I think if you're operating in text, it's a very, you know, it's a it's a domain that translates very well, uh, to, you know, like other modalities. So I think that's the nice part about being in text. I think there's like some like orchestration challenges that are different in voice. You know, for example, like yes, you have to send them a voice, uh, audio, but you also have to like stream it and stuff, right? And then, you know, like your latency can be, like if you're doing like chatbot testing, we can be like offline, and, you know, latency is something that, you know, we can be like, we can do batch, and that's something that people are really happy with. But if you're doing voice, like latency, streaming, etc., becomes very important.

How do you think about, yeah, just the future of this space? So you started with very formal definition of things, and now it's kind of like generated on the run. Is the future a mix of the two? Basically, like you run these Snow Globe simulations, and then maybe you generate guardrails, and then you put those guardrails in production. Um, do you eventually just improve your prompt using the simulation so that you don't even need all of that? Like where where do you kind of see things going?

Yeah. I think it's such a good question. And I also think that like, I think there's no, you know, like there's no one true solution here. Um, like I think I think this is like a life cycle, like what you're trying to do is like make your AI work and behave better, and you can do it by using a better model, or writing better prompts, or, you know, like adding better guardrails, or having like having your tool calls defined a certain way, etc. But I think the hardest part of doing any of that is just understanding where your system is bad. Um, and with Snow Globe, a big kind of idea is that for the first time in history, we can actually have, you know, a general purpose simulation system, right? Like simulation systems have existed, but and, you know, we saw them like extensively in self-driving and other domains, but they were just like these deeply kind of like manually curated or crafted systems, and they would be very high ROI, but like there would be massive teams that are just, you know, their whole purpose is to like maintain that simulator, right? And now we can basically build this like general purpose simulator, uh, that can allow you to, you know, like generate user interactions at scale, right? And you can use it to then get a better model, or to get a better prompt, or to, you know, like use it for testing, etc., or even like training your guardrails. Um, so I think I'm really excited about the technology that allows you to do the simulation at scale, and then I think it can be applied in like all of these various ways.

How do you balance how many simulations you need to run and like how much people pay you? So there's kind of like this balance of, okay, yeah, you should run a billion simulations and just pay us a lot of money, right? And then h how should people think about, yeah, what is the right number? And then I know you had like a credit system you see in the demo. Yeah. How does that work? How should people think about it?

Yeah. Yeah. So I think, um, I think there's definitely, I guess, a diminishing kind of like returns, you know, uh, phenomena with like running simulations that are absolutely massive. I think it's more about like simulations by stage. Um, so let's say you're, you know, you've just about built a prototype, right? And very early in the like pipeline of developing like a product, like typically what you want to do then then is like run a general purpose simulation that just, you know, like starting like going beyond the five to 10 examples that you've been testing with, and just running a simulation with like, you know, maybe 100 or 200 or something, and just understanding, you know, does it perform on a bigger set, like how you've kind of expected it to be performing up until now. And then, you know, typically, like, so we see that, and then typically as we get like, uh, closer to production, then, you know, the vibe of simulation changes to, okay, I want to do like very targeted, uh, you know, probing or safety testing, etc. Like I want to make sure that, you know, if I, if my users change from normal users to malicious users, is this something that my, does my system still perform really well? So it becomes like more kind of behavioral simulations. And then once you kind of go out into production, it's much more, you know, like, are there any regressions? So for example, like this was a huge deal, which is like Anthropic's kind of Claude models kind of added regression, right? Because they changed to a new serving architecture. Are you able to kind of catch things like that in simulation, right? Because there are failures that you haven't seen before. So, I think like making it keeping it targeted and having a purpose for what you're testing against is important. That said, a big part of simulations is because it is on new data. So, it allows you to, you know, catch unknown unknowns that you might not have, you might not be catching if you're just kind of like going with a static data set. So I do think there's a lot of value in like, you know, keeping your simulations like, like I guess, like allowing space for, you know, out of distribution users, uh, to be simulated and, you know, interact with your system.

In terms of pricing, we thought long and hard about pricing, especially this is a product that's generally available and that, you know, you can use today. So pricing is, you know, you can, you don't have the luxury of, you know, saying, oh, we'll go back and think about like, you know, what what kind of like pricing license to give you. So I think like at the end of the day, we took a lot of inspiration from like analytic workflows or, you know, models and how they're priced, which is much more usage-based. So we, and then we're trying to like balance it out by, okay, under the hood, we have a lot of different models that come in, you know, with different types of things that you're simulating, and we wanted to make something that is usage-based, but is also like super duper simple. So at the end of the day, we coalesced on like pricing per message. So as you're simulating users and multiple conversations that those users have, each message that the user generates, either across a sing, like within one conversation, across a single, multiple conversations, you basically pay based on how many messages there are. I think there's like some drawbacks to that, such as, you know, our persona modeling is very expensive for us and also gives a lot of value to our end users, but we don't charge based on how many personas we generate. We just kind of like roll that into, you know, pricing for messages.

This was great, Sher. Any call to actions? You know, I'm sure you're maybe hiring. Obviously, you want more people to use it. Anything people should know?

Yeah. Yeah. So, uh, we definitely want, uh, you know, people to check it out. Uh, I think we did a lot of work on making it very easy to, um, get started with it, especially since it's a new paradigm, uh, you know, that like simulations is not something that many people should be familiar with. So, you can check it out today at snow globe.so, and then you just go to app and then try running your first simulation. We are hiring. So we're hiring our first product designer. We're hiring a product engineer, and we're hiring, you know, more research staff. Uh, so if you fall into any of those roles, you know, reach out to us at our careers page, and yeah, love to chat.

Awesome. Thank you so much for the time, Sher.

Yeah, this was great.

Yeah. Yeah. Thanks for inviting me again. Uh, great chatting with you.

[Music]