📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

The Secret Ingredient to AI Adoption Nobody Talks About | James Evans (Head of AI at Amplitude)

The AI Ready Show55:34

Transcription

I think AI like makes everyone dangerous at other people's jobs. And so there's definitely been some moments where it's like PM will be like, "Oh, like I had an idea for a feature and like here's a prototype. Like I feel like it's ready to go. Like what do you think?" And the designer's like, "Oh man, like thanks bud. Like I'll take it from here." Sort of, you know?

Um I think the best motivator is seeing a peer do better than you because they're using a new tool. Like that makes you think like, "Oh man, I'm like falling behind." And I think if you're not seeing other people use those tools, it can be hard to appreciate their impact.

I think the prototyping has like lessened the need for PRDs. Reading AI generated PRDS just like feels dystopian to me.

[Music]

Autosklls helps teams go from AI curious to AI ready. We run tailored workshops to uncover where AI has the biggest impact in your business. Whether that's automating repetitive tasks, cutting hours out of your daily workflows, or choosing the right mix of tools and custom solutions. It's practical, hands-on, and built to get your team using AI for work that matters. Book your free 30inut AI readiness audit at autoskills.ai.

[Music]

Today I'm joined by James Evans who AI initiatives at Amplitude. James previously co-founded Command AI which Amplitude acquired about a year ago. This conversation explores how established software companies are evolving with AI, especially those handling massive customer data sets. As organizations struggle with complex data while building AI features, analytics tools face unique adoption challenges. We discuss how analytics is changing and what this means for product builders. James shares Amplitude's vision to transform analytics from manual dashboards to proactive insights, what he calls anomaly detection v2, combining traditional methods with LLMs to identify patterns and suggest actions. He also explains how session replay data might eventually replace complex taxonomies, creating a Tesla vision moment for analytics. We cover personalization challenges in AI products and how Amplitude teams use their Slack-based analytics assistant, Moda, which saw 10X usage after moving from a standalone site to Slack. For anyone building digital products curious about how AI will transform analytics and user insights, this episode offers perspectives from someone at the forefront of this evolution. I hope you enjoy the episode.

James, welcome to the show.

Thanks for having me. Big fan.

My pleasure. Well, to start off, can you give a quick introduction of yourself, your background, and what led you to Amplitude?

For sure. So, um, yeah, I always struggle with this because I really don't feel like I have much of a background. I was co-founder of a company originally called Command Bar, rebranded to Command AI three weeks before being acquired. So, either the best or worst rebrand in history. The original premise of the company was to make software easier to use with natural language prel not preLM but but preLM being, you know, the universal technology that it is today. So premise was basically that it's weird that software has become so much more ubiquitous and so much more powerful since the early days, but like usability hasn't really changed. Like we're still stuck, you know, trying to find the right page or like looking into help docs and chat doesn't didn't really seem at the time the right pattern. Other patterns like, you know, popups didn't really seem like the right pattern. So we said, let's embed a natural language text box in every every product and let users sort of just describe what they're trying to do and then let the software react. Then LLM sort of like dropped from the sky. Um, and we actually ate a slice of humble pie and said, actually, you know what, now it does seem like chat is the right form factor. And so, um, we started leaning into that, um, and ironically in product messaging as well. And so that's what we were doing, uh, until we were acquired by Amplitude about a year ago. So Amplitude, I think people probably know it. It's like a digital analytics company. And so now the Command AI products are part of Amplitude. And now I work in Amplitude. I work on, you know, the Command AI stuff. I work on AI broadly and I also lead our, um, all of our non-analytics products. So experimentation, session replay. Yeah.

And can you tell us in your own words, like, why that partnership or why that acquisition made so much sense for Amplitude to acquire what was newly rebranded as Command AI?

I know I tried to, I was thinking about in like the acquisition docs, I was think we bought the command.ai domain name for an ungodly sum and I was like, I wanted to do the weiwork thing where like I carved it out as like I was the owner of command.ai and so you know, I could do what I wanted with it, but it didn't, didn't make sense. Basically, um, Amplitude, uh, for like 10 years was just dominating in digital analytics and like, as I think a pretty natural consequence of their growth was go from being just the insight part to also the, you know, the action part so that you can like tweak a digital experience and they were already doing that with some capabilities like experimentation, feature flagging, experimentation for example. For us, Command AI was the action part without the insight part. Like we built some insight, um, capabilities to help you understand like where should you deploy a nudge or how should you configure your your chatbot. Um, but we always were used in tandem with an analytics solution, like often Amplitude. And so we basically became, you know, part of the action side of, um, Amplitude. And so now we, we, you know, we always knew at some point we would have to build some type of analytics and it was very daunting cuz like it's a huge space and so much, you know, work has gone into making those tools great. Um, and so we ended up taking the, you know, partner with a great analytics, uh, tool approach versus like building our own. It's been great. Like, I, I know there's like a spectrum of kind of outcomes for tech acquisitions. Um, and this wasn't like the biggest in the world. We were like 40, 50 people when we were acquired, but yeah, so far so good. So yeah, and it makes a ton of sense.

And, you know, it's been interesting following Amplitude. Amplitude has been investing very heavily in the AI space in the form of acquisitions of startups like Command AI and also other AI native analytics startups. I'm curious what is like the big difference prej chat GPT and post chat GPT, if we can define that as like the moment that LLMs became broadly popularized in regards to digital analytics. How does digital analytics change, if at all, or is it the same thing just applied in a different context?

It's a good question. Like, I don't think we, I'll give you our current hypothesis, but like, I don't think we've nailed it, frankly, and like, I don't think anyone in the space has nailed it. We have two beliefs. The biggest problem with analytics that Amplitude has had to contend with for 12, 13 years is that it's quite effortful to do. And the premise of Amplitude is like, oh, it should be easier to be close to the data than in a BI tool, you know, like PowerBI or Tableau or whatever. And I think does a pretty good job of that, but it's still like a lot of effort. The customers who get the most value out of AMP. We have customers who spend like 20 hours a week in the tool and they get like immense value from it. But like, obviously not all of our customers just structurally can spend that much time in the tool. And so you can't build with the assumption, I think oftentimes people make the mistake of building the assumption that like, if your tool is good, people will spend a lot of time in it. When in reality, like people do a lot of stuff and being in like in a web app is not, um, always not able to spend, you know, tens of hours a week in it. So on the analytic side specifically, we will, the promise of stuff like this is like, English is the universal interface, text is the universal interface, like I should just be able to chat with my data. And we are definitely working on versions of that where a lot less of analytics will be charts and filtering and these like constraining your thoughts into what the GUI is allowing you to do. I think the GUI still has a, there's like an interesting search versus first browse dynamic with analytics where actually the GUI can help you sort of structure your own thinking about data. So I don't think it's going to go away completely, but I do think there will be a lot more chatting with your data and we're working on that with tools like MCP. Um, and I think a step beyond that is like, what's the purpose of analytics? It's to like notice when things happen and take action. That's why Amplitude has added actions like acquiring us. And so I think the direction we're very excited about getting to soon is instead of like chatting with your data is like co-piloting, right? Like you're still in the tool, it's just easier. Maybe you spend less time to get to an insight. Would be way cooler if it's like the tool is just doing analytics for you and pinging you when something interesting has happened and maybe it's even pinging you with a suggestion of what action to take based on the observation that's been made in the background. That's much more like, you know, sell the work, not just the software. So I think we're transitioning from being like analytics software to being analytics provider in the background. Our agent agentic product which we announced in June is sort of our answer to that. So the idea is like anyone who's spending time in Amplitude instead can configure these background agents to monitor things that they care about. It's kind of like anomaly detection and then you can also wire it up to the actions you're using, whether it's, you know, guides, chat, etc., to get suggestions of like what you should do, um, in response to that change. So that's sort of how we see it today, but it's still early.

What's a maybe an end-to-end example of that where you might set up an agent to monitor for things like maybe a higher drop-off at a certain stage of the product onboarding process and then certain actions to address that? Like what might it look like to configure something like that in Amplitude today?

Yeah, that's that's a great example. There's always like the product example and then there's like the website example. So let's take a product example. Self-serve software onboarding is like, I think the easiest to think about. And what's crazy is that like, it's been cool. Amplitude works with like really big, has a lot of really large customers. And you would think, I think there's always, there's this idea in tech that like the biggest companies are running like all of the experiments and are like squeezing like all the juice out of personalization. My experience is that that's like very much not the case. And we have very large customers who don't even run experiments on a per region or per language basis. So aren't even like taking advantage of potential differences in behavior among, you know, different regions or languages. So imagine you configure an agent that says, hey, I want you to monitor, I'm define a funnel with you. So we've got a funnel, we've got a definition of like what we want our new users to do. So maybe it's like, we always use fintech examples. So maybe it's like, create an account and then it's, you got to do KYC, so you got to give us your social security number. A lot of people drop off when that happens and then you get through, then you create an account and then maybe you need to connect your bank account and then now you can send a payment and once you've sent a payment, you know, you're likely to retain and give us money. So, uh, we configure that funnel and then the agent is going to monitor across, you know, all the the slices of your user base that you care about for how that funnel is going to evolve and it's looking for two things. One is it's looking for deviations. So that's like something broke or, you know, there was a change in in sentiment around giving you know, your social security number or something like that. Um, and then the second thing is opportunity. So we bake a lot of benchmarks into these agent templates. And so the idea is it's also looking for, it's kind of running and come up with ideas in the background and potentially suggesting experiments for like, hey, I think we could do better for this audience for, you know, the social security page if we use different language or if we change up the interface. Do you want me to run that experiment? So that's an example of it's like pushing an idea to you the same way if you like hired someone, you're like, your job is to optimize this funnel, come up with stuff. So frankly, like on the back end, a lot of what these do is we, they come up with ideas and then they, they try to reach confidence in an idea being good enough to push to the user. And so you can think about it as like every hour, they're like reaching into the bag of things they might want to change. Look at some sessions, look at the data and consider like, oh, do I think this is a reasonable idea to suggest to the user?

And I, I love this idea of proactive AI. It's something I know a lot of companies are thinking about. Rather than building AI agents that are available self-serve for users to tap into, you have agents that are proactively identifying opportunities for improvement in some area or another and then in your case automating the process of addressing some of those maybe opportunities or deficiencies. What are some unique challenges that you may have faced in building some of these proactive features? And also, what does it look like to build some of these proactive features? Is it very much classic supervised machine learning anomaly detection type stuff that as you mentioned, or are LLMs useful for this step as well?

Such a good question. Wow. I'll highlight two. One is I really think it is helpful in the in the proactive analytics AI, you know, with AI space or concept, I think it's really helpful to think about it as V2 of anomaly detection. What's a problem with anomaly detection? Chatty alerts. People lose trust or silence or turn off alerts if they're like showing up all the time and not useful. So that has been a huge design challenge for us is like, and also people in different ideas of like what useful is. And so we actually, on the whole, like have a pretty high bar. They're not super chatty. They're probably not chatty enough, frankly. They're probably like missing opportunity because in our onboarding funnel, like we need you not to turn off the agent to get value from it. So we start out with a really high bar of what useful is when you're outside of the app and we're actually like pushing it to you in Slack or or email or whatever. So, so less noise and making sure that you're capturing signal rather than having a lot of noise and also capturing a lot of signal because the noise can potentially lead to churn of your own product. Basically, the, to get more specific, um, we don't want an agent to alert you. Obviously, there's a large magnitude change, like it's always going to alert you, just like with anomaly detection. Um, we don't want an agent to alert you to a small fluctuation unless it believes it has a reasonable hypothesis as to why the behavior has occurred. One of the huge advantages we have is, um, we have, uh, sessions. So if a customer configures it, the agent doesn't just have access to the events that a customer has sent to Amplitude, but they can, you know, see what the user did, subject to all the kind of privacy knobs and dials that we have in our session replay tool. And what that allows us to do is allows us to say like, okay, here's the part of my funnel. Now I'm going to watch like stuff that happened before, stuff that happened after, see if I can detect like a change in behavior that I think is contributing to this.

Well, that's interesting. It's kind of like, I think it's actually just really easy to, the litmus test for usefulness is imagine it's not an AI, imagine it's a human. Imagine you hire someone and they're like pinging you all the time with like, oh, hey, this changed, this, and you're always just going to be like, like, why? Like, don't, don't just ping me if you notice something.

Yeah, exactly. Like, you refresh the dashboard, like, congrats, like, do the work and then ping me again. If it's really important, go ahead, flag to me. But if it's not, do that extra step of build the context, build the hypothesis, maybe even build a suggestion for, oh, I think we should put a guide here. I think we should change the language. And so that is, um, we, and to your question around how LLMs play in, um, it's basically a blocker for the for the, um, notification to be sent is, you know, you do a reasoning step of, hey, do you think you have a solid hypothesis? Is there a concrete action that the recipient could take based on this information?

Okay, that's super interesting. So, anomaly detection V2 is combining some of those traditional techniques with a filtering mechanism using the reasoning capabilities of LLMs to actually sift through and identify anomalies that are worth addressing because there's a hypothesis behind them and a reason to sort of believe that this is a meaningful anomaly to address and there's a meaningful course of action that you can take to address that anomaly.

It's basically been really hard to distinguish. I think you've had to do kind of advanced ML stuff to distinguish like meaningful anomaly from within, you know, the bounds of of user sets. And LLMs, I think make that all provided you have the right context on what's potentially driving the anomaly. Again, for us, that's the sessions. It's way easier to have a sense for, is this random or is this causal in some way?

Got it. Okay. So, I'm sure a lot of your users are now experimenting with building AI tools themselves. And these AI tools are generally, or these AI products are generally a different breed of products because they're non-deterministic. They take in a variety of different inputs and their outputs are non-deterministic. You can't necessarily guarantee a response. And that's a feature, not a bug. That's part of the magic of LLMs is that they can handle and personalize based on a variety of different types of inputs. How does analytics change to accommodate tools that are a lot more fluid and unpredictable, like generated apps?

Exactly. Yeah. Um, I'm still looking for the answer to this one. Um, I think I'll give you, I'll give you my early sense. First of all, I think the jury is still out. There aren't a lot of those apps. I think there the jury is still out on like how there's a spectrum from, you know, you've got, you build a software product and it's the same for everyone to, oh, it's, you know, maybe you personalize it with overlay like popups, or maybe there's, you know, different, there's an onboarding quiz at the beginning and there's, you know, personalization based on that, all the way to like N of 1 generated UI. Um, we obviously have a lot of apps at the extreme end of not personalized. We have a decent amount that are semi-personalized. Very few that are, I actually don't really know of any that are like fully generated. Um, and so we are, like, house view on generated N of 1 personalization is that there are some areas where it makes sense. I think support chat is a great place where we've already embraced like N of 1 personalization, but for UI, like the purpose of UI and why I think we won't end up in a world where all apps collapse to a command prompt is it structures your thinking about what you can do with the software. And so our customers who ask us like, hey, should we be looking into like generated UI, our take is like, it's actually a lot harder to reason about and it's a lot harder to iterate on if you know you're wild while you're naively generating, you know, different stuff for every user. Um, I do think we'll end up with a lot more components that are at least generated with some guardrails in a more personalized way. And I think the solution is probably eval, like you're defining success and the, there's a meme out there, like everything is RL, everything is reinforcement learning. Like, we've been doing this in analytics for a long time, like trying to define success and there's been a spectrum of adoption. Like, there's a bunch of companies out there that I think are not super clear on what all their funnels are and what a successful session looks like versus an unsuccessful session. If you're trying to generate UI, you have to do that because otherwise you have nothing with which to like optimize the generation. So eval, like having a moment, uh, right now, and I think they're like going to be necessary for, um, any sort of generated UI to take off.

It, it makes a lot of sense because for analytics, you need to have some level of structure in order to detect patterns, right? It's, it's, uh, in order to understand what the different steps in the funnel are and what an anomaly might be. And eval, in essence, are an effort to introduce structure to the unpredictable outputs that can be generated from these AI tools so that you can have at least some metrics to back up how good or bad or to classify what is being generated by the AI. So it makes a lot of sense. And I know you also published a blog post on the Amplitude blog around personalization and what is too much personalization, and we'll make sure to link that link to that in the show notes. Is this the ethos of that blog post, or are there any other insights from that post that you want to elaborate on?

Yeah, that's basically, that's, yeah, that's spot on. Um, uh, the, I think people are excited about the prospect of perfect person. Everyone's always been excited about the prospect of perfect personalization. And a general thing I believe in our space is that even before AI, like there are so many places where basic best practices, I think it's often easy to forget. Like, you're on tech Twitter, you know, you're only interact with like YC startups, I think it's easy to feel like all the low-hanging fruit has been squeezed out of like software design and then try to make like a doctor's appointment. Thank God, there's all these like AI for healthcare startups that are like trying to make this easier, but like, they clearly, you know, aren't taking advantage of all like the basic stuff we've learned about how to build software over the last 20 years. And so one of the gists of that article is like, there's, there's the mid-level personalization stuff, like chat, like just asking users about themselves, users or visitors, still really valuable and there, it's, I think it's going to continue because like I said before, helps you reason about your users. It's really hard to reason about a hundred different, you know, users, all who are special snowflakes. It's a lot easier to reason about, oh, these are the users broken down by region. And these are the users broken down by sophistication, etc. And I think that leads to better thinking about what new products and services you can build for your users. So that's the gist is like, yes, now we can do N of one, super exciting, but like, don't forget that there's a ton of value to be created with like category-level personalization as well. And LLMs can help with that. LLMs, they already are with chat. We do, uh, LLM-generated guides. This is, you know, an issue near and dear to my heart. Weird corner of software, but most pop-ups are really annoying because they're not, we have learned that pop-ups are not likely to be helpful because they are very one-size-fits-all. You, you create a pop-up to like introduce a new feature or like get someone to use something and it's not like, you don't expect it's like, oh, something I should pay attention to. You can do N of one generated pop-ups. That's something we're actually working on right now. You can already use AI to generate pop-ups that are personalized for a bunch of your audiences. That's not N of 1, but it's a hell of a lot better than what people do today because you could do it. It just takes a long time to create pop-ups for all those groups. And now you can.

Yeah, I mean, you could really boil down all of the challenges that occur in developing with LLMs comes down to the amount of unpredictability that you're dealing with. And the more personalization that you include, the more unpredictability that you add into the product as a byproduct. It may increase the value, but it also increases the challenges associated with making it work consistently and reliably. And I like your framework of describing it more as a gradient rather than a, hey, either you have a product that is not personalized or that is hyper-personalized. There's a middle ground where you can benefit from some structure, and that structure enables more robust analytics and tracking while still getting some of the benefits of that personalization. So I'm going to shift gears just a little bit and talk about the broader challenge of data analysis and using AI to augment data analysis. So AI is very clearly having a moment and it has been having a moment. It's been a long moment, uh, in software engineering. Yeah, in general, it's just been having a long moment, but especially in software engineering, right? And the amount of, uh, code that is AI-generated, um, within organizations, um, has been increasing, uh, as a proportion for many organizations, very quickly. And there's a lot of predictions from folks like Dario at Anthropic, which are aggressive predictions, like 90% of code will be AI-generated by the end of year. I think you said today that happened at Anthropic and he said, "Just because we didn't fire 90% of the engineers doesn't mean we weren't right about our prediction." Interesting. Okay. Interesting. Yeah, I was, I was seeing some bubbling up on Twitter about that prediction and, uh, I think a lot of people were secretly hoping that that was not going to pan out to be true. But, uh, I mean, you can clearly read the tea leaves and we are absolutely heading towards that direction. But in the case of data analysis, the challenge is, I would say, much more formidable in using AI to augment and automate data analysis. What makes data analysis so much harder to automate than software engineering?

What would you say it boils down to? Just way higher, higher dimension. Um, so like if you think about, uh, trying to predict, like our substrate is a generic event, which is basically like a gener, a JSON structure, and every customer, like, has a different JSON structure to the taxonomy. Um, and so like, this gets talked about like ad nauseam at Amplitude, which is, you know, it's, you can't out, you can't use AI to sneak your way out of a bad taxonomy. Now, there's a couple, there's a couple solution tracks to this. You can try to use AI to come up with a better taxonomy. That's something we're working on. So really basic example, you know, we have auto-capture and so we can suggest to events that seem important and we can suggest descriptions for those events. That's one thing we can do. Uh, data quality for anyone who's thought about this is as much, is much more about ongoing maintenance than it is, uh, just the upfront definition. So you create some conflicting events, we should tell you that. And we're starting, we have product called Data Assistant that does that. There's a big drop off in events relative to historical averages that doesn't seem consistent with what's happening to the product. Sessions are really helpful for this. So imagine, like, simple example, one of our customers, um, deleted an event accidentally for a subgroup, um, for like, their critical conversion event. And so it looked like conversion was tanking, but if you go and look at the sessions, it's like, oh, people, there's no change in behavior. So it's an instrumentation issue. So today we can flag that too. You should argue we should just be able to fix that for you in your codebase, um, by making that change for you. So that's number two. Um, number three, I think is the most exciting, um, potential, you know, leapfrog moment in analytics is, do we need a taxonomy at all if we have the lossless representation of a user session, which is again, the, um, comes from our session replay capability. So my view is that we're probably going to experience like a Tesla vision moment soon where we, if you think about like event-based taxonomy is kind of like large, expensive, hard to maintain, very structured, and very, very useful when you get it right, versus just watch the, just watch the entire session and just use that to try to reason about the user session. Maybe suggest critical events on the basis of what you're seeing in those sessions. I think that is probably the best path to getting around the everyone has to build and maintain a good taxonomy problem to unleashing, um, AI on a customer's data because today, um, if you, uh, just try to do that, there's a bunch of companies where you're not going to get anything useful because the taxonomy is really hard to read. And a fourth bonus problem is that you might have a good taxonomy, but, uh, if you want a human to engage with it, they have to understand it. And so one of the clearest pieces of feedback we got when we started releasing AI-powered analytics capabilities is it was set up so that people could like ask it about events and they would get information back about events. And the clearest feedback we got was people would write in things like they would describe a situation and then rely on Amplitude to figure out what the right event is describing that situation. The users themselves like don't, you know, don't have any context in their taxonomy. And so we had to do a bunch of work and we're currently doing this work to make it so that the first thing an either an agent or just natural language interface in Amplitude is good at is helping you pick the right events and potentially at that moment, like flagging to you if there's an issue with your taxonomic. Chile are events are very high dimensional. Um, the, there's a wide spectrum in quality of taxonomies today. M. Maintaining them is a problem. And maybe sessions are all you need and we won't be have to talk about taxonomies, uh, going forward.

Yeah. And it goes into this idea of the importance of the semantic layer as well. When you're dealing with such high dimensions, having an understanding of what those dimensions represent gives not only people but also LLMs a much more pointed way of reasoning about and exploring those different dimensions in a way that makes sense rather than just simple column names or simple title descriptions. Um, but the idea of using session replays as a catch-all to circumvent the need to develop very detailed taxonomies, keep them up to date is an interesting one. Is Amplitude doing anything around that today? Is there anything available to users to be able to leverage that video data to do analytics work?

Not today. Well, we have session replay as like a kind of traditional capability. Um, but we are, yeah, we're working on being able to leverage some of that for like augmenting or potentially replacing your, um, data quality. It's kind of like another analogy is, um, kind of like Netflix and the, the idea that you could just like stream huge movies like over the internet, like at one point seemed like kind of crazy. Now the idea of like, you know, doing analysis over thousand, hundreds of thousands of hours of like user behavior, I think a lot of people look at that and like, oh, seems intractable. Turns out, by the way, that like session replays aren't actually like replays. They're, you know, much lower dimension than if you try to do AI with actual like videos of sessions. That is really hard and really slow and really expensive. Sessions, anyone who works in space knows session replays are actually like much lower dimension clickstream data, basically, or representations of the DOM that you can turn back into, um, a session. So, uh, I think it's a lot more just like Netflix, the internet caught up and Netflix could start streaming. Like, I think it's a lot more tractable than people give it credit for. It's still early. Like, I don't know if it's a 2026 thing, but I think it's very unlikely we're sitting here in 2030 and we're still talking about, oh man, like taxonomies are really brittle and like people have a hard time picking the right event to use and I've got four versions of the same event. Like, that's just it. It really seems like we can move beyond those problems.

And what are the gaps that we would fill in by 2030 to get to that Tesla vision, I guess, benchmark? Is it more compute? Is it better reasoning capabilities with these models?

I think it's more like basic ML. Um, the thing about sessions is they're really nice. That's the nice thing about events is when you start, they're very, um, you pick what you care about and you only have what you care about and you define like, a user added an account, a user added something to the cart, a user clicked unsubscribe, like, it's just what you care about. Um, sessions are super messy. Like, did the user hover over this thing because they like, they're moving their hand, or does it indicate confusion? Like, you can just go down a million. So I think the, the balance is probably going to be, how do we start with sessions, form a basis of what's important, report that back to you today, human, maybe in the future, agent, to like confirm, hey, yeah, what you're seeing there is my, you know, buying flow or my happy path or whatever. Um, I think there is an element here, probably again, like going back to eval, like training. If you have an, you have to have a representation today of what good looks like. And today people represent that with events. And then I want X% of people who did the add to cart event to do the checkout event. Um, I think in a session world, it's much more just like how, um, anyone who's working on browser use model today, have these, have these eval environments to define like how to do stuff. I think we're going to have to have that and some ability to train, uh, any sort of session-based analytics tool about like, here's what good user behavior looks like, here's what bad user behavior looks like by just showing. Um, and I'm not exactly sure how that plays out today. Like, obviously you've got companies like Merore that are like making gajillions of dollars in creating these data sets for, um, for the foundation model companies. Like, clearly it's not tractable to create the same amount of data for every software company. Um, and so I think it's going to be on us to help figure out like, how do we start with sessions so you don't have to do any instrumentation, report back to you what we think is important, get some feedback, and then on an ongoing basis, you're not saving like all the information or doing analysis over all the information, you're sucking out the useful stuff from sessions automatically.

It's a really cool, compelling vision of what a tool like Amplitude can become, and I'm super excited for that. Let's maybe move the focus from how you're using AI in product to how AI is being used internally across your company and your team. What are some of the ways that folks at Amplitude are using AI to augment their own work and save time, maybe improve quality, or whatever it might be? What are some of those most interesting use cases that come to mind?

Um, I would say the most popular are the least interesting or at least like contrarian today. So like engineering, everyone uses Cursor, um, and on, uh, we're building our own, you know, support chatbot, so like we'll have that too. Um, I think one of the probably the most interesting development on this front has been we developed, we have like an internal version of our MCP server called, um, Moda, that like plugs into like all our like internal databases. Historically, like we've been pretty bad, I think at, um, making this is ironic, but like making, you know, customer, we have our own analytics tool and then there's data that exists in other products, you know, like Salesforce data and stuff like that. I think we've done a pretty bad job internally of like enabling everyone to have access to that data. And we recently hooked up a bunch of data sources to this internal MCP product and people can ask like, hey, how many people that buy sess buy the analytics also buy session replay? How many accounts? Oh, how many of those are spending more than a million dollars with it? That was like really hard to answer before because it required, it required understanding of our product taxonomy, which is actually good and like, thankfully, like easy to use, not that would be really embarrassing if that not the case. It also requires understanding of like our Salesforce taxonomy, um, and that was hard. That's hard for like a product person, like, oh, I'm going to Salesforce, it's supposed to look like, are one of these five SKUs, like attached to the customer? That's what defines whether they buy session replay. You know, once you reach a certain scale as a company, like your Salesforce taxonomy just gets blown up and it's like really confusing, just because there's all these like one-off discounts and like promotions and stuff like that that get represented in weird ways. Um, this internal tool is really good at that. It's just really good at like figuring out, I mean, we give it some understanding of these internal taxonomies and so it's enabled, it's ironic because the purpose of Amplitude is to get people closer to the data. I think we, we do a great job with that on the product side internally, and this has done a a great job of extending that to like other other business systems. It's in Slack. Usage 10xed once we brought it into Slack. Um, we had it as like a another like an internal site before. Um, and so now like a lot of Slack is people, you know, it's kind of cool. You'll be in a thread and someone will be like arguing about like, oh, I think the persona for this product is different than the persona of this product. And then they'll be like, at Moda, like, can you, you know, verify my understanding? And it doesn't always get it right, but, uh, it's a lot. I think it's created this vibe. People, I think, gave up on the idea of like verifying questions like this with data. And it is right enough of the time that I think it has encouraged people to be more data-driven because it feels more accessible. It's like just a query away.

Oh, I love that. And one of the things that you touched on there is something that we hear across so many different companies is a challenge of pulling from various data sets and enabling those data sets for these different use cases. And, and, uh, that might be similar to MODA. Some of the challenges you've already laid out. It's dealing with the different taxonomies and maybe very complex taxonomies of some of the tools that you're connecting to. What are some of the other challenges that may be involved in making a tool like MODA as useful as possible?

I think the Slack thing is sounds so simple, but I actually think the dynamic at play is, you know, how Midjourney like was weird because it like took off in Discord and like, oh, a weird consequence of that is that you can see other people's prompts. Um, I think the same is true. I think that's why bringing Moda to Slack was more valuable. It's not because it's just in Slack, it's because it's public. And so if I see you send something, I'm like, "Oh, whoa, I didn't know it could do that." And so I think that's why it's been, uh, such a good decision to put it into Slack because it shows you like, I think ultimately, like the driver, I think a lot of companies, ourselves included, like a year ago or whatever, were basically like, everyone has to use AI now because it's like going to make you more productive. Um, and I think it's kind of some people will just take to it, but other people are just gonna be, just they're happy, they're productive, like they don't want to introduce new things. I think the best motivator is seeing a peer like do better than you because they're using a new tool. Like that makes you think like, oh man, I'm like falling behind. And I think if you're not seeing other people use those tools, it can be hard to appreciate their impact. And so you might see like, oh, suddenly an engineer is like more productive because they're using Cursor, but that's different than like watching you send a query to Moda. And so I think like the, a healthy amount of like peer competition has been like emotionally a huge factor in encouraging people to use AI. Another one is AI prototyping. Um, been encouraging like everyone on the product team, you know, if you have an idea, don't over-prototype. Like if you can just do it in a Scallet draw, just do it in a Scallet draw. But like, if it's complicated, then do a prototype. Um, and then, you know, in a product meeting, if someone brings out a cool prototype and it's like, oooh and ahh, like, you don't want to be the guy that's like never getting oooh and ahh.

I, I love that. Great lesson there in just enabling the public usage of these tools to improve adoption of those tools. And the, the Slack form factor lends really nicely to that. You mentioned prototyping tools, you mentioned Cursor for the engineering org. Are there any other tools that have seen a ton of adoption and ROI internally?

I think the prototyping has like kind of lessened the need for PRDs. I also just think like maybe it's just a me thing, but reading AI-generated PRDs just like feels dystopian to me. Um, and so we really haven't, uh, I'm sure some people use them, but we haven't seen like widespread adoption of those. Um, other things, uh, yeah, writing in general, I think has probably been, um, not super adopted. I'd say it's like engineering support is huge. Um, you know, obviously dogfooding there. What's the support team using? Like, what types of tools? Uh, so we have our own, we're building our own product. Um, so that was actually a Command AI product called Copilot. Uh, and now it's, um, in the process of being, uh, rebuilt as Amplitude Assistant.

Amazing. Are you able to share details on Amplitude Assistant?

Yeah, basically it's, um, I mean, it's, you can see the, the precursor, uh, in the Command AI version on our website, but basically, um, it's a chatbot that lives in the bottom right. And the premise kind of goes back to the start of our conversation around one of the things that was exciting, I think, to Amplitude about our company is that everyone, AI for support, clearly amazing because there's a lot of like, you know, easy answers and that degree of custom optimization that LLMs can offer is like really helpful versus getting a generic like IVR answer or having to wait to talk to a human. Um, one of the issues with chatbots that, um, we found is that they are often like devoid of context on who the user is and what they've been doing. So chatbots are kind of weird because they're like in, in the website, but they're like a portal that a different team has to like talk to the user. They're not typically owned by like product. It's not like part of the or marketing if it's a website. It's not like part of the digital product. It's like this kind of other thing on the side. And so the opportunity we saw was like, well, we know we had able to like know a ton about the user context or the visitor context from like what they did before they did after. Um, and so we can use that to personalize, like we remember like what you were doing when you came to ask the question, like maybe you asked a question about a feature and we saw you use it incorrectly and so we can personalize the answer accordingly. But the biggest impact of this, I think, is going to be on resolution, um, classification. So I have a huge pet peeve with how products measure Rex, measure resolution today. So if you look at like how a bunch of our competitors measure resolution, they look at deflection, which

Is like, did you escalate to a human? I think this is like the worst definition because you can come into my store, I could tell you to [ __ ] off and you leave the store, and I've deflected you from further talking to me. You know, like, terrible definition.

Um, so what we can do is we can see what did you go on to do in the website or the product, and does that, did, does that uh imply that we answered your question correctly? E-commerce, you asked a question about a product. Did you go on to buy that product? Like, that's like actual success definition. So we're tying, we're able to tie the chat experience to the product outcome in a way that I don't think anyone else can. Super cool.

And is Amplitude Assistant something that lives on the customer's site for their users to be able to access? And would they, would they use it to ask specific product questions? Is that's the use case? Any, any question? It's a support. It's a support chatbot. Like, our view is that there's, there's one bot you get. There's, there aren't going to be a bunch of natural language interfaces, you know, sprinkled throughout the product. It's like, there's probably going to be one in the bottom right. It's going to be the one you turn to when you have a question about anything you need help. Um, and our product is one of those. It will, will live in the bottom right. Um, and be the place you go to ask a support question, ask a product question, anything.

And, and going back to the work that you did at Command AI, is there a neat keyboard shortcut to activating it? Good question. We haven't, uh, we haven't landed on it. Command K is the natural choice, but for, for nostalgia's sake, but it's a good question. I'll get back to you. Yeah, I, I'm person, I'm a huge fan of of uh, of of interfaces like that. Command K in tools like Superhuman, and I'm starting to see Command P a lot. Command P for some reason is is starting to emerge in some tools.

Um, how are you seeing the way that PMs, analysts, and engineers them working together? I guess a work dynamic between different folks, sort of designers into the mix. Has that evolved as well as a result of these new tools? Yeah, totally. It's such a good question. It's like they've definitely bleed into each other more now. Like, I think AI like makes everyone dangerous at other people's jobs. Um, and so there's definitely been some moments where it's like the PM will be like, "Oh, like I had an idea for a feature and like here's a prototype. Like I feel like it's ready to go. Like, what do you think?" And the designer's like, "Oh man, like thanks bud. Like, I'll take it from here." sort of, you know, um, and then vice versa, like everyone can kind of code now. Um, so you can like hop into Cursor and like just get stuff done. So there's definitely been situations where, you know, you see those tweets where it's like, "If I could join Apple for a day, I would like fix this bug and then leave." Like, there are definitely situations where now if you're like a PM or a designer, you just like fix stuff um that you had to previously rely on an engineer for. So, I think our quality, like quality of life ship fixes have probably like improved as a result. Um, everyone could always do the PM's job, so that one's not like a big surprise, but like, yeah, if you want to create a PRD now, you can. Um, now the challenge with that is like, I think some companies have taken the approach where it's like, we're just going to have engineers, and engineers are going to do certainly do the products, maybe even do the designing because especially now they're they're more dangerous to those things. We haven't taken that approach yet. Um, I don't know if we will. And so I think the, the fun challenge is like, how do you still have like clear boundaries? So it's like, you don't want everyone like running loose in the codebase. Uh, and frankly, like I don't know if we've, I feel like we're in the era where like, there's going to be, you had the, there was this concept of like the product engineer um that was popular like two years ago. I think what that showed was like new capabilities, new skill sets are creating like hybrid roles, like that's kind of like a designer engineer. U those are like unicorns, like can't hire enough of those people. I think there are probably like fewer than 200 of those like in the world. Um, but I think we're kind of waiting and seeing like what's the next, what are the new hybrid roles that are going to emerge? Like maybe it's a, you know, product, and I just don't think we've quite gotten there yet. And so the challenge is, how do you give people clear ownership of their lanes when everyone can kind of like partially do their job?

Yeah. Well, design engineers, I'm seeing are a very hot commodity right now. And you hinted at the PRD, the value of a PRD maybe decreasing because people are able to build such high fidelity prototypes so quickly. Having strong design skills and a sense for good UI UX and being able to turn that into code, I think is incredibly powerful. Is that something you're seeing as well as designers being very empowered?

Oh, totally. I mean, I think today's generation of tools, like I think Figma makes is maybe getting good enough, but like the issue with the prototyping, I think it's just like wild inconsistency, frankly. If you go to Amplitude today, like there's a lot of inconsistency already, and I think the idea of like, anyone can do design now has the risk, like we've just seen a lot of prototypes that are like, "Oh man, now there's like 30 different ways of selecting an event." You know, um, it's not like a hard, that shouldn't be that hard of a problem to like make sure you use the design system. Um, this is like an example, I think something that you could get wrong if you suddenly have all, you know, all engineers and all product people like now doing design is, oh, maybe you don't get as much consistency, maybe we don't care. Like especially in a world of generated um UI, like maybe thumb's important, maybe it's just a tool thing, like making sure that the prototyping tool you use like has access to your design system. But I think we're kind of in like the wild west era where like things like this are really obvious problems that don't seem that hard, but like haven't been figured out yet. So I do expect in the next, you know, year. I'm not going to be here saying like, "Oh, our prototypes result in inconsistent design."

Great. Great. So, just to close, James, uh, since we're running on time here, what are you most excited about over the next year or so as it pertains to AI at Amplitude?

H um, a lot of our the cool stuff that we've talked about, like natural language analytics, background agents, using sessions to you know, potentially leap frog taxonomy, like a lot of this stuff is still kind of in like labs mode um with early adopters. And so I'm just really excited to, I think the people who can benefit most are probably the companies that are not the ones that are going to early adopt. And so just like getting it out there. I think the other thing is it's just like such a fun time to work in tech because every like six months, you know, there's like totally new toys you get to play with in the form of these tonation models. But like, you know, there's beyond LLMs, like we haven't figured out ways of, you know, using the new video models to like do cool stuff, or like even the new, the the world building models, which I think are like the coolest thing ever. Um, and so I just feel like it's a really fun time to be in tech because if you think about it, before this, like before LLMs, what was like the biggest technical innovation and like SAS that we could take advantage of? Maybe like React and like single page applications. Pretty cool, but like so much less cool than than this generation of stuff. So I'll close. Uh, you mentioned the the AI engineering Daario prediction that 90% of code would be written by AI. When I heard that, all I could think of was the biology prediction that Bitcoin is going to a million, like that they seemed like similar levels of audacious to me. And like Bitcoin is 100K. I'm not saying it didn't do well, but like it's far away from a million. The fact that like whether it's like 70% or 80% or 90%, like we kind of got there with engineering is like totally mind-blowing to me. Such a cool time.

James, anything to plug as we close?

Nah, Amplitude. Check out our stuff. See if you know, give me feedback. Let me know if I'm talking out of my ass. Or I appreciate you taking your time and uh, and thanks again for for joining the show.

[Music]

Autosklls helps teams go from AI curious to AI ready. We run tailored workshops to uncover where AI has the biggest impact in your business. Whether that's automating repetitive tasks, cutting hours out of your daily workflows, or choosing the right mix of tools and custom solutions. It's practical, hands-on, and built to get your team using AI for work that matters. Book your free 30-minute AI readiness audit at autoskills.ai.

[Music]