📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Investors Called This Entrepreneur's Idea 'Ridiculous.' Now They've Raised $100 Million

Forbes25:20

Transcription

Hi everyone. I'm so young. I'm one of the co-founders of 12 Labs, where I lead our go-to-market efforts.

Uh, today I'm just going to talk about video. And, uh, video is interesting because it's often a forgotten data format when it comes to how we understand, utilize, and analyze. Uh, but I don't think anyone here can reject the idea that we use video every day. It's how we communicate. It's how we, you know, tell stories. It's how we record our lives now. Um, it's how we work together.

And so, 12 Labs, as a company, we were founded, uh, I think four and a half years ago, and we started thinking about what can be done to be able to contextually comprehend videos at scale that allows us to better utilize, uh, all of the knowledge, insight, and stories that live inside video. So, we'll go into that.

Uh, very simply, video dominates the way that, you know, we live today. It's 90% of all the data in the world. It's going to increase. Um, and just as an example, here are some of the industries where video is so important. Um, video-centric applications, video-centric workflows. Um, of course, media, sports, entertainment is a really big place where the most valuable stories that we love to watch, uh, that we love to share, uh, are captured and created. Um, advertising also, um, told in video is incredibly powerful. Uh, we look at the creator economy, um, law enforcement, how do we stay safe, uh, be able to manage evidence, do investigations, um, healthcare, automotive, um, enterprise knowledge, all in the form of video.

But here's the problem, which was, uh, even if we know that video is powerful, the technologies to be able to understand video used to rely on frame understanding. So, if we assume maybe a video might have, every second, maybe let's assume 24 to 30 frames, you're extracting a frame and you're analyzing what objects are in a frame. Uh, sometimes we can even do, because speech recognition is very commoditized, we can even extract transcripts from video and analyze the dialogue and analyze, uh, the language part of video.

But neither is fully sufficient because if you think about the way we see, hear, and understand the world, we're actually, it's actually temporal, right? We see things that are changing, moving. Uh, we hear sounds that are not necessarily in language form. If I clap, that's a sound. If I say something, that's also dialogue. Um, but also sound. So, how do you combine all of these things, put it in the context of time, and be able to, almost like a human, contextually comprehend what is happening inside video so that we're able to use, uh, the context? We can simply, very easily, search for things and scenes and moments in video. Uh, we can ask questions of what's happening inside and almost be able to communicate and freely use all of that context that is available to us because we're storing a lot of it.

Um, here's also another challenge. Uh, when we had started the company, at the time, LLMs were still very nascent, but even LLMs at the time, um, and even today, uh, are not, they can't see the whole picture. Uh, language is very different as a data format from video. If you think about the way that words are spoken or said, uh, every word has a meaning. There's a reason why we say each word. Uh, video is very redundant in nature, and it's multimodal, right? Uh, it has all of these different inputs, and you have to combine them together and understand everything temporally.

And so, traditional ways of analyzing video, um, we talked about frame-based computer vision, there's analyzing speech, um, but actually, what we found is that most customers, most enterprises that are video-centric, a lot of the ways that, as consumers, we find video for ourselves is manual. Uh, so you actually go back to footage, and as you're watching it, you take notes, you tag it, and a lot of organizations actually store spreadsheets of tags. Uh, often times, they're video file-based tags. You're not even getting, uh, in-scene, in-video, temporal context. Uh, you're usually storing, you're taking notes on what's happening, and that's really what you're searching through, uh, years later when you're needing to better utilize those assets.

Uh, I'll walk through some examples of a case study later on on what that means. And so, if we were to describe things that we need to find, let's say you're a sports organization, you're a sports team, and you're trying to find when there's a touchdown happening, maybe there's a crowd, the fans applauding, um, maybe there's a really cool logo in the background while something really exciting is happening during a game. How do you actually describe and find that exact segment so that you can create a highlight from it? Um, or so that you can, uh, show it to fans.

And the same problem and challenge exists for every organization that has video. Whether you're trying to go through petabytes of, uh, evidence to be able to do an investigation and or write reports. Uh, whether you're trying to create and repurpose content from older shows. Uh, it's all the same problem of how do you understand footage at scale? Um, in traditional ways of tags, you would not be able to do this. You would not be able to describe something you're looking for and be able to find it exactly when you need to. There is no Control-F for video.

And so, the technology, uh, that 12 Labs builds is multimodal video understanding. Uh, we built foundation models that allow you to do semantic search, uh, and retrieval across vast amounts of multimodal data, in this case, video. Uh, you also are able to communicate and converse with video. So, this is an example of a query that you could search that you can run: "receiver catches deep ball for touchdown," and it would go across all of those moments that you have and pinpoint that that exact scene.

You could also, uh, because the AI has created a memory of all of the footage, all of the videos that you have, you're also able to reason with it. So, you can ask, "Describe what happens in this clip. What's funny about this video? Why would an audience find it interesting?" Uh, "Create a highlight from it," and be able to extract response, intelligent responses from it based on the AI's understanding of your content.

This is a way of, um, so this is an overview of how a video foundation model will work. In the case of 12 Labs, we have two models. We have our Moringo model, and think of Moringo, uh, as a way of creating this first memory of video. So, the AI will see, hear, and kind of understand and almost create this memory, and we store these in the form of vector embeddings and, uh, create vector embeddings that capture all of the context of video. And these, this memory, these embeddings allow you to very easily search, uh, and find exact moments, uh, across your library.

Then we have our second model, which is called Pegasus. We like to name our models after horses. In this case, it's a fictional horse, but, um, Pegasus allows you to then reason on top of the memory that Moringo has created. And so, these two work together to enable a holistic, uh, set of use cases, uh, whatever type of use case you have, uh, for your video.

So, one of the industries where video is super powerful, as we know, is, uh, where content is created and stored, um, media and entertainment, um, advertising, and how we, you know, tell stories. This is, I believe, also being captured in video. Um, so, let's take a few examples of overall processes of, of content management. Uh, so, one, when you're creating content, when you're going from raw footage to a story that can be told to many, that's actually a, it's, in some ways, it's a search problem. Uh, you have a narrative in your mind as a creative. How do you actually go across all of your, the raw footage you have, uh, extract all of the moments that are relevant, and be able to put them together so that you can very quickly iterate, um, on different stories and be able to, you know, uh, put that together? And that's a problem that not only individual creators would have, uh, but also the largest organizations and the largest, uh, production houses, would have too.

Uh, distribution, um, when we get recommended content, uh, from platforms, are they actually personalized? Is it actually understanding what I like to watch? Because what I said I like to watch when I signed up for something could be very different from my actual behavior. And so, uh, recommending, being able to help me as a user discover content that's more relevant to myself that I actually enjoy watching. Um, all of these are recommendation and discovery that can be powered through enhanced content understanding.

Uh, I think a lot of us are probably finding, uh, shorter-form content on social media that we're, you know, spending a lot of time scrolling through. I know scrolling through. I know I am. Content repackaging and monetization. How do you actually repurpose all of these incredible stories that have been created in the past into more digestible and consumable formats that is more personal? And because you can do this at scale, it becomes more personalized, um, way for me to consume as a viewer. Uh, these are all challenges that we find, uh, challenges, or opportunities, as you can call it, opportunity that our customers have that, uh, enhanced contextual intelligent video understanding is able to help with.

And so, for 12 Labs, we build the AI that can understand video, um, and we provide the technology in forms of APIs that allow our customers and partners to be able to build on top of the technology that we have. So, really quickly, I'm going to go through an example, um, of a customer case study. So, Maple Leaf Sports Entertainment is, uh, one of the largest sports entities in Canada. Um, if anyone's here from Canada, uh, they have the NBA team, Toronto Raptors, Maple Leafs, which is an NHL team. And just like many other content or media organizations, the challenge is after every game, um, and there's also a huge historical archive there as well, uh, how do you actually access the footage that you own, the assets you own, and have so that you can very quickly put together stories for fans? And the demand for content is just exponentially growing, and fans want to see something that's more personalized to them. And that means you have to find a way to be able to scale the way you access your own footage library and to be able to piece these stories together, uh, so that you can distribute them in a more personalized manner.

And so, in the case of Maple Leaf Sports, they actually did something incredibly cool, which is take video understanding technology and build an agentic workflow that can do prompt-based highlight creation. So, if I wanted to create a highlight, uh, around Vince Vince Carter's 30th anniversary, I could actually, as a creative, um, in a meeting room with all of our marketing team, I could actually just type in a prompt or describe things that I have ideas of in my head. And instead of everyone leaving the meeting and coming back a week later, uh, we could actually visually see what those highlights look like in real time as we're talking about it. And so, the time it takes to create highlights, um, and sports content would take, could take from 16 hours to days. And now you're cutting that down to nine minutes. And that's not just an efficiency. It's about how do we do more as an organization. Um, and it's an incredibly innovative, uh, use case where we saw last year one of the first use cases of true, uh, agent workflows built on top of video understanding that actually allows you to scale as an organization and to grow the way that you tell stories. And it's something that we're seeing increasingly more of, uh, from not just sports organizations, but from every, could be content creators, um, more established legacy media and studios as well.

Um, I'm going to walk through on time, but here's a quick example of said technology. [Music] [Music] And so, when you describe things that you're trying to look for, uh, whoever, whether you're in a part of the organization or a partner, >> you'd be able to find that exact second or moment. And in this case, now we're taking the memory that the AI has created, and we're creating, we're reasoning on top of it. We're asking, "Create hashtags, descriptions." You can go as far as to say, "Let's create highlights and chapters for me," um, and ask very specific questions, uh, about the footage you have. Yeah, I'm going to skip over the rest.

Um, yeah, and different, we, 12 Labs is a horizontal platform, meaning the video, uh, as a foundation model, the AI can comprehend any type of footage, whether it's animation, could be sports, uh, could be your dashcam footage. Um, in this case, I'm, I'm going to just wrap up the case study on Maple Leaf Sports Entertainment and why it's so exciting. One is, of course, time savings and efficiency. How do you take the best talent you have, the creative talent you have, and be able to say, you know, save their time and be able to do things more efficiently? Um, but the more exciting thing is how do you do more of that? And so, how do you build an engine that optimizes itself? Um, how do you build a creative engine, um, you know, that you can actually scale engagement, uh, through with fans? And third is, of course, personalization. Um, how do you recreate things that really resonate with specific audiences? Um, and it's the, it's really the content machine, the engine that helps you do that. Um, so, yeah, that was a little bit about video understanding, um, and 12 Labs. Hope that was interesting. Thank you.

>> Now joining So Young on stage, please welcome back moderator Zoya Hassan.

>> Back.

>> Thank you so much for that presentation. You must be tired of all the talking. Please drink your water. Um, it's really fascinating stuff, you know, especially like, as someone non-technical at all, like it's really just fascinating all the things you can do with AI that I couldn't ever imagine. But I want to, you know, humanize the technology a little bit more and, uh, talk more about 12 Labs and your story. So, can you tell me a little bit about when you guys were founding this company, between 24 and 26 years old, all you co-founders? I want to hear, I want to hear a little bit about the founding days and what it took for you guys to get to where you are, the nitty-gritty details.

>> Yeah, absolutely. So, uh, we have five co-founders. Uh, I'm on the business development and the go-to-market side, but my co-founders actually, uh, met working together, um, as researchers and scientists building AI to be able to understand vast amounts of data from language, image. They couldn't quite tackle video because there was no technology for video understanding at the time. So, that's really where the idea came from. Um, I met our CEO and co-founder, my co-founder and CEO, Jay, in high school. So, uh, we met in the ninth grade in boarding school. Um, and we bonded. And so, before we started, we started, like I think all of us have been doing this for more than half of our adult lives now. But, um, we were incredibly good friends. We trusted each other, uh, both in terms of what they could do, the potential, but also, um, as humans, um, as well.

>> Yeah. Are you guys still good friends?

>> We're all good friends. I'm glad to hear that. Yeah.

>> Uh, 12 Labs has raised over $100 million. Was that challenging at all? Especially in the early rounds.

>> Yeah. So, as a startup, when we had first started, uh, it was, think 2020, we started building the tech. We officially incorporated in 2021. And when we had started, uh, LLMs were super nascent at the time too. I don't think, I don't know if GPT-3 was out yet. Uh, and so the concept of, "Hey, we want to build general-purpose foundation models. They are multimodal," and multimodal means AI can understand sound, visual, and video. Uh, and we want to build APIs that allow any developer or enterprise to access capabilities of these models, sounded ridiculous to investors. Um, and so, I remember like, we weren't even talking to investors at the time because we had a very strong conviction of what we wanted to do and why, which was, we just want to build, um, models that, you know, anyone can build on top of. And this was meant to be general-purpose, but that's not typically not the advice that startup founders get. Startup founders get the advice of, "Hey, find a very narrow problem, build an application, completely verticalize first, and then expand." Um, but AI is very different today. AI has, you know, horizontal, to build infrastructural technology, uh, a lot of it has to be, uh, horizontal.

>> And so, our first ever investor, uh, was Index Ventures, actually. Um, and the way they, we got their attention was our CTO at the time was getting a lot of, Aiden was, he's Forbes 30 Under 30, by the way. I didn't qualify anymore. But, um, so Aiden, uh, we got a lot of questions around, "Hey, you, you guys are like these young, uh, founders in their 20s, and you're like 12 people, who were 12 people at the time, 12 people, uh, cool stuff, but I'm sure every other hyperscaler, Microsoft, etc., are trying to do, do exactly what you're trying to do, and they have unlimited resources." And so, at the time, our CTO wanted to prove something, right? So, we had, uh, $50k in compute credits from AWS that they give to startup founders. And so, he, uh, made, made a really big bet, and it was a gamble for us at the time. We had no funding. So, he took $50k in startup credit, spent it all on training the baby, baby version of what our models are today, and, uh, that was Moringo at the time, and built, uh, AI that could search, retrieve, uh, from video. And he entered the ICCV VQA Challenge, which is one of the most prestigious, uh, conferences and competitions in the AI, like video space, uh, hosted by Microsoft. They're essentially competing against Microsoft's state-of-the-art model and won first place. And that's how, uh, our first investors, who had a strong thesis around the future of multimodal technology, that's how we first got connected. And then we also got to meet Dr. Orfee, and she was also, she also became an investor and, um, advisor to the company. And, yeah.

>> Incredible. What about landing your first customer? How did you pitch them?

>> Yeah, so, um, video is really, if you think about it, it's ubiquitous. It's in every industry. And so, if you're trying to land your first customer, uh, it means that's exciting, but it's also really challenging because it means anyone could be your customer. So, we knocked on many, many doors. Um, and I can't mention this by name, but, uh, one of our very first customers, uh, we actually had, we went to a conference just to meet him and grab him as he was walking off the stage from a panel. And he was managing all of the content for one of the largest sports leagues in the world. He was actually presenting, coincidentally, about the challenge of tagging, managing tags, and having this huge organization be able to tap into all of the archives, uh, of video assets they have, and how inefficient or challenging that was. And, uh, we had like all of their, like, we found videos from their organization on YouTube. We made a customized demo for him. As he was coming off of stage, we grabbed him, did a demo for him on our iPad. Um, and he, we were super lucky. We did many of these types of things. But when you find the right person at the right organization with the right problem, um, they will understand that you're really early as a company and as a product. And actually, we signed our first POC with him for $5,000. And really, the $5,000 was to, you know, it's to make sure we can actually go through a legal process. Um, but what we were doing was, we were learning, and we learned so much from him around the processes of media management, uh, and, uh, kind of use them as a design partner to be able to build our product and technology around them. Um, and kind of, we were at the same time, we're exploring all of these other potential customers around how do we, you know, is this applicable to other prospects and the larger market? Um, but if you can do that while you're working with, you know, the first few, uh, customers who really believe in you, not just because of what you have now, but because of the future of, like, what you're trying to create, then it becomes an incredible partnership. And, uh, we still work with them today, and they're at different, you know, organizations now. Um, but, you know, they bring you, them bring you with them as well.

>> Yeah, absolutely. So, what I heard from that, guys, is that if you would like to grab her as she's walking off stage and pitch your product, she's open to it. Um, what should this crowd know about building an AI? Especially, you know, what's the opportunity for small startups to be really successful when there's such big players in the AI world that are dominating, you know, whether that's OpenAI, that's Anthropic, like they dominate the scene, but startups are obviously necessary. And that might be collaboration, not competition, right? So, any thoughts on that?

>> Uh, you have to really understand, I think, your, your problem space really, really well. If you think about hyperscalers and how large companies work, they have so many products, and sometimes, even within their organizations, you might have a lot of silos. And for startups, the biggest strength that we still, we have, and we, we protect and we grow, we need to grow as we grow as a company, is how do you minimize and shorten the life, the communication cycle between what you're learning from market and how that gets back into product and be able to create the flywheel. Uh, and if we're talking about AI, that would be a flywheel between how do you go from market, um, to product, to data, to, to research, and back again. And, uh, video is interesting because video isn't just about the model. It's also about everything that has to do with processing video at scale, really efficiently. Um, and so for us, that was a huge moment that we had, the speed, the focus, the prioritization. Um, and that's something that, uh, large companies simply cannot do. And if you've raised money as founders, then you're spending all of that capital on solving this one problem. No other organization is going to spend, invest everything they have into solving that single problem. So, um, I think it's being able to really vocalize, like, what problem you're solving really well and just go after that specific thing. Um, because really, that's that's your, that becomes your strength.

>> No, that's really great advice. My last question for you, and arguably the most important one. Why is there an Eleven Labs and a Twelve Labs? And who came first?

>> Twelve Labs came first.

>> Okay.

>> Almost, I think, a year. Um, but we, we work, we know the Eleven Labs folks really well. Actually, a couple years ago, there was like a, something that started on Twitter. Um, same question. We get this all the time. And so, we did a hackathon at Shack 15 in SF called 23 Labs. We called it the 23 Labs Multimodal AI Hackathon. Um, we're actually doing another one in a couple weeks in New York. This one's going to be focused on advertising and like multimodal AI, um, to power like advertising technology. Um, but yeah, amazing. I know we got some audience questions, but we're unfortunately out of time. Are you open to, you know, people asking you questions when you're out in the crowd later?

>> Yes, absolutely.

>> Okay, catch it then, guys. Thank you so much, So Young.

>> Thank you, everyone.

>> Thank you all.