Transcription
Open source AI is booming, according to Hugging Face CEO Clem Dong. He's seen the same story play out again and again. Companies start out on Frontier APIs, but as they scale, the cost pushes them towards open-source models.
So today, instead of our usual Friday news roundup, we're turning things over to Rebecca Balon, who talked to Clem about why the open versus closed fight matters so much to the AI industry and why he's worried about the possibility that a handful of big companies could end up controlling everything. Stick [music] around.
Welcome back to Equity, TechCrunch's flagship podcast about the business of startups. I'm Rebecca Balon, and today we're joined by Clem Dong, the co-founder and CEO of Hugging Face. Clem, welcome to the show.
>> Thanks for having me. Super happy to be here.
>> Yeah, I'm really excited to talk to you. I mean, you've had a pretty incredible career. You're almost like you're you've got a little bad boy status, right? You like turned down opportunities at Google and eBay earlier in your career, you know, trying to trying to stick to startups and open tech and then, you know, founded Hugging Face and it's been it's taken off, right? You've raised nearly $400 million and I'm kind of curious, you know, I brought you on today because there's been so much talk about open uh open source again, right? Like it's just feels like there's these peaks where open source comes into everyone in the mainstream's attention. So when you look at Hugging Face today, I'm curious like what is a data point that best captures how much open source AI um has changed in the last year?
>> Yeah, there are a couple of data points. First, kind of like the the the volume and the number of models and data sets that are being shared on the platform is is quite uh astonishing. I think now there's a new repository created every 7 seconds on the platform. So that's almost three million public models, one million public data sets that have been shared on the on the platform. So it shows a little bit a different picture than the, you know, one model to rules them all and everyone talking just about one model. I think in reality, we're seeing more and more companies using a lot of different models with a lot of specialized customized models for for their need for their specific use cases. Uh, I think we've seen also uh in terms of adoption by enterprises. Now we have uh half the Fortune 500 uh using Hugging Face and so using using either their their own private models or or open source models. Uh so we've really been uh seeing kind of like a trend in the past past few months as more people get interested in uh open source models.
>> So so you're saying enterprises are using it. Are they actually deploying open models in production or are they still experimenting? Like how far away from the mainstream?
>> Yeah, a lot of them are are deploying them uh in production. I would even argue that the the typical kind of like flow is that companies start by using frontier APIs maybe at the beginning, you know, to experiment to launch a new feature and then when they really hit production and they hit scale, the cost is starting to be too big really with with frontier models. So that's usually when they uh switch to uh to open source models to to power their their workload. So and I and I suspect this trend will will continue. Maybe in a few years, kind of like uh frontier models will be for experimenting and some really kind of like um high value tasks and then most of the production workloads will actually be powered either by private models within companies or or by open source models.
>> Mmm, yeah, that would be pretty devastating for companies like OpenAnthropic which are desperately trying to increase their the amount of um like usership that they have right.
>> Yes and no, because uh I mean I I still think they can build very very valuable companies by, you know, being kind of like at the frontier level for example for for reasoning or or for some some tasks. Uh, ultimately kind of like AI will be so big that I think there's a world where both an OpenAI and Anthropic are valuable companies and also most of the workloads are are powered by uh other types of of models.
>> So I'm not uh I'm not too worried for them and, you know, uh they're probably going to be the most most valuable companies like somewhere by the end of this year or or next year. Yeah. Yeah. So they they're going to be fine. I'm I'm not I'm not worried about them.
>> Yeah. There seems to be this kind of perfect storm of factors that's bringing open source back into the public attention, right? I mean, you mentioned tokens, right? We recently saw Alex Karp, Palantir's Alex, Alex Karp, ranting about how uh token infrastructure being used by big labs like OpenAI, Anthropic um is, you know, criminal and he's pushing for more open models. Um, and then so we've got the token issue and then on the other side, we've got again the conversation around the Chinese models catching up and then within all of that, there was this absolute cluster of uh the Trump administration limiting the release of private AI models. What stands out to you as kind of the thing that's pushing it the most?
>> Well, I think the point of Alex Karp that, you know, companies want more control basically and and transparency in the systems they use is the one that resonates the most with me because obviously it's something that we've been saying for for a while, which kind of really makes sense. Like, you know, if you if you're an AI company or technology company, you don't want to outsource your core capabilities, AI to another they are kind of like a black box API that that you don't control, that you don't have any visibility on, that you don't really kind of like uh have any any sort of ownership. Uh, so this kind of like idea that companies need to own AI and own models instead of renting them and outsourcing them to to someone else. Uh, is is cap like the thing that uh I hear the most from uh from companies and and customers and and community members these days, which uh which makes a lot of sense in my opinion.
>> I mean, not to mention like if you own your own model, you are not at risk of if the government decides, we're going to shut this down for safety reasons, that you're kind of left in the lurch there, right?
>> Yeah. Yeah. Yeah. It sounds like a more sustainable way of of building building the technology. By the way, uh much closer to what we've seen with software, right? Like uh I mean that's that's how software has always been built with like everyone being able to write their own code uh and build their own software stack instead of kind of like delegating that or outsourcing that to to other companies. But if you consider that AI is kind of like the next generation of software or software 2.0, I think this approach makes much more sense.
>> Does that require like is there enough talent of people that are able to implement these um models into businesses? I mean, larger enterprises sure, but I mean across the board.
>> I think so because uh I mean what we're seeing at Hugging Face is that we had a few hundred thousand users like uh like three years ago, three four years ago uh and now we have 16, 17 million AI builders using using the platform. What we've seen is that a lot of the software engineers, especially with uh agents, are now able and it's becoming easy for them to uh train models, run models themselves, optimize models themselves. So I think uh with with kind of like agents, we're seeing that it's becoming easier and easier for software engineers to uh run their own models, optimize their own models, train their own models. And we're seeing that like across the board, not only from, you know, smaller startups or like very kind of like AI native startups, but also uh in enterprises.
>> Interesting. And so are they are they just kind of using like h how how would uh a startup, let's say if they want to work on their own model, like how would they use Hugging Face to do that? Because I know that you're I feel like Hugging Face is in an interesting place right now. It's like you're part GitHub for AI, but you're also kind of like becoming a little bit AWS in terms of services, right? So, how how how might they use um yeah,
>> Hugging Face?
>> Yeah, it's it's a very kind of like a modular platform. So, it depends a lot on your on your skills, on the structure of your team, on your goals and things like that. But most users really start from kind of like an open-source based model, right? So, maybe they're going to use like a Llama 2 for for a generative workloads, open AI, open GPT uh for for some other task, and NVIDIA, NeMo, any any sort of model uh and kind of like deploy them directly on their own infrastructure uh and basically start running workloads. That's usually kind of like the the starting point and then progressively you see teams kind of like uh wanting to do some optimization, particularly for example if they have uh compute uh constraints or if they have cost constraints or speed uh con constraints. Um, so they they can start kind of like optimizing these models, optimizing these weights for their specific use cases and then they'll at some point start to post-train these models to basically be more accurate for their specific use cases. Um, and that's kind of like usually uh usually the process. You start from really off-the-shelf solutions and then you end up by really uh controlling a lot of the workloads yourself and building a lot of the systems yourself, which creates actually your differentiation from other companies and other organizations, right? Like you you want to build these skills of like building AI systems better than your competitor, and that's what's going to differentiate you in the long in the long run.
>> Yeah, that's a good point. Now my brain is is stuck on um one of the models you mentioned. You mentioned Llama um 2, which to take it back to kind of the the geopolitics of it all. Um, so this is one of the Chinese models that's been getting a lot of attention for its amazing agentic capabilities recently. And Hugging Face's own spring 2026 report says that Chinese models accounted for most of the downloads, right? Like 41%. So, China's surpassing um the US monthly uh and in overall downloads. What are some other findings about how Chinese models are doing on the platform versus US models and what do you think this says about the state of open source right now?
>> In my opinion, this is a very big challenge because in an ideal world, I think we would want more of the open-source models, especially that are used in the US, to be shared by American companies instead of Chinese Chinese companies. Um, so I know that a lot of organizations in the US are working toward this goal, right? Like you have Nvidia that has I've I've told has become like the the king of American open source AI lately by sharing a lot of very interesting data sets, a lot of very powerful models themselves, like uh like NeMo. Uh, and there are a lot of startups like RC Reflection. >> There are more and more organizations, I feel like in the US that are sharing in open source, but we need uh we need much more because if you if you think of it, kind of like open source is kind of like the foundational stack for the rest of the AI stack. Um, and I think uh you you'd want kind of like every country to have kind of like some sort of sovereignty on on each uh parts of the stack. So I think that would be uh much much better to have a world where a lot of the open source used in the US is actually created by American organizations.
>> Yeah. Well, are you are you seeing like a lot of American organizations using Chinese open source models? Like is there not, you know, some kind of a stigma against that?
>> No. No. Uh, we're seeing a lot of them using uh Chinese open source, right? Some of them famously shared about it, right? Cursor uh talked about how uh their models were built uh on top of Chinese open source. You have Brian Chesky from from Airbnb that has been very vocal about uh about uh open-source AI. Uh, the majority of the scale-ups in the US that are using open source are now using open open weights from from China. Uh, also uh we don't talk about them a lot anymore, but uh all academia. So if you go to Stanford, if you go to Harvard, um, because the only way to really learn, study, do research on AI is to have uh open source and open weights, right? You can't really study an API because it's a complete black box. >> Right? >> And so all the all academia, all the research community is basically um using Chinese Chinese open source.
>> What what's the risk of that though? Like why is it why does it matter? Maybe they're just better at open source, right? Like what's the what's the problem?
>> The So there are a couple of challenges. Um, you know, uh the main one I would say is that open source is kind of like both the foundation and a very uh strong accelerating factor for AI in general. Uh, in my opinion, the reason why the US is ahead now is because from 2016 to 2022, 2023, the US was super open with open research, open source AI, everyone collaborating and sharing with each other, right? The famous example is the T5 GPT came out from Google sharing in open source transformers, right? And so, um, open source creates in a way the conditions for your AI leadership. Sure.
>> And so almost automatically uh if China keeps leading in open source, keeps sharing all this research openly in China, it creates kind of like this accelerating development of the field. And I wouldn't be surprised if as a result, China starts to lead AI in general. Uh, probably next year or or the year after.
>> Well, I mean, what would you say to people? What would you say to people who argue that China's open source is only improving so well because they're really good at distillation attacks and, you know, copying the homework of closed frontier models?
>> I I would say it's very reductive uh and very simplistic, because Chinese China has some of the best uh developers and researchers in AI. Now, um, we've we know distillation to be a very, very small factor in the ability to create good good models. It's a practice that everyone is doing, including people, including companies in the US. So if it was as easy just to do distillation to to get good at building AI models, there would be many other countries and including in the US where we would be much better in open source AI. Uh, the reality is that um, they have really, really good research teams in China um doing really well and taking a a much more open and collaborative approach to AI than in the US. And that's that's why they're they're successful.
>> What what about like the risks? Like is it riskier? Are open models riskier because they're harder to control? Right? Like obviously Trump's the Trump administration had limited the release of Claude's um sorry Anthropic's Mythos and Fable and then also OpenAI's um GPT 5.6 um due to cybersecurity concerns etc. With open models, you know, you're seeing open models catch up and as a result, there's a lot more cybersecurity attacks because they kind of, you know, while the models aren't Mythos level or Fable level, uh, necessarily, they are good enough and they have fewer guardrails to stop them from um carrying out these kinds of attacks. So, how do you balance, you know, safety with access?
>> Yeah. So historically, open source has always been less dangerous than kind of like some closed source secret kind of like behind closed doors initiatives. Uh, and the reason why is because it's more transparent. Uh, and so it's easier to understand the capabilities of it and to create kind of like mitigations for them. For example, for defenders to patch the cybersecurity risks uh that they know open source models can can do. The reality is that these risks already exist by the fact that models uh just are here uh in the frontier labs right, and already distributed to organizations through APIs. The reality also is that uh guardrails or APIs are very shallow and quite quite ineffective. I think that's that's something that we've seen in the past past few weeks. Like you can put some guardrails uh and have a feeling that it's safe, but the reality is that it's very, very easy to jailbreak them. It's uh, you know, possible to steal the weights. Uh, and if the these kinds of capabilities start to uh happen and to be possible in one lab, it's very probable that uh other labs will be able to also replicate the same thing, you know, not only in the US but but in China.
>> So is the argument less that there should be guardrails or?
>> I I think the argument is that you don't really make it safe by keeping it behind closed doors for just a few players. You actually make it more dangerous because you create asymmetry of power and asymmetry of capabilities between some actors and uh that have access and can't use can steal can kind of like use in a malicious way these weights and other people who can't defend themselves. The way you make the world safer in my opinion is by by leveling up the playing fields uh and really kind of like uh creating transparency on these models and giving them both for like the attackers and the defenders and obviously making it harder and illegal to to do the attacks, right? Like it's it's the same as if you take like other pieces of of of technology. Um, you know, uh you you don't really kind of like prevent risks by making it like uh illegal to share some some material to create kind of like dangerous things. You you make it illegal to kind of like uh use these kind of like tools.
>> Right? With this approach, you uh enable kind of like innovation, you enable competition, you enable job creation, you don't create monopolies, right? Uh, in in my opinion, the biggest risk in AI is concentration of power.
>> Right? >> We we mentioned, you know, some of the AI companies becoming the most uh valuable companies in the world uh very, very soon. In my opinion, they're also becoming the most powerful companies in the world. You've seen that for example with the their interaction with the Department of War, right? If if you would have told me a few months ago that an AI company could be in a situation of power compared to the American Department of War, I would have told you that's crazy, but it's actually what's what's happening. And so one of the, in my opinion, one of the biggest risks in AI is that you end up in a world where a few companies are completely dominating AI, getting to an amount of power and an amount of wealth that you've never seen before. It's basically similar to if there was just one or two companies being able to do software. Um, and in my opinion, this is this is the real dangerous scary scenario. And if you don't do anything to fight that, >> if you actually just enable a few companies to build frontier AI and if you kind of like create the the regulatory environment that allows them to do it and no one else, uh, you you end up in my opinion in in a very, very scary and dangerous world.
>> Yeah. So, so do you think that there's um moves that the US government should be making to um, you know, I think that the Trump's AI plan kind of hand-waved at open source, but I don't really know that much has been done. Is there anything like in particular that you would point out as something that you'd like to see um the US government do to support open open source development?
>> Yeah, public support um is kind of like very important right now because uh yeah, for for some reason uh in the US every time we talk about open source now, we we talk about how it's unsafe, which is is kind of weird uh and it creates this weird counter incentive for anyone to actually do open source versus what it used to be where open source was and open research was really celebrated as something that, you know, companies doing this were actually putting, you know, collective contributions ahead of profit maximizing. You know, like I I remember 10 years ago, if you had a company that was sharing their research, sharing open source, people would be cheering for it, which in my opinion is the right thing to do versus right now when the company is doing that, people are questioning it and be like, oh, but is it's safe. Should they should they be doing that? So I think if the the US administration can continue to really uh uh show support for open source, contribute also, like we have uh interestingly, we've had a lot of uh uh American organizations, American public organizations uh contributing to open source AI. For example, a few days ago, the I think it's called the National Design Agency uh released an open-source model for PII detection on on Hugging Face. Um, so, uh, public organizations can contribute a lot to open data for for open source, which is still kind of like a big uh big bottleneck. So these are are some of the things that I think uh could be interesting for them to do to to support and uh and foster open science and open source AI more in the US.
>> Yes. Yeah. I think uh open source maybe needs a little bit of a of a rebrand or makeover for for certain people. But you know, there is a risk when when you have open data sets and sharing open data sets, right? Like uh Hugging Face yourself, you're embroiled in a recent lawsuit, right? Evox Productions, don't know who they are. They're suing um Hugging Face alongside Stability and Runway, claiming that for Hugging Face specifically, I think the the claim is about you guys hosting data sets that have copyrighted images. I'm I'm curious, you know, that that lawsuit's ongoing, so you probably can't comment on it too much, but how do you think about legal risk as a platform that's hosting other people's other companies models and data sets? Does this change how you vet or moderate content at all?
>> So of course uh, you know, we we follow all all regulation and and kind of like follow all the the rules, all the legal kind of like uh um things that that we need to to follow as a as a platform. It's been important for us uh right from the start and we actually did a lot of initiatives to kind of like give more legal clarity to to the field. For example, a few years ago, we introduced a new type of license that uh that allows open source models, open weight models to have kind of like more um more clarity on the kind of use cases they can they can be used for. You know, again, I think the the challenge and the tradeoff is uh between kind of like uh doing it in in private and doing it in in public and how how different this is. Uh, the reality is that we know now that a lot of the closed source uh labs have been actually, you know, basically using the whole web uh without any sort of kind of like copyright.
>> Yeah, I guess with open source, you don't need a subpoena, right? You can you can get it. Well, I mean, with open source, not open weight, a lot of the time open weight models, you can't actually see the data sets. So >> yeah, so making them public instead of like um, you know, keeping it private behind closed doors uh gives more attribution. It it allows people to actually know what is used and what is not being used. Um, and and also it's kind of like uh used in a different way. Like uh if you look at how fair use has been built has been designed for for copyright, there's always this balance between obviously, you know, giving attribution and protecting creation, but also allowing innovation to to continue, giving tools for people to uh to learn and to get education, right? So, for example, you can use copyrighted material to if you're a teacher, right? To to teach uh kids, right? Which which makes a lot of sense, right? You don't want to make kind of like teachers having having to pay for everything they they use or they teach uh if it's for like the public public good. And you see the same thing for for open source, right? If some data and some data sets are shared in open source for everyone to use for free, >> you know, not for profit. >> Yeah. >> Um, this is very different than, you know, a lab using that to make billions of dollars of revenue without any public contributions.
>> This is America, Clem. What are you talking about? [laughter] [gasps]
>> All right. Well, look, okay, just to quickly pivot because I think we can we can go into the the benefits of open source all day.
>> I can talk about all about all of that for hours.
>> Yeah. And I could listen for hours, honestly. Um, but so one thing I want to ask you about since like you're clearly like I know I made a joke, this is America, Clem, but like, you know, speaking of the way American companies run things versus the way you're doing things at Hugging Face. One thing that I thought was really interesting was yeah, you've raised 400 million um, but not like for three years, right? You haven't done a round in three years and I think you also turned down a huge investment from Nvidia last year, right? So I'm c like how are you thinking about fundraising in this environment? Like you have become such a huge part of the AI infrastructure at the moment, yet you're not following the Silicon Valley, you know, fundraise at all costs rules.
>> Yeah, we've always taken a little bit of a original unique unique approach to things. We feel like we're building kind of like a a platform for a community uh and they're trusting us with kind of like uh sharing their their data and their models on the platform. So we have some sort of kind of like a long-term responsibility to to them. So we've always taken kind of like a bit of a conservative approach to things and not necessarily kind of like maximizing short-term revenue but instead kind of like focusing on on long-term sustainability. Compared to most AI startups uh and companies uh we're quite capital efficient uh in the sense that that we don't need, you know, billions of dollars of compute to to run. Um, we kind of like uh close to profitability. We we just recently started to touch the money that we raised three years ago. So we we're in like a in a position where we optimize more for kind of like long-term sustainability of the company than kind of like a short-term, you know, profit or fundraising u maximization. Uh, and and we we're pretty happy to to be in this this position and I think it's quite quite aligned with what we're building, right? We like become kind of like the the storage and collaboration platform for AI builders, right? We have kind of like strong network effects on on the platform. Uh, but also we are a platform, so we have to create kind of like uh, you know, like 100 times more value than uh we would be creating if we weren't a platform and and capture like one person, 2% of of this value. So these these things also take take time and take kind of like a long-term long-term approach to to it, similar to kind of like a social network if you if you think of it.
>> Yeah.
>> But uh but also I think in the long run, they're um quite unique and uh quite quite interesting. I think we're ending up with this approach on a on a a quite strategic and interesting position in the field, right? Like uh we're not in a very highly competitive position. We're more kind of like in a unique position where we can keep kind of like creating value for for the community and for AI builders.
>> Now when I'm thinking about value and, you know, capital and where it's flowing, right? You you have kind of a bird's eye view over I guess like where capital is flowing, but also what are some underrepresented opportunities? Like where where on Hugging Face, like what kinds of data sets um are forming in certain industries that you're not seeing capital going to at the same rate that they're, you know, joining Hugging Face.
>> Yeah. This is very very big disparity. That's why a few few months ago someone asked me if they if we were in a AI bubble and I answered that we were probably in a LLM API bubble, but definitely not in an AI bubble because there are a lot of domains, topics that are underinvested. Um, for example, uh, local AI, right? Like the ability to actually run AI on your phone, on your laptop, on your own data center rather than kind of like running it on the cloud. Uh, you know, there aren't a lot of like companies investments. Uh, there.
>> Is that a hardware issue?
>> Well, I think it's a lot of it is kind of like an investor's kind of like mimetic behavior issue where uh, you know, investments.
>> Yeah, a lot of the investments kind of like focus on the few of the very hot and and kind of like common topics that everyone is is talking about. Another one is obviously, you know, biology, chemistry, uh, all these domains have seen very, very little investment compared to uh uh text LLM APIs in the past past few years. Um, and so there there are a lot to to do that too. And then of course, there's robotics, right? Like you have Hugging Face has Reachy. I think I see one behind you. Is that
>> Yeah, I have a couple of them behind me.
>> Is that Reachy Mini?
>> Yeah, it's a couple of like the first first iterations of uh of Reachy Mini.
>> So okay. So then is robotics where open source has a big advantage because like no one company can collect all the physical data?
>> Uh, there are a couple, yeah, differences between kind of like robotics and the rest of AI as you mentioned. I think data is not going to be only just more important, but also more difficult. For example, when we look at the robotics data sets on Hugging Face, they are huge. Uh, they're really massive just because, you know, video, image data sets are really much much harder to work on than text data sets. Uh, we we're starting to to talk to people who are hosting on Hugging Face, you know, petabytes sized, you know, data sets for for robotics. Uh, the the second important thing also is I think the the trust uh issue with robotics. You know, I have a couple of like Reachy Minis at home. Uh, and I have like babies, right? And so when when I think about kind of like having a robot at home that interacts with my environment, interacts with my family, interacts with my privacy, I think it's even scarier than for the rest of AI to have like a black box system just controlled by a few organizations. Especially if these organizations' CEO is is kind of like not the most stable >> person in the world. And so I think for robotics, even more for the rest of AI, you you need more transparency. You need more open source to have a lot of different companies competing. You need to understand how it's working, why it's working like that.
>> Yeah, that's a really good point. I hadn't thought of that, right? Like you're going to have a robot in your house. You're like, "Okay, well, I'd like to know what's going on underneath the hood. I'd like to know what you're collecting about me. Like when you're are you actually off when I say you're off, etc." Um, so open source is probably a less scary, you know, version of
>> And you want choices, you know, it's it's the same uh it's the same for AI in general. Like like a world where you have one or two choices is is a very scary world because you give up some of your agency, you give up some of your ability to kind of like uh decide and and reward different things. So you want you want open source for competition to kind of like really empower not just one or two companies, but hundreds, thousands, tens of thousands of of companies to be able to build different things and and give people choices.
>> Ah, I feel like I want to have a whole second episode with you about robotics. So maybe you can join us again sometime. But in the meantime, we have gone over, we are out of time. Um, Clem, where can our listeners connect with you online?
>> Uh, Twitter or LinkedIn, you know, it's like usually the best the best way to uh follow a little bit what what we're doing. Uh, and and people can reach out to me there too.
>> Okay, great. Well, thank you so much for joining. This has been great. Uh, to our listeners, you can find me on Twitter and LinkedIn as well. Uh, you can find Equity at EquityPod on X and Threads. Talk to you next time. Equity is hosted by TechCrunch Senior Reporters and produced by Terresa Locolo with editing by Cal. Subscribe on YouTube or wherever you [music] get your podcasts and find out what's next at techcrunch.com/events. Thanks so much for listening and we'll talk to you next time.