📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

The State of AI 2025: Insights from Nathan Benaich at Air Street Capital #stateofai #ai

The Research and Applied AI Summit - RAAIS25:20

Transcription

Hey everyone, I'm Nathan Benes, founder and general partner of Air Street Capital. And after a couple months of work, I'm super excited that today is a launch day for our state of AI report covering 2025.

I've been working in AI for over a decade now. Uh went to grad school doing binformatics and cancer research and I've seen the early days of AI software companies, AI research, and I can really say that this year has been a monumental achievement for the AI community.

Now, this report, if you haven't encountered it before, is really meant to be an analysis of AI research, industry, politics and safety, generative AI usage, focusing on the last 12 months. Our goal is to inform conversation around AI and its implications for the future. And for that reason, we publish it as an open access document that we really hope that you use in your everyday work. It's available now at state of.ai and seen in major publications.

And importantly, it's the effort of the community. We have reviewers from big tech companies, from academic groups, um from policy groups and startups who have really tried to keep us honest this year. There's been a couple dozen of them that have contributed. And to all of you who've published uh and achieve amazing things that we cover in the report, thank you for that.

So, in this presentation, I'm going to try and run through a couple of my favorite slides from this year's report. The overall document is about 300 and so slides, so it's pretty chunky. It's definitely grown year on year and that's really just because there's been so much happening in AI over the last 12 months.

So really one of the major stories in research is yes I know 12 months has gone by and open AI seems to be consistently producing the most intelligent models at the frontier. Now according to various benchmarks whether it's artificial analysis epoch and others open AI models like GPT5 are continually at the frontier but that gap seems to be narrowing and there's been a really fastm moving and evolving ecosystem in the open source community which we'll dig into right now and that story has really been a flipping of meta lama models which were originally extremely popular to their strategy shifting a little bit and China's open source community really stepping into the fold. And in particular, I wanted to highlight Quen, which is produced by Alibaba. And in particular, through these analysis, you can see the number of downloads that Quen models have achieved on HuggingFace and other portals has really skyrocketed. And these models are really accessible. They come in sizes that are easy for anybody to use. And as a result of that, they are kind of the default for the open source community.

Now, one other major area in research over the last 12 months has been um the evolving way that we're building RL uh into our models. And so, RL is really about how does a model learn from its experience? How can it take actions in an environment? We've gone over the last few years from very simple environments like binary outcome video games, win loss to sort of fuzzy matching environments to rubrics and now increasingly to verifiable rewards where a model can take actions in an environment and we know definitively yes that action is good or not good. And those actions can be taken across long horizons of of time. And so what what does that really bring? One of the uh major achievements in the last year has been a variety of labs uh hitting the international math olympiad gold medal performance uh using a augmented mathematics. This kind of achievement would have been magic a couple of years ago and it's really astounding how quickly progress has happened.

What's also cool is that it's not just the models that are getting smarter. Experts are getting smarter using the models. So here we have examples from Deep Minds Alpha Zero, which teaches chess grandmasters brand new concepts that have actually applied in gameplay and actually improve their game. AI agents have been an increasing part of the scientific method, which is generally all about consuming large amounts of research, formulating hypotheses, running experiments, analyzing those experiments, reframing the hypothesis, doing more experiments, and kind of completing that recursive loop. But instead of humans just doing all of that, they're increasingly involving AIS to to help them along with it as well. And um this particular work is nent, but it's borne fruit. Uh it's discovered novel gene candidates for disease, novel pathways, and um and we expect a lot more progress to come in the next 12 months.

Indeed, models that were trained originally on natural language have moved of course into other domains like in biology. And um now we see empirical evidence that the very same scaling law of nature that got us all excited about language models and text is actually also present in biological sequences. And here it's the task of protein um sequence prediction.

Another area that's really taken the research area by storm is um the physical AI narrative. So this is really about moving AI away from just a computer onto mobile robots and how can they learn the environment around them and um kind of exhibit the same reasoning characteristics that we see in really advanced language models. And so we've seen this idea of chain of thought uh where a model is explaining its sort of step-by-step reasoning to solve a task move into the robotic space in this in the concept of chain of action where a model is outputting intermediate steps along a plan before another model executes the action.

On the more developer side of uh research, uh this model context protocol that was released last year sort of became this sort of interoperable USBC for AI tools. The general idea is that a model can through the MCP connect to all sorts of data sources and use those data sources to solve increasingly complex problems. Like for example, we'd love to have agents that can search our inbox, search our drives, uh and interoperate with all sorts of other software products. And we've seen MCP which was originally authored by Anthropic really become uh one of the standards. Now of course this also exposes some interesting cyber security uh risks that are evolving and that we'll get into a bit later.

Now focused on industry the sort of big meta point here is AGI is gone. Now we're all about super intelligence. Uh every major technology executive that has been pursuing an AI agenda has sort of adopted this narrative. Um, it doesn't seem that long ago that no one really knew what AGI meant. And now I think increasingly fewer people know what super intelligence means. But it's exciting.

Indeed, the fight at the frontier is pretty ruthless. When you look at various uh benchmarks, whether it's LM Marina or artificial analysis, the turnover of top models at the frontier is very fast. But there are two that really stick out. One is uh Google DeepMinds models Gemini and the other is is OpenAI. And what's pretty cool is that we're getting more capabilities for less. There it wasn't that long ago that we used to believe these large models would consume so much uh power, so much money, so much resources that when would they ever become profitable. And I think that kind of capability to cost ratio is becoming more and more encouraging.

The other thing that's incredibly encouraging is that these companies that are making use of general purpose models are actually making incredible revenues. Here we show data from RAMP which looks at the overall adoption of AI across their customer base growing very very fast particular uptick since January this year. These AI products are increasingly sticky. The cohorts actually look pretty impressive improving year on year and the uh overall contract size is growing materially too.

Indeed, AI first startups, which are uh ones that again productize generative AI through novel web apps and consumer experiences, are seeing their revenues accelerate very quickly and indeed they actually um accelerate far faster and perform better than top cortile peers in other sectors.

A major story as you'll remember in February was this big deepseek freakout moment when the NASDAQ saw a huge draw down. Nvidia saw a huge draw down as a result of this paper which came out describing a potential reasoning system that may have been trained for much much smaller amounts than what Frontier Labs in the US were talking about. Now of course these numbers were overestimated. Indeed, they didn't describe the entire research effort. And uh what resulted after that was uh everybody realizing this in public markets and then piling back into the thesis that cheaper intelligence produces more demand which requires more chips which drive more uh usage and uh on we go. Otherwise termed the Jevans paradox.

And so it's really astounding to see just how far we've come. It was really only in 2015 and there were some experiments before that but really 2015 where we saw some evidence that more GPUs, more compute led to models converging a lot faster. And 10 years later, we're in the situation where AI is really becoming the most important, the most uh largely focused on uh scientific and engineering effort of our generation with gigawatt scale capacity data centers being built up, hundreds of billions of dollars if not trillions of dollars allocating towards creating these computational resources to produce uh intelligence that more and more of us want to consume in our day-to-day.

This has given rise to this narrative around sovereign AI where every country effectively wants to have access to the creation of intelligence. Um and we've documented all these um investments across different nation states whether it's in the EU, the UK, Canada, etc. And the sums are really gargantuan.

A major choke point that we outline in the deck is around energy and power. Um clearly the US is trying to um soften a lot of its policies to encourage the creation of data centers and the spinning up of new power sources which is um you know particularly sensitive topic in many states. Meanwhile, China is really blazing ahead. They've added far more capacity over the last few years. Their effective operating revenue margin is a lot higher and uh and their regulations are looser.

Now we dive back into the topic of Nvidia and just how powerful and important it is as a company to just AI progress in general but also to this increasingly large buildout. Uh in in the deck we look at the number of papers that make use of Nvidia every year and in particular several chipsets that the company makes. And here we show that actually H100 and H200s are increasingly being used as those clusters get developed. Meanwhile, older chips are falling off uh the curve.

One of our favorite analysis last year was just like how how important were the competitors to Nvidia and just how lucrative of of an investment would have been had you invested in these companies versus Nvidia. And here we've refreshed that analysis and effectively shown that the return on investing all the dollars that were put into competitors instead into Nvidia would have yielded an astronomically larger return. And so while there are competitors here, it's still very very clear that Nvidia is number one. And Nvidia on its own has become an increasingly important player in the capital markets stepping in to fund uh new AI companies, new labs, um but also a consortium of neoclouds that absorb its capacity and spin up cloud services specifically for uh for deep learning workloads.

On the politics side, there's been a lot of news in the US. Of course, President Trump's administration announced the action plan, which is this grand strategy, including a 100 different policies that are proposed to ensure that the US really stands at the forefront of innovation. Um that plan includes a lot of things that I won't summarize here but one of the more important ones is this concept of a US technology stack where the government wants to export all areas of the stack from hardware to software to tools abroad so that basically international community builds on it. It's actually leading in open source and really trying to revitalize that effort and rolling back a bunch of regulations that the prior administration put in place.

Now back into this AI stack. This policy is really about shifting away from broad diffusion controls and export controls and things of that nature into more of an export-led strategy. And this is around endorsing this American AI stack effectively to counter China's digital Silk Road playbook. There's been a lot of zigzagging around these export control topics as I think the administration is just trying to test out what is effective and what isn't. Nvidia's lobbying budgets has grown as a result of this because first the commerce secretary announced tougher controls and then companies were allowed to export and then they were allowed to export if they uh paid a tax to the US government and then the Chinese government stopped uh importing uh of foreign chips and we just been zigzagging ever since.

And interestingly, the US government has pursued a bunch of pretty unusual partnerships with the private sector where that's taking a 10% stake in Intel or effectively taking a 15% cut of sales on foreign chip sales from AMD and Nvidia and China. Uh the US government being given a golden share uh in US steel and you know the state of state of AI regulation is pretty patchworky. There's been over a thousand AI related bills that were introduced across states with about 10% of them actually becoming laws. And as you might imagine, AI companies are not particularly happy about this.

Another major thing is these international AI governance conferences and AI safety institute networks that were instituted after the main safety conference in the UK at Bletchley have seemingly pretty much lost steam. Um, in the last year there was a few conferences that the US didn't even show up to. And importantly, JD Vance said at the Paris AI action summit that the future of AI is not going to be won by hand ringing about safety.

So back in Europe, who's afraid of the AI act? Um, this act has been in place for many years and is starting to actually get implemented. In August, there were rules around general purpose AI codes of practice. But interestingly, the tone around enforcement of the overall act has really been watered down as it's increasingly clear that accelerating on AI is probably the most important economic um driver of growth and the US is really steaming ahead while Europe is pulling the brakes. And China of course is accelerating ahead too no matter the cost. Zinping did signal on all hands that it's really all hands on deck and told ministers to redouble their efforts on AI and this is despite huge debt levels. The CCP is allocating 10% increase in science and technology spend on AI. And if you look at how aggressively their companies have been producing open source models that are getting adopted by the community, I think the strategy is very very worth paying attention to.

Over in the valley, of course, uh as AI companies get big and scale fast, there's been this issue of acquisitions from the FTC, our outgoing head. And this really gave rise to these reverse acquires. We were all expecting a sort of trump bump to M&A as the president is very pro technology which hasn't really happened. So we've seen an acceleration of these reverse acquires where big companies will acquire the IP or the team or the data from a company and leave behind a remain co. The largest of these deals is scale AI.

On the jobs front, there's been some pretty interesting and potentially concerning action where if we look at entry- level jobs versus experienced workers, it really looks like hiring for entry level has dropped quite significantly. Whereas experienced workers still have this tacet knowledge which is important for enterprises to to hold on to.

Over in safety now, there's been a real changing of the tides here where many labs that were set up for the purposes of pursuing safety and um and stopping existential risk of AI have sort of pretty materially have shifted their u their tone and uh the current US administration also has been diluting um various safety related topics. Various companies have missed self-imposed deadlines for the purposes of actually pushing forward models and um I think this is something that is rapidly evolving and is actually still important to keep investing into.

So here if we look at just how much money has been spent on external AI safety testing, it's probably on the order of $130 million and that compares to the overall budget of of AI across all these major labs and big companies of close to hundred billion. And of course there are some structural conflicts of interest here where internala safety teams are really trying to do their best to effectively balance the race for commercialization with producing safe artifacts that you know everybody everyday people can use. Uh at the same time those are the individuals that are closest to the lab. So they have the most influence but but again they're under like tough pressures.

It's not really all about the money. many external orgs lack uh the ability to attract this talent and um we just need to be like spending a lot more time on this. This is also because the state of incidents is going up. There's been more and more issues and um and announcements of LLM misuse. They're a little bit innocuous for now, but they could become more serious later on. Indeed, there have been issues of North Korean state actors that are infiltrating other governments using language models that are of course like um pretty capable at um influencing people's behavior.

Cyber security uh capabilities have been increasing at a rapid pace. Meter research shows that AI task completion capabilities double every seven months across general domains. And one researcher even seemed to um find that those capabilities would increase every 5 months. And so these are really tasks that are relatively easy for now, but they become more and more complicated. And based on this trajectory, it looks like we might have a fairly significant cyber security issue that we need to tackle.

The other issue in safety is really these u model demonstrations which suggest fragility and alignment. Uh there are examples here where uh models fake alignment effectively. Uh examples where they uh are aware that they're in evaluation and therefore uh behave differently. Um there issues around self-preservation or behaviors that look like self-preservation and whether those are actually really deeply embedded in the model and easy to unwind or can be unwound.

But at the end of the day like there is good amount of um energy and momentum going into interpretability. There's been a novel a set of novel methods for tracing just exactly how language models work figuring out the activations which really shifts a focus from features to bundles of features that interact with one another. And you know the entire industry has talked about hallucinations. Hallucinations that makes models less reliable. This general narrative has declined over time as we've sort of shifted away from broad hallucination classification of responses and much more to the token level hallucination detection. And this is important because more detailed and granular interpretability really allows us to create much more robust systems.

And back to this topic of faking alignment. This is uh a pretty interesting and concerning uh example where um we actually see conflicting training objectives during training uh can prevent modification of the behavior. Um but you can actually revert this behavior when it's unmonitored. Um this is like an evolving topic and particular shout out to Anthropic who's been doing some exceptional work here. Uh there's other examples of fine-tuning that can unlock these cartoon villain p villain personas where uh models basically learn a broader latent concept like behave like a villain and then that surfaces across completely unrelated prompts.

But at the end of the day it's really exciting to see large labs come together and actually test evaluations on each other's models completely independently. For example, Enthropic reported that 03 looked as well or better aligned than clawed on most axes.

Now, interesting. This year, we decided to run a survey uh of AI practitioners. We uh accumulated about 1,200 responses, which I think is one of the largest uh AI usage surveys out there from practitioners. And this data, I'll present a couple of slides here, but it'll be also open access for anybody to use.

So across this data set we found that actually 95% of people out of those,200 use AI both at work and their personal lives. And pretty interestingly 76% of them pay out of their own pockets. Quite a few people are actually paying over $200. Uh about 10% of them actually and 70% of the people report that their organization's budgets for generative AI have actually grown in the past year.

Importantly, we asked about what are the barriers to scaling. And most people would say it's the upfront time to configure systems to make sure they work reliably. Answering questions around data privacy and a lack of expertise and controls and integrations. So these are all things that are effectively teething problems that can be solved over time.

We asked what were the most surprising moments that people had with AI. And uh of course the most popular one was coding. Um frequently saw as a surprise. People were pretty amazed about the complexity of applications that could be developed autonomously with AI. This aligns to this idea of long range task solving which has increased over the last 12 months. People were also really surprised about video, image and audio generation and uh you can really understand why these products are truly magical when it comes down to the most frequently used generative AI use cases in enterprise also coding content generation documentation all these tasks which are um pretty run-of-the-mill for for AI agents today. Chai GBT was certainly the most popular followed by Claude and Gemini and Perplexity. Interestingly, Deep Seek wasn't super far away from Grock.

There's a lot more predictions that I uh that I would love to put in the report. Uh but we decided to uh evaluate the ones from last year. Uh most notably, we predicted that an AI app or website produced by just one person with no coding ability would go viral. Indeed, this happened and in fact, one person who did this even sold his company for almost $und00 million. Uh we just we also predicted that Frontier AI labs would implement meaningful changes to data collection practices after uh big trials. In fact, Anthropic had a landmark $ 1.5 billion settlement with authors around how they acquired books. Uh we also predicted that open source alternatives to OpenAI's 01 would surpass them on various benchmarks and we certainly saw that in January with uh Deepsec R1.

So for the next 12 months we make a range of predictions. uh I won't run through all of them, but one is that we believe open-ended agents will make a meaningful scientific discovery. I would say Nobel Prize, but the uh the cycle of the Nobel Prize is a bit longer than 12 months. Uh last year's Alphafold Nobel Prize was probably one of the fastest in history. Uh we also think a Chinese lab will overtake a US lab uh at the frontier on a major leaderboard and uh perhaps that Trump would issue an unconstitutional executive order to ban state AI legislation as the nation really accelerates towards super intelligence.

Now if you enjoyed this I really encourage you to follow all of our other writing which we produce on Airhe press. It's a range of news, technical analysis, research, policy, uh, and every, uh, month analysis of all these topics in the form of a newsletter called Guide to AI. I also encourage you to come to our various global events where we focus on bringing best practices from the very best uh, individuals in the space. We're really about creating serendipity and the most connections to help accelerate your career, spark new ideas, and um, and help, you know, everybody generate a positive future with AI. On an annual basis, we run our uh research and applied AI conference in London in June and a variety of meetups uh across the US and Europe.

And so with that, it's been a real pleasure to take a couple minutes to run through this year's data report. I'm Nathan Bes, founder of Airirstream Capital. And uh please check it out at state of.ai and let us know what you think. Thanks.