📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

How AI Models Steal Creative Work — and What to Do About It | Ed Newton-Rex | TED

TED15:08

Transcription

Translator: Hani Eldalees

The technology and vision behind generative AI are amazing, but stealing the world's creators' work to build it is not. There are three key things AI companies need to build their models, and three key resources – people, compute, and data. Namely, engineers to build the models, GPUs to run the training process, and data to train the models on. AI companies spend enormous amounts on the first two, sometimes a million dollars per engineer and up to a billion dollars per model. But they expect to take the third resource, training data, for free.

Currently, many AI companies are training on creative work they haven't paid for or even asked permission to use. This is unfair and unsustainable. But if we re-architect our training data and license it, we can build a better AI ecosystem that works for everyone, both the AI companies themselves and the creators, without whose work these models wouldn't exist.

Most AI companies today do not license the majority of their training data. They use web scrapers to find as much content as possible, download it, and train on it. They are often very secretive about what they train on, but what is clear is that training on copyrighted work without a license is rampant. For example, when the Mozilla Foundation looked at 47 large language models released between 2019 and 2023, they found that 64 percent of them were trained, in part, on Common Crawl, a dataset that includes copyrighted works, such as newspaper articles from major publications. Another 21 percent did not disclose enough information to know which of the two methods. Training on copyright without a license has quickly become the standard in most of the generative AI industry.

But this training, this unlicensed training on creative work, has serious negative consequences for the people behind that work. And that's for the simple reason that generative AI competes with its own training data. This is not the narrative AI companies like to portray. We like to talk about democratization, about allowing more people to be creative. But the fact that AI competes with its own training data is inevitable. A large language model trained on short stories can create competing short stories. An AI image model trained on stock photos can create competing stock photos. An AI music model trained on music licensed for TV shows can create competing music for TV show licensing. These models, however imperfect, are so fast and easy to use that this competition is inevitable.

And this is not just theory. Generative AI is still very new, but we are already seeing the exact kind of impacts you would expect in a world where generative AI competes with its own training data. For example, the famous director Ram Gopal Varma recently said he would use AI music in all his future projects. In fact, there are multiple reports of people starting to listen to AI music instead of human-produced music, and recently, an AI song reached number 48 on the German charts. In all these cases, AI music is competing with songs it was trained on.

Or take Kelly McCreary. Kelly is an artist from Nashville. For 10 years, they made enough money selling their work that art was their full-time income. But in 2022, a dataset that included their work was used to train a popular AI image model. Their name was one of many that a large number of people used to create art in the style of specific human artists. Kelly's income dropped by 33 percent almost overnight. Illustrators around the world tell similar stories, being undercut by AI models that they have reason to believe were trained on their work.

The freelance platform Upwork wrote a white paper where they looked at the impacts they saw in the job market for generative AI. They looked at how job postings on their platform had changed since the introduction of ChatGPT, and sure enough, they found exactly what you would expect, which is that generative AI had reduced demand for freelance writing tasks by 8 percent, which rises to 18 percent if you only look at what they call lower-value tasks. So our raw data, along with the individual stories we hear, all align with the logical assumption: "Generative AI competes with the work it was trained on." It's so fast and easy to use, it's inevitable, and it competes with the people behind that work.

Creators are now arguing that this training is illegal. The legal framework of copyright gives creators the exclusive right to license copies of their works, and AI training involves making copies. Here, in the US, many AI companies argue that AI training falls under the copyright exception for fair use, which allows unlicensed copying in a limited set of circumstances, such as creating a parody of a work. Creators and rights holders strongly disagree, saying there is no way this narrow exception can be used to legitimize the mass exploitation of creative work to create automated competitors to that work.

And for the record, I completely agree. Of course, this question has not yet been tested in courts, and there are currently about 30 lawsuits ongoing brought by rights holders against AI companies, which will help address this question. But this will take some time, and creators are suffering from what they consider unfair competition right now. So they propose a solution that has been used and worked before – licensing. If a commercial entity wants to use copyrighted work, whether it's to manufacture goods or build a streaming service, it licenses that work.

AI companies now have a host of reasons why this doesn't apply to them. There's the fair use legal exception I've already mentioned. There's also the argument that because humans can train on copyrighted work without a license, AI should be allowed to do so. But this is a claim that is hard to justify. Artists have been learning from each other for centuries. When creating, you expect others to learn from you. You learn from a host of sources, from other art to textbooks to taking lessons. Much of this you or someone else paid for, supporting the entire ecosystem.

In generative AI, commercial entities valued in the millions or billions of dollars are extracting as much content as possible, often against creators' wishes, for no compensation, and making multiple copies along the way – which are subject to copyright law – to create a highly scalable competitor to what they are copying. Capable, to the point that there are AI image generators estimated to produce 2.5 million images a day and AI song generators producing 10 songs a second. To say that human learning and AI training are the same thing and should be treated the same way is preposterous.

AI companies also argue that licensing their training data would be impractical. They use so much training data that individual payments to each data creator would be minuscule. But this is true of many content licensing markets. Content creators still want to be paid, even if it's a small amount. AI companies argue that they simply use too much data for licensing to be feasible. But it's very hard to believe that in a world where there are a host of datasets you can access with permission. You can license data from media companies. There were 27 major deals between AI companies and rights holders just last year, and that's not to mention smaller, unreported deals. There are marketplaces for training data where you can get more data. You can augment this with data in the public domain – that is, that has no copyright, such as the 500 billion-word Common Corpus dataset. You can further augment this with synthetic data, that is, data generated by the AI model itself, which typically has no copyright. So there are many options available to you if you want to build your model without infringing copyright.

But the strongest evidence that it's possible to license all your data is that there are many companies that are already doing it. I know, because I've done it myself. I've worked in what we now call generative AI for over a decade, and last September, my team at Stability AI released an AI music model that was trained on licensed music. A number of other companies have done the same, and I founded Fairly Trained to highlight this fact and these companies. Fairly Trained is a nonprofit that certifies generative AI companies that do not train on copyrighted work without a license. We launched in January of this year, and we have already certified 18 companies.

Now these companies take a variety of approaches to licensing their training data. We have an AI voice model trained on licensed individual voices. We have an AI music model licensed with over 40 music catalogs. We have a large language model trained only on public domain data, mostly government documents and records. We have companies that have paid upfront fees for their data. We have companies that share their revenue with data providers. There is no single answer to the exact details of how one of these licensing deals works. The beauty of licensing is that both parties can come together and figure out what works for them. And this is happening more and more now.

You will hear that the requirement to license training data somehow stifles innovation, and that only large AI companies can afford these huge upfront licensing fees. But in reality, it's the small startups that are going to the trouble of licensing all their data, and they are doing it, often, without exorbitant upfront licensing fees, but using models like revenue shares. And there's another key aspect to licensing your training data. All this training on copyrighted work is forcing publishers to shut off access to their content. The Data Source Initiative looked at 14,000 commonly used websites in AI training datasets, and they found that over a one-year period, looking only at the highest-value domains for AI training, the number that were restricted by opt-outs or terms of service rose from three percent to between 20 and 33 percent. The web is being progressively shut down due to unlicensed training. This is bad for new AI models, for newcomers to the market, but also for everyone – researchers, consumers, and others, who benefit from an open internet.

It should come as no surprise that the general public does not agree with AI companies about what they can train their models on. One poll from the AI Policy Institute, in April, asked people about the common practice among AI companies of training on publicly available data. This is data that is openly available online, it includes a lot of copyrighted work, such as news articles, and often pirated media. 60 percent of people said that should not be allowed versus only 19 percent who said it should be allowed. The same poll went on to ask whether AI companies should compensate data providers. 74 percent said yes, and only 9 percent said no. Time and time again, when we ask the public these questions, they show support for requirements around permission and payment, and a rejection of the idea that something being publicly available somehow makes it fair game. And the people who make the art that society consumes feel the same way.

We launched today the "Statement on AI Training," which is a short, simple open letter that simply states: "The unlicensed use of creative works to train generative AI poses a significant and unfair threat to the livelihoods of the people behind those works, and should not be permitted." This has already been signed by 11,000 creators worldwide, including Nobel Prize-winning authors and Oscar-winning actors and composers. And if you agree with this sentiment, I encourage you to sign it today at aitrainingstatement.org.

What this statement and prior statements make clear is that these artists, these creators, view unlicensed training on their work by generative AI models as completely unfair and potentially catastrophic for their professions. So if you are an advocate for unlicensed AI training, just remember that the people who wrote the music you listen to and the books you read probably disagree.

So where does this leave us? Well, for now, many of the world's artists, writers, musicians, and creators outright hate generative AI. And we know, in their own words, that one of the reasons is that we train on their work without asking them. But it doesn't have to be this way. The AI industry and the creative industries can be mutually beneficial and should be. But for this mutually beneficial relationship to emerge, we have to start from a position of respecting the value of the works being trained on and the rights of the people who made them.

I am not arguing that all AI development should be stopped. I am not arguing that AI shouldn't exist. What I am arguing is that the resources used to build generative AI should be paid for. Licensing is hard work. It will slow you down in the short term, but you will eventually get to the exact same place – models with the same capability and power – and you will do so without forcing the world's publishers to tear down their gates and destroy the commons, and without alienating the world's creators against you.

So I hope more AI companies will follow the lead of those we have certified at Fairly Trained, and license all their training data. I hope employees at these companies will ask their employers to do so. And I hope everyone who uses generative AI will ask about their favorite models that they were trained on. There is a future where generative AI and human creativity can coexist, not just peacefully, but symbiotically. It's been a difficult start, but it's not too late to change course. Thank you. (Applause)