Skip to content
← All articles
Music News
12,320,916 YouTube tracks are already training an AI. Look yourself up.
Photo · Background: Tom Majric via Pixabay. Treatment: BCKSTG.

12,320,916 YouTube tracks are already training an AI. Look yourself up.

By BCKSTG EditorialLast reviewed:

12,320,916 YouTube tracks are public in LAION-DISCO-12M, the dataset feeding AI music models. Search yours. Read what Google argued in court.

For the last eighteen months, AI labs have been pulling music in from wherever it played. Some announced their datasets. Others let theirs surface online on their own. The rule is simple. If your music is published publicly, someone treats it as material to train a model on.

The biggest receipt sitting in plain sight is called LAION-DISCO-12M. The methodology started as a 2023 NeurIPS paper, DISCO-10M, from a team at ETH Zürich. The German nonprofit LAION ran the same recipe at larger scale and put the result on Hugging Face in November 2024. A list of 12,320,916 songs referenced by YouTube ID, with title, artist, view count, and duration for each. About 91 years of music. Roughly 750 MB of parquet files, licensed Apache 2.0.

DISCO-12M isn’t the only public music dataset in the wild. SLEEPING-DISCO 9M uses the same recipe at a smaller scale. The Atlantic’s AI Watchdog tool searches across both, plus others. If you’re not in one, you could be in another.

The dataset doesn’t contain the audio. It contains the receipts that point to the audio on YouTube. Anyone with the IDs can pull the songs and train a model on them. The distinction matters for lawyers. For the artist whose songs are in the list, it matters less. The track is already flagged.

Search The Atlantic’s AI Watchdog for your name

Who takes the hit

Independent artists. Major labels have legal departments. You don’t.

DISCO-12M was pulled from YouTube, which is also the distribution platform, discovery engine, and de facto release valve for most independent music in the first place. There isn’t really an opt-out.

The Atlantic’s AI Watchdog beat has been tracking these releases as they land, and their tool searches across multiple datasets.

What Google just argued in court

The case has a backstory worth knowing. In November 2023, Google DeepMind launched Lyria and credited four researchers as core contributors. Within months all four left Google, founded Udio, and got sued by the RIAA in June 2024. Udio settled with Universal Music Group in October 2025, agreeing to rebuild on licensed music only. Google watched the people who built its music AI get sued for exactly that, then released Lyria 3 anyway on February 18, 2026, inside the Gemini app with its 750 million monthly users.

That timeline is what the lawsuit hangs on. Per the complaint’s footnoted citations to Google’s own published research (chiefly the MusicLM paper from Google Research), at least 44 million clips and 280,000 hours of music were used to train Google’s music models, with the actual training set substantially larger.

On Monday, June 8, 2026, Google filed a motion to dismiss the class-action lawsuit a group of independent artists brought against it in the US District Court for the Northern District of Illinois (Kogon v. Google LLC, No. 1:26-cv-02582, filed March 6, 2026 by Loevy & Loevy). The suit accuses Google of using their songs without permission to train Lyria 3, its music-generation model. Music Business Worldwide reported the filing on June 10, 2026, in a story by Mandy Dalugdug.

Google’s central argument fits in one line. When the plaintiffs uploaded their songs to YouTube, they granted Google a license that covers, among other things, AI training. If the court agrees, that same argument becomes the template for anyone training a model on music uploaded to a platform with broad terms of service. Not just for DISCO-12M. For everything that comes after.

The legal discussion can take months. The material part will not. The songs are already in the list.

What parts of the fine print on platforms like YouTube should independent artists be reading the closest, so they don’t accidentally hand over permissions that cover AI training, on a service that is borderline mandatory for a music career?

A Lawyer’s Take

“The first thing any artist needs to know about these SLEEPING-DISCO and LAION-DISCO datasets is that YouTube was not involved in any way. These were non-profit research groups who assembled open-source metadata to be used in the research and development of generative AI models by downstream users.

The YouTube / Lyria 3 lawsuit raises another interesting issue. Modern artists who want to give themselves the best chance of succeeding and growing a fan base don’t feel like they have a choice. They need to be where fans are regardless of how bad the terms are. For AI-training purposes, see what each platform’s policies and terms say, and where possible, update your settings so you can opt out of training.”

Ryan Schmidt, Esq. Music attorney. Represents artists, songwriters, and producers on record deals, publishing agreements, and catalog transactions. Former recording and touring artist.

A Producer’s Take

"If your biggest fear is that somebody can type a prompt into a computer and replace your entire artistic output, your competition is not AI. Your competition is being boring."

Gino The Ghost. 6x Grammy winner. CA7RIEL & Paco Amoroso, Saweetie, Nathy Peluso.
Via @ginotheghost / ginotheghost.com. A hot take worth chewing on.

Look yourself up

Use The Atlantic’s AI Watchdog tool. It searches LAION-DISCO-12M, SLEEPING-DISCO 9M, and other public AI music training datasets in one place. Type your artist name there.

We briefly hosted our own mirror of LAION-DISCO-12M while reporting this story. A small site like ours can’t responsibly host a 793,000-artist index for the long haul, and a single-dataset search was always going to cover less ground than The Atlantic’s. Theirs is the better tool. Use it.

Before the Lyria 3 case produces any legal answer, there are three concrete moves.

One. Document. Screenshot The Atlantic’s lookup with the date. If a class action moves forward, that count is useful evidence.

Two. Ask the host to pull you. The Hugging Face dataset page carries a community report button and a contact email. Removal is handled there, not at LAION or Google.

Three. Watch the case. Whatever the court decides about the YouTube license will set the template for the next dataset, not this one.

FAQ

What is LAION-DISCO-12M?

LAION-DISCO-12M is a public dataset published on Hugging Face in November 2024 by LAION, a German nonprofit best known for compiling the image datasets behind Stable Diffusion. The dataset lists 12,320,916 songs by YouTube video ID, with title, artist name, view count, and duration recorded for each track. About 91 years of continuous audio if you played it back to back. The package weighs roughly 750 MB across parquet files, licensed under Apache 2.0 so anyone can use, modify, and redistribute it commercially without paying the artists whose songs it indexes. The dataset itself doesn’t contain audio files. It contains the YouTube IDs that point to the songs, which is enough for any AI lab with a download script to pull the audio and train a generative music model on it. The methodology started as a 2023 NeurIPS paper called DISCO-10M from researchers at ETH Zürich; LAION ran the same recipe at larger scale.

Was my music used to train AI?

Open The Atlantic’s AI Watchdog tool, type any spelling variation of your artist name that appears on YouTube, and you will see how many tracks appear across multiple AI music training datasets along with sample track titles. A match doesn’t prove a specific AI company trained their model on your songs. The datasets are publicly hosted on Hugging Face and openly used by the AI research community, but no AI lab is required to disclose which datasets it trained on. What a match does confirm: your work is in a packaged, downloadable list specifically built for AI music training.

How do I get my music removed from DISCO-12M?

Removal is handled at the dataset host, Hugging Face, not at LAION (the nonprofit that compiled it) or Google (which hosts the underlying YouTube videos). Open the LAION-DISCO-12M dataset page on Hugging Face, scroll to the community section, and submit a community report flagging the rows that contain your tracks. There is also a contact email listed on the dataset page for direct removal requests. Document your case first by screenshotting your AI Watchdog lookup result with the date visible. That timestamp is potentially useful evidence if a class-action lawsuit moves forward and you join as a plaintiff. Removal at Hugging Face stops further distribution of your YouTube IDs through that dataset, but doesn’t pull the audio off YouTube itself, and doesn’t undo any training that already happened. Other public music datasets exist; you may need to submit takedowns at each separately.

What is Google’s motion to dismiss the Kogon case actually arguing?

Google argues that when the plaintiffs uploaded their songs to YouTube, they accepted YouTube’s terms of service, and those terms granted Google a license that covers, among other things, training AI models on the uploaded content. If the court agrees, the same argument becomes the template for any company training a model on music uploaded to any platform with broad terms of service. Not just DISCO-12M, but everything that comes after. The Northern District of Illinois case is Kogon v. Google LLC, case number 1:26-cv-02582, filed March 6, 2026 by the Chicago law firm Loevy & Loevy. The motion to dismiss was filed June 8, 2026 and first reported by Music Business Worldwide on June 10. A ruling could take months. The plaintiffs include eleven named independent artists representing a proposed class of similarly situated artists.

What does this mean for independent artists who depend on YouTube?

YouTube is the distribution platform, discovery engine, and de facto release valve for most independent music. The platform’s terms of service are written by Google and updated unilaterally; uploading a track is treated as legal acceptance of whatever those terms currently say. If the court accepts Google’s reading in Kogon, opting out of AI training and staying on YouTube as a career platform start to look mutually exclusive. The choice isn’t really a choice when leaving means losing the discovery surface that fed the career in the first place. That is the actual stakes question the artists in Kogon are testing: whether a one-sided clickwrap agreement is enough legal cover for a $2 trillion company to extract a downstream commercial product from work uploaded by independent creators. The answer will set the template for the next dataset, and the one after that.

Open The Atlantic’s AI Watchdog

← Read moreBCKSTG Playback

Got a story for Playback? Send the angle our way.

Pitch a Story