Industry update

Your Music May Be in These AI Training Datasets. Here Is How to Check.

The Atlantic published four searchable databases covering 21 million tracks used to train AI music models. Any artist can search right now to see if their recordings were scraped without permission.

Bradley J Simons
Bradley J Simons
Updated July 15, 2026
Editorial review due September 15, 2026

Short answer

The Atlantic's AI Watchdog lets artists search four music datasets covering about 21.2 million tracks shared among AI developers. Search your artist name, titles, and ISRCs, save any matches, and verify the underlying dataset. A match shows availability to developers, not proof that a specific company trained on the recording.

What happened?

The Atlantic reporter Alex Reisner published an investigation on June 14, 2026 and made four music datasets searchable for artists. Two of the datasets each hold about 100,000 recordings. The other two are far larger: one contains approximately 9 million tracks, the other roughly 12 million.

Scale of the four databases (tracks identified, June 2026)
Dataset A12 million tracks
Dataset B9 million tracks
Dataset C~100,000 tracks
Dataset D~100,000 tracks

The databases include recordings from major artists alongside tens of thousands of lesser-known independent musicians. The Atlantic found that three are distributed as lists of links to music on YouTube or Spotify. The Free Music Archive collection includes audio and metadata gathered from Creative Commons releases.

Why this matters for independent artists

Before this database existed, an artist who suspected their music had been scraped had no practical way to check. Suspicion is not evidence. What The Atlantic tool gives you is a specific, searchable record of which tracks appeared in four datasets available to developers.

That is a different thing from proof of infringement. The major labels and the Recording Academy have been building their copyright cases against Suno and Udio (filed June 2024) without this kind of direct lookup tool. Independent artists behind the class action against Google over Lyria 3 filed in March 2026 face the same challenge: showing that a specific work was used, not just that large amounts of copyrighted music were swept in. A database entry is a starting point for that argument, not a concluded case.

The Atlantic also notes that a dataset match does not show which developer downloaded it or which tracks a developer selected for a model run. For most independent artists, the practical value right now is preserving a verifiable record that the work appeared.

How to check if your music is in the databases

Search The Atlantic tool

The Atlantic published its searchable databases alongside the investigation. You can search by artist name, song title, or ISRC. No account is required. If you find your work, note the dataset name and the specific track information so you have a record.

What to do with a result

Finding your recording in the database does not mean you should immediately take legal action. It means you have evidence worth preserving. If you are already in contact with a music rights attorney or a collective action related to AI training, share what you found. If not, keep a record of what appeared and check back as the legal landscape develops.

Your ISRC is the most precise search term. If you do not know the ISRC for a recording, your distributor or publishing administrator should have it on file, and it should appear on any streaming platform listing for the track.

What is still unclear?

Open questions

The databases show what was available and shared among AI developers, not necessarily what was ingested and used for training. AI companies can argue their model was not trained on a specific track even if that track appears in a shared dataset. How courts weigh that distinction is still being worked out in the active cases against Suno, Udio, and Google. There is also the question of what legal options are available to independent artists who are not part of the existing class actions and whose music appears in the databases. That is not settled. What is clear is that the tool gives you a factual starting point that did not exist a week ago.

Sources

Frequently asked questions

How can I check whether my music is in an AI training dataset?

Search The Atlantic's AI Watchdog with your artist name, track titles, and ISRCs. Try alternate spellings and collaborators, then save the dataset name and matched record for anything you find.

Does a dataset match prove an AI company trained on my song?

No. A match shows that the recording or a link to it appeared in a dataset available to developers. It does not prove that a named company downloaded that dataset or included the track in a model run.

Do these music datasets contain audio files?

The structure varies. LAION-DISCO-12M says it contains public YouTube links and metadata, while the Free Music Archive dataset includes downloadable audio and metadata under Creative Commons licenses.

What should an artist save after finding a match?

Save the search result, dataset name, track identifiers, URL, access date, and screenshots. Keep the original recording, release metadata, registration records, and ownership documents with that evidence.

Related Velveteen guides

Related Velveteen tools

Artist-facing updates

Get music industry updates without the noise

Short notes on platform changes, royalty issues, and release marketing moves that actually affect independent artists.

Improve this page

Was this useful? Send a signal or flag a correction.