Your Music May Be in These AI Training Datasets. Here Is How to Check.
Reviewed August 25, 2026: The Atlantic's searchable music-dataset investigation remains available. A match documents dataset inclusion, not proof that a company trained on the recording.
Short answer
Reviewed August 27, 2026: The Atlantic’s AI Watchdog remains a reference for four music datasets covering about 21.2 million tracks. Search artist names, titles, and ISRCs, preserve matched records, and verify the named dataset. Inclusion shows that material appeared in a dataset, not that a particular company trained on it.
What happened?
The Atlantic reporter Alex Reisner published an investigation on June 14, 2026 and made four music datasets searchable for artists. Two of the datasets each hold about 100,000 recordings. The other two are far larger: one contains approximately 9 million tracks, the other roughly 12 million.
The databases include recordings from major artists alongside tens of thousands of lesser-known independent musicians. The Atlantic found that three are distributed as lists of links to music on YouTube or Spotify. The Free Music Archive collection includes audio and metadata gathered from Creative Commons releases.
Reviewed August 25, 2026: the searchable investigation and LAION dataset description remain available, and no source reviewed for this update changes the core limitation. A result documents inclusion in a dataset; it does not prove a particular model used it.
Why this matters for independent artists
Before this database existed, an artist who suspected their music had been scraped had no practical way to check. Suspicion is not evidence. What The Atlantic tool gives you is a specific, searchable record of which tracks appeared in four datasets available to developers.
That is a different thing from proof of infringement. The major labels and the Recording Academy have been building their copyright cases against Suno and Udio (filed June 2024) without this kind of direct lookup tool. Independent artists behind the class action against Google over Lyria 3 filed in March 2026 face the same challenge: showing that a specific work was used, not just that large amounts of copyrighted music were swept in. A database entry is a starting point for that argument, not a concluded case.
The Atlantic also notes that a dataset match does not show which developer downloaded it or which tracks a developer selected for a model run. For most independent artists, the practical value right now is preserving a verifiable record that the work appeared.
How to check if your music is in the databases
Search The Atlantic tool
The Atlantic published its searchable databases alongside the investigation. You can search by artist name, song title, or ISRC. No account is required. If you find your work, note the dataset name and the specific track information so you have a record.
What to do with a result
Finding your recording in the database does not mean you should immediately take legal action. It means you have evidence worth preserving. If you are already in contact with a music rights attorney or a collective action related to AI training, share what you found. If not, keep a record of what appeared and check back as the legal landscape develops.
Your ISRC is the most precise search term. If you do not know the ISRC for a recording, your distributor or publishing administrator should have it on file, and it should appear on any streaming platform listing for the track.
What is still unclear?
Open questions
The databases show what was available and shared among AI developers, not necessarily what was ingested and used for training. AI companies can argue their model was not trained on a specific track even if that track appears in a shared dataset. How courts weigh that distinction is still being worked out in the active cases against Suno, Udio, and Google. There is also the question of what legal options are available to independent artists who are not part of the existing class actions and whose music appears in the databases. That is not settled. What is clear is that the tool gives you a factual starting point that did not exist a week ago.
Sources
Frequently asked questions
How can I check whether my music is in an AI training dataset?
Search The Atlantic's AI Watchdog with your artist name, track titles, and ISRCs. Try alternate spellings and collaborators, then save the dataset name and matched record for anything you find.
Does a dataset match prove an AI company trained on my song?
No. A match shows that the recording or a link to it appeared in a dataset available to developers. It does not prove that a named company downloaded that dataset or included the track in a model run.
Do these music datasets contain audio files?
The structure varies. LAION-DISCO-12M says it contains public YouTube links and metadata, while the Free Music Archive dataset includes downloadable audio and metadata under Creative Commons licenses.
What should an artist save after finding a match?
Save the search result, dataset name, track identifiers, URL, access date, and screenshots. Keep the original recording, release metadata, registration records, and ownership documents with that evidence.
Related Velveteen guides
Related Velveteen tools
Get music industry updates without the noise
Short notes on platform changes, royalty issues, and release marketing moves that actually affect independent artists.
Was this useful? Send a signal or flag a correction.