Research · The Verge ·

The Atlantic reveals music tracks used in AI training datasets

The Atlantic released a searchable database revealing the vast scale of unlicensed music used to train AI models, exposing gaps between what artists licensed and what AI companies actually used.

Based on reporting by The Verge — analysis by dalili

The Atlantic has created a searchable database showing millions of music tracks that were freely available in datasets used to train AI models—despite not being licensed for that purpose. The investigation highlights a critical gap in how AI training data was sourced: while many datasets were technically public, the artists behind them never consented to their use in AI training.

The scale is staggering. Billions of tracks from Spotify, YouTube, and other sources ended up in training datasets with no compensation to artists. The database lets users search by artist name to see if their work was included without permission, uncovering a systemic issue in how AI companies accessed training data.

For the music industry, the findings reinforce calls for clearer licensing frameworks around AI training. The publication marks a shift toward transparency—showing the gap between how AI is imagined (trained on ethical, licensed data) and how it was actually built (with whatever data was cheapest and most available).

Key takeaways

  • Millions of unlicensed music tracks used in AI training
  • Artists uncompensated for intellectual property
  • Highlights data transparency gap in AI industry

Why it matters

AI training transparency is crucial as legal battles over data rights escalate. The Atlantic's database is a landmark showing the gap between artist consent and AI company practice, pushing toward clearer licensing norms.