The Atlantic has created a searchable database showing millions of music tracks that were freely available in datasets used to train AI models—despite not being licensed for that purpose. The investigation highlights a critical gap in how AI training data was sourced: while many datasets were technically public, the artists behind them never consented to their use in AI training.
The scale is staggering. Billions of tracks from Spotify, YouTube, and other sources ended up in training datasets with no compensation to artists. The database lets users search by artist name to see if their work was included without permission, uncovering a systemic issue in how AI companies accessed training data.
For the music industry, the findings reinforce calls for clearer licensing frameworks around AI training. The publication marks a shift toward transparency—showing the gap between how AI is imagined (trained on ethical, licensed data) and how it was actually built (with whatever data was cheapest and most available).