Article image
Cover: generated with Midjourney, edited in Photoshop.
 

Suno’s Songs Put the Training Set on Trial

Generated tracks have become evidence in a ruling that could move AI music from extraction to licensing.

markus brinsa 13 august 13, 2026 9 9 min read create pdf website all articles

Verified Sources

On July 31, Munich Regional Court ruled that Suno was not entitled to use music represented by GEMA. The first-instance decision may be appealed, and the full written reasoning matters. But the case has already made one shift difficult to reverse.

AI music companies have long described training as an upstream technical process and generation as a downstream creative act. GEMA’s case connected those two stages.

It argued that the six works at issue had entered Suno’s model without authorization, and that the outputs made the model’s relationship to those works audible.

That is the deeper significance of the ruling. Copyright’s AI problem is no longer confined to an abstract fight over whether an enormous training corpus can be inspected, traced, or characterized after the fact. In music, a model may produce evidence against its own data practices. The output is not merely the product being sold. It can become an evidentiary route back into the system that produced it.

For Suno, the immediate consequences include a revenue disclosure order and damages whose amount has yet to be determined. For the wider market, the harder question is what happens when the distinction between input and output stops protecting the economics of unlicensed training.

Six Works, One Much Larger Argument

GEMA filed its suit against Suno in January 2025. It concerned the musical works behind six songs: “Atemlos,” “Daddy Cool,” “Rasputin,” “Big in Japan,” “Forever Young,” and “Mambo No. 5.” Lyrics were not part of the case before the Munich court.

The court’s March hearing background is unusually revealing. It said it was undisputed that Suno’s model had been trained on the six works in question. It also recorded GEMA’s allegation that Suno had used stream-ripping techniques to extract music from YouTube while bypassing the platform’s Rolling Cipher, a protection designed to prevent downloads.

Many AI copyright disputes begin with a vast corpus and no practical way for a rights holder to show whether a particular work entered it. Here, the dispute was more concentrated. The parties were arguing over specific compositions, specific training, and specific outputs.

GEMA documented 176 prompts concerning “Atemlos,” 124 for “Big in Japan,” 14 for “Daddy Cool,” 12 for “Forever Young,” eight for “Rasputin,” and four for “Mambo No. 5.” Its prompts included original lyrics, a desired musical style, and the title of the work. They did not include instructions for melody, harmony, rhythm, or arrangement.

That distinction is central. If the user supplied the musical expression at issue, an AI company can plausibly cast the output as a user-authored reconstruction. If the prompt supplied only identifying material and the system returned recognizable compositional features, the model itself becomes much harder to treat as a neutral conduit.

GEMA’s claim was that the outputs demonstrated memorization of protected works and, therefore, unlawful reproduction within the model as well as further infringement through the generated audio. Suno denied that the works were protected or recognizable in the outputs. It argued that training data is neither contained nor stored in the model; its weights and parameters, Suno said, encode generalized patterns rather than recoverable copies. It also argued that the relevant training was protected by U.S. fair use and that GEMA’s targeted prompting interrupted the causal chain.

Those are not peripheral defenses. They are close to the operating theory of generative AI: models learn patterns, users make requests, outputs are new. The ruling does not settle every version of that argument in every jurisdiction. It does something more commercially consequential. It establishes that this defense cannot simply be asserted at a high level when recognizable output and a documented training relationship are sitting in the same case.

The Output Became the Record

The AI industry has often treated provenance as an issue of scale. A company may acknowledge that it trained on material available on the open internet, while arguing that the size and transformation of the corpus make individual works legally or technically irrelevant.

Suno’s own California disclosure says its music models are trained on tens of millions of publicly available music files and related metadata. It also says those datasets may include both public-domain works and songs subject to intellectual-property protection. That disclosure is more candid than much of the sector’s public language, but it still describes training at a level that leaves the relationship between particular inputs and particular outputs unresolved.

GEMA’s case attacked that gap.

A generated track that is recognizably close to a protected composition is not a dataset inventory. It does not establish the complete path by which a file was acquired, processed, weighted, or surfaced in every instance. Nor does every resemblance establish copying.

Music has conventions, genres reuse structures, and similarity questions can turn on fine distinctions of expression and influence.

Yet outputs can change the economics of proof. They provide a rights holder with something that broad transparency commitments often do not: an observable result that can be compared to a known work, tested repeatedly, and placed before a court. The question is no longer only whether a company can explain its model architecture. It is whether it can explain why a model produces a particular result when the user has not asked it to supply the relevant melody, harmony, rhythm, or arrangement.

That is a more uncomfortable question for a system trained at web scale. It pushes firms from generic claims about statistical learning toward model-specific evidence about memorization, filtering, safeguards, and provenance.

The litigation therefore reaches beyond the old argument over whether weights are copies. A model need not function as a searchable archive for an output to create legal exposure. The practical issue is whether a protected work can be materially recovered through the behavior of the system.

Licensing Is Becoming Infrastructure

GEMA’s strategy is not limited to enforcement. It is building a commercial alternative to the data practices it is challenging.

In July, shortly before the Suno ruling, GEMA launched PLAI by GEMA, a curated music dataset for AI-tool providers. The offering bundles audio files, metadata, author rights, and master rights. GEMA says the package includes roughly 57,000 works, about 178,000 audio files, and more than 60 genres. Klangio, a Karlsruhe-based music-transcription company, became its first customer.

This is the part of the story that matters most for the market. GEMA is not simply arguing that a model should pay after it has been trained. It is attempting to turn cleared training data into an input product: one that can be selected, documented, priced, and supplied at scale.

That changes the strategic frame. For years, many generative AI companies treated rights clearance as a drag on the speed and breadth of model development. The assumption was that data abundance offered a structural advantage: the internet was large, content was available, and the legal question could be managed later.

A rights-cleared dataset makes the opposite claim. It says the durable advantage may come from knowing what a model was trained on, what rights travel with that data, and what uses are commercially supportable after deployment.

That does not mean a curated dataset can instantly replace the diversity of the open web. It cannot. The data products that emerge from this moment will involve tradeoffs in breadth, genre, language, cost, and creative utility. A model trained on licensed production music is not automatically a substitute for one trained on the accumulated history of recorded culture.

But the market is no longer choosing between frictionless data and no data. It is choosing between two kinds of infrastructure: opaque abundance that carries litigation and reputational risk, or governed supply that carries explicit cost and constraints.

That is a normal commercial choice in mature industries. Cloud providers pay for power, studios clear music, publishers acquire rights, and financial firms build compliance systems before offering products at scale. AI has been unusual because its core material often arrived without an agreed market structure. The Suno ruling adds pressure to build one.

Europe Is Turning Provenance Into a Business Question

The European Union’s AI Act does not decide the copyright merits of the Suno case. It does, however, create an environment in which providers of general-purpose AI models must adopt policies to comply with Union copyright law, including rights reservations, and publish a sufficiently detailed summary of training content. The Act also requires technical documentation about training and testing to be available to regulators.

These requirements do not make every model transparent in the way a rights holder might want. They preserve room for trade secrets and do not convert a public summary into a complete training log. Still, they alter the baseline.

Training-data provenance is moving from a question that can be deferred to litigation toward a question that regulators, downstream partners, investors, and commercial customers may reasonably ask before deployment. A model provider that cannot explain its data practices faces more than a lawsuit. It faces a procurement problem.

This is especially acute in music because the sector already has institutions designed to administer rights at scale. GEMA represents composers, lyricists, and publishers; record labels and other rights holders control other relevant interests; platforms have long operated through negotiated licensing arrangements. The machinery is imperfect and frequently contentious, but it exists.

Generative AI arrived with a very different economic model. It sought to turn the accumulated work of creators into a general-purpose capability, then sell access to the capability at marginal cost.

The legal dispute is now forcing a confrontation between that model and the cultural economy it draws from.

The central question is not whether AI music should exist. GEMA itself says it does not seek to prevent the use of its repertoire by AI systems. The question is who controls the terms under which cultural material becomes training infrastructure, and who receives a share when that infrastructure produces commercial value.

The Pressure Will Move Beyond Music

Music is an unusually clear testing ground because its structure can make resemblance legible. A melody, harmonic progression, rhythm, and arrangement can be compared. A rights holder can generate outputs, document prompts, prepare musicological analysis, and place a recognizable result before a court.

Other creative domains will not always be as tractable. Text, images, code, voice, film, and design each raise different questions about style, expression, transformation, and market substitution. But the broader pattern is portable.

AI companies will face more claims that combine two forms of evidence: a showing that protected work entered the training pipeline and a showing that the deployed system can reproduce, retrieve, or closely echo aspects of that work.

The litigation will not be limited to the philosophical question of whether a model “stores” data. It will increasingly turn on what the model demonstrably does.

That has implications for model design. Firms will invest more heavily in provenance records, dataset governance, model evaluations for memorization, output safeguards, rights-reservation systems, and licensing partnerships. Those investments will not be equally affordable for every company. Large incumbents may be better positioned to absorb the cost, negotiate catalog deals, and maintain legal teams. Smaller developers may find that licensing is both a route to legitimacy and a new barrier to entry.

The danger is a market in which only the largest technology companies and the largest rights owners can afford to participate. The opportunity is a more disciplined system in which creators, publishers, tool builders, and specialized dataset providers can negotiate over value before a model reaches mass scale.

That outcome is not guaranteed. A licensing market can become concentrated, expensive, and exclusionary. It can privilege catalog owners over independent creators. It can produce technical compliance that satisfies paperwork while obscuring real questions of compensation and cultural power.

But the old equilibrium is weakening. Suno’s case shows that a company cannot rely indefinitely on the claim that training is invisible while its outputs remain audible.

A New Cost of Building

The Munich ruling does not end the legal conflict around AI training. It does not resolve fair use in the United States, define every text-and-data-mining exception, or establish a universal standard for output similarity. Its legal force will depend on the appellate process and on the reasoning eventually available to the market.

Its strategic force is already clearer. AI companies are being asked to treat training data as a governed production input rather than a free resource that disappears into a model.

Rights holders are being pushed to prove not only that their work was taken, but that the system’s behavior reveals a continuing commercial relationship to that work. Regulators are beginning to require documentation that turns provenance into an operational obligation.

For music, the next phase will be fought over license terms, catalog access, output controls, revenue allocation, and the cost of proving that a model has learned too much from a particular body of work.

The unresolved pressure is whether the companies that built their advantage through scale can adapt before courts and regulators turn the cost of unlicensed training into a recurring cost of doing business.

About the Author

Markus Brinsa writes about AI failure, enterprise risk, governance, and the structural shifts underneath them — the through-line being the gap between AI governance on paper and what systems actually do at runtime. He created Chatbots Behaving Badly, a publication and podcast investigating real incidents in which AI systems gave bad advice, were manipulated, or failed in ways that mattered. He is the Founder & CEO of SEIKOURI Inc., an international strategy firm that gives enterprises and investors human-led access to pre-market AI — and converts first looks into rights and rollouts that scale. Access creates possibility. Rights create leverage. Scale turns early advantage into durable position. The two halves are the same work from opposite ends: SEIKOURI gets clients to AI early and makes sure what they deploy holds up once it's running. Thirty years bridging technology, strategy, and cross-border growth across the U.S. and Europe. I close the gap between what leaders expect AI to do and what it actually does in the wild.

brinsa.com
©2026 copyright by markus brinsa | brinsa.com™