By Sara Doczi, EFA Intern
AI regulation has been one of the most prominent issues of recent years, spanning copyright, algorithmic bias, high-risk AI decisions, transparency, child safety, and non-consensual deepfakes. But regulation moves far slower than AI development does, leaving consumers and businesses with little choice but to fight AI giants using existing legal frameworks…tools that were never built for this fight.
Within the sphere of AI training data, copyright has become one of the most important legal frameworks used by creatives to fight tech giants. But copyright can be a double-edged sword. As recent legal battles show, using existing copyright laws that do not explicitly address the issue of AI may not be the most effective way of protecting one’s own intellectual property.
How copyright law defines harm in AI cases
Many people feel uncomfortable, even violated, by the idea of an AI using material they created as training data. In a society where legal and judicial systems function as the primary arbiters of justice, legal action can feel like the only way forward. But most modern international law is not built to recognise and quantify feelings of discomfort or unease. Within the legal sphere, “harm” is defined through narrow, specific metrics tied to a person or their property.
So people have had to find ways of translating that harm into a language the law understands, and copyright might become the primary vessel for doing so. There have been a wave of copyright lawsuits against major AI companies in recent months, one of the most notable being Anthropic’s landmark settlement with more than 300,000 authors, including Charles Graeber, the award-winning journalist behind The Good Nurse.
The USD $1.5 billion (roughly AU $2.1 billion) settlement might look like a win for those authors. The reality is bleaker: not only does this amount shrink significantly when we consider average payout, but more importantly, the courts did not, strictly speaking, find it illegal for Anthropic to train models like Claude on copyrighted material, only that doing so via pirated copies crossed a line. This follows the precedent set by a major ruling last year, when a US federal judge sided with Meta against a group of authors.
Meta’s case is a clear illustration of how “harm” operates very differently in a legal sense than it does in everyday understanding. Meta won partly because the authors couldn’t show that the company’s use of their work had actually hurt sales of their original books. Legal frameworks built around products tend to look strictly at markets: the harm the law recognises is one where a product is damaged – through, for example, unfair competition and being tied directly to sales – rather than any intrinsic recognition of the distress caused to the author.
AI Training Data: Fair Use Has Limits
The US “fair use” doctrine, with its notably vague boundaries, plays a key role here too. The doctrine allows unlicensed use of copyrighted material in situations that advance freedom of expression – comment, criticism, parody, teaching, reporting, or scholarship. In practice, Australian copyright law, a common test is whether a use is “transformative” of the original work. That’s exactly how Meta won its June 2025 case against Richard Kadrey and other authors: the judge found AI training to be “highly transformative.”
Australia’s copyright laws are stricter than US ones, with the Copyright Act 1968 (Cth) having a “fair dealing” exception rather than a “fair use” one. Instead of employing the vague doctrine that guides American copyright cases, Australian fair use law allows for the use of copyrighted material only if it fits in a specific, closed list of purposes. Still, the US fair use doctrine is incredibly relevant, as a large number of AI models are trained in the United States, meaning that US laws apply to them.
But courts outside the US can still offer a glimmer of hope for those worried about the doctrine being exploited. At the end of July, the Munich Regional Court ruled against AI music generator Suno in its case brought by the German collecting society GEMA. The ruling matters because it shows companies can’t rely on US fair use protections worldwide. The court found that Suno’s output, simply by being accessible in Germany, constituted copyright infringement there, regardless of where the training itself took place.
It’s a strange irony that a doctrine meant to encourage engagement with art and culture is now being used to shield tech giants in court. As both the Anthropic and Meta cases show, there’s a troubling pattern in these rulings: judges tend to take issue specifically with training on pirated material, rather than with the broader act of using someone’s intellectual property without consent in the first place.
AI Training Is Destroying Rare Books
Recent reporting from The Guardian has surfaced something even more unsettling than the basic question of how AI is trained. Unsealed court filings revealed earlier this year that Anthropic had been “optimising” its training process by buying second-hand physical books, slicing off their spines, and mass-scanning the pages.
This practice, known as “destructive scanning”, is especially alarming when it comes to rare or antique books as the copies being bought and destroyed may be among the last surviving copies of a given edition. Worse, the practice is difficult to track, since the companies rarely do the bulk purchasing themselves. Instead, they go through third-party “book recycling” firms, which have no obligation to disclose who’s on the receiving end of their large, seemingly random second-hand orders.
Destructive scanning is one more example of how hard it is to bring legal action over training data, the process is obscure by design, and the law’s narrow definition of harm limits the scope of what can even be challenged.
Can Copyright Protect Creators From AI?
Right now, copyright cases are one of the sharpest tools artists and creators have to avoid becoming little more than a vast dataset waiting to be mined. But the courts, particularly in the US, remain fixated on piracy, licensing, fair use, and market competition, which sharply limits how effectively anyone can actually fight back against tech giants.
Copyright law, after all, was built for issues of a market-based society: to protect a person’s ability to make money from their own work, and remains limited to it. That function still matters enormously. It’s what stands between creatives and a future where AI-generated books, articles, designs, and music simply put them out of work.
Copyright law was never designed to grapple with something even bigger and more existential: what it means for human-made work to be absorbed, often without consent, into the creation of large, loosely regulated, and morally unresolved large language models. The judgement of AI-generated art must be incredibly layered, especially when considering cases that fall outside of the individual-centred nature of copyright cases discussed above.
The fact that anyone can prompt an AI image generator to create “Aboriginal art” and pass it off as authentic can not only be traumatic for the rightful and traditional owners of these art styles, as Amy Allerton, a Gumbaynggir and Bundjalung woman and founder of an Indigenous design company points out, but can also become detrimental to protecting cultural heritage.
Copyright, ultimately, is incredibly limited by its focus on individuals and markets. Proper ex ante AI regulation must address these wider issues, ensuring clear and strong protections against the bastardisation of human-made creative property; protections that move beyond existing legal frameworks built to address only individual, human subjects.
Image credit: Unsplash/Markus Winkler
Related Items:
- EFA Board Welcomes New Intern: Sara Doczi 4 August 2026
- World IP Day 2026: Copyright, Creativity, and the… 24 April 2026
- EFA Condemns Government's "Opportunity-First" AI… 3 December 2025
- EFA Slams Trump-Style AI Deregulation Agenda 7 August 2025
- #roboNDIS, Data Privacy and Security Failures:… 29 April 2025