By – Ishita Jain
Abstrarct
This article analyses the question of whether the training of generative AI on copyrighted content constitutes a violation of intellectual labour rights through expropriation or a legally permitted transformative use, and its implications on the economics of intellectual labour.The article discusses the varied American jurisprudence on AI training, particularly Thomson Reuters v. Ross Intelligence, Bartz v. Anthropic, and Kadrey v. Meta. While these cases generally recognised AI training as a transformative use, they differed significantly on questions of market harm and the extent to which the training data could be shown to originate from copyrighted works.The article then shifts focus to India, where the exhaustive list of fair dealing exceptions under Section 52 of the Copyright Act, 1957 makes the pending Delhi High Court decision in ANI Media v. OpenAI is an important one for the Global South.
Introduction
Every big language model relies on an enormous pile of human creation: novels, news stories, code, photographs, poems; everything was harvested at a scale that no library could contain. The people who created it have, by and large, never been consulted or compensated. It is now up to the law of intellectual property to determine whether this represents the biggest robbery in the history of art, or a legitimate transformation of knowledge freely available. The courts of three continents answer that question every day, and they disagree.
This essay identifies the nature of that disagreement, the dividing doctrine, the first decisions in the United States, their billion-dollar settlement, and India’s forthcoming test of a 1957 law, arguing that the issue is one of economics: the very existence of intellectual labor as a business enterprise.
The Doctrinal Fault Line
AI proponents maintain that the training process constitutes a non-expressive usage of copyrighted material since after the training process, the system does not hold any readable copies of the books, but instead it holds a statistical relationship between the words and the ideas. According to them, this kind of analysis is transformative and hence qualifies as fair use. In response, critics point out that for such statistics to be generated, the full copyrighted material must be copied during the process and this would compete with the original owners.
What American Courts Have Said
Judgments started trickling in from 2025, and they quickly split. The first one came from a Delaware court, where it was established that the use by Thomson Reuters of Westlaw headnotes to train an artificial intelligence legal research tool was not a fair use of the materials, because the resulting product was a market substitute rather than a transformative work. Almost instantly, in Bartz v. Anthropic,I, Judge William Alsup held that training AI models on lawfully acquired books could constitute a highly transformative use under copyright law, while making it clear that the use of pirated “shadow libraries” could not be justified. Just two days later, Judge Vince Chhabria in Kadrey v. Meta ruled that Meta’s AI would result in market flooding and thus be a copyright infringement.
A growing consensus indicates that AI training may qualify as a transformative use; however, the source of the training data and the extent of market harm remain the decisive considerations.. Chhabria’s “market dilution” theory states that the sheer weight of a glut of machine-made copies can outweigh any amount of transformation, is the case that creators will make against OpenAI, Google and Stability AI.
The Indian Test
Although India does not recognise a broad fair use doctrine like the United States, the debate remains highly relevant. As a major producer and consumer of digital content in the Global South, India’s approach to AI training and copyright is likely to have significant implications for the future of intellectual property governance. Section 52 of the Copyright Act, 1957 provides some exceptions, namely fair dealing for purposes such as private study, criticism, and news reporting, but, in the case of Super Cassettes v. Chintamani Rao, it was held by the Delhi High Court that the list provided in the statute was exhaustive. The absence of an adaptive and expansive test of four factors applicable to technology brings forth the case of ANI Media v. OpenAI, filed in November 2024. ANI asserts that OpenAI has “scraped” its news content, used the same in its outputs, and even created fictional news content bearing the name of ANI. OpenAI maintains that the process of training and storage of data by its model is conducted outside India and it learns only the pattern and not the expression. Justice Amit Bansal has reserved his judgment after thirty-two hearings in 2026.
The Labour Crisis Beneath The Doctrine
This is an economic problem of who gets the credit for intellectual work. Three years go into the creation of a novel; ten thousand novels will train a machine learning algorithm to create a fairly decent copy within one second, at near zero marginal costs. The problem is one of scale, when a single person cannot negotiate with the system having trillions of parameters, and the opt-outs proposed by the industry break down the copyright regime, which was never an opt-out system. The position of the creators of algorithms cannot be overlooked either: Humans have always studied other works, and nobody ever had to ask for permission to read something.Strict liability could concentrate AI development in the hands of a few large corporations able to afford licensed datasets, limiting research and innovation. It may also encourage AI development to shift towards jurisdictions with specific text and data mining (TDM) exceptions, such as the European Union, Japan, and Singapore.
Conclusion: Beyond the Binary
The truth of the matter is that “theft or fair use” is a false dichotomy, and the settlement now coming into view recognises this. Pirated training data is no longer defensible on legal grounds. Transparency is becoming mandatory: The EU AI Act’s disclosure requirements and California’s Training Data Transparency Act, which took effect in January 2026, compel developers to disclose what they trained on. A market in licenses is developing, from agreements between publishers and AI laboratories to the collective licensing systems modeled after the music industry. The function of copyright specified with the Constitution is not to halt technological advance but to make creation economically rational. Ultimately, the issue is not whether machines can learn from human effort; they already do. The question is whether the humans who made these machines possible, will benefit from their creation.
About the Author
I am Ishita Jain, a law student pursuing BCom LLB (Hons) in my 3rd year , with a strong interest in technology law, digital rights, public policy, and the legal challenges created by emerging technologies. Through legal research and writing, I aim to examine how technological progress can be made more accountable, inclusive, and fair for individuals and communities.
Image Source : https://lawsblog.london.ac.uk/2024/02/05/on-the-use-of-copyright-works-as-training-data-for-artificial-intelligence/

