By – Ishita Jain
Abstract
Artificial Intelligence is powered by vast amounts of human-generated content, yet the creators of this data often receive neither consent, recognition, nor compensation. This article examines how AI training raises concerns relating to digital labour, copyright, privacy, intellectual property, and transparency, while arguing that ethical AI development must be grounded in consent, accountability, and respect for human creativity.
Introduction: The invisible labour behind artificial intelligence:
Modern AI systems rely on vast amounts of human-generated data including articles, images, code, and social media content to learn and produce sophisticated outputs. This creates a paradox: while millions of people collectively provide the knowledge that powers AI, they are treated merely as users rather than contributors and receive neither recognition nor compensation for their role.
Humans as data: Reducing Creativity Into Raw Material:
The most frightening aspect about the artificial intelligence economy is that human expression can so easily be translated into mere “data.” The moment something becomes “data,” it starts sounding neutral, technical, and devoid of any ownership. However, data never comes into existence on its own accord. Poems do not consist of merely data, but rather someone’s imagination. Paintings are not merely data, but someone’s creativity. Legal articles cannot just simply become data, because they are someone’s research work. Even our own life stories are more than just plain data.
AI technologies develop themselves using massive amounts of data, which include various types of human creations , including texts, images, videos, codes, websites, and many other things.
Learning from pre-existing knowledge is absolutely fine. Society advances by learning, getting inspired, and improving. Students learn from books, lawyers from previous judgments, and artists from observing other artists. Unlike AI, when an individual studies a book, they are simply acquiring knowledge rather than reproducing it at scale to compete with its creator.A corporation storing millions of books, posts, images, and articles to train a successful AI system for financial gain is making profits on a grander scale.
AI learning from humans is not the central concern; the real issue is that AI uses human-created knowledge to produce content and services that increasingly compete with the very people who created them. Across writing, translation, and programming, this creates a cycle where human creativity trains AI systems, companies profit from the outputs, and workers are forced to adapt to the changing technological landscape.
The new digital labour: Unpaid and unseen
In the industrial era, labour was visible: workers used machinery in factories under direct supervision. In the digital age, however, activities such as creating content, writing reviews, sharing tutorials, and tagging information generate value while concealing the labour behind them. This highlights the emergence of an unpaid digital workforce—not because users intentionally worked for AI companies, but because their online contributions were scrapped to train AI models without compensation. As a result, artists became training data, writers became sample material, and programmers became coding models.Another category of workers involved in AI that remains hidden are data laborers, annotators, and content moderators. There is also another group of hidden workers behind AI, including content moderators, data annotators, and human trainers, whose efforts are essential to the functioning of these systems. They review and classify data, evaluate chatbot responses, filter harmful content, and correct errors generated by AI models. Their labour helps make AI systems more accurate, efficient, and safe. Yet, their contribution remains largely invisible. While the world sees the “smart” AI system, it rarely sees the people who sorted, annotated, and refined the vast datasets on which it depends.
AI is thus far from being an artificial phenomenon. AI uses the language and creativity of humans, contains human bias and prejudice, displays the brilliance of human intelligence, has human errors, and involves human labor. This is why the AI machine looks like an independent one; its creators have been stripped of its visibility.
Consent, Ownership and the Legal Gap:
A key concern is the absence of meaningful consent. While users may make their content publicly available, this does not imply consent for it to be scraped and commercially used to train AI systems. Access does not equal appropriation artists, writers, and programmers share their work for public use, not for its incorporation into proprietary AI models.Copyright law is particularly relevant to the issues surrounding AI training. While copyright protection generally extends only to the expression embodied in an author’s work, AI systems are often trained on millions of texts, images, and other creative works rather than directly reproducing a single copyrighted work. This creates complex questions regarding ownership, fair use, and unauthorized exploitation of creative content. Privacy law is equally important where personal data is involved.
However, the AI training debate extends beyond personal information and encompasses creative works, non-personal data, publicly available online content, community knowledge, and copyrighted material. In India, the Digital Personal Data Protection Act, 2023 represents an important step towards regulating personal data, yet existing legal frameworks such as the Information Technology Act, 2000 and the Copyright Act, 1957 were not designed to address the challenges posed by large-scale AI training models. Consequently, significant legal uncertainties remain regarding consent, transparency, and the use of human-created content in the development of artificial intelligence. Thus, current regulations are clearly not ready to meet the challenges of the emerging AI economy. The legal framework is still based on the ideas of users, owners, platforms, and employees. However, AI introduces another actor, the human contributor whose work is mined without any relationship to the company profiting from their efforts.
Data colonialism and human dignity:
AI training also raises concerns of digital colonialism, where knowledge, language, culture, and human creativity are extracted by powerful corporations without equitable returns. For India, a major producer of multilingual digital content, effective AI governance is essential to ensure its citizens benefit from AI rather than becoming uncompensated sources of training data.
Conclusion: Humanity as the teacher of the machine:
AI has reinvented the concept of labour. No longer restricted to the office, the factory, or traditional employment, labour can take place in a caption, in a review, in a paragraph, in a picture, in a translation, in a legal comment, or simply in a conversation online . Humanity has been training AI for years now, but only now do we begin to realize the price of that training.
The debate is not about rejecting AI but ensuring that the humans whose knowledge powers it are protected through transparency, consent, recognition, and fair compensation. After all, those who taught AI should not be excluded from its benefits.
About the Author
I am Ishita Jain, a law student pursuing BCom LLB (Hons) in my 3rd year , with a strong interest in technology law, digital rights, public policy, and the legal challenges created by emerging technologies. Through legal research and writing, I aim to examine how technological progress can be made more accountable, inclusive, and fair for individuals and communities.
Image Source : google.com/url?q=https://enterpriseai.economictimes.indiatimes.com/news/industry/revolutionizing-the-workforce-how-human-ai-collaboration-is-redefining-jobs/129891041&sa=D&source=docs&ust=1784789974789577&usg=AOvVaw1zX-CBoWccl6XxdocIEgOu

