By — Ishita Jain
Abstract
Today, artificial intelligence systems are trained on enormous amounts of human-generated data, such as photos, writings, personal data, artworks, social media and other online content. However, this poses the following question: does a person have a right to forbid an AI system from using his/her data for its training purposes? The answer to this question is far from being simple. On one hand, copyright law regulates some types of creative work, privacy law governs the use of personal data, and consent serves as the means of control over the use of data by individuals. On the other hand, AI systems are being trained by using automated scraping in bulk, which makes individual consent not only impractical but hardly enforceable either. This article considers the developing tension between AI development and individual control over data in the light of copyright, privacy and consent laws.
Introduction: When Your Data Becomes Someone Else’s Training Material
Unlike human beings, artificial intelligence learns differently. Training of generative AI relies on very large datasets that contain books, images, articles, artworks, websites and other types of information. The scale of such training implies the emergence of a very problematic reality: what an individual writes, photographs or posts on social media might be used to train a commercial AI model without any knowledge of the individual himself/herself.The question of whether people can opt out of AI training seems straightforward enough.
In India, there are certain copyright and privacy laws that help in restricting the use of an individual’s information in AI training. However, these laws include DPDP Act 2023 and Copyright Act 1957; neither of which has been framed specifically for generative AI. Thereby, posing a significant lack of measures in handling AI training.
Consent: Did We Ever Really Say Yes?
Consent is one of the fundamental principles for achieving individual control over data. In accordance with GDPR, the consent should, as a rule, be informed and voluntary. But the training of the AI happens at an enormous scale, involving billions of examples collected from many thousands of sites, and obtaining consent from all users becomes unfeasible in such conditions. Individuals have rights for access, erasure and objection to processing. According to Article 21, a person may, in certain cases, object to the processing of his/her data on the grounds of legitimate interest. The problem, however, persists: the right to erasure provided by law does not guarantee that the AI model will be able to “forget” the processed data.
Copyright: Is AI Training Copying?
Another method of legal opposition against the use of AI in training lies in copyrights, which protect original expressions but not the underlying ideas, facts, names or experience of any individual. Here the main question in case AI companies copy copyrighted works in their training programs is whether this constitutes an infringement or qualifies as an exception such as the fair use doctrine.
Such instances have resulted in notable lawsuits like The New York Times v. OpenAI, Andersen v. Stability AI, and Thomson Reuters v. Ross Intelligence, where courts examine whether copying of works by commercial entities for AI training purposes is legally justified. Nevertheless, there are definite limitations in using copyrights in this regard, since various kinds of personal data, such as faces, names, or even facts, are not eligible for copyright protection.
Privacy: Your Information, Your Control?
Privacy laws provide more coverage since they protect personal information as well as artistic creations. In India, the Digital Personal Data Protection Act, 2023 has provisions for consent, data processing and rights of the individual, indicating a trend of more personal control over personal data.
The problem with this is that it becomes more complicated if public information is used for a new purpose through AI technologies. A person may share a picture of him/herself on Instagram or an article online without having in mind the use of the information by a company developing AI technology.It becomes obvious that the old model of notice and consent does not work here. The mere existence of a privacy policy means nothing.
The Emerging “Right to Remain Untrained”
The right to remain untrained would move further beyond the existing privacy and copyright laws. It would recognize the capacity of an individual to state that: “You can access my information, but you cannot use it for training purposes for your AI systems.”
This right could be implemented in many ways. AI systems creators could be required to respect the machine-readable opt-out options. Website owners could send technical signals that their website is not supposed to be used for training purposes. The individual could receive an accessible way to object to the use of personal information for training purposes. The AI systems developers could also be obliged to make a disclosure of what kind of categories and sources were used for training.However, while the EU AI Act makes some progress toward transparency and imposes obligations on creators of general purpose AI systems related to copyright compliance and the summary of training data, transparency doesn’t mean individual control over one’s data.
The Bigger Problem: Can We Actually Unlearn?
The greatest hurdle will be technical. If a person manages to show that their data has been improperly used in AI training, then erasing the impact of that data from the trained model is another issue altogether.Unlike a mere deletion of files, AI cannot “unlearn” individual data. The cost of re-training the models is high, whereas machine unlearning is still under research. Thus, any significant right to be untrained will need both legal and technical means to track and delete data from AI.
Conclusion: From Permission to Participation
It is not just whether or not the AI can learn from the human’s data. It is about the person’s right to control what happens to their data and information. There are some legal protections through copyright, privacy, and consent, but none of these grants the right to stop AI training completely.
The discussion thus needs to move away from the legal vs. illegal use of AI training and to data autonomy. People should be given a choice in terms of allowing their data and creative work to become part of AI systems. The innovation by AI should go on but not at the expense of the individual’s rights.
About The Author:
I am Ishita Jain, a law student pursuing BCom LLB (Hons) in my 3rd year, with a strong interest in technology law, digital rights, public policy, and the legal challenges created by emerging technologies. Through legal research and writing, I aim to examine how technological progress can be made more accountable, inclusive, and fair for individuals and communities.
Image Source: https://lawbhoomi.com/ai-and-database-protection-laws/

