AI and Data: A Legal and Ethical Dilemma | ChatGPT OpenAI | OpenAI Chat | Chat OpenAI | Turtles AI

AI and Data: A Legal and Ethical Dilemma
How Big Tech Is Exploiting Data Without Consent, and Proposals for a More Fair, Transparent Future
Editorial Team23 December 2024

 

Big tech companies, in their exploitation of data to train AI models, raise significant legal and ethical questions. Although existing laws are not effective enough, radical proposals such as placing AI models in the public domain are emerging. This article explores the main dilemmas and possible solutions.

Key points:

  • Tech companies ignore laws to train AI with personal data.
  • Financial penalties are not enough to stop illicit behavior.
  • The “fruit of the poisoned tree” model could be applied to LLMs.
  • The proposal to make AI models public offers a radical and useful solution to combat data abuse.

Over the past few years, the evolution of AI models, particularly those based on large language models (LLMs), has raised questions about privacy, data protection, and compliance with the law. Major technology companies around the world, led by OpenAI and others, have developed advanced AI using vast amounts of data from the Internet, including personal data of users. However, these practices have often been carried out without the necessary consent of individuals, raising concerns about data misuse. Despite some regulatory initiatives and fines for copyright, privacy, and security violations, Big Tech continues to ignore regulations, partly because the economic penalties imposed are insufficient to deter giants with virtually unlimited resources. In this context, a radical initiative emerges as a possible solution: the proposal to make AI models public domain, as a way to redress the balance and counteract the abuse of power by technology companies.

The problem is that large AI models, like OpenAI’s GPT-4, are built on vast amounts of data from a wide range of sources, many of which have never authorized the use of their information. This is similar to criminal law, where there are strict rules about how evidence must be collected. A legal concept known as “poison fruit” means that if evidence is collected illegally, it cannot be used in court. Some argue that a similar logic should be applied to LLMs: if models are trained on illegally collected data, then they should be eliminated as “poisoned fruit” of the technology. However, this position raises moral and environmental challenges, as training such models requires massive energy resources and infrastructure. A recent analysis found that training GPT-4 consumed as much energy as about 4,500 homes over a 100-day period, with a devastating environmental impact that varies depending on the energy sources used.

This leads us to another important aspect of the debate: environmental ethics. While there is a justification for calling for the elimination of AI models if they were obtained through illicit practices, destroying them would have an environmental impact as damaging as the training process itself. The energy consumption related to the training of these systems is so high that it is a central concern, especially when considering the environmental cost of all the resources involved. Therefore, while the concept of the "fruit of the poisoned tree" is legally valid, it is not sustainable economically and ecologically.

At this point, a radical idea emerges that could solve both the legal and environmental problem: making AI models public domain. Models, built on data that belongs to everyone, should be considered common goods. If a model has been trained using data that belongs to millions of people, it might be right that the model, once developed, is no longer controlled by a single entity, but made available to the community. This would not only eliminate illicit profits for technology companies, but would also give everyone the opportunity to benefit from these tools. Furthermore, such a solution could be an effective response to violations of privacy rights, forcing companies to be more transparent and respectful of the law. The idea is that if a company breaks the law by collecting data, its models should be made publicly available, making it impossible to profit from illicit practices.

However, this proposal presents practical difficulties, especially on a legal level. Current laws do not explicitly provide for the possibility of forcing companies to make their AI models public, nor is there a clear mechanism that could force a company like OpenAI to release a product developed with the use of data collected without consent. In Europe, the GDPR has created advanced data protection legislation, but it has no direct powers to order the release of AI models into the public domain. Furthermore, technology companies may seek to avoid any kind of sanctions through legal forum shopping, choosing more favorable jurisdictions and further complicating the enforcement of the laws.

Despite these difficulties, the idea of ​​public access to AI models represents a challenge to a system that continues to privilege the economic interests of large companies. If we really want to ensure that technology works for the common good and not for the profit of a few, it will be necessary to address global cooperation between countries, which allows the application of international treaties similar to those already existing in other sectors, to enforce users’ rights and prevent Big Tech from continuing to operate with impunity.

If existing legislation is not enough to stop these practices, we will need to think of innovative solutions, and the proposal to make AI models public could be one of the most promising avenues.

However, the road to a world in which AI models are used in an ethical, transparent and human rights-respecting manner is still long and full of legal and practical obstacles.