Generative AI and Legal Risks: OpenAI’s Sora Case | Chat GPT gratis | ChatGPT app | OpenAI italiano | Turtles AI
Using undisclosed data to train AI raises complex legal issues. OpenAI faces controversy over the provenance of the dataset with the launch of Sora. Potential risks include copyright, trademark, and likeness viola.
Key Points:
- Lack of transparency: OpenAI did not disclose the exact source of the training data.
- Legal implications: Training may include copyrighted content.
- Probabilistic models: AI can faithfully reproduce protected data.
- Evolving industry: Legal disputes will shape the future of the creative industry.
OpenAI is at the center of heated debate over the legality of training Sora, its innovative video-generating AI. The recently launched Sora stands out for its ability to produce high-quality clips, from text or visual prompts, up to 20 seconds long, adapting to different resolutions and aspect ratios. However, the mystery surrounding the source of the data used to train the model raises important questions. While OpenAI has acknowledged using “publicly available” data and licensed content from platforms like Shutterstock, it is unclear whether footage from video game walkthroughs or Twitch streams was included without permission. This lack of clarity has led legal experts to speculate on possible copyright infringement.
Sora appears to be able to emulate iconic settings and styles from well-known video games. From near-faith replicas of Super Mario Bros. to scenes reminiscent of shooter games like Call of Duty and Counter-Strike, the model’s ability to generate such specific content suggests that training may have involved protected material. Furthermore, the presence of details that recall famous streamers, such as Auronplay or Pokimane, further fuels speculation about the nature of the dataset. OpenAI has tried to limit the generation of protected content, implementing filters that prevent explicit references to registered titles. However, the model seems able to circumvent these restrictions through indirect descriptions.
The central issue concerns the possible use of unauthorized material during the training phase. Experts point out that gameplay videos have multiple levels of copyright protection: from the content of the game itself to unique creations by players, including any contributions generated by users. For games like Fortnite, which allow the customization of maps and scenarios, intellectual property rights multiply, increasing the legal risk for those who use such materials without a license.
Generative models, by their probabilistic nature, learn patterns from the data provided during training. This feature, although powerful, carries the risk of reproducing elements almost identical to the original ones, raising concerns about possible copyright infringements. The AI industry is already facing lawsuits related to content generation. Microsoft, OpenAI, and others have been sued for alleged infringement related to reproducing copyrighted works, and similar cases involve companies like Midjourney and Stability AI. Content creators and rights holders are demanding more protections, arguing that indiscriminate training of AI models violates fair use principles.
The example of Google Books, which prevailed in a similar case, shows that courts may recognize transformative value in AI models. However, even in this scenario, end users who consume recognizable outputs could face legal liability. Sora’s ability to emulate assets like textures, voices, and animations introduces additional trademark and likeness rights risks. One particularly problematic application would be the use of models like Sora to generate real-time interactive games, which could mimic proprietary content.
Legal and technological developments in this space will continue to redefine the relationship between AI and intellectual property, charting an important path for the future of the creative industry.
