AI training with the controversial data repository “The Pile” is back in court with a lawsuit from Chicken Soup for the Soul, LLC, accusing almost every major tech company of piracy. The problem is that Apple denies using it to train Apple Intelligence. Artificial intelligence is a term that has almost lost all meaning due to its application to everything.
See also: Apple blocks updates for AI “vibe coding” apps

In this sense, it appears that a guidance has mistakenly included Apple, while it has previously denied the use of the referenced dataset.
According to a lawsuit by Chicken Soup for the Soul, LLC, Apple, Meta, xAI, Google, Anthropic, OpenAI, Perplexity , and NVIDIA are infringing copyright by training their respective AI tools on a dataset known as “The Pile.” Although this dataset is filled with proprietary content, such as YouTube subtitle files, it was not used by Apple to train Apple Intelligence.
This lawsuit comes at a time when artificial intelligence is at the center of technological development, with many companies investing significant amounts in the development and training of their systems. “The Pile” is a widely used dataset that includes a large variety of data, from scientific articles to literary works, and has been used by many companies to train their models.
See also: Apple Sports: Allows you to watch NCAA March Madness in real time

However, the use of such data has raised concerns regarding copyright and the ethics of using it without permission. Apple, for its part, has stated that it has not used “The Pile” to train its systems, emphasizing that the lawsuit is unfounded. The company claims that the training of Apple Intelligence is based on data that has been obtained legally and with respect for copyright.
This case highlights the challenges that technology companies face in their effort to develop advanced AI systems, while also needing to ensure that data usage is conducted in a legal and ethical manner.
See also: Apple fixes WebKit vulnerability in iOS, iPadOS, and macOS

The outcome of the lawsuit could have significant impacts for the industry, determining how companies can use data to train their models in the future. While artificial intelligence continues to evolve, data management and protection of intellectual property rights remain critical issues that must be addressed by all involved parties.
