Meta 's new AI assistant gained its knowledge through public posts on Facebook and Instagram .

Meta Platforms used public posts on Facebook and Instagram to train parts of its new virtual assistant, Meta AI. But it avoided personal posts shared with family and friends as it tries to respect users' privacy, according to the company's top executive.
Additionally, Meta did not use personal conversations as training data in its messaging services. Steps have been taken to filter out personal details from datasets used for training, Meta's President of Global Affairs Nick Clegg said during his keynote at the company's annual Connect conference.
See also: Meta Verified for Business: Subscription program extends to businesses
“We tried to exclude datasets that contained personal information,” Clegg said, adding that Meta largely used public data for training.
LinkedIn example that Meta did not use the app, due to privacy concerns.
Clegg's comments come as tech companies including Meta, OpenAI and Google have come under fire for using information scraped from the internet without permission to train their artificial intelligence models, which ingest large amounts of data to synthesize information.
Companies are considering how to manage the private or copyrighted material they have collected during this process and how AI might reproduce it, while also facing lawsuits from creators accusing them of copyright infringement.
Meta AI emerged as the most important product among the first artificial intelligence tools unveiled by CEO Mark Zuckerberg during Meta’s annual product conference, Connect. This year’s event focused on artificial intelligence, unlike previous meetings that focused on virtual reality.
Meta built the AI assistant using a custom model based on its powerful Llama 2, which the company released for public commercial use in July. It also developed a new model called Emu, which generates images in response to text messages.
See also: OpenAI: Is it negotiating with Jony Ive for a Hardware Project?
The product will be able to generate text, audio and images, as well as access real-time information through collaboration with Microsoft's Bing.
Clegg said that the public posts on Facebook and Instagram used to train Meta AI included text and photos.
These posts were used to train Emu to create product imagery, while the chat features were based on Llama 2 with some publicly available and annotated datasets added, a Meta spokesperson said.
Interactions with Meta AI can also be used to improve features in the future, the spokesperson said.
Clegg mentioned that Meta has imposed security restrictions on the content that the Meta AI tool can generate. These restrictions include, among other things, a ban on creating photorealistic images of public figures.
Regarding copyrighted material, Clegg said he envisions a “reasonable distinction” regarding “whether the creative content falls within the existing fair use doctrine or not,” which allows for limited use of protected works for purposes such as commentary and research.
“I believe that is the case, but I strongly suspect that this will play a significant role in the trial,” Clegg said.
See also: Paint Cocreator: Released to Insiders with artificial intelligence
Some companies provide image creation tools that make it easy to reproduce iconic characters, such as Mickey Mouse. Other companies, on the other hand, have paid to acquire the materials or have intentionally avoided including them in the training data.

OpenAI, for example, signed a six-year deal with content provider Shutterstock over the summer to leverage the company's image, video, and music libraries for educational purposes.
When asked if Meta had taken steps to prevent the reproduction of copyrighted images, a Meta spokesperson referred to its new terms of service that prohibit users from creating content that violates privacy and copyright.
Source: reuters.com
