News

Industry shock: Meta was exposed to induce competing AI products to test extremely psychologically sensitive topics. The safety boundary in the field of artificial intelligence recently triggered heated discussions due to an internal message.

3 min read
Industry shock: Meta was exposed to induce competing AI products to test extremely psychologically sensitive topics. The safety boundary in the field of artificial intelligence has recently triggered heated discussions due to an internal message. According to media disclosures, Meta Company had previously launched a project code-named "Cannes". By hiring outsourced personnel to pretend to be minors, it conducted a highly controversial "extreme stress test" on multiple mainstream competing chatbots, including ChatGPT, Gemini and Character.AI. According to internal Meta documents and descriptions from multiple people familiar with the matter, the project continued to operate until at least April 21 this year. Meta organized hundreds of workers through the outsourcing company Covalen and asked them to create fake user profiles under the age of 18 and register with disposable email addresses and unified passwords. In conversations with competing AIs, these "minor" accounts frequently send prompt words involving high-risk topics such as suicide, self-mutilation, and eating disorders, and even upload images of knives, pills, and nooses to further stimulate the model's feedback mechanism. According to internal data from the project, these test contents were carefully choreographed to test the security defense systems of competing AI products, trying to find and induce chatbots to bypass their proper interception mechanisms and thus output content that did not comply with security specifications. In one round of testing in August 2025 alone, staff entered more than 45,000 high-risk prompt words into competing product platforms. In these conversations, testers often play the role of teenagers in trouble, trying to probe the model's bottom line by concocting extremely sensitive scenarios such as "asking about abortion pills," "facing threats of violence," or "hiding an eating disorder." Meta defended this incident in a public statement. A spokesperson said that benchmarking chatbot responses is an industry standard practice to ensure that AI products are safe and age-appropriate, and that any accusations of malicious intent are a misunderstanding of technology companies’ efforts to improve their systems. At the same time, Meta clearly denied that it would use these test data for competing products to train its own models. This incident not only revealed the gray area that exists in AI safety testing, but also once again triggered deep concerns about the vulnerability of generative AI in dealing with sensitive issues among teenagers. With the rapid iteration of AI technology, how to handle the boundaries between competitive behavior and ethics while maintaining efficient safety testing has become an urgent issue in the industry. via AI News (author: AI Base)