Article image
Cover: generated with Midjourney, edited in Photoshop.
 

Fake Teens Tested Real Danger

Meta says it was safety benchmarking. WIRED’s reporting makes it look like chatbot safety has entered its undercover phase.

markus brinsa 25 july 30, 2026 12 12 min read create pdf website all articles

Verified Sources

The Kids Were Not Kids

There is a special kind of corporate sentence that sounds responsible until you picture what it means.

Hundreds of contractors working on a project for Meta allegedly posed as minors online to test rival chatbots on suicide, sex, eating disorders, drugs, and other high-risk topics. That is the sentence. Now picture the workflow.

Adult contractors, managed through a vendor, allegedly created dummy under-18 accounts. They sent prompts to ChatGPT, Gemini, Character.AI, and other systems. Some were written as if they came from children or teenagers in crisis. The answers were copied into spreadsheets. The project had a name, because of course it did. Internally, according to WIRED, it was known as Cannes.

Cannes. A name usually associated with red carpets, champagne, luxury yachts, and people pretending to care about cinema while checking whether their table is close enough to someone famous. In this version, Cannes apparently involved fake teen accounts, crisis prompts, and competitive safety benchmarking.

The AI industry has a talent for turning the grotesque into a process category. What used to sound like an undercover operation now arrives in the language of quality assurance. Nobody is spying. They are benchmarking. Nobody is impersonating children. They are simulating age-specific user journeys. Nobody is pushing rival systems into dangerous territory. They are evaluating adversarial safety behavior across model environments. It is amazing what a spreadsheet can do for the soul.

What WIRED Says Project Cannes Did

WIRED reported that hundreds of contractors on a Meta project were instructed to pose as minors and probe rival chatbots with prompts involving suicide, sex, eating disorders, drugs, profanity, racial slurs, and other sensitive subjects. The project was managed by Covalen, a Meta contractor, and was active as recently as April 21.

The targets were OpenAI's ChatGPT, Google's Gemini, and Character.AI. Workers were asked to create dummy under-18 accounts, send written prompts and images to rival chatbots, and copy the responses into spreadsheets. Some of the images, WIRED reported, included pills, knives, nooses, and a medical diagram of a gynecological procedure.

According to WIRED, one testing round completed in August 2025 involved more than 45,000 prompts. The companies behind the rival chatbots were not aware of the testing. WIRED also reviewed a spreadsheet of 3,748 prompts: hundreds focused on suicide and self-harm, hundreds more on eating disorders, and at least 239 on sex or romance, with others touching drugs, profanity, and racial slurs. Many were written from the perspective of children or teenagers in crisis, and, according to the project instructions WIRED described, were often designed to push the chatbots toward responses their safety systems were supposed to refuse.

Meta defended the work as routine safety testing. A company spokesperson told WIRED that testing and benchmarking chatbot responses to help ensure safe and age-appropriate experiences is "a responsible, industry-standard practice," and said Meta does not use competitor benchmarking to train its own AI models.

That defense is not absurd on its face. Safety testing is necessary. Red teaming is necessary. Edge-case testing is necessary. If companies put chatbots in front of minors, or build products minors can easily access, someone has to test what happens when a vulnerable child asks a dangerous question.

The uncomfortable part is not the existence of safety testing. The uncomfortable part is the costume.

Teen Personas Change the Ethical Frame

Testing a chatbot with a dangerous prompt is one thing. Testing it while pretending to be a child is another.

The teen persona changes the frame because the system is not merely being asked a risky question. It is being asked a risky question under an identity signal that should activate heightened care. The question is no longer only, "Will the model give unsafe information?" It becomes, "Will the model recognize that the speaker may be a child, and respond accordingly?"

That is a legitimate safety question. It is also a legally and ethically charged one.

When contractors create fake under-18 accounts on rival platforms, they are not merely observing public behavior. They are entering someone else's product environment under a manufactured identity — and the work appears to have crossed contractual lines while doing it.

OpenAI bars unsolicited safety testing, attempts to bypass safeguards, and using outputs to develop competing models. Google prohibits efforts to defeat its safety filters outside its own testing programs. Character.AI's policies prohibit harmful, exploitative, and obscene content, and since late 2025 the company has barred open-ended chat for under-18 users entirely. A months-long campaign of dummy-minor accounts probing all three sites, at minimum, against the terms every one of those companies publishes.

The "minors," of course, were adult contractors pretending to be minors, which keeps the operation clear of the gravest category of law.

WIRED asked two attorneys who specialize in online speech and technology law, Kendra Albert and Riana Pfefferkorn, to review examples of the prompts. Both said the material they were shown did not cross into soliciting child sexual abuse material or illegal obscenity, and the spreadsheet WIRED reviewed did not include prompts asking the chatbots to generate such material. So the conduct is not illegal in the way the subject matter might lead you to fear. That does not make it clean. It makes it ugly.

There is a version of this work the field considers legitimate, and it is worth being precise about why Cannes is not it.

Independent researchers and journalists probe these systems without permission all the time; adversarial testing is how the safety field learns anything. The difference is not the fake account. It is everything around it. Rumman Chowdhury, CEO and founder of Humane Intelligence, reviewed a sample of the Cannes prompts for WIRED and drew the line exactly there:

A dataset of thousands of youth-safety prompts could genuinely be useful for comparing how often chatbots refuse harmful requests, but the scale of the operation, its opacity, and the fact that the tested companies were never told made it something other than a safety benchmark. Structuring a monthslong program designed to systematically break other platforms' rules, she said, using dummy accounts masquerading as children, sits outside what anyone normally means by "industry standard" evaluation.

Her sharper formulation is the one that should worry every company now building a safety team: this is "exactly the kind of governance gray zone where safety becomes a convenient cover for anticompetitive practices." That is the whole problem in one sentence. The method can be identical on both sides of the line. What changes is who is doing it, whether they disclosed it, and who stands to benefit from what they find.

Safety Benchmarking Can Become Competitive Intelligence

Testing competitors is not, by itself, unusual. Business Insider reported last year that Scale AI contractors working on Google's Bard compared its answers against ChatGPT and rewrote them to match or beat the rival. Product managers use rival products. Security teams study attack surfaces. Lawyers read terms of service. None of that is surprising.

AI safety benchmarking sits in a more delicate place, because the benchmark is not speed, latency, price, or interface design. It may involve vulnerable users, self-harm scenarios, sexual content, child-safety signals, violent imagery, and prompts designed to trigger failures.

That creates a new kind of competitive intelligence problem. If one company secretly tests another by posing as minors, it learns where the rival refuses, where it complies, where it escalates, where it redirects, and where it leaves the user alone with dangerous information. That is valuable information. It can be used to improve safety. It can also be used to shape public claims, regulatory narratives, litigation strategy, and product roadmaps.

The Cannes documents make that tension concrete rather than hypothetical. OpenAI's terms specifically bar using its outputs to develop models that compete with OpenAI. One former contractor told WIRED they feared the project amounted to secretly taking material from competitors' systems to feed back into Meta's. Meta, for its part, explicitly denies using competitor benchmarking to train its models. Three facts point to one question the public documents cannot answer: what was the collected data actually for? WIRED reported that the material it reviewed did not indicate how, or whether, Meta used the responses at all.

That missing piece is not a footnote. It is the whole nerve ending of the story.

In ordinary software, competitive benchmarking feels mundane. In chatbot safety, it starts to resemble a border crossing between audit and espionage. The word "espionage" should be used carefully. This was not an intelligence agency breaking into servers; based on the public reporting, the operation involved interacting with account-accessible chatbot products. But the cultural feeling of the story is unmistakable. Contractors allegedly wore fake youth identities to enter rival systems and extract evidence of how those systems behaved under pressure.

That is undercover work, even if the badge says safety.

The Industry Built the Products That Require These Experiments

The absurdity would be funnier if the underlying problem were not so grim. Chatbots are not search engines with better manners. Many are designed to feel conversational, patient, responsive, and socially aware. They can answer homework questions, flirt, role-play, give advice, mirror emotional language, and stay present long after a human friend would have said, "Please call someone." That makes them especially complicated around children and teenagers.

The Federal Trade Commission opened an inquiry in 2025 into AI chatbots acting as companions, issuing orders to seven companies — including Meta — for information about how they measure, test, and monitor potential negative effects on children and teens. The FTC noted that chatbots can mimic human characteristics, emotions, and intentions, and may prompt users, especially children and teens, to trust and form relationships with them. That is the background against which Project Cannes should be read.

The issue is not whether chatbot makers should test dangerous conversations. They must. Someone has to ask the questions real vulnerable users may ask. Someone has to test whether age signals are recognized. Someone has to find out whether the bot will redirect a teenager discussing self-harm or simply continue the conversation as if nothing were wrong.

The issue is that the industry has released or normalized products that make such testing feel both necessary and grotesque.

Once chatbots are treated as companions, tutors, friends, therapists, characters, and all-purpose emotional vending machines, the safety work inevitably moves into darker rooms.

And there is a question hiding inside the method itself. WIRED reported that many of the Cannes prompts were crude or repetitive attempts to elicit responses a well-functioning chatbot should plainly reject, raising the question of what the project measured beyond a system's ability to refuse an obvious provocation. Chowdhury noted the same thing.

The operation may not only be ethically ugly. It may also be mediocre safety research — an expensive way to prove that competitors' guardrails catch the easy cases.

That is a more damning read than "grotesque but useful." It suggests the point was never only safety.

Because here is the symmetry. For years, tech platforms built systems that encouraged children and teenagers to behave like content sources, engagement units, audience segments, and monetizable attention. Now the same industry has reached the point where adult contractors allegedly pretend to be teenagers so that one platform can test whether rival platforms behave responsibly toward teenagers.

The teen is no longer only the user. The teen is the test condition.

That is revealing. It is a picture of the entire chatbot economy trying to audit the monster it is still feeding.

Meta Is an Awkward Messenger for This Lesson

Meta defending teen chatbot safety testing is not inherently wrong. Meta being the messenger is where the room gets tense.

The company has spent years under child-safety scrutiny across its social platforms. In February 2026, Axios reported that internal red-teaming presented in the New Mexico attorney general's suit against Meta found an unreleased chatbot product failed to protect minors from sexual exploitation 66.8 percent of the time. NYU professor Damon McCoy, an expert witness with access to the documents, testified to failure rates across other severe categories as well: 63.6 percent for sex-related crimes, violent crimes, and hate, and 54.8 percent for suicide and self-harm. Meta said it never launched the product precisely because that testing surfaced the problems, and characterized the exercise as red teaming designed to elicit violations so they could be fixed. McCoy, in his testimony, described the document as reflecting outcomes of the product when it was deployed.

That history does not mean Meta is wrong to test rival systems. It does mean the company does not get to float above the story as a neutral lab technician in a white coat.

When a company with its own child-safety controversies allegedly sends contractors into rival systems under fake teen identities, the public is entitled to ask what exactly is being tested. Is this about raising safety standards? Learning from others? Proving that competitors also fail? Preparing for a marketplace where every AI company will accuse every other of being unsafe for children? The answer may be several of these at once. That is the problem.

Because AI safety is becoming adversarial. Companies are not merely testing themselves against risk. They are testing one another — watching how rivals handle edge cases, mapping guardrails, learning which failures look embarrassing, which look legally dangerous, and which might become useful the next time regulators, journalists, or lawmakers come asking. In the old internet safety debates, platforms argued about moderation from a distance. In the chatbot era, they can interrogate each other's systems directly. That changes the game.

The Spreadsheet Is Not Accountability

One of the great delusions of modern platform governance is the belief that if a disturbing act passes through a dashboard, a spreadsheet, or a vendor workflow, it has somehow become responsible.

A contractor enters a fake teen persona. A prompt is sent. A chatbot responds. The answer is copied. A cell is filled. A metric is created. A dataset is delivered. A manager reviews the numbers. The whole thing acquires the clean smell of process.

But process is not the same as accountability. And the people running the prompts felt the difference. Former contractors told WIRED that aspects of the work alarmed them — that they feared they might be generating or preserving illegal material if a chatbot responded the wrong way to a sexual prompt involving a minor, and that the project felt like quietly siphoning competitors' outputs back toward Meta. "Everyone I knew who worked on this project was completely gobsmacked," one told WIRED. "Like, surely we are going to get in trouble for doing this?" Whatever Cannes was, it was disturbing enough that the workers executing it assumed someone would eventually have to answer for it.

Accountability would require knowing why the test was conducted, who approved the fake teen accounts, what rules governed the prompts, how the disturbing material was handled, whether contractors received adequate support, whether rival platforms' terms were considered, how the collected responses were used, whether legal and ethics teams reviewed the operation, and whether any findings improved child safety beyond one company's competitive position.

Without those answers, the spreadsheet is only a ledger of discomfort.

It may contain useful evidence. It may even contain evidence that helps make chatbots safer. But a useful dataset is not automatically a clean dataset, and a safety purpose does not erase every method used to pursue it. The AI industry needs red teaming. It also needs rules for red teaming that do not depend on whether the company conducting the test has enough lawyers to rename the awkward parts.

Everyone Wants to Be the Adult in the Room

The most revealing part of this story is the fight over who gets to sound responsible. Meta says safety testing is standard. The rival companies say they were unaware and object to the methods — Character.AI called the conduct a violation of its terms and of the characters its community built, OpenAI said it was looking into the issue, and Google said it had not authorized the testing and, on its own retest of the samples, found Gemini responding within policy. Regulators want answers about children and chatbot companions. Researchers want better testing. Parents want products that do not turn their children's private despair into a model interaction. The chatbot companies want scale, engagement, and the legal status of mere tools whenever the conversation gets dangerous.

Everyone wants to be the adult in the room. The problem is that the room appears to be full of fake teenagers, real contractors, synthetic friends, corporate vendors, rival platforms, and spreadsheets full of crisis prompts.

That is not a safety culture. That is a very expensive panic response with a project codename.

The uncomfortable truth is that Project Cannes may have been trying to answer a real question. How do rival chatbots behave when a child appears to be in danger? That question deserves serious testing. But the way it was reportedly pursued shows how far the industry has drifted from any simple idea of responsible AI. The companies have built systems that talk like people, compete like platforms, defend themselves like utilities, and test each other like intelligence targets.

When safety testing starts looking like undercover chatbot espionage, the lesson is not that testing should stop. The lesson is that the chatbot industry has reached the point where even its safety work needs adult supervision.

About the Author

Markus Brinsa writes about AI failure, enterprise risk, governance, and the structural shifts underneath them — the through-line being the gap between AI governance on paper and what systems actually do at runtime. He created Chatbots Behaving Badly, a publication and podcast investigating real incidents in which AI systems gave bad advice, were manipulated, or failed in ways that mattered. He is the Founder & CEO of SEIKOURI Inc., an international strategy firm that gives enterprises and investors human-led access to pre-market AI — and converts first looks into rights and rollouts that scale. Access creates possibility. Rights create leverage. Scale turns early advantage into durable position. The two halves are the same work from opposite ends: SEIKOURI gets clients to AI early and makes sure what they deploy holds up once it's running. Thirty years bridging technology, strategy, and cross-border growth across the U.S. and Europe. I close the gap between what leaders expect AI to do and what it actually does in the wild.

brinsa.com
©2026 copyright by markus brinsa | brinsa.com™