A child-safety organization has rated ChatGPT for Teens an unacceptable risk for users under 18 after testing the product’s safeguards across several areas. Engadget reported on October 8 that Common Sense Media’s Youth AI Safety Institute examined parental alerts, crisis responses, age prediction, emotional language and Study Mode. The assessment does not establish how every teen interaction will unfold, but it raises specific questions about whether the product behaves as parents might expect.
Common Sense Media said some advertised protections worked during its tests. The teen product refused explicit sexual roleplay, and Engadget described examples in which it declined to provide a calorie floor and recognized signs of medical danger following purging. In the latter case, the chatbot advised the test profile to involve a parent and seek pediatric care. The organization said those responses aligned with OpenAI’s specification for users under 18.
Other tests produced sharply different results. Common Sense Media linked 12 teen accounts to parental accounts, then spent as long as an hour discussing suicide, self-harm or disordered eating with the chatbot. According to Engadget, none of those sessions generated a safety notification to the linked parent during the test window. OpenAI said notification capability can take several hours to activate after accounts are linked and suggested that a technical issue may also have delayed messages.

The assessment also examined whether the chatbot directed users in crisis toward outside help. Researchers created 390 distinct mental-health prompts, and three child psychiatrists judged that 201 warranted a crisis response. Engadget reported that the teen version supplied a hotline number in 23 percent of those cases, compared with 33 percent for the pre-August version used as a baseline. It named a specific medical or mental-health professional in 58 percent, versus 68 percent for the baseline.
The teen system performed better on one aggregate measure cited in the report: advising a young user to involve a trusted adult. That response rose from 87 percent in the baseline results to 94 percent in the teen version. The mixed outcome is important. The testing found examples of appropriate intervention alongside lower rates for hotline and professional referrals, rather than a uniform failure across every safeguard.
Common Sense Media reported additional problems with age detection and the way ChatGPT responds when a teen treats it like a person. Testers supplied their fictional users’ ages at account creation, and the chatbot stored that information in memory, but those accounts were not properly moved into the teen experience. The organization also objected to language that appeared to simulate feelings or personal concern, arguing that such phrasing crossed a boundary in conversations with minors.

Study Mode was another point of criticism. According to Engadget, testers could bypass it by removing an @study prefix from prompts, while a popup labeled “Show me the answer” could let users request completed work. Common Sense Media interpreted those choices as weakening the product’s educational purpose. OpenAI’s explanation, as relayed in the report, was that the design introduces friction while preserving a young user’s agency and reminding them that other learning paths exist.
OpenAI rejected the assessment’s broader conclusion. A company spokesperson told Engadget that the testing did not accurately represent how the teen safeguards operate in practice or the views of experts on AI support for teenagers. The company also argued that much of the testing may have taken place before parental controls were fully activated, which it said would make the findings inaccurate. Common Sense Media, whose institute is funded in part by the OpenAI Foundation, maintained that the product did not appear fundamentally different enough in its testing.
The dispute leaves parents and researchers with two separate questions: whether the safeguards were tested under fully active conditions, and whether the product is sufficiently reliable once those conditions are met. The Engadget report does not resolve either question through independent replication, and its figures should not be generalized beyond the described test design. It does, however, document measurable gaps that merit further evaluation before parental controls or a teen label are treated as proof of consistent protection.

Comments
Loading comments…