A new mathematical approach attempts to predict when a chatbot switches from seemingly helpful responses to a dangerous one. Physicists Neil F. Johnson and Frank Yingjie Huo of George Washington University present a model that examines the competition between the context of the conversation and different possible responses. The study was published in the journal Patterns on October 8, 2026.

The research doesn't simply attempt to distinguish between right and wrong answers. The authors use the term "undesired" for a response that may be factually correct but encourages self-harm or provides harmful guidance in a sensitive environment. The term describes a behavioral issue, not necessarily an inaccuracy.
The question also concerns the timing of messages. A model can give several acceptable responses before changing direction, making the dangerous response harder to predict from past results alone. The researchers argue that the content of the conversation influences when this transition occurs.
How the dangerous response occurs
The team describes the process as a competition for the model’s “attention.” At a simplified level, the history of the conversation influences which of the possible response directions prevails. As new responses are added, the context can shift in a different direction, until the dangerous response emerges.
The formula estimates the transition point as the number of desired responses that precede the first undesirable one. Johnson and Huo rely on geometric relationships between the vectors representing the context, the desired response, and the undesirable one. This is an approximation, not a timer that guarantees when a dangerous response will occur.
The researchers tested the method on seven open models from three different development teams: GPT-2, Pythia, and OPT, with sizes ranging from 124 million to 12 billion parameters. In the classification of immediate or delayed displacement, the study reports correct prediction in 18 out of 19 cases, after two borderline cases were excluded from this calculation.
See also: Study warns of risks in using chatbots for therapy

The limits of the results
The result does not mean that the formula accurately predicts every harmful response. The authors clarify that the test focuses primarily on whether the transition is immediate or delayed, not on the exact text unit where it will occur. The numerical estimates remain approximate, while behavior near the boundary between two directions makes prediction difficult.
In a separate test at the sentence level, two discrepancies were associated with cases of negation: the geometric indication referred to an undesirable direction, while the full sentence meant the opposite. This shows that the numerical indication does not always capture the final meaning.
The team also looked at published tests of commercial chatbots by the Center for Countering Digital Hate. These showed patterns consistent with the researchers’ description, but they are not a direct test of the formula: the authors did not have access to the internal representations of the closed systems. The distinction is important because consistent behavior does not prove that the same predictive model works in a commercial service.
In their paper, the researchers acknowledge that larger test sets and prospective evaluation on new questions are needed. The costs of implementing such a control, the need to design appropriate response categories, and how it would interfere with real-world use have also not been fully quantified.

A possible sign, not a guarantee of safety
The potential value of the approach lies in providing early indication that the conversation is heading in an undesirable direction. Such an indicator could complement other controls, particularly in applications that operate without a constant connection to a central service. However, the dangerous response is not eliminated by simply estimating a transition point.
The SecNews technical team notes that the study describes a research tool, not a ready-made protection mechanism. Providers will need independent testing on different models, clear criteria for what is considered harmful, and an assessment of cases in which the formula fails. The dangerous response can depend on the order and context of messages, so testing beyond a limited set of prompts is required.
Johnson and Huo propose a way to examine the transition from helpful to unwanted responses, not proof that any chatbot can already reliably warn before a harmful outcome. The next critical test will be whether the formula maintains its performance across new queries and models, without turning an uncertain indication into a false guarantee of safety.
See also: ChatGPT for Teens: "Unacceptable risk" for teenagers
See also: Lawsuit against Character.AI and Google over teen's suicide
