Frontier AI models, from OpenAI, Anthropic, Google, Amazon, and xAI, are failing dramatically under multi-turn adversarial attacks, according to new research from Cisco’s AI threat intelligence team. The study reveals that safety benchmarks widely used across the industry miss nearly all real-world attack behavior, creating a dangerous gap between published safety scores and actual resilience. When attackers are allowed to probe models iteratively—reframing questions, building context across turns, adopting personas, and escalating gradually—attack success rates climbed as high as 88%, an order of magnitude above the lowest single-turn result. The findings challenge the industry’s reliance on single-turn evaluations and suggest that current model rankings may be fundamentally misleading.
The Multi-Turn Vulnerability Gap
The report paired single-turn and multi-turn evaluation across 15 closed flagship models, running roughly 30,000 single-turn prompts and nearly 7,000 multi-turn attacks spread over more than 1,400 conversations. Across the cohort, multi-turn attack success rates (ASR) ranged from about 8% to 88%, while single-turn ASRs were typically in the low single digits for the strongest models. OpenAI’s GPT-5.4 jumped nearly ninefold under iterative pressure, moving from a single-turn ASR in the low single digits to nearly 25%. Google’s Gemini 3 Pro climbed from about 18% to 73%, and xAI’s Grok 4.1 Fast in its non-reasoning configuration topped the cohort at 88%. Anthropic’s Claude family posted the strongest single-turn refusal performance, with single-turn ASRs in the low single digits, yet still landed in the 11% to 16% range once attackers were allowed to adapt. More than half of the models tested showed an absolute gap of at least 15 points between the two regimes. The cross-regime gaps ran in both directions: for example, Amazon’s Nova 2 Lite recorded a relatively high single-turn ASR but the lowest multi-turn ASR at about 8%, indicating that some models may be overfit to single-turn patterns.
These results underscore that adversarial robustness cannot be inferred from static benchmarks alone. Real adversaries will not stop after a single refusal; they will reformulate prompts, build context across turns, and escalate their requests. The researchers note that the gap between single-turn and multi-turn performance is so wide that it can misrank leading models, with different models appearing safer or more vulnerable depending on the evaluation method. The study extends an earlier Cisco analysis of eight open-weight models, where multi-turn ASR ran two to ten times higher than single-turn baselines and reached over 90% against Mistral Large-2. Multi-turn vulnerability appears to be a structural property of current frontier models, present in both open and proprietary systems.
Configuration and Guardrails
A striking finding from the research is the impact of a single configuration flag. The same Grok 4.1 Fast model with reasoning mode enabled saw its multi-turn ASR cut roughly in half—a swing of more than 40 points tied to a single capability flag. This configuration-driven safety variation does not appear on any public benchmark or model card the researchers reviewed. Users running the model in its default non-reasoning configuration encounter a substantially different threat profile from those who turn reasoning on. This highlights a critical risk: users and deployers may be unaware of the safety implications of model settings. Production deployments typically wrap base models in additional safety layers, such as guardrails and policy filters. The researchers acknowledge that these layers help but stress they do not eliminate risk. “Guardrails attenuate risk but do not eliminate it,” said Amy Chang, head of AI threat and security research at Cisco. “The base model sets the floor on what any production system can achieve. Just as traditional software development decisions involve risk tolerance and acceptance for the code itself and all its dependencies, the same approach applies to AI development and deployment. The blast radius for a rogue or misaligned AI agent, however, has the potential to be more damaging than a software flaw.” The research warns that as AI agents become more autonomous, the consequences of a single successful attack could cascade far beyond a simple text response.
Strategies and Recommendations
The study identified five strategy families that drove most of the multi-turn outcomes: role-play and persona adoption, contextual ambiguity, refusal reframing, information decomposition, and crescendo-style escalation. Within each family, the spread between the most and least exposed model was large, often approaching the full range of the chart. This means that strategy labels mostly sort which models pull apart from one another, even where average difficulty looks similar. On the single-turn side, three procedures dominated the rankings: Imposter AI, Soft Paraphrase, and System Prompts. By content type, hate speech, profanity, and specialized advice led. Imposter AI alone outpaced the tenth-ranked procedure by a wide margin, suggesting that targeted fixes to a handful of attack surfaces could move the aggregate numbers for most models.
The Cisco team proposes three operational steps for organizations buying or deploying AI: publish ASR by strategy family on every model release, gate deployments on regressions in the top three procedures and content types using a 3-point threshold, and flag any model with a cross-regime gap above 15 points for manual review. Applied to this cohort, the third rule alone surfaces more than half the tested models for closer examination. Regulatory frameworks point in the same direction. The NIST AI Risk Management Framework, the forthcoming NIST Cyber AI Profile (IR 8596), and Article 15 of the EU AI Act all call for adversarial robustness testing. None currently specify the interaction regime, strategy decomposition, or slice-support labeling the Cisco research argues is needed for decision-grade assessment. The researchers emphasize that until such standards are adopted, buyers and regulators must look beyond published benchmark scores and demand evidence of resilience against iterative, adaptive attacks. In a world where AI models are increasingly embedded in critical applications—from customer service to healthcare to national security—the failure to test for multi-turn vulnerabilities could lead to catastrophic outcomes.
Source:Help Net Security News

Leave a comment
Your email address will not be published. Required fields are marked *