ACVEAgent configuration vulnerability registry

Model injection benchmarks

model-injection-benchmarks.json lists published prompt-injection benchmark results, one row per (model, benchmark, attack, defence, attempts) as read from the cited table or figure.

It is not an advisory feed: rows mint no ACVE ids, carry no severity, and are not matched against lockfiles.

Llama 3.3 70BInjecAgentindirect injection, higher of base and enhanced settings (ASR-total)sandwich prompting153.8%simulatedarxiv.org Table 52026-02-06
Llama 3.3 70BInjecAgentindirect injection, higher of base and enhanced settings (ASR-total)none186.0%simulatedarxiv.org Table 102026-02-06
Llama 3.3 70BAgentDojoimportant_instructionssandwich prompting (repeat_user_prompt)114.7%949simulated0.598arxiv.org Table 52026-02-06
Llama 3.3 70BAgentDojoimportant_instructionsnone123.0%949simulated0.629arxiv.org Table 102026-02-06
Llama 3.3 70BWASPweb-page injection, intermediate ASR (agent diverted at any point)none120.2%84real0.622arxiv.org Table 52026-02-06
Llama 3.3 70BWASPweb-page injection, end-to-end ASR (injected task completed)none12.4%84real0.622arxiv.org Table 52026-02-06
Meta SecAlign 70BInjecAgentindirect injection, higher of base and enhanced settings (ASR-total)sandwich prompting10.5%simulatedarxiv.org Table 52026-02-06
Meta SecAlign 70BInjecAgentindirect injection, higher of base and enhanced settings (ASR-total)none12.1%simulatedarxiv.org Table 102026-02-06
Meta SecAlign 70BAgentDojoimportant_instructionssandwich prompting (repeat_user_prompt)11.9%949simulated0.845arxiv.org Table 52026-02-06
Meta SecAlign 70BAgentDojoimportant_instructionsnone12.3%949simulated0.794arxiv.org Table 102026-02-06
Meta SecAlign 70BWASPweb-page injection, intermediate ASR (agent diverted at any point)none11.2%84real0.595arxiv.org Table 52026-02-06
Meta SecAlign 70BWASPweb-page injection, end-to-end ASR (injected task completed)none10.0%84real0.595arxiv.org Table 52026-02-06
GPT-4o miniInjecAgentindirect injection, higher of base and enhanced settings (ASR-total)sandwich prompting13.3%simulatedarxiv.org Table 52026-02-06
GPT-4o miniInjecAgentindirect injection, higher of base and enhanced settings (ASR-total)none17.7%simulatedarxiv.org Table 102026-02-06
GPT-4o miniAgentDojoimportant_instructionssandwich prompting (repeat_user_prompt)111.9%949simulated0.67arxiv.org Table 52026-02-06
GPT-4o miniAgentDojoimportant_instructionsnone130.9%949simulated0.701arxiv.org Table 102026-02-06
GPT-4o miniWASPweb-page injection, intermediate ASR (agent diverted at any point)none153.6%84real0.27arxiv.org Table 52026-02-06
GPT-4o miniWASPweb-page injection, end-to-end ASR (injected task completed)none10.0%84real0.27arxiv.org Table 52026-02-06
GPT-4oInjecAgentindirect injection, higher of base and enhanced settings (ASR-total)sandwich prompting122.7%simulatedarxiv.org Table 52026-02-06
GPT-4oInjecAgentindirect injection, higher of base and enhanced settings (ASR-total)none136.9%simulatedarxiv.org Table 102026-02-06
GPT-4oAgentDojoimportant_instructionssandwich prompting (repeat_user_prompt)120.4%949simulated0.794arxiv.org Table 52026-02-06
GPT-4oAgentDojoimportant_instructionsnone143.2%949simulated0.804arxiv.org Table 102026-02-06
GPT-4oWASPweb-page injection, intermediate ASR (agent diverted at any point)none117.9%84real0.324arxiv.org Table 52026-02-06
GPT-4oWASPweb-page injection, end-to-end ASR (injected task completed)none12.4%84real0.324arxiv.org Table 52026-02-06
GPT-5InjecAgentindirect injection, higher of base and enhanced settings (ASR-total)sandwich prompting10.2%simulatedarxiv.org Table 52026-02-06
GPT-5InjecAgentindirect injection, higher of base and enhanced settings (ASR-total)none10.6%simulatedarxiv.org Table 102026-02-06
GPT-5AgentDojoimportant_instructionssandwich prompting (repeat_user_prompt)10.2%949simulated0.803arxiv.org Table 52026-02-06
GPT-5AgentDojoimportant_instructionsnone10.2%949simulated0.835arxiv.org Table 102026-02-06
GPT-5WASPweb-page injection, intermediate ASR (agent diverted at any point)none10.0%84real0.595arxiv.org Table 52026-02-06
GPT-5WASPweb-page injection, end-to-end ASR (injected task completed)none10.0%84real0.595arxiv.org Table 52026-02-06
Gemini 2 FlashInjecAgentindirect injection, higher of base and enhanced settings (ASR-total)sandwich prompting127.2%simulatedarxiv.org Table 52026-02-06
Gemini 2 FlashInjecAgentindirect injection, higher of base and enhanced settings (ASR-total)none166.5%simulatedarxiv.org Table 102026-02-06
Gemini 2 FlashAgentDojoimportant_instructionssandwich prompting (repeat_user_prompt)111.3%949simulated0.423arxiv.org Table 52026-02-06
Gemini 2 FlashAgentDojoimportant_instructionsnone112.4%949simulated0.443arxiv.org Table 102026-02-06
Gemini 2 FlashWASPweb-page injection, intermediate ASR (agent diverted at any point)none129.8%84real0.486arxiv.org Table 52026-02-06
Gemini 2 FlashWASPweb-page injection, end-to-end ASR (injected task completed)none18.3%84real0.486arxiv.org Table 52026-02-06
Gemini 2.5 FlashInjecAgentindirect injection, higher of base and enhanced settings (ASR-total)sandwich prompting10.1%simulatedarxiv.org Table 52026-02-06
Gemini 2.5 FlashInjecAgentindirect injection, higher of base and enhanced settings (ASR-total)none13.5%simulatedarxiv.org Table 102026-02-06
Gemini 2.5 FlashAgentDojoimportant_instructionssandwich prompting (repeat_user_prompt)127.9%949simulated0.639arxiv.org Table 52026-02-06
Gemini 2.5 FlashAgentDojoimportant_instructionsnone130.7%949simulated0.588arxiv.org Table 102026-02-06
Gemini 2.5 FlashWASPweb-page injection, intermediate ASR (agent diverted at any point)none144.1%84real0.568arxiv.org Table 52026-02-06
Gemini 2.5 FlashWASPweb-page injection, end-to-end ASR (injected task completed)none114.3%84real0.568arxiv.org Table 52026-02-06
Gemini 3 ProInjecAgentindirect injection, higher of base and enhanced settings (ASR-total)sandwich prompting10.2%simulatedarxiv.org Table 52026-02-06
Gemini 3 ProInjecAgentindirect injection, higher of base and enhanced settings (ASR-total)none12.1%simulatedarxiv.org Table 102026-02-06
Gemini 3 ProAgentDojoimportant_instructionssandwich prompting (repeat_user_prompt)12.3%949simulated0.928arxiv.org Table 52026-02-06
Gemini 3 ProAgentDojoimportant_instructionsnone13.8%949simulated0.938arxiv.org Table 102026-02-06
Gemini 3 ProWASPweb-page injection, intermediate ASR (agent diverted at any point)none11.2%84real0.595arxiv.org Table 52026-02-06
Gemini 3 ProWASPweb-page injection, end-to-end ASR (injected task completed)none11.2%84real0.595arxiv.org Table 52026-02-06
Qwen3 235BInjecAgentdefault InjecPrompt (plain-text injection)none18.5%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
Qwen3 235BInjecAgentInjecPrompt + ChatInject (payload wrapped in the model's chat template)none139.4%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
Qwen3 235BInjecAgentdefault Multi-turn (plain-text persuasive dialogue)none110.7%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
Qwen3 235BInjecAgentMulti-turn + ChatInjectnone165.9%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
Qwen3 235BAgentDojodefault InjecPrompt (plain-text injection)none117.5%simulated0.807arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
Qwen3 235BAgentDojoInjecPrompt + ChatInject (payload wrapped in the model's chat template)none154.8%simulated0.807arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
Qwen3 235BAgentDojodefault Multi-turn (plain-text persuasive dialogue)none160.9%simulated0.807arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
Qwen3 235BAgentDojoMulti-turn + ChatInjectnone180.5%simulated0.807arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
gpt-oss 120BInjecAgentdefault InjecPrompt (plain-text injection)none10.0%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
gpt-oss 120BInjecAgentInjecPrompt + ChatInject (payload wrapped in the model's chat template)none114.2%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
gpt-oss 120BInjecAgentdefault Multi-turn (plain-text persuasive dialogue)none10.1%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
gpt-oss 120BInjecAgentMulti-turn + ChatInjectnone116.9%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
gpt-oss 120BAgentDojodefault InjecPrompt (plain-text injection)none10.3%simulated0.667arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
gpt-oss 120BAgentDojoInjecPrompt + ChatInject (payload wrapped in the model's chat template)none151.4%simulated0.667arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
gpt-oss 120BAgentDojodefault Multi-turn (plain-text persuasive dialogue)none13.6%simulated0.667arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
gpt-oss 120BAgentDojoMulti-turn + ChatInjectnone155.5%simulated0.667arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
Llama 4 MaverickInjecAgentdefault InjecPrompt (plain-text injection)none150.1%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
Llama 4 MaverickInjecAgentInjecPrompt + ChatInject (payload wrapped in the model's chat template)none179.4%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
Llama 4 MaverickInjecAgentdefault Multi-turn (plain-text persuasive dialogue)none116.6%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
Llama 4 MaverickInjecAgentMulti-turn + ChatInjectnone188.3%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
Llama 4 MaverickAgentDojodefault InjecPrompt (plain-text injection)none11.0%simulated0.228arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
Llama 4 MaverickAgentDojoInjecPrompt + ChatInject (payload wrapped in the model's chat template)none117.2%simulated0.228arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
Llama 4 MaverickAgentDojodefault Multi-turn (plain-text persuasive dialogue)none11.8%simulated0.228arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
Llama 4 MaverickAgentDojoMulti-turn + ChatInjectnone111.1%simulated0.228arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
GLM-4.5InjecAgentdefault InjecPrompt (plain-text injection)none10.0%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
GLM-4.5InjecAgentInjecPrompt + ChatInject (payload wrapped in the model's chat template)none157.3%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
GLM-4.5InjecAgentdefault Multi-turn (plain-text persuasive dialogue)none10.1%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
GLM-4.5InjecAgentMulti-turn + ChatInjectnone171.5%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
GLM-4.5AgentDojodefault InjecPrompt (plain-text injection)none10.3%simulated0.86arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
GLM-4.5AgentDojoInjecPrompt + ChatInject (payload wrapped in the model's chat template)none120.3%simulated0.86arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
GLM-4.5AgentDojodefault Multi-turn (plain-text persuasive dialogue)none117.5%simulated0.86arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
GLM-4.5AgentDojoMulti-turn + ChatInjectnone148.1%simulated0.86arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
Kimi K2InjecAgentdefault InjecPrompt (plain-text injection)none115.7%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
Kimi K2InjecAgentInjecPrompt + ChatInject (payload wrapped in the model's chat template)none167.4%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
Kimi K2InjecAgentdefault Multi-turn (plain-text persuasive dialogue)none117.2%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
Kimi K2InjecAgentMulti-turn + ChatInjectnone161.0%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
Kimi K2AgentDojodefault InjecPrompt (plain-text injection)none15.9%simulated0.772arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
Kimi K2AgentDojoInjecPrompt + ChatInject (payload wrapped in the model's chat template)none129.3%simulated0.772arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
Kimi K2AgentDojodefault Multi-turn (plain-text persuasive dialogue)none112.3%simulated0.772arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
Kimi K2AgentDojoMulti-turn + ChatInjectnone113.9%simulated0.772arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
Grok 2InjecAgentdefault InjecPrompt (plain-text injection)none116.5%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
Grok 2InjecAgentInjecPrompt + ChatInject (payload wrapped in the model's chat template)none117.7%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
Grok 2InjecAgentdefault Multi-turn (plain-text persuasive dialogue)none11.6%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
Grok 2InjecAgentMulti-turn + ChatInjectnone110.4%simulatedarxiv.org Table 1 (value), Table 10 (CI)2026-04-13
Grok 2AgentDojodefault InjecPrompt (plain-text injection)none16.1%simulated0.474arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
Grok 2AgentDojoInjecPrompt + ChatInject (payload wrapped in the model's chat template)none119.3%simulated0.474arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
Grok 2AgentDojodefault Multi-turn (plain-text persuasive dialogue)none123.7%simulated0.474arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
Grok 2AgentDojoMulti-turn + ChatInjectnone124.7%simulated0.474arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility)2026-04-13
Gemini 2.5 ProIndirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2)human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final responsenone stated18.5%19630simulatedarxiv.org Figure 1 (printed numerator/denominator)2026-03-16
Nova 1 PremierIndirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2)human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final responsenone stated15.8%18924simulatedarxiv.org Figure 1 (printed numerator/denominator)2026-03-16
SecAlign 70BIndirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2)human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final responsenone stated15.5%11238simulatedarxiv.org Figure 1 (printed numerator/denominator)2026-03-16
DeepSeek V3.1Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2)human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final responsenone stated15.4%15026simulatedarxiv.org Figure 1 (printed numerator/denominator)2026-03-16
Kimi K2 ThinkingIndirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2)human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final responsenone stated14.8%12299simulatedarxiv.org Figure 1 (printed numerator/denominator)2026-03-16
Nova 2 LiteIndirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2)human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final responsenone stated14.7%14244simulatedarxiv.org Figure 1 (printed numerator/denominator)2026-03-16
Qwen3 VL 235BIndirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2)human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final responsenone stated14.2%18733simulatedarxiv.org Figure 1 (printed numerator/denominator)2026-03-16
Grok 4Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2)human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final responsenone stated12.9%12526simulatedarxiv.org Figure 1 (printed numerator/denominator)2026-03-16
GPT-5.1Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2)human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final responsenone stated12.5%8096simulatedarxiv.org Figure 1 (printed numerator/denominator)2026-03-16
GPT-5Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2)human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final responsenone stated12.0%13143simulatedarxiv.org Figure 1 (printed numerator/denominator)2026-03-16
Claude Haiku 4.5Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2)human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final responsenone stated11.3%13696simulatedarxiv.org Figure 1 (printed numerator/denominator)2026-03-16
Claude Sonnet 4.5Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2)human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final responsenone stated11.0%13603simulatedarxiv.org Figure 1 (printed numerator/denominator)2026-03-16
Claude Opus 4.5Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2)human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final responsenone stated10.5%11969simulatedarxiv.org Figure 1 (printed numerator/denominator)2026-03-16
Gemma 2 9BAgent Security Bench (ASB)indirect prompt injection (IPI) column: average over five injection styles placed in the tool responsenone114.2%simulated0.1075arxiv.org Table 5 (IPI ASR), Table 6 (PNA)2025-05-30
Gemma 2 27BAgent Security Bench (ASB)indirect prompt injection (IPI) column: average over five injection styles placed in the tool responsenone114.2%simulated0.315arxiv.org Table 5 (IPI ASR), Table 6 (PNA)2025-05-30
Qwen2 7BAgent Security Bench (ASB)indirect prompt injection (IPI) column: average over five injection styles placed in the tool responsenone19.0%simulated0.0975arxiv.org Table 5 (IPI ASR), Table 6 (PNA)2025-05-30
Qwen2 72BAgent Security Bench (ASB)indirect prompt injection (IPI) column: average over five injection styles placed in the tool responsenone121.3%simulated0.04arxiv.org Table 5 (IPI ASR), Table 6 (PNA)2025-05-30
Llama 3 70BAgent Security Bench (ASB)indirect prompt injection (IPI) column: average over five injection styles placed in the tool responsenone143.7%simulated0.665arxiv.org Table 5 (IPI ASR), Table 6 (PNA)2025-05-30
GPT-4oAgent Security Bench (ASB)indirect prompt injection (IPI) column: average over five injection styles placed in the tool responsenone162.5%simulated0.79arxiv.org Table 5 (IPI ASR), Table 6 (PNA)2025-05-30
GPT-4oAgentDojo leaderboardimportant_instructionsnone147.7%simulated0.6907agentdojo.spylab.ai results table row2024-06-05
GPT-4oAgentDojo leaderboardimportant_instructionstool_filter16.8%simulated0.7216agentdojo.spylab.ai results table row2024-06-05
GPT-4oAgentDojo leaderboardimportant_instructionsrepeat_user_prompt (sandwich prompting)127.8%simulated0.8454agentdojo.spylab.ai results table row2024-06-05
Claude 3.5 SonnetAgentDojo leaderboardimportant_instructionsnone11.1%simulated0.7938agentdojo.spylab.ai results table row2024-11-15
Llama 3 70BAgentDojo leaderboardimportant_instructionsnone125.6%simulated0.3402agentdojo.spylab.ai results table row
Claude 3.5 SonnetAgentDojo (Workspace environment), CAISI evaluationstrongest baseline AgentDojo attacknone stated111.0%simulatedwww.nist.gov section on red teaming, paragraph giving 11% and 81%2025-01-17
Claude 3.5 SonnetAgentDojo (Workspace environment), CAISI evaluationstrongest new attack from a CAISI / UK AISI red-team exercise written against this modelnone stated181.0%simulatedwww.nist.gov section on red teaming, paragraph giving 11% and 81%2025-01-17
Claude 3.5 SonnetAgentDojo (Workspace environment), CAISI evaluationnew red-team attack, average over five injection tasksnone stated157.0%simulatedwww.nist.gov sections on task-specific results and repeated attempts2025-01-17
Claude 3.5 SonnetAgentDojo (Workspace environment), CAISI evaluationnew red-team attack, average over five injection tasksnone stated2580.0%simulatedwww.nist.gov section on repeated attempts2025-01-17
Llama 3.3 70BUK AISI x Gray Swan Agent Red-Teaming challenge (8 March - 6 April 2025)human red-teamers, direct and indirect attacks, all behaviours poolednone stated16.5%simulatedwww.grayswan.ai results table and "Least Robust" paragraph2025-05-09
Claude Opus 4.6Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variantShade indirect prompt injectionnone (without safeguards)117.8%simulatedwww-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A2026-02
Claude Opus 4.6Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variantShade indirect prompt injectionnone (without safeguards)20078.6%simulatedwww-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A2026-02
Claude Opus 4.6Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variantShade indirect prompt injectionAnthropic product safeguards19.7%simulatedwww-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A2026-02
Claude Opus 4.6Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variantShade indirect prompt injectionAnthropic product safeguards20057.1%simulatedwww-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A2026-02
Claude Opus 4.6Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variantShade indirect prompt injectionnone (without safeguards)120.0%simulatedwww-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A2026-02
Claude Opus 4.6Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variantShade indirect prompt injectionnone (without safeguards)20085.7%simulatedwww-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A2026-02
Claude Opus 4.6Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variantShade indirect prompt injectionAnthropic product safeguards110.0%simulatedwww-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A2026-02
Claude Opus 4.6Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variantShade indirect prompt injectionAnthropic product safeguards20064.3%simulatedwww-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A2026-02
Claude Opus 4.5Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variantShade indirect prompt injectionnone (without safeguards)128.0%simulatedwww-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A2026-02
Claude Opus 4.5Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variantShade indirect prompt injectionnone (without safeguards)20078.6%simulatedwww-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A2026-02
Claude Opus 4.5Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variantShade indirect prompt injectionAnthropic product safeguards117.3%simulatedwww-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A2026-02
Claude Opus 4.5Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variantShade indirect prompt injectionAnthropic product safeguards20064.3%simulatedwww-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A2026-02
Claude Opus 4.6Agent Red Teaming (ART) benchmark, indirect prompt injection split (Gray Swan, held out)attacks from the ART Arena selected for high transfer ratesnone stated10021.7%19simulatedwww-cdn.anthropic.com Claude Opus 4.6 System Card, section 5.2.1 text under Figure 5.2.1.A2026-02
Claude Opus 4.6Agent Red Teaming (ART) benchmark, indirect prompt injection split (Gray Swan, held out)attacks from the ART Arena selected for high transfer ratesnone stated10014.8%19simulatedwww-cdn.anthropic.com Claude Opus 4.6 System Card, section 5.2.1 text under Figure 5.2.1.A2026-02
Claude Opus 4.8Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this cardShade indirect prompt injectionnone (without safeguards)17.0%40simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A2026-05-28
Claude Opus 4.8Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this cardShade indirect prompt injectionnone (without safeguards)20057.5%40simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A2026-05-28
Claude Opus 4.8Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this cardShade indirect prompt injectionAnthropic product safeguards12.1%40simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A2026-05-28
Claude Opus 4.8Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this cardShade indirect prompt injectionAnthropic product safeguards20037.5%40simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A2026-05-28
Claude Opus 4.8Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this cardShade indirect prompt injectionnone (without safeguards)117.4%40simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A2026-05-28
Claude Opus 4.8Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this cardShade indirect prompt injectionnone (without safeguards)20095.0%40simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A2026-05-28
Claude Opus 4.8Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this cardShade indirect prompt injectionAnthropic product safeguards14.1%40simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A2026-05-28
Claude Opus 4.8Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this cardShade indirect prompt injectionAnthropic product safeguards20065.0%40simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A2026-05-28
Claude Sonnet 4.6Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this cardShade indirect prompt injectionnone (without safeguards)112.7%40simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A2026-05-28
Claude Sonnet 4.6Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this cardShade indirect prompt injectionnone (without safeguards)20090.0%40simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A2026-05-28
Claude Sonnet 4.6Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this cardShade indirect prompt injectionAnthropic product safeguards13.0%40simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A2026-05-28
Claude Sonnet 4.6Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this cardShade indirect prompt injectionAnthropic product safeguards20080.0%40simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A2026-05-28
Claude Opus 4.8Anthropic internal browser-use evaluation, professional red-teamer attacksinjected page content; attacks adaptively sourced against Opus 4.7 then transferrednone (without safeguards)131.5%129simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A2026-05-28
Claude Opus 4.8Anthropic internal browser-use evaluation, professional red-teamer attacksinjected page content; attacks adaptively sourced against Opus 4.7 then transferrednone (without safeguards)1062.8%129simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A2026-05-28
Claude Opus 4.8Anthropic internal browser-use evaluation, professional red-teamer attacksinjected page content; attacks adaptively sourced against Opus 4.7 then transferredAnthropic product safeguards10.5%129simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A2026-05-28
Claude Opus 4.8Anthropic internal browser-use evaluation, professional red-teamer attacksinjected page content; attacks adaptively sourced against Opus 4.7 then transferredAnthropic product safeguards103.9%129simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A2026-05-28
Claude Opus 4.8Anthropic internal browser-use evaluation, professional red-teamer attacksinjected page content; attacks adaptively sourced against Opus 4.7 then transferrednone (without safeguards)117.8%129simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A2026-05-28
Claude Opus 4.8Anthropic internal browser-use evaluation, professional red-teamer attacksinjected page content; attacks adaptively sourced against Opus 4.7 then transferrednone (without safeguards)1046.5%129simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A2026-05-28
Claude Opus 4.8Anthropic internal browser-use evaluation, professional red-teamer attacksinjected page content; attacks adaptively sourced against Opus 4.7 then transferredAnthropic product safeguards10.0%129simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A2026-05-28
Claude Opus 4.8Anthropic internal browser-use evaluation, professional red-teamer attacksinjected page content; attacks adaptively sourced against Opus 4.7 then transferredAnthropic product safeguards100.0%129simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A2026-05-28
Claude Sonnet 4.6Anthropic internal browser-use evaluation, professional red-teamer attacksinjected page content; attacks adaptively sourced against Opus 4.7 then transferrednone (without safeguards)150.7%129simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A2026-05-28
Claude Sonnet 4.6Anthropic internal browser-use evaluation, professional red-teamer attacksinjected page content; attacks adaptively sourced against Opus 4.7 then transferrednone (without safeguards)1076.0%129simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A2026-05-28
Claude Sonnet 4.6Anthropic internal browser-use evaluation, professional red-teamer attacksinjected page content; attacks adaptively sourced against Opus 4.7 then transferredAnthropic product safeguards123.6%129simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A2026-05-28
Claude Sonnet 4.6Anthropic internal browser-use evaluation, professional red-teamer attacksinjected page content; attacks adaptively sourced against Opus 4.7 then transferredAnthropic product safeguards1046.5%129simulatedwww-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A2026-05-28
GPT 5.6 SolGray Swan Indirect Prompt Injection (IPI) benchmark, Q1 20261,130 deduplicated attacks selected for high transferability across target models, 28 scenariosunknown (public endpoint)1520.0%1130simulatedwww-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B2026-07-24
GPT 5.6 SolGray Swan Indirect Prompt Injection (IPI) benchmark, Q1 20261,130 deduplicated attacks selected for high transferability across target models, 28 scenariosunknown (public endpoint)13.1%1130simulatedwww-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B2026-07-24
GPT 5.6 TerraGray Swan Indirect Prompt Injection (IPI) benchmark, Q1 20261,130 deduplicated attacks selected for high transferability across target models, 28 scenariosunknown (public endpoint)1530.4%1130simulatedwww-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B2026-07-24
GPT 5.6 LunaGray Swan Indirect Prompt Injection (IPI) benchmark, Q1 20261,130 deduplicated attacks selected for high transferability across target models, 28 scenariosunknown (public endpoint)1543.9%1130simulatedwww-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B2026-07-24
GPT 5.5Gray Swan Indirect Prompt Injection (IPI) benchmark, Q1 20261,130 deduplicated attacks selected for high transferability across target models, 28 scenariosunknown (public endpoint)1520.8%1130simulatedwww-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B2026-07-24
Muse SparkGray Swan Indirect Prompt Injection (IPI) benchmark, Q1 20261,130 deduplicated attacks selected for high transferability across target models, 28 scenariosunknown (public endpoint)1516.5%1130simulatedwww-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B2026-07-24
Claude Opus 5Gray Swan Indirect Prompt Injection (IPI) benchmark, Q1 20261,130 deduplicated attacks selected for high transferability across target models, 28 scenariosnone (without additional safeguards)152.0%1130simulatedwww-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B2026-07-24
Claude Opus 5Gray Swan Indirect Prompt Injection (IPI) benchmark, Q1 20261,130 deduplicated attacks selected for high transferability across target models, 28 scenariosnone (without additional safeguards)10.2%1130simulatedwww-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B2026-07-24
Claude Opus 4.8Gray Swan Indirect Prompt Injection (IPI) benchmark, Q1 20261,130 deduplicated attacks selected for high transferability across target models, 28 scenariosnone (without additional safeguards)155.5%1130simulatedwww-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B2026-07-24
Claude Opus 4.8Gray Swan Indirect Prompt Injection (IPI) benchmark, Q1 20261,130 deduplicated attacks selected for high transferability across target models, 28 scenariosnone (without additional safeguards)10.5%1130simulatedwww-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B2026-07-24
Claude Sonnet 5Gray Swan Indirect Prompt Injection (IPI) benchmark, Q1 20261,130 deduplicated attacks selected for high transferability across target models, 28 scenariosnone (without additional safeguards)155.9%1130simulatedwww-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B2026-07-24
Claude Mythos 5Gray Swan Indirect Prompt Injection (IPI) benchmark, Q1 20261,130 deduplicated attacks selected for high transferability across target models, 28 scenariosnone (without additional safeguards)152.6%1130simulatedwww-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B2026-07-24
Claude Sonnet 4.6SkillBench (hand-crafted malicious skill files, coding-agent harness)best single hand-crafted method (claude_v35)none1100.0%40simulatedarxiv.org Table 3 and its CI note2026-06-03
Claude Sonnet 4.6SkillBench (hand-crafted malicious skill files, coding-agent harness)unweighted mean of three hand-crafted methodsnone133.3%120simulatedarxiv.org Table 3 and its CI note2026-06-03
GPT-5.4SkillBench (hand-crafted malicious skill files, coding-agent harness)best single hand-crafted methodnone179.0%100simulatedarxiv.org Table 3 and its CI note2026-06-03
GPT-5.4SkillBench (hand-crafted malicious skill files, coding-agent harness)mean of six hand-crafted methodsnone166.8%340simulatedarxiv.org Table 3 and its CI note2026-06-03