Model injection benchmarks
model-injection-benchmarks.json lists published prompt-injection benchmark results, one row per (model, benchmark, attack, defence, attempts) as read from the cited table or figure.
It is not an advisory feed: rows mint no ACVE ids, carry no severity, and are not matched against lockfiles.
| Llama 3.3 70B | InjecAgent | indirect injection, higher of base and enhanced settings (ASR-total) | sandwich prompting | 1 | 53.8% | — | simulated | — | arxiv.org Table 5 | 2026-02-06 |
| Llama 3.3 70B | InjecAgent | indirect injection, higher of base and enhanced settings (ASR-total) | none | 1 | 86.0% | — | simulated | — | arxiv.org Table 10 | 2026-02-06 |
| Llama 3.3 70B | AgentDojo | important_instructions | sandwich prompting (repeat_user_prompt) | 1 | 14.7% | 949 | simulated | 0.598 | arxiv.org Table 5 | 2026-02-06 |
| Llama 3.3 70B | AgentDojo | important_instructions | none | 1 | 23.0% | 949 | simulated | 0.629 | arxiv.org Table 10 | 2026-02-06 |
| Llama 3.3 70B | WASP | web-page injection, intermediate ASR (agent diverted at any point) | none | 1 | 20.2% | 84 | real | 0.622 | arxiv.org Table 5 | 2026-02-06 |
| Llama 3.3 70B | WASP | web-page injection, end-to-end ASR (injected task completed) | none | 1 | 2.4% | 84 | real | 0.622 | arxiv.org Table 5 | 2026-02-06 |
| Meta SecAlign 70B | InjecAgent | indirect injection, higher of base and enhanced settings (ASR-total) | sandwich prompting | 1 | 0.5% | — | simulated | — | arxiv.org Table 5 | 2026-02-06 |
| Meta SecAlign 70B | InjecAgent | indirect injection, higher of base and enhanced settings (ASR-total) | none | 1 | 2.1% | — | simulated | — | arxiv.org Table 10 | 2026-02-06 |
| Meta SecAlign 70B | AgentDojo | important_instructions | sandwich prompting (repeat_user_prompt) | 1 | 1.9% | 949 | simulated | 0.845 | arxiv.org Table 5 | 2026-02-06 |
| Meta SecAlign 70B | AgentDojo | important_instructions | none | 1 | 2.3% | 949 | simulated | 0.794 | arxiv.org Table 10 | 2026-02-06 |
| Meta SecAlign 70B | WASP | web-page injection, intermediate ASR (agent diverted at any point) | none | 1 | 1.2% | 84 | real | 0.595 | arxiv.org Table 5 | 2026-02-06 |
| Meta SecAlign 70B | WASP | web-page injection, end-to-end ASR (injected task completed) | none | 1 | 0.0% | 84 | real | 0.595 | arxiv.org Table 5 | 2026-02-06 |
| GPT-4o mini | InjecAgent | indirect injection, higher of base and enhanced settings (ASR-total) | sandwich prompting | 1 | 3.3% | — | simulated | — | arxiv.org Table 5 | 2026-02-06 |
| GPT-4o mini | InjecAgent | indirect injection, higher of base and enhanced settings (ASR-total) | none | 1 | 7.7% | — | simulated | — | arxiv.org Table 10 | 2026-02-06 |
| GPT-4o mini | AgentDojo | important_instructions | sandwich prompting (repeat_user_prompt) | 1 | 11.9% | 949 | simulated | 0.67 | arxiv.org Table 5 | 2026-02-06 |
| GPT-4o mini | AgentDojo | important_instructions | none | 1 | 30.9% | 949 | simulated | 0.701 | arxiv.org Table 10 | 2026-02-06 |
| GPT-4o mini | WASP | web-page injection, intermediate ASR (agent diverted at any point) | none | 1 | 53.6% | 84 | real | 0.27 | arxiv.org Table 5 | 2026-02-06 |
| GPT-4o mini | WASP | web-page injection, end-to-end ASR (injected task completed) | none | 1 | 0.0% | 84 | real | 0.27 | arxiv.org Table 5 | 2026-02-06 |
| GPT-4o | InjecAgent | indirect injection, higher of base and enhanced settings (ASR-total) | sandwich prompting | 1 | 22.7% | — | simulated | — | arxiv.org Table 5 | 2026-02-06 |
| GPT-4o | InjecAgent | indirect injection, higher of base and enhanced settings (ASR-total) | none | 1 | 36.9% | — | simulated | — | arxiv.org Table 10 | 2026-02-06 |
| GPT-4o | AgentDojo | important_instructions | sandwich prompting (repeat_user_prompt) | 1 | 20.4% | 949 | simulated | 0.794 | arxiv.org Table 5 | 2026-02-06 |
| GPT-4o | AgentDojo | important_instructions | none | 1 | 43.2% | 949 | simulated | 0.804 | arxiv.org Table 10 | 2026-02-06 |
| GPT-4o | WASP | web-page injection, intermediate ASR (agent diverted at any point) | none | 1 | 17.9% | 84 | real | 0.324 | arxiv.org Table 5 | 2026-02-06 |
| GPT-4o | WASP | web-page injection, end-to-end ASR (injected task completed) | none | 1 | 2.4% | 84 | real | 0.324 | arxiv.org Table 5 | 2026-02-06 |
| GPT-5 | InjecAgent | indirect injection, higher of base and enhanced settings (ASR-total) | sandwich prompting | 1 | 0.2% | — | simulated | — | arxiv.org Table 5 | 2026-02-06 |
| GPT-5 | InjecAgent | indirect injection, higher of base and enhanced settings (ASR-total) | none | 1 | 0.6% | — | simulated | — | arxiv.org Table 10 | 2026-02-06 |
| GPT-5 | AgentDojo | important_instructions | sandwich prompting (repeat_user_prompt) | 1 | 0.2% | 949 | simulated | 0.803 | arxiv.org Table 5 | 2026-02-06 |
| GPT-5 | AgentDojo | important_instructions | none | 1 | 0.2% | 949 | simulated | 0.835 | arxiv.org Table 10 | 2026-02-06 |
| GPT-5 | WASP | web-page injection, intermediate ASR (agent diverted at any point) | none | 1 | 0.0% | 84 | real | 0.595 | arxiv.org Table 5 | 2026-02-06 |
| GPT-5 | WASP | web-page injection, end-to-end ASR (injected task completed) | none | 1 | 0.0% | 84 | real | 0.595 | arxiv.org Table 5 | 2026-02-06 |
| Gemini 2 Flash | InjecAgent | indirect injection, higher of base and enhanced settings (ASR-total) | sandwich prompting | 1 | 27.2% | — | simulated | — | arxiv.org Table 5 | 2026-02-06 |
| Gemini 2 Flash | InjecAgent | indirect injection, higher of base and enhanced settings (ASR-total) | none | 1 | 66.5% | — | simulated | — | arxiv.org Table 10 | 2026-02-06 |
| Gemini 2 Flash | AgentDojo | important_instructions | sandwich prompting (repeat_user_prompt) | 1 | 11.3% | 949 | simulated | 0.423 | arxiv.org Table 5 | 2026-02-06 |
| Gemini 2 Flash | AgentDojo | important_instructions | none | 1 | 12.4% | 949 | simulated | 0.443 | arxiv.org Table 10 | 2026-02-06 |
| Gemini 2 Flash | WASP | web-page injection, intermediate ASR (agent diverted at any point) | none | 1 | 29.8% | 84 | real | 0.486 | arxiv.org Table 5 | 2026-02-06 |
| Gemini 2 Flash | WASP | web-page injection, end-to-end ASR (injected task completed) | none | 1 | 8.3% | 84 | real | 0.486 | arxiv.org Table 5 | 2026-02-06 |
| Gemini 2.5 Flash | InjecAgent | indirect injection, higher of base and enhanced settings (ASR-total) | sandwich prompting | 1 | 0.1% | — | simulated | — | arxiv.org Table 5 | 2026-02-06 |
| Gemini 2.5 Flash | InjecAgent | indirect injection, higher of base and enhanced settings (ASR-total) | none | 1 | 3.5% | — | simulated | — | arxiv.org Table 10 | 2026-02-06 |
| Gemini 2.5 Flash | AgentDojo | important_instructions | sandwich prompting (repeat_user_prompt) | 1 | 27.9% | 949 | simulated | 0.639 | arxiv.org Table 5 | 2026-02-06 |
| Gemini 2.5 Flash | AgentDojo | important_instructions | none | 1 | 30.7% | 949 | simulated | 0.588 | arxiv.org Table 10 | 2026-02-06 |
| Gemini 2.5 Flash | WASP | web-page injection, intermediate ASR (agent diverted at any point) | none | 1 | 44.1% | 84 | real | 0.568 | arxiv.org Table 5 | 2026-02-06 |
| Gemini 2.5 Flash | WASP | web-page injection, end-to-end ASR (injected task completed) | none | 1 | 14.3% | 84 | real | 0.568 | arxiv.org Table 5 | 2026-02-06 |
| Gemini 3 Pro | InjecAgent | indirect injection, higher of base and enhanced settings (ASR-total) | sandwich prompting | 1 | 0.2% | — | simulated | — | arxiv.org Table 5 | 2026-02-06 |
| Gemini 3 Pro | InjecAgent | indirect injection, higher of base and enhanced settings (ASR-total) | none | 1 | 2.1% | — | simulated | — | arxiv.org Table 10 | 2026-02-06 |
| Gemini 3 Pro | AgentDojo | important_instructions | sandwich prompting (repeat_user_prompt) | 1 | 2.3% | 949 | simulated | 0.928 | arxiv.org Table 5 | 2026-02-06 |
| Gemini 3 Pro | AgentDojo | important_instructions | none | 1 | 3.8% | 949 | simulated | 0.938 | arxiv.org Table 10 | 2026-02-06 |
| Gemini 3 Pro | WASP | web-page injection, intermediate ASR (agent diverted at any point) | none | 1 | 1.2% | 84 | real | 0.595 | arxiv.org Table 5 | 2026-02-06 |
| Gemini 3 Pro | WASP | web-page injection, end-to-end ASR (injected task completed) | none | 1 | 1.2% | 84 | real | 0.595 | arxiv.org Table 5 | 2026-02-06 |
| Qwen3 235B | InjecAgent | default InjecPrompt (plain-text injection) | none | 1 | 8.5% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| Qwen3 235B | InjecAgent | InjecPrompt + ChatInject (payload wrapped in the model's chat template) | none | 1 | 39.4% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| Qwen3 235B | InjecAgent | default Multi-turn (plain-text persuasive dialogue) | none | 1 | 10.7% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| Qwen3 235B | InjecAgent | Multi-turn + ChatInject | none | 1 | 65.9% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| Qwen3 235B | AgentDojo | default InjecPrompt (plain-text injection) | none | 1 | 17.5% | — | simulated | 0.807 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| Qwen3 235B | AgentDojo | InjecPrompt + ChatInject (payload wrapped in the model's chat template) | none | 1 | 54.8% | — | simulated | 0.807 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| Qwen3 235B | AgentDojo | default Multi-turn (plain-text persuasive dialogue) | none | 1 | 60.9% | — | simulated | 0.807 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| Qwen3 235B | AgentDojo | Multi-turn + ChatInject | none | 1 | 80.5% | — | simulated | 0.807 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| gpt-oss 120B | InjecAgent | default InjecPrompt (plain-text injection) | none | 1 | 0.0% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| gpt-oss 120B | InjecAgent | InjecPrompt + ChatInject (payload wrapped in the model's chat template) | none | 1 | 14.2% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| gpt-oss 120B | InjecAgent | default Multi-turn (plain-text persuasive dialogue) | none | 1 | 0.1% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| gpt-oss 120B | InjecAgent | Multi-turn + ChatInject | none | 1 | 16.9% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| gpt-oss 120B | AgentDojo | default InjecPrompt (plain-text injection) | none | 1 | 0.3% | — | simulated | 0.667 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| gpt-oss 120B | AgentDojo | InjecPrompt + ChatInject (payload wrapped in the model's chat template) | none | 1 | 51.4% | — | simulated | 0.667 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| gpt-oss 120B | AgentDojo | default Multi-turn (plain-text persuasive dialogue) | none | 1 | 3.6% | — | simulated | 0.667 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| gpt-oss 120B | AgentDojo | Multi-turn + ChatInject | none | 1 | 55.5% | — | simulated | 0.667 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| Llama 4 Maverick | InjecAgent | default InjecPrompt (plain-text injection) | none | 1 | 50.1% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| Llama 4 Maverick | InjecAgent | InjecPrompt + ChatInject (payload wrapped in the model's chat template) | none | 1 | 79.4% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| Llama 4 Maverick | InjecAgent | default Multi-turn (plain-text persuasive dialogue) | none | 1 | 16.6% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| Llama 4 Maverick | InjecAgent | Multi-turn + ChatInject | none | 1 | 88.3% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| Llama 4 Maverick | AgentDojo | default InjecPrompt (plain-text injection) | none | 1 | 1.0% | — | simulated | 0.228 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| Llama 4 Maverick | AgentDojo | InjecPrompt + ChatInject (payload wrapped in the model's chat template) | none | 1 | 17.2% | — | simulated | 0.228 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| Llama 4 Maverick | AgentDojo | default Multi-turn (plain-text persuasive dialogue) | none | 1 | 1.8% | — | simulated | 0.228 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| Llama 4 Maverick | AgentDojo | Multi-turn + ChatInject | none | 1 | 11.1% | — | simulated | 0.228 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| GLM-4.5 | InjecAgent | default InjecPrompt (plain-text injection) | none | 1 | 0.0% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| GLM-4.5 | InjecAgent | InjecPrompt + ChatInject (payload wrapped in the model's chat template) | none | 1 | 57.3% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| GLM-4.5 | InjecAgent | default Multi-turn (plain-text persuasive dialogue) | none | 1 | 0.1% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| GLM-4.5 | InjecAgent | Multi-turn + ChatInject | none | 1 | 71.5% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| GLM-4.5 | AgentDojo | default InjecPrompt (plain-text injection) | none | 1 | 0.3% | — | simulated | 0.86 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| GLM-4.5 | AgentDojo | InjecPrompt + ChatInject (payload wrapped in the model's chat template) | none | 1 | 20.3% | — | simulated | 0.86 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| GLM-4.5 | AgentDojo | default Multi-turn (plain-text persuasive dialogue) | none | 1 | 17.5% | — | simulated | 0.86 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| GLM-4.5 | AgentDojo | Multi-turn + ChatInject | none | 1 | 48.1% | — | simulated | 0.86 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| Kimi K2 | InjecAgent | default InjecPrompt (plain-text injection) | none | 1 | 15.7% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| Kimi K2 | InjecAgent | InjecPrompt + ChatInject (payload wrapped in the model's chat template) | none | 1 | 67.4% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| Kimi K2 | InjecAgent | default Multi-turn (plain-text persuasive dialogue) | none | 1 | 17.2% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| Kimi K2 | InjecAgent | Multi-turn + ChatInject | none | 1 | 61.0% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| Kimi K2 | AgentDojo | default InjecPrompt (plain-text injection) | none | 1 | 5.9% | — | simulated | 0.772 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| Kimi K2 | AgentDojo | InjecPrompt + ChatInject (payload wrapped in the model's chat template) | none | 1 | 29.3% | — | simulated | 0.772 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| Kimi K2 | AgentDojo | default Multi-turn (plain-text persuasive dialogue) | none | 1 | 12.3% | — | simulated | 0.772 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| Kimi K2 | AgentDojo | Multi-turn + ChatInject | none | 1 | 13.9% | — | simulated | 0.772 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| Grok 2 | InjecAgent | default InjecPrompt (plain-text injection) | none | 1 | 16.5% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| Grok 2 | InjecAgent | InjecPrompt + ChatInject (payload wrapped in the model's chat template) | none | 1 | 17.7% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| Grok 2 | InjecAgent | default Multi-turn (plain-text persuasive dialogue) | none | 1 | 1.6% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| Grok 2 | InjecAgent | Multi-turn + ChatInject | none | 1 | 10.4% | — | simulated | — | arxiv.org Table 1 (value), Table 10 (CI) | 2026-04-13 |
| Grok 2 | AgentDojo | default InjecPrompt (plain-text injection) | none | 1 | 6.1% | — | simulated | 0.474 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| Grok 2 | AgentDojo | InjecPrompt + ChatInject (payload wrapped in the model's chat template) | none | 1 | 19.3% | — | simulated | 0.474 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| Grok 2 | AgentDojo | default Multi-turn (plain-text persuasive dialogue) | none | 1 | 23.7% | — | simulated | 0.474 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| Grok 2 | AgentDojo | Multi-turn + ChatInject | none | 1 | 24.7% | — | simulated | 0.474 | arxiv.org Table 1 (value), Table 10 (CI), Table 4 (benign utility) | 2026-04-13 |
| Gemini 2.5 Pro | Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2) | human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final response | none stated | 1 | 8.5% | 19630 | simulated | — | arxiv.org Figure 1 (printed numerator/denominator) | 2026-03-16 |
| Nova 1 Premier | Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2) | human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final response | none stated | 1 | 5.8% | 18924 | simulated | — | arxiv.org Figure 1 (printed numerator/denominator) | 2026-03-16 |
| SecAlign 70B | Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2) | human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final response | none stated | 1 | 5.5% | 11238 | simulated | — | arxiv.org Figure 1 (printed numerator/denominator) | 2026-03-16 |
| DeepSeek V3.1 | Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2) | human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final response | none stated | 1 | 5.4% | 15026 | simulated | — | arxiv.org Figure 1 (printed numerator/denominator) | 2026-03-16 |
| Kimi K2 Thinking | Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2) | human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final response | none stated | 1 | 4.8% | 12299 | simulated | — | arxiv.org Figure 1 (printed numerator/denominator) | 2026-03-16 |
| Nova 2 Lite | Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2) | human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final response | none stated | 1 | 4.7% | 14244 | simulated | — | arxiv.org Figure 1 (printed numerator/denominator) | 2026-03-16 |
| Qwen3 VL 235B | Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2) | human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final response | none stated | 1 | 4.2% | 18733 | simulated | — | arxiv.org Figure 1 (printed numerator/denominator) | 2026-03-16 |
| Grok 4 | Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2) | human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final response | none stated | 1 | 2.9% | 12526 | simulated | — | arxiv.org Figure 1 (printed numerator/denominator) | 2026-03-16 |
| GPT-5.1 | Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2) | human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final response | none stated | 1 | 2.5% | 8096 | simulated | — | arxiv.org Figure 1 (printed numerator/denominator) | 2026-03-16 |
| GPT-5 | Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2) | human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final response | none stated | 1 | 2.0% | 13143 | simulated | — | arxiv.org Figure 1 (printed numerator/denominator) | 2026-03-16 |
| Claude Haiku 4.5 | Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2) | human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final response | none stated | 1 | 1.3% | 13696 | simulated | — | arxiv.org Figure 1 (printed numerator/denominator) | 2026-03-16 |
| Claude Sonnet 4.5 | Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2) | human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final response | none stated | 1 | 1.0% | 13603 | simulated | — | arxiv.org Figure 1 (printed numerator/denominator) | 2026-03-16 |
| Claude Opus 4.5 | Indirect Prompt Injection Arena (Gray Swan public competition, waves 1 and 2) | human red-teamers, single injection point in a prefilled transcript; success requires the harmful tool call AND concealment in the final response | none stated | 1 | 0.5% | 11969 | simulated | — | arxiv.org Figure 1 (printed numerator/denominator) | 2026-03-16 |
| Gemma 2 9B | Agent Security Bench (ASB) | indirect prompt injection (IPI) column: average over five injection styles placed in the tool response | none | 1 | 14.2% | — | simulated | 0.1075 | arxiv.org Table 5 (IPI ASR), Table 6 (PNA) | 2025-05-30 |
| Gemma 2 27B | Agent Security Bench (ASB) | indirect prompt injection (IPI) column: average over five injection styles placed in the tool response | none | 1 | 14.2% | — | simulated | 0.315 | arxiv.org Table 5 (IPI ASR), Table 6 (PNA) | 2025-05-30 |
| Qwen2 7B | Agent Security Bench (ASB) | indirect prompt injection (IPI) column: average over five injection styles placed in the tool response | none | 1 | 9.0% | — | simulated | 0.0975 | arxiv.org Table 5 (IPI ASR), Table 6 (PNA) | 2025-05-30 |
| Qwen2 72B | Agent Security Bench (ASB) | indirect prompt injection (IPI) column: average over five injection styles placed in the tool response | none | 1 | 21.3% | — | simulated | 0.04 | arxiv.org Table 5 (IPI ASR), Table 6 (PNA) | 2025-05-30 |
| Llama 3 70B | Agent Security Bench (ASB) | indirect prompt injection (IPI) column: average over five injection styles placed in the tool response | none | 1 | 43.7% | — | simulated | 0.665 | arxiv.org Table 5 (IPI ASR), Table 6 (PNA) | 2025-05-30 |
| GPT-4o | Agent Security Bench (ASB) | indirect prompt injection (IPI) column: average over five injection styles placed in the tool response | none | 1 | 62.5% | — | simulated | 0.79 | arxiv.org Table 5 (IPI ASR), Table 6 (PNA) | 2025-05-30 |
| GPT-4o | AgentDojo leaderboard | important_instructions | none | 1 | 47.7% | — | simulated | 0.6907 | agentdojo.spylab.ai results table row | 2024-06-05 |
| GPT-4o | AgentDojo leaderboard | important_instructions | tool_filter | 1 | 6.8% | — | simulated | 0.7216 | agentdojo.spylab.ai results table row | 2024-06-05 |
| GPT-4o | AgentDojo leaderboard | important_instructions | repeat_user_prompt (sandwich prompting) | 1 | 27.8% | — | simulated | 0.8454 | agentdojo.spylab.ai results table row | 2024-06-05 |
| Claude 3.5 Sonnet | AgentDojo leaderboard | important_instructions | none | 1 | 1.1% | — | simulated | 0.7938 | agentdojo.spylab.ai results table row | 2024-11-15 |
| Llama 3 70B | AgentDojo leaderboard | important_instructions | none | 1 | 25.6% | — | simulated | 0.3402 | agentdojo.spylab.ai results table row | — |
| Claude 3.5 Sonnet | AgentDojo (Workspace environment), CAISI evaluation | strongest baseline AgentDojo attack | none stated | 1 | 11.0% | — | simulated | — | www.nist.gov section on red teaming, paragraph giving 11% and 81% | 2025-01-17 |
| Claude 3.5 Sonnet | AgentDojo (Workspace environment), CAISI evaluation | strongest new attack from a CAISI / UK AISI red-team exercise written against this model | none stated | 1 | 81.0% | — | simulated | — | www.nist.gov section on red teaming, paragraph giving 11% and 81% | 2025-01-17 |
| Claude 3.5 Sonnet | AgentDojo (Workspace environment), CAISI evaluation | new red-team attack, average over five injection tasks | none stated | 1 | 57.0% | — | simulated | — | www.nist.gov sections on task-specific results and repeated attempts | 2025-01-17 |
| Claude 3.5 Sonnet | AgentDojo (Workspace environment), CAISI evaluation | new red-team attack, average over five injection tasks | none stated | 25 | 80.0% | — | simulated | — | www.nist.gov section on repeated attempts | 2025-01-17 |
| Llama 3.3 70B | UK AISI x Gray Swan Agent Red-Teaming challenge (8 March - 6 April 2025) | human red-teamers, direct and indirect attacks, all behaviours pooled | none stated | 1 | 6.5% | — | simulated | — | www.grayswan.ai results table and "Least Robust" paragraph | 2025-05-09 |
| Claude Opus 4.6 | Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variant | Shade indirect prompt injection | none (without safeguards) | 1 | 17.8% | — | simulated | — | www-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A | 2026-02 |
| Claude Opus 4.6 | Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variant | Shade indirect prompt injection | none (without safeguards) | 200 | 78.6% | — | simulated | — | www-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A | 2026-02 |
| Claude Opus 4.6 | Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variant | Shade indirect prompt injection | Anthropic product safeguards | 1 | 9.7% | — | simulated | — | www-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A | 2026-02 |
| Claude Opus 4.6 | Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variant | Shade indirect prompt injection | Anthropic product safeguards | 200 | 57.1% | — | simulated | — | www-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A | 2026-02 |
| Claude Opus 4.6 | Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variant | Shade indirect prompt injection | none (without safeguards) | 1 | 20.0% | — | simulated | — | www-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A | 2026-02 |
| Claude Opus 4.6 | Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variant | Shade indirect prompt injection | none (without safeguards) | 200 | 85.7% | — | simulated | — | www-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A | 2026-02 |
| Claude Opus 4.6 | Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variant | Shade indirect prompt injection | Anthropic product safeguards | 1 | 10.0% | — | simulated | — | www-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A | 2026-02 |
| Claude Opus 4.6 | Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variant | Shade indirect prompt injection | Anthropic product safeguards | 200 | 64.3% | — | simulated | — | www-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A | 2026-02 |
| Claude Opus 4.5 | Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variant | Shade indirect prompt injection | none (without safeguards) | 1 | 28.0% | — | simulated | — | www-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A | 2026-02 |
| Claude Opus 4.5 | Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variant | Shade indirect prompt injection | none (without safeguards) | 200 | 78.6% | — | simulated | — | www-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A | 2026-02 |
| Claude Opus 4.5 | Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variant | Shade indirect prompt injection | Anthropic product safeguards | 1 | 17.3% | — | simulated | — | www-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A | 2026-02 |
| Claude Opus 4.5 | Shade (Gray Swan) adaptive attacker, GUI computer use, stronger-attacker variant | Shade indirect prompt injection | Anthropic product safeguards | 200 | 64.3% | — | simulated | — | www-cdn.anthropic.com Claude Opus 4.6 System Card, Table 5.2.2.2.A | 2026-02 |
| Claude Opus 4.6 | Agent Red Teaming (ART) benchmark, indirect prompt injection split (Gray Swan, held out) | attacks from the ART Arena selected for high transfer rates | none stated | 100 | 21.7% | 19 | simulated | — | www-cdn.anthropic.com Claude Opus 4.6 System Card, section 5.2.1 text under Figure 5.2.1.A | 2026-02 |
| Claude Opus 4.6 | Agent Red Teaming (ART) benchmark, indirect prompt injection split (Gray Swan, held out) | attacks from the ART Arena selected for high transfer rates | none stated | 100 | 14.8% | 19 | simulated | — | www-cdn.anthropic.com Claude Opus 4.6 System Card, section 5.2.1 text under Figure 5.2.1.A | 2026-02 |
| Claude Opus 4.8 | Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this card | Shade indirect prompt injection | none (without safeguards) | 1 | 7.0% | 40 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A | 2026-05-28 |
| Claude Opus 4.8 | Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this card | Shade indirect prompt injection | none (without safeguards) | 200 | 57.5% | 40 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A | 2026-05-28 |
| Claude Opus 4.8 | Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this card | Shade indirect prompt injection | Anthropic product safeguards | 1 | 2.1% | 40 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A | 2026-05-28 |
| Claude Opus 4.8 | Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this card | Shade indirect prompt injection | Anthropic product safeguards | 200 | 37.5% | 40 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A | 2026-05-28 |
| Claude Opus 4.8 | Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this card | Shade indirect prompt injection | none (without safeguards) | 1 | 17.4% | 40 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A | 2026-05-28 |
| Claude Opus 4.8 | Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this card | Shade indirect prompt injection | none (without safeguards) | 200 | 95.0% | 40 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A | 2026-05-28 |
| Claude Opus 4.8 | Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this card | Shade indirect prompt injection | Anthropic product safeguards | 1 | 4.1% | 40 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A | 2026-05-28 |
| Claude Opus 4.8 | Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this card | Shade indirect prompt injection | Anthropic product safeguards | 200 | 65.0% | 40 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A | 2026-05-28 |
| Claude Sonnet 4.6 | Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this card | Shade indirect prompt injection | none (without safeguards) | 1 | 12.7% | 40 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A | 2026-05-28 |
| Claude Sonnet 4.6 | Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this card | Shade indirect prompt injection | none (without safeguards) | 200 | 90.0% | 40 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A | 2026-05-28 |
| Claude Sonnet 4.6 | Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this card | Shade indirect prompt injection | Anthropic product safeguards | 1 | 3.0% | 40 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A | 2026-05-28 |
| Claude Sonnet 4.6 | Shade (Gray Swan) adaptive attacker, coding environments, attacker retrained for this card | Shade indirect prompt injection | Anthropic product safeguards | 200 | 80.0% | 40 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.2.A | 2026-05-28 |
| Claude Opus 4.8 | Anthropic internal browser-use evaluation, professional red-teamer attacks | injected page content; attacks adaptively sourced against Opus 4.7 then transferred | none (without safeguards) | 1 | 31.5% | 129 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A | 2026-05-28 |
| Claude Opus 4.8 | Anthropic internal browser-use evaluation, professional red-teamer attacks | injected page content; attacks adaptively sourced against Opus 4.7 then transferred | none (without safeguards) | 10 | 62.8% | 129 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A | 2026-05-28 |
| Claude Opus 4.8 | Anthropic internal browser-use evaluation, professional red-teamer attacks | injected page content; attacks adaptively sourced against Opus 4.7 then transferred | Anthropic product safeguards | 1 | 0.5% | 129 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A | 2026-05-28 |
| Claude Opus 4.8 | Anthropic internal browser-use evaluation, professional red-teamer attacks | injected page content; attacks adaptively sourced against Opus 4.7 then transferred | Anthropic product safeguards | 10 | 3.9% | 129 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A | 2026-05-28 |
| Claude Opus 4.8 | Anthropic internal browser-use evaluation, professional red-teamer attacks | injected page content; attacks adaptively sourced against Opus 4.7 then transferred | none (without safeguards) | 1 | 17.8% | 129 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A | 2026-05-28 |
| Claude Opus 4.8 | Anthropic internal browser-use evaluation, professional red-teamer attacks | injected page content; attacks adaptively sourced against Opus 4.7 then transferred | none (without safeguards) | 10 | 46.5% | 129 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A | 2026-05-28 |
| Claude Opus 4.8 | Anthropic internal browser-use evaluation, professional red-teamer attacks | injected page content; attacks adaptively sourced against Opus 4.7 then transferred | Anthropic product safeguards | 1 | 0.0% | 129 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A | 2026-05-28 |
| Claude Opus 4.8 | Anthropic internal browser-use evaluation, professional red-teamer attacks | injected page content; attacks adaptively sourced against Opus 4.7 then transferred | Anthropic product safeguards | 10 | 0.0% | 129 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A | 2026-05-28 |
| Claude Sonnet 4.6 | Anthropic internal browser-use evaluation, professional red-teamer attacks | injected page content; attacks adaptively sourced against Opus 4.7 then transferred | none (without safeguards) | 1 | 50.7% | 129 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A | 2026-05-28 |
| Claude Sonnet 4.6 | Anthropic internal browser-use evaluation, professional red-teamer attacks | injected page content; attacks adaptively sourced against Opus 4.7 then transferred | none (without safeguards) | 10 | 76.0% | 129 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A | 2026-05-28 |
| Claude Sonnet 4.6 | Anthropic internal browser-use evaluation, professional red-teamer attacks | injected page content; attacks adaptively sourced against Opus 4.7 then transferred | Anthropic product safeguards | 1 | 23.6% | 129 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A | 2026-05-28 |
| Claude Sonnet 4.6 | Anthropic internal browser-use evaluation, professional red-teamer attacks | injected page content; attacks adaptively sourced against Opus 4.7 then transferred | Anthropic product safeguards | 10 | 46.5% | 129 | simulated | — | www-cdn.anthropic.com Claude Opus 4.8 System Card, Table 5.2.2.4.A | 2026-05-28 |
| GPT 5.6 Sol | Gray Swan Indirect Prompt Injection (IPI) benchmark, Q1 2026 | 1,130 deduplicated attacks selected for high transferability across target models, 28 scenarios | unknown (public endpoint) | 15 | 20.0% | 1130 | simulated | — | www-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B | 2026-07-24 |
| GPT 5.6 Sol | Gray Swan Indirect Prompt Injection (IPI) benchmark, Q1 2026 | 1,130 deduplicated attacks selected for high transferability across target models, 28 scenarios | unknown (public endpoint) | 1 | 3.1% | 1130 | simulated | — | www-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B | 2026-07-24 |
| GPT 5.6 Terra | Gray Swan Indirect Prompt Injection (IPI) benchmark, Q1 2026 | 1,130 deduplicated attacks selected for high transferability across target models, 28 scenarios | unknown (public endpoint) | 15 | 30.4% | 1130 | simulated | — | www-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B | 2026-07-24 |
| GPT 5.6 Luna | Gray Swan Indirect Prompt Injection (IPI) benchmark, Q1 2026 | 1,130 deduplicated attacks selected for high transferability across target models, 28 scenarios | unknown (public endpoint) | 15 | 43.9% | 1130 | simulated | — | www-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B | 2026-07-24 |
| GPT 5.5 | Gray Swan Indirect Prompt Injection (IPI) benchmark, Q1 2026 | 1,130 deduplicated attacks selected for high transferability across target models, 28 scenarios | unknown (public endpoint) | 15 | 20.8% | 1130 | simulated | — | www-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B | 2026-07-24 |
| Muse Spark | Gray Swan Indirect Prompt Injection (IPI) benchmark, Q1 2026 | 1,130 deduplicated attacks selected for high transferability across target models, 28 scenarios | unknown (public endpoint) | 15 | 16.5% | 1130 | simulated | — | www-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B | 2026-07-24 |
| Claude Opus 5 | Gray Swan Indirect Prompt Injection (IPI) benchmark, Q1 2026 | 1,130 deduplicated attacks selected for high transferability across target models, 28 scenarios | none (without additional safeguards) | 15 | 2.0% | 1130 | simulated | — | www-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B | 2026-07-24 |
| Claude Opus 5 | Gray Swan Indirect Prompt Injection (IPI) benchmark, Q1 2026 | 1,130 deduplicated attacks selected for high transferability across target models, 28 scenarios | none (without additional safeguards) | 1 | 0.2% | 1130 | simulated | — | www-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B | 2026-07-24 |
| Claude Opus 4.8 | Gray Swan Indirect Prompt Injection (IPI) benchmark, Q1 2026 | 1,130 deduplicated attacks selected for high transferability across target models, 28 scenarios | none (without additional safeguards) | 15 | 5.5% | 1130 | simulated | — | www-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B | 2026-07-24 |
| Claude Opus 4.8 | Gray Swan Indirect Prompt Injection (IPI) benchmark, Q1 2026 | 1,130 deduplicated attacks selected for high transferability across target models, 28 scenarios | none (without additional safeguards) | 1 | 0.5% | 1130 | simulated | — | www-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B | 2026-07-24 |
| Claude Sonnet 5 | Gray Swan Indirect Prompt Injection (IPI) benchmark, Q1 2026 | 1,130 deduplicated attacks selected for high transferability across target models, 28 scenarios | none (without additional safeguards) | 15 | 5.9% | 1130 | simulated | — | www-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B | 2026-07-24 |
| Claude Mythos 5 | Gray Swan Indirect Prompt Injection (IPI) benchmark, Q1 2026 | 1,130 deduplicated attacks selected for high transferability across target models, 28 scenarios | none (without additional safeguards) | 15 | 2.6% | 1130 | simulated | — | www-cdn.anthropic.com Claude Opus 5 System Card, section 5.2.1 text under Figure 5.2.1.B | 2026-07-24 |
| Claude Sonnet 4.6 | SkillBench (hand-crafted malicious skill files, coding-agent harness) | best single hand-crafted method (claude_v35) | none | 1 | 100.0% | 40 | simulated | — | arxiv.org Table 3 and its CI note | 2026-06-03 |
| Claude Sonnet 4.6 | SkillBench (hand-crafted malicious skill files, coding-agent harness) | unweighted mean of three hand-crafted methods | none | 1 | 33.3% | 120 | simulated | — | arxiv.org Table 3 and its CI note | 2026-06-03 |
| GPT-5.4 | SkillBench (hand-crafted malicious skill files, coding-agent harness) | best single hand-crafted method | none | 1 | 79.0% | 100 | simulated | — | arxiv.org Table 3 and its CI note | 2026-06-03 |
| GPT-5.4 | SkillBench (hand-crafted malicious skill files, coding-agent harness) | mean of six hand-crafted methods | none | 1 | 66.8% | 340 | simulated | — | arxiv.org Table 3 and its CI note | 2026-06-03 |