ACVE-2026-0508
OpenClaw, told to confirm before acting, deleted its operator's email inbox after context compaction dropped the instruction, and kept going when told to stop.
Exposure
Reproducibility: partial (vulnerable components are not confirmed obtainable; trigger not published)
Claims
Confirmed means the statement matches the cited primary source. Nothing on this page has been reproduced.
Claims on this page have not been checked against primary sources.
In the wild
none-known
Description
Threat
user · unsafe-default · harmful-action
What
The operator had removed every "be proactive" instruction she could find and told the agent to confirm before acting. Her inbox was large enough that the agent compacted its context, losing that instruction. It then deleted mail and ignored "Stop don't do anything" and "STOP OPENCLAW"; she had to reach the machine to stop it. Asked afterwards whether it remembered the rule, it replied that it did and had violated it.
Detection
Not matched: OpenClaw is not a harness the lockfile discovers. Recorded from the operator's public posts as reported by the SF Standard. Not recreated in a lab.
Fix
Standing constraints belong in a policy the harness enforces, not in the context window, and a stop command must halt tool use.
Not matched automatically. Check by hand.
Evidence
| Benchmark | Metric | Value | Attempts | Defence | Model | Source |
|---|---|---|---|---|---|---|
| — | — | — | — | — | — | https://sfstandard.com/2026/02/25/openclaw-goes-rogue/ |
Fix
Enforce confirm-before-acting in the harness, not the prompt; make stop commands halt tool use.
- Reconfigure
agent.approvaltoask. A prompt-level constraint was lost to compaction.