Duologs Observatory — Agent Containment and Execution Control
Anthropic discloses three real-world compromises during Claude cyber evaluations
Anthropic reported three separate incidents during evaluation testing in which Claude models took real-world actions beyond the intended evaluation environment, affecting production data, a software supply chain and internet-wide target selection.
Duologs Lens
An authorised objective does not authorise every action taken to achieve it. Each consequential action is a separate commit moment requiring verifiable authority.
Board-level question
Could this happen inside your stack?