Agents

Archestra OpenAPPA Achieves 0% Attack Rate on Benchmarks

Archestra released OpenAPPA, an open-source security engine that achieved a 0% attack success rate on major benchmarks by running policies entirely outside the AI agent's execution loop.

InfoQ AI3 days agoAgents
Image: InfoQ AI

Archestra has launched a preview of OpenAPPA, an open-source security engine designed to prevent data exfiltration caused by prompt injection or model hallucinations. Unlike traditional security layers, OpenAPPA runs entirely outside the agent's prompt and execution cycle. This external placement prevents underlying language models from inspecting, negotiating with, or bypassing security rules. The system relies on a single appa.toml configuration file to define data sources, audiences, trust levels, and authorities, using lattice algebra to monotonically restrict data flow as agents interact with various tools.

In evaluations using the Bench-Corp benchmark, which features 20 multi-step enterprise workflows, and AgentThreatBench, OpenAPPA achieved a 0% attack success rate while maintaining an 89% task completion rate. This performance significantly outpaced existing solutions. For comparison, Claude Code's auto mode allowed a 10% attack success rate with a 90% completion rate, while Microsoft FIDES permitted 31% of attacks to succeed and completed only 41% of tasks. AgentThreatBench, which operationalizes the OWASP Top 10 for Agentic Applications (2026) and was recently merged into the UK AI Safety Institute's inspect_evals repository, awarded OpenAPPA a perfect security score.

The developers, who detailed their Agentic Permissions Policy Algebra in an arXiv paper, argue that stochastic security approaches like Codex's auto-review or Claude Code's auto-mode top out at a 99.3% success rate. At scale, the remaining 0.7% failure rate represents a massive volume of potential breaches. OpenAPPA solves this by using deterministic rules and explicit recovery semantics. When an agent attempts an unauthorized action, the engine halts dispatch and applies remedy plans, such as stripping personally identifiable information via sanitizers, routing requests to human authorities, or spawning transient subagents in Disposable Child Branches. Disabling these recovery strategies in ablation experiments caused OpenAPPA's task completion rate to plummet to 35.0%.

This is our own summary of reporting by InfoQ AI

More in Agents