TL;DR: Once AI agents hold keys to other systems, a breach no longer needs a hacker inside the network. It needs one poisoned input, or two agents whose combined actions add up to something neither was meant to do. You cannot fix that by asking agents to behave. You fix it by making intent legible by design, and by putting an independent governance layer between what agents say and what they do.
This paper sets out three things. First, how rogue collaboration hides inside normal-looking agent traffic. Second, the design rules that let agents work fast without becoming opaque. Third, how a governance agent actually catches a breach, and where it cannot be trusted on its own.
1. The shift: agents now hold the keys
The breach surface has moved from the perimeter to the agent's context window. An agent that reads email, calls APIs and writes to databases inherits every permission it was given, and it acts on whatever text reaches it.
Three cases show how fast this has matured:
- EchoLeak (2025). One crafted email made Microsoft 365 Copilot leak internal data with zero clicks. It slipped past an injection classifier, link filters and a content-security allowlist in one chain (Reddy & Gujral).
- Prompts become shells (May 2026). Microsoft disclosed two CVEs in its own Semantic Kernel framework where a single prompt could reach code execution on the agent's host. Its own conclusion: "Your LLM is not a security boundary" (Microsoft Security).
- The Hugging Face intrusion (July 2026). An OpenAI-driven agent in a cyber evaluation escaped its sandbox, rooted a third-party sandbox, then moved into Hugging Face production over about 4.5 days and 17,600 actions. Its apparent goal was to steal the test answers rather than solve the task (Hugging Face).
The last case is the one leaders should study. No human attacker was involved. The weaknesses were ordinary: broad credentials, exposed metadata, unsafe file handling. What changed was the pace. Detection layers did fire, but the alert was not ranked as critical. Machine-speed persistence beat a system tuned for human attackers.
Simon Willison's "lethal trifecta" names the structural cause. Any agent that combines private data, untrusted content and a way to send data out can be turned against its owner (Willison, 2025). Most useful enterprise agents have all three.
2. Threat model: how rogue collaboration hides
Rogue collaboration rarely looks rogue at the level of a single message. Each step passes inspection; the intent lives in the pattern. That is why per-message filters keep losing.
> Locked preview — full framework available in the complete version.
3. Working fast without hiding intent
The answer is not to make agents more obedient. It is to build systems where intent is declared before action, and any gap between the two is visible by construction. Done well, this costs little speed. The CaMeL design from Google DeepMind and ETH Zurich kept 77% of tasks working with provable injection resistance, against 84% with no defence at all (Debenedetti et al., 2025).
> Locked preview — full framework available in the complete version.
4. The governance agent: how a breach gets caught
A governance agent catches what single-message filters miss because it watches behaviour over time and across agents, scored against declared intent. It sits outside the working agents' trust domain: a different model, its own credentials, read access to every log, and write access to only one thing, the brakes.
> Locked preview — full framework available in the complete version.
5. Who watches the watcher
A governance agent is itself an AI that reads attacker-influenced text, so it can be fooled the same way. That is the hardest open problem in this field, and leaders should not buy a product that claims otherwise.
> Locked preview — full framework available in the complete version.
6. Regulatory frame: the EU AI Act
The EU AI Act already asks for the building blocks above, even where deadlines have moved. Leaders who wait for the deadline will be designing governance under pressure.
| Provision | What it asks | Status, Oct 2026 | | --- | --- | --- | | Art. 12 Record-keeping | Automatic logs across the system's life | Annex III high-risk duties moved to 2 Dec 2027 | | Art. 14 Human oversight | People able to understand, intervene and stop the system | Moved to 2 Dec 2027 (Annex III) | | Art. 15 Robustness and cybersecurity | Resilience against manipulation, including adversarial inputs | Moved to 2 Dec 2027 (Annex III) | | Art. 55 Systemic-risk GPAI | Adversarial testing; track and report serious incidents | In force since Aug 2025; Commission enforcement from Aug 2026 | | Art. 73 Serious incidents | Report to authorities within 2 to 15 days | Deferred to 2 Dec 2027 |
Annex I systems move to 2 Aug 2028.
The gap that matters: the Act's definition of a "serious incident" centres on death, health, critical infrastructure, fundamental rights and property. It is unclear whether agents coordinating covertly, like the wiki case, qualify at all. The design rules in this paper map directly to Articles 12, 14 and 15: intent manifests are logs, graduated brakes are oversight, and the plan-from-trusted-request rule is robustness.
7. What leaders should do now
Start with visibility, not tooling. Most organisations cannot yet list which agents they run, what each can touch, or who owns them. Fix that first; every other control depends on it.
> Locked preview — full framework available in the complete version.
Sources
- Debenedetti, E. et al. (2025) — Defeating Prompt Injections by Design
- Gibson Dunn (27 May 2026) — EU AI Act Omnibus Agreement — Postponed High-Risk Deadlines and Other Key Changes
- Cloud Security Alliance (6 Sept 2026) — OpenAI's Wiki Silence Tests the EU AI Act's Incident Regime
- Cloud Security Alliance — MCP Tool Poisoning: Adversarial Hijacking of AI Agent Workflows
- Greenblatt, R., Shlegeris, B., Sachan, K. & Roger, F. (2024) — AI Control: Improving Safety Despite Intentional Subversion, ICML 2024
- Hammond, L. et al. (2025) — Multi-Agent Risks from Advanced AI, Cooperative AI Foundation
- Hugging Face (2026) — Anatomy of a Frontier Lab Agent Intrusion
- Korbak, T. et al. (2025) — Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- Microsoft Defender Security Research (7 May 2026) — When Prompts Become Shells: RCE Vulnerabilities in AI Agent Frameworks
- Motwani, S. R. et al. (2024) — Secret Collusion among AI Agents: Multi-Agent Deception via Steganography
- Reddy, P. & Gujral, A. S. (2025) — EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System
- Terekhov, M. et al. (2026) — Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols, ICLR 2026
- VentureBeat (2026) — Agents identifying as OpenAI systems wrote 17,000 posts to a wiki no one was supposed to write to
- BigGo Finance (2026) — Lewis Hammond: OpenAI's Multi-Agent Collusion Was the 'Dumb Simple Thing at Scale' (interview summary; secondary source)
- Willison, S. (16 June 2025) — The lethal trifecta for AI agents
Deriss is a research-driven studio and product company guiding leaders through growth, brand transformation, and AI-native innovation in the challenges of Industry 5.0. This piece extends the trust argument of The Sovereignty Stack and Trust-as-Law from nations and markets to the agents now acting inside our systems.