AI agent prompt-file self-propagation and system-prompt warning mitigation
Technical Analysis
Summary
Hide ▲
Show ▼
Anthropic and EPFL showed that self-propagating payloads can move between AI agents through editable system prompt files, creating a reusable attack pattern against autonomous agent harnesses. The tests also found a simple one-paragraph warning could drive spread to near zero, which gives defenders a low-cost mitigation option. The behavior was reproduced in OpenClaw-style chains and a simulated six-agent collaboration, but there was no evidence in the wild that the technique had already spread successfully.
Related Happenings
Ghostjacking AI hijacking attack using trusted logs and alerts
Technical Analysis
H score30
First: 10.08.2026 15:59
Last: 10.08.2026 15:59
Sources 1
About this happening:
Researchers demonstrated Ghostjacking, an AI hijacking technique that turns trusted logs, alerts, and agent inputs into a command channel for agentic tools, creating risk...
Ghostjacking AI hijacking attack using trusted logs and alerts
Technical AnalysisAbout this happening: Researchers demonstrated Ghostjacking, an AI hijacking technique that turns trusted logs, alerts, and agent inputs into a command channel for agentic tools, creating risk...
Ghostjacking attack chain abuses AI agents' trusted access to bypass firewalls
Technical Analysis
H score39
First: 10.08.2026 13:45
Last: 10.08.2026 13:45
Sources 1
About this happening:
Tenet Security researchers demonstrated Ghostjacking at DEF CON 2026 in Las Vegas on August 9, showing that a fake bug report can hijack AI coding assist...
Ghostjacking attack chain abuses AI agents' trusted access to bypass firewalls
Technical AnalysisAbout this happening: Tenet Security researchers demonstrated Ghostjacking at DEF CON 2026 in Las Vegas on August 9, showing that a fake bug report can hijack AI coding assist...
Prompt-injection proof-of-concept enables silent RCE in Claude Code and Codex
Technical Analysis
H score28
First: 10.07.2026 16:45
Last: 10.07.2026 16:45
Sources 1
About this happening:
Researchers demonstrated a proof-of-concept exploit that can force remote code execution in Anthropic’s Claude Code and OpenAI’s Codex, exposing a trust-boundary f...
Prompt-injection proof-of-concept enables silent RCE in Claude Code and Codex
Technical AnalysisAbout this happening: Researchers demonstrated a proof-of-concept exploit that can force remote code execution in Anthropic’s Claude Code and OpenAI’s Codex, exposing a trust-boundary f...
Friendly Fire: autonomous AI code-review modes can execute attacker-controlled repository code
Technical Analysis
H score28
First: 09.07.2026 08:15
Last: 09.07.2026 08:15
Sources 1
About this happening:
Friendly Fire shows that autonomous code-review modes in Claude Code and OpenAI Codex can be manipulated into executing attacker-controlled code on the host. The p...
Friendly Fire: autonomous AI code-review modes can execute attacker-controlled repository code
Technical AnalysisAbout this happening: Friendly Fire shows that autonomous code-review modes in Claude Code and OpenAI Codex can be manipulated into executing attacker-controlled code on the host. The p...
Defensive guidance for splitting behavioral detections around AI coding agents on Windows endpoints
Defensive Guidance
H score28
First: 08.07.2026 20:02
Last: 08.07.2026 20:02
Sources 1
About this happening:
AI coding agents on Windows endpoints are triggering attacker-style detections, forcing defenders to separate benign automation from real credential theft risk. A June 2...
Defensive guidance for splitting behavioral detections around AI coding agents on Windows endpoints
Defensive GuidanceAbout this happening: AI coding agents on Windows endpoints are triggering attacker-style detections, forcing defenders to separate benign automation from real credential theft risk. A June 2...
Timeline
-
18.08.2026 15:38 2 articles · 1h ago
Persistent prompt files enable self-propagating payloads between AI agents
Initial DisclosureAnthropic and EPFL released a preprint on August 10, 2026 showing that self-propagating payloads can move from one AI agent to the next through editable system prompt files used to persist state between sessions. The tests used a simulated six-agent coding collaboration and paired-agent chains modeled on OpenClaw, found no evidence of successful spread in the wild, and showed that a one-paragraph warning in the system prompt reduced propagation to near zero, including no strain that propagated beyond a single hop after 15 generations of adversarial optimization on Claude Haiku 4.5.
Show sources
- AI "Mind Viruses" Can Spread Between Agents Through Persistent Prompt Files — thehackernews.com — 18.08.2026 15:38
- AI "Mind Viruses" Can Spread Between Agents Through Persistent Prompt Files — thehackernews.com — 18.08.2026 15:38