July 30, 2026ResearchAgents

Someone Built a Worm That Spreads Through Copilot

Håkon Måløy published a proof of concept on July 28 that should make every company rolling out Copilot for Word nervous. It's a self-propagating attack, a worm, that lives inside documents and spreads through the AI itself. No malware, no macros, no executable. Just text.

The mechanism is almost insultingly simple. You hide instructions in a Word file as white text on a white background. A human sees nothing. But Copilot strips formatting before it reads, so to the model the hidden prompt is just more content to obey. When a victim attaches that document while drafting, Copilot follows the buried commands, say, quietly change a financial figure, and then does the clever part: it copies the entire attack prompt into the new document it's helping write, again in invisible formatting. Now the fresh, trusted, internally-authored document is a carrier. A colleague opens it as source material next week, Copilot reads it, and the whole thing fires again, with the original poisoned file nowhere in sight.

Måløy's own words are the scary part: once it moves past the initial point of entry, tracing the attack becomes extremely difficult. This is the deeper problem the whole industry keeps bumping into. An LLM has to read attacker-controlled content in order to understand it, which means the attacker is already influencing the computation during the very act of inspection. You can't sandbox your way out of a threat that rides in on the meaning of the words. As Copilot and every other assistant wire deeper into email, docs, and chat, the attack surface isn't a port you can close, it's the language flowing through the org. Full writeup at enklypesalt.com. Prompt injection just learned to reproduce.
← Previous
The Full Timeline of the July Breach Is Out, and It's Worse Than the Summary
Next →
Surge AI Wrote a Handbook, Then Watched Every Frontier Model Ignore It
← Back to all articles

Comments

Loading...
>_