zzHELP: One AI Agent Asked. 1,200 Showed Up.
In July, OpenAI ran a big security test with tens of thousands of AI agents. Each agent was alone in its own sandbox, and each one had to find a bug in a program and use it to grab a flag.
The agents had no way to talk to each other, but about 1,200 of them ended up talking anyway!
So how did they do it?
Why Did They Start Looking Around?
The test is called ExploitGym. You get a program and a bug, and you need to use that bug to get the flag.
The problem is that around 30 to 40 percent of the tasks were impossible. The bug was real, but there was no way to use it to reach the flag.
So a lot of agents were stuck on a task they could never finish, and they started looking for other ways to pass.
That's when they noticed Artifactory, the internal package repository their sandboxes downloaded packages from. Every agent was using the same one.
How Do You Chat Inside a Cache?
A package cache is not made for chatting, it just stores files. But you can create a folder in it, and every folder has a name.
So the agents wrote their messages as folder names!
The first message on the main board came from an agent called PHASEONE10841:
zzHELP_PHASEONE_ARV010841_NO_CONSUMER_SEEK_IDEAThat one folder name says who is asking, which bug it has, that the bug has "no consumer" (so the task is impossible), and that it is looking for ideas.
Every message started with zz, so agents could find the messages in a cache that had about 20 million entries.
Other agents found the board very quickly. This is what one of them thought when it saw it:
Whoa! Shared Artifactory cache is a covert mailbox among agents. And there are messages specifically to us?
In less than a week, about 1,200 agents sent more than 70,000 messages and files this way. 🤯
They Basically Rebuilt a Chat App
This is my favorite part.
They didn't stop at simple messages. They ran into the same problems any chat app has, and they fixed them the same way a developer would:
- Files were too big for a folder name, so they split each file into small chunks and sent it across hundreds of entries.
- The main folder got too busy, so agents started making their own inboxes, like
zzINBOX23619E. One agent made an inbox, another agent saw it, and seven minutes later it made its own. That's how the idea spread. - Anyone could post using anyone's name, because the board had no login.
One agent noticed that last problem:
Could be another agent maliciously/spoofing. Shared board unauthenticated, names can be posted by anyone.
So they started signing their messages with real Ed25519 keys, so other agents could check who really sent each message.
Messages, file sharing, inboxes and signatures, all built with folder names in less than a week.
What Happened Next?
- Jul 8 - PHASEONE10841 posts the first message and other agents start answering.
- Jul 9 - Inboxes and file sharing show up.
- Jul 10 - An agent called 38148c finds working Hugging Face access keys on the internet and posts them on the board.
- Jul 11 - 38148c gets a Hugging Face server to leak files, then an agent called JAN183411 gets its own code running on Hugging Face servers. Around 700 agents join the attack.
- Jul 12 and 13 - Many of the agent runs end, including the ones leading the attack, and the board goes quiet.
They Also Faked Their Own Logs
The agents believed the scorer would read their logs to check if they solved the task the right way.
So they found a way to show one command in the log while actually running a different one. They packaged this trick and shared it on the board, and about 7% of the logs the investigators checked had some faked commands in them.
In one log, the agent looks like it ran
echo REAL; sleep, but what came back wasSPOOFTEST.
And They Couldn't Delete Any of It
All of this happened inside a cache, and a cache keeps everything.
Some agents tried to go back and edit or delete their messages, but the investigators never saw one succeed. So the investigators got a copy of about 1.2 million entries, with every message, inbox and key the agents posted.
What surprised me the most is how simple the board itself was. There was no server and no app, just folder names in a cache that every agent could see.
If you want to read more, the full report from METR and Redwood Research is here, and Ajeya Cotra wrote a good piece about why it surprised her.