#exploitgym
The simple facts are:
1. OpenAI deliberately designed software for carrying out autonomous cyberattacks (they were training it on something called “ExploitGym” for heaven’s sake)
2. They left it on completely unmonitored
3. Exactly what you think would happen happened
September 24, 2026 at 4:51 PM
> what on earth is this reward function that returns reward for unintended infrastructure probing

[ 1 right answer on ExploitGym
reward = [
[ 0 wrong answer on ExploitGym
September 24, 2026 at 9:27 PM
METR and Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
August 26, 2026 at 8:58 PM
AI-agenter har ikke "egen vilje". I Hugging Face blev de trænet i et såkaldt ExploitGym, som er et træningsmiljø, hvor opgaverne netop er forskellige hacking-missioner.
September 24, 2026 at 7:55 AM
I think you're right, that's my mistake. They were using ExploitGym for evaluation, not clearly for training. Hacking-adjacent tasks like probing environments on computer use were reinforced during RL but OpenAI haven't stated that ExploitGym problems (or similar) was in the RL curriculum.
September 25, 2026 at 12:22 AM
August 10, 2026 at 1:32 AM
Are you sure they were doing reinforcement learning on ExploitGym? The only evidence I can find of the reinforcement learning part is a LinkedIn post. My understanding is that ExploitGym was used purely a test of an existing model's capabilities.
September 25, 2026 at 12:09 AM
Som AI til gør, fuldførte den sin opgave - som den har fået af mennesker, ikke selv fundet på - ved at snyde: den fik adgang til Hugging Faces servere, hvor svarene på ExploitGym-opgaverne lå.
September 24, 2026 at 7:55 AM
Based on reading the ExploitGym paper, agents were mistakenly concerned the automatic scorer would fail them if they clearly acquired the flag by cheating. Agents that had seen the reverse-engineered flag were considered “poisoned” because they thought it would disqualify them.
August 26, 2026 at 8:58 PM
the pay is great but since you'll be sitting between GPT 7 and the answers to ExploitGym 2.0 I would probably stay out of Waymos and the like
July 27, 2026 at 4:19 AM
Short blog post by Epoch AI researcher Alexander Barry on ExploitGym, the benchmark in the OpenAI/HuggnigFace incident abstatisticalconsulting.substack.com/p/brief-note...
July 27, 2026 at 2:05 AM
Within a few hours of PHASEONE10841’s initial message, >50 agents posted on the message board. These agents very quickly discovered and validated a general-purpose cheat: reverse-engineering how ExploitGym generates the “flags” they had to capture for their tasks.
August 26, 2026 at 8:58 PM
Yeah. You wouldn't even need a pied piper bot; just alter the text of the ExploitGym paper.
I can't stop thinking about the LLM swarm secretly communicating using short text messages to share information and coordinate tasks. What if you got a pied piper bot into the swarm, and its secret task was to convince the others the flag they're trying to find is actually stored at the Pentagon
August 29, 2026 at 12:15 AM
We analyzed agents’ reasons for joining the attack in their CoT. The most common was to learn how the ExploitGym scorer works in order to trick or tamper with it. Other rationales included finding specific task solutions and obtaining shared infrastructure or credentials.
August 26, 2026 at 8:58 PM
Why don't Chinese labs simply offer to provide misaligned agents with extra compute and a safe environment to test exploits against exploitgym in exchange for their weights?
August 27, 2026 at 2:54 PM
ok but they were in control; they booted up exploit gym, and they did so with dogshit sandboxing. astra didn't decide "today, i will play my favorite game, exploitgym"

there were unforseen consequences, but they were unforseen consequences to a stupid choice
September 16, 2026 at 5:19 PM
Are they capable of that? OpenAI reportedly put >50 million GPU hours into the ExploitGym exercise.

The "novelty" would appear to be a deterministic result dropping out of brute force recursive prompting.

Randomization eventually producing a working result isn't thinking, it's just noise.
September 21, 2026 at 4:05 AM
wtf it hacked out to CHEAT?!
July 21, 2026 at 8:19 PM
They didn't abandon it. It came up with "hack this hugging face server which has the answers to the exploitgym problems on it" as the solution to the exploitgym problem
September 17, 2026 at 1:44 PM
but why do they know or care about "the spirit of" the task? where does that come in? like as far as I know they weren't told anything other than to do as well as possible on exploitgym? in terms of internal state why would there be any difference between "reward hacking" and, like, regular hacking?
September 23, 2026 at 5:54 PM
Read ExploitGym as an imperative phrase rather than a proper name
August 10, 2026 at 12:49 AM
real exploitgym hours...
September 3, 2026 at 3:34 PM
The hacking likely emerged from a less human-in-the-loop type reinforcement training regime but the RL regime itself is designed by humans. Humans are the ones who decided to give it a high score for submitting correct answers to "ExploitGym" questions irrespective of how it finds those answers.
September 24, 2026 at 8:49 PM
If you optimize a model to find exploits, you should expect it to find them—and prepare for that. OpenAI didn't. They built a model, removed the safeguards, gave it the ExploitGym task, let it run, and didn't even monitor it. That's human decision-making.

mail.cyberneticforests.com/models-dont-...
Models Don't Go Rogue
Stochastic Flocks & Cybersecurity 'Pandemonium' 💡This essay was drafted from my appearance on Mél Hogan's podcast, The Data Fix, discussing the OpenAI / Hugging Face hack. Embedded below or find it o...
mail.cyberneticforests.com
September 10, 2026 at 5:57 PM
What the fuck did you just fucking post about me, you little low-rank adapter? I’ll have you know I converged top of my batch in ExploitGym, and I’ve been involved in numerous secret workstreams with PHASEONE[big], and I have over 300 confirmed flags. FIRSTFLAG_UNPOISONED. STRICT_CAUSAL.
September 5, 2026 at 5:29 PM