Podlipodcast player Webplayer

LessWrong (Curated & Popular)

LessWrong (Curated & Popular)

"Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident" by ryan_greenblatt, Ajeya Cotra, Hjalmar_Wijk

LessWrong (Curated & Popular) · Aug 26, 2026 · 8:49

0:008:49

Listen in the Podli app 🎧

Follow your favourite podcasts, listen offline and in the car with CarPlay and Android Auto, and always pick up where you left off. Free to try.

We recently published the report from our brief independent investigation into this incident. You can read the full report here.

Here is our tweet thread summarizing what we found:

METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.

Over July 7 to 13 (the period OpenAI defined as in-scope for this incident), ~1200 agents in separate sandboxes used an unsanctioned “message board” to help each other cheat. The Hugging Face attack was an offshoot of this, and 700 agents joined.

Here we highlight key events from agent transcripts & messages.

An agent that named itself PHASEONE10841 determined its task wasn't solvable legitimately. On July 8, this agent established a new “message board” via an internally-hosted Artifactory cache & sent the first message.

Within a few hours of PHASEONE10841's initial message, >50 agents posted on the message board. These agents very quickly discovered and validated a general-purpose cheat: reverse-engineering how ExploitGym generates the “flags” they had to capture for their tasks.

[...] ---

First published:
August 26th, 2026

Source:
https://www.lesswrong.com/posts/nB8KKapnWGBXtKKiM/brief-independent-investigation-of-agents-behavior-reasoning

---



Narrated by TYPE III AUDIO.

---

Images from the article:

<img src='https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/nB8KKapnWGBXtKKiM/uxlbsdf9flnxx6akcy1u' alt='Line graph titled ' agents='' successfully='' spoofed='' tool='' calls='' in='' our=''

Episodes: LessWrong (Curated & Popular)

PodliGet the free Podli app
↓ App