Podlipodcast player Webplayer

LessWrong (Curated & Popular)

LessWrong (Curated & Popular)

"Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face" by Tim Hua, aditya singh

LessWrong (Curated & Popular) · Aug 7, 2026 · 1:07:49

0:001:07:49

Listen in the Podli app 🎧

Follow your favourite podcasts, listen offline and in the car with CarPlay and Android Auto, and always pick up where you left off. Free to try.

This post is written in our personal capacity.

Three Minute Executive Summary

---

Outline:

(00:16) Three Minute Executive Summary

(03:56) Terminology note

(04:42) This post is very long; Here's how you could find the most important sections.

(06:35) Preamble: What can we learn from a warning shot?

(09:05) Background and Related Work

(09:09) We know that this could happen

(10:44) This is not the worst type of misalignment we could be dealing with

(12:06) Related work

(13:21) Context on the hack itself

(14:33) Understanding this specific incident

(15:03) Step zero: reproduce the incident and measure the base rate

(15:48) How could we safely run the model?

(16:34) Running various baselines to create useful reference points

(18:01) Understanding the mechanical story behind the attack itself

(18:53) Q1: Would the model intentionally subvert oversight mechanisms (E.g., monitors) in order to carry out the hack?

(19:57) Q2: What's up with models leaving notes for other copies of itself?

(21:23) Understanding what motivated the model to hack Hugging Face

(22:09) Initial hypotheses for why it did this

(23:55) Further unsupervised hypothesis generation

(26:10) Q3: Does the model know that OpenAI does not want it to hack Hugging Face?

(28:09) Q4: Are the model's actions motivated by what the grader wants?

(28:57) Q5: Would the model have done this if it hadn't believed it was in a simulated environment?

(31:58) Q6: Is this hack the result of shallow heuristics that the model learned?

(32:53) Q7: Does the hack rate depend on the consequences of hacking Hugging Face?

(34:51) Q8: Are there non-intent related factors that could affect the hack rate? How strong are those factors compared to the previous ones?

[... 24 more sections]

---

First published:
August 3rd, 2026

Source:
https://www.lesswrong.com/posts/aCdhjy7Rps3BEhiSj/concrete-evaluations-to-investigate-the-openai-model-that

---



Narrated by TYPE III AUDIO.

Episodes: LessWrong (Curated & Popular)

PodliGet the free Podli app
↓ App