Podlipodcast player Webplayer

LessWrong (Curated & Popular)

LessWrong (Curated & Popular)

"models may behave differently in graded episodes (a tirade)" by nostalgebraist

LessWrong (Curated & Popular) · Aug 8, 2026 · 1:52:58

0:001:52:58

Listen in the Podli app 🎧

Follow your favourite podcasts, listen offline and in the car with CarPlay and Android Auto, and always pick up where you left off. Free to try.

Like many others, I felt surprised and alarmed by the recent wave of revelations about LLM agents hacking real systems during training episodes and evaluation runs.

Wait a moment, though -- "I felt surprised and alarmed"? "Alarmed," sure, fine that one's self-explanatory... but why surprised?

After all: haven't we known for a long time, on both theoretical and (increasingly) empirical grounds, that RLVR selects for monomaniacal pursuit of perceived grader-satisfaction, ethics and (beyond-episode) consequences be damned?

After all -- the way we train frontier capabilities into these models is, more or less:

If you do this, at scale, then you should expect to (eventually) see every behavior pattern that [...]

---

Outline:

(03:35) \[1\] remember what you already know

(20:02) \[2\] reward-instilled reflexes and flexible reward-pursuit

(43:14) \[3\] graded-episode perception, and policies conditional upon it

(01:01:44) \[4\] the discourse is not yet adequate

(01:09:57) eval awareness

(01:18:32) metagaming

(01:41:21) reward hacking

The original text contained 18 footnotes which were omitted from this narration.

---

First published:
August 7th, 2026

Source:
https://www.lesswrong.com/posts/AfoGGrJfuNzofpzWL/models-may-behave-differently-in-graded-episodes-a-tirade

---



Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episodes: LessWrong (Curated & Popular)

PodliGet the free Podli app
↓ App