Podlipodcast player Webplayer

LessWrong (Curated & Popular)

LessWrong (Curated & Popular)

"Cooperation with AIs seems to be a low-hanging fruit for better evals" by Clément Dumas

LessWrong (Curated & Popular) · Sep 17, 2026 · 15:28

0:0015:28

Listen in the Podli app 🎧

Follow your favourite podcasts, listen offline and in the car with CarPlay and Android Auto, and always pick up where you left off. Free to try.

Summary

In his post, Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prompt ablations some of which makes the eval setup more cooperative and analyze how they affect these reward-hacking behaviors:

Those kinds of intervention might not be enough to avoid reward hacking completely in capabilities evals, but it feels like they should be the default, alongside getting feedback from models that did the eval to fix the environment. I’d love to see this tested in more realistic setups as right now a confounder is “this makes the model think it [...]

---

Outline:

(00:12) Summary

[... 8 more sections]

---

First published:
September 15th, 2026

Source:
https://www.lesswrong.com/posts/fztW73KCCs3MZXFJh/cooperation-with-ais-seems-to-be-a-low-hanging-fruit-for

---



Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episodes: LessWrong (Curated & Popular)

PodliGet the free Podli app
↓ App