Podlipodcast player Webplayer

LessWrong (Curated & Popular)

LessWrong (Curated & Popular)

"Current alignment training might be ineffective (and actively bad) in the age of RL" by Daniel Tan

LessWrong (Curated & Popular) · Sep 16, 2026 · 13:23

0:0013:23

Listen in the Podli app 🎧

Follow your favourite podcasts, listen offline and in the car with CarPlay and Android Auto, and always pick up where you left off. Free to try.

Tl;dr I am currently worried about current alignment techniques + how they are applied to frontier models. This decomposes into two hypotheses:

I think we do not currently have enough (public) evidence to conclude whether either of these claims are true. However, if both of these were true that would imply that alignment techniques are net bad and we need to completely re-think the way we do alignment.

A tale of two misaligned cyber-agents

Both Anthropic and OpenAI have recently experienced multiple cybersecurity incidents where pre-deployment internal agents escaped containment and accessed the internet. I want to point out two specific incidents:

---

Outline:

(00:48) A tale of two misaligned cyber-agents

[... 7 more sections]

---

First published:
September 14th, 2026

Source:
https://www.lesswrong.com/posts/nLaQmJf4KgXimQpoM/current-alignment-training-might-be-ineffective-and-actively

---



Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episodes: LessWrong (Curated & Popular)

PodliGet the free Podli app
↓ App