Listen in the Podli app 🎧
Follow your favourite podcasts, listen offline and in the car with CarPlay and Android Auto, and always pick up where you left off. Free to try.
There has been much discussion recently around whether a large portion of alignment research is net negative. Without endorsing or refuting them, the basic arguments here are:
- Prosaic alignment of models is becoming a bottleneck for capabilities.
- Therefore improving the prosaic alignment of models enables faster capabilities advances, which bring us closer to RSI.
- It is unlikely these prosaic alignment methods remain sufficient during the RSI loop, and so this work brings us closer to doom.
- Furthermore, dealing with these more prosaic failures reduces the likelihood of a warning shot of sufficient magnitude to cause a slowdown which would prevent RSI.
On the basis of this argument, some urge alignment researchers at AGI companies to quit outright. But quit to do what? Missing from this exchange so far has been a discussion of opportunity costs. If you aren’t going to do (technical) work on “Alignment” – either inside or outside of an AGI company – what should you work on?
In this post, I outline a contrast between “Alignment Engineering” – the dominant model for what “working on alignment” looks like (inside labs, and in the field as a whole) with “Misalignment Science”. I begin by characterising [...]
---
Outline:(01:51) "Alignment Engineering"
(06:22) AI Safety and the ML tradition
(07:42) Implicit work trials
(09:36) "Misalignment Science"
(14:56) Conclusion
(15:39) Postscript: Iterating ourselves into oblivion
The original text contained 4 footnotes which were omitted from this narration. ---
First published:
October 5th, 2026
Source:
https://www.lesswrong.com/posts/FogmcDHA6AdMGukum/alignment-engineering-vs-misalignment-science
---
Narrated by
TYPE III AUDIO.