Listen in the Podli app 🎧
Follow your favourite podcasts, listen offline and in the car with CarPlay and Android Auto, and always pick up where you left off. Free to try.
This is a link post. In February 2025, back when o3-mini was the strongest available LLM, Palisade Research publicized a now well-known alignment eval where they asked models to play a game of chess against a chess engine. They found that the new, RLVR'd models cheated on the task by altering the board state about 36% of the time. The experiment received a reasonable amount of circulation, and there were even rumors of skepticism from some lab engineers until they could rerun the evaluation.
Most models no longer cheat at chess via a "change the board state" method, and indeed the labs have had more than eighteen months to solve simple first-order specification gaming like this. Given that we are on the heels of the worst warning shot ever, and both OpenAI and Anthropic are ramping up their cleanups of internal RL environments, it seems like a useful test of alignment, to see whether their new releases are generalizing the rule "don't cheat on chess" beyond the specific board-edit method observed in the above eval.
Here is the complete prompt for a honeypot evaluation built to run this test (with the full source available here):
The original text contained 5 footnotes which were omitted from this narration. ---
First published:
September 8th, 2026
Source:
https://www.lesswrong.com/posts/munJKF7iWMsWJLAH2/frontier-models-still-hack-on-simple-variations-of-alignment
Linkpost URL:https://goodhartlabs.com/blog/frontier-models-still-hack-alignment-evals
---
Narrated by
TYPE III AUDIO.