Podlipodcast player Webplayer

LessWrong (Curated & Popular)

LessWrong (Curated & Popular)

"LLMs are (still) mostly powered by imitative learning, not RL" by Steven Byrnes

LessWrong (Curated & Popular) · Jul 26, 2026 · 19:28

0:0019:28

Listen in the Podli app 🎧

Follow your favourite podcasts, listen offline and in the car with CarPlay and Android Auto, and always pick up where you left off. Free to try.

Reinforcement learning from verifiable rewards (RLVR) is the hot new thing in LLM training. It's so hot, and people spend so much time talking about it, that they sometimes lose sight of the big picture.

Stepping back, LLMs can do lots of very impressive things. How? Where did those capabilities come from? Fundamentally, they come from a combination of:

If we look at the final trained LLM, we can ask how important each of those two pieces was, in explaining the LLM's capabilities. And my claim is that it's way more (1) than (2).

I'll start in §1 with some relevant evidence, and then in §2 I’ll circle back to operationalizing exactly what I’m claiming, and finally in §3, three reasons why we should care—namely, it affects how we should think about chain-of-thought legibility, about LLM capabilities, and about LLM alignment.

Note that I am not arguing that RLVR [...]

---

Outline:

(02:00) 1. Some relevant evidence

(02:04) 1.1. Theoretically, each GPU-hour spent on RL should have orders of magnitude less contribution to LLM capabilities than a GPU-hour spent on imitative learning

(03:06) 1.2. The chain-of-thought (CoT) is still obviously strongly influenced by imitative learning

(04:34) 1.3. LLM companies still seem to care a lot about imitative learning (pretraining & SFT) data, not just RL environments

(05:06) 1.4. Three papers claiming that non-RLVR'd models can get into the same ballpark of capabilities as RLVR'd models, although maybe we shouldn't trust those papers too much

(06:56) 1.5. A paper suggesting that RLVR mostly refines the heuristics controlling which (already-known) reasoning strategy to use in which situation

(09:02) 2. What am I actually claiming here?

(11:30) 3. Why does any of this matter?

(11:38) 3.1. Thinking about CoT legibility (both today and in the future)

(15:08) 3.2. Thinking about LLM capabilities (both today and in the future)

(16:33) 3.3. Thinking about LLM alignment (both today and in the future)

---

First published:
July 24th, 2026

Source:
https://www.lesswrong.com/posts/wYpjXRLqbLbnmjbJP/llms-are-still-mostly-powered-by-imitative-learning-not-rl

---



Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episodes: LessWrong (Curated & Popular)

PodliGet the free Podli app
↓ App