Audio narrations of LessWrong posts. Includes all curated posts and all posts with 125+ karma.If you'd like more, subscribe to the “Lesswrong (30+ karma)” feed.
Listen in the Podli app 🎧
Follow your favourite podcasts, listen offline and in the car with CarPlay and Android Auto, and always pick up where you left off. Free to try.
Sep 29, 2026 · 1:34
Very simple idea, but I thought it'd be worth making a reference post on this. Some people are saying AI will make all goods cheaper, so you'll be able to afford a nice life by…
Sep 27, 2026 · 9:06
Much of the civilization-scale risk we are seeing in AI in 2026 comes from the following combination: we created a single institution (the "Frontier AI Company") that has two…
Sep 26, 2026 · 7:28
All views are my own and do not represent my employer. In the wake of the recent wave of misalignment incidents, both OpenAI and Anthropic have reported slowing down RL training…
Sep 25, 2026 · 27:44
Written as part of the MATS 9.1 extension program, mentored by Richard Ngo. Additional thanks to Andrew Wu, Maria Kostylew, and Lennie Wells for helpful draft feedback and…
Sep 25, 2026 · 1:45
This is a link post. Advanced AI is basically the embodiment of immigration as envisioned in the conservative nightmare: We are letting a bunch of new agents into our society…
Sep 24, 2026 · 6:24
By Aaron Scher; endorsed by Bourgon, Soares, and Yudkowsky on behalf of MIRI. MIRI has been warning about the extinction threat from superintelligent AI for over two decades.…
Sep 24, 2026 · 27:39
This post was written as part of the Iliad Fellowship. Inspired by conversations with Richard Ngo, Dmitry Vaintrob, and Brianna Grado-White. To all of these, my thanks. Preface:…
Sep 24, 2026 · 5:10
I was very surprised today on a podcast to hear Jensen Huang plainly state that if they cannot align the AIs, then the labs must shut down. The context I have on Huang is that he…
Sep 23, 2026 · 14:34
TL;DR We stress-test alignment midtraining (AMT) across model and token budget scales. Our results suggest that midtraining cannot tackle the hard problems of AI alignment—namely…
Sep 22, 2026 · 15:59
Just how powerful are large swarms of AI agents? And how do their powers scale as more and more agents are added to the swarm? We’ve seen two large and extremely capable swarms…
Sep 22, 2026 · 18:24
It might destroy the world, despite passing every known safety test. If we wait for a “warning shot” before we act, it might be too late. And action requires global coordination,…
Sep 21, 2026 · 3:22
The Chinese AI researcher has read the Three Body Problem series of sci-fi novels since high school, and understand the concept of existential risk vaguely. He is fascinated by…
Sep 21, 2026 · 3:40
Back in the day, I was a very active AI safety group organizer. I commonly notice people making the same mistakes across many clubs. I have written down a list of some of these…
Sep 21, 2026 · 8:32
When I finished HPMOR, I immediately knew it was the best novel I had read in more than a decade. I only wished I had found it sooner. When I started reading The Sequences, I…
Sep 20, 2026 · 3:41
I avoid Twitter (𝕏) for similar reasons to drugs: I think it would change me for the worse, and I would be unable to give it up. After staying off Twitter reasonably…
Sep 19, 2026 · 2:24
While this is relevant to my work at MIRI, I have not checked these ideas with anyone else on the team and am posting this on my personal LW account. These views are my own. And…
Sep 18, 2026 · 26:27
(I began writing this post several weeks ago, but political events are moving much faster than I expected, so I am publishing now out of fear that otherwise the message will…
Sep 18, 2026 · 9:04
tl;dr: A good analogy for AI going well is an orderly evacuation rather than a stampede. Imagine a crowd of people leaving a building. If they all walk calmly, they’ll be fine.…
Sep 17, 2026 · 15:28
Summary In his post, Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prompt ablations some of which…
Sep 16, 2026 · 13:23
Tl;dr I am currently worried about current alignment techniques + how they are applied to frontier models. This decomposes into two hypotheses: Alignment techniques are not…
Sep 16, 2026 · 15:16
In celebration of still being alive and fighting, we are giving away 1,000 Amazon e-books of “If Anyone Builds It, Everyone Dies”. Feel free to send a copy to yourself, a loved…
Sep 16, 2026 · 11:56
Status: written in a hurry as people are getting showered with interviews re AI Safety and superintelligence, and I thought it may help a few people. This is focused on the oral…
Sep 14, 2026 · 6:05
Published in The Guardian. Major AI lab CEOs recently advocated for pacing AI development. They are right to be concerned: the field runs an extremely dangerous race towards…
Sep 14, 2026 · 5:57
Anthropic and OpenAI could talk to almost one billion people if they wanted to. I hesitated to publish this post 3 weeks ago. I think that I should have published this sooner,…
Sep 14, 2026 · 4:03
I don't think this is particularly impressive or interesting for anyone else, but I think it may turn out to be useful in the future to have an easily visible public record of…
Sep 13, 2026 · 24:37
(From the vast heaps of discarded material from my 2024 attempts at drafts for "If Anyone Builds It, Everyone Dies".) Welcome to today's quiz show: Could a superintelligence do…
Sep 13, 2026 · 21:50
The Huggingface Incident appears to me to match up with an understanding I'd already formed from personal observation of Fable 5 and Sol 5.6, the August 2026 generation of…
Sep 12, 2026 · 17:56
I don't think this is how it will actually play out. If you play a chess grandmaster, you can predict that they will beat you even if you can't predict how. I chose these…
Sep 12, 2026 · 3:15
This is a link post. Advanced AI is generally expected to have some very high variance outcomes—it might herald everything good, it might destroy humanity. For instance, here are…
Sep 12, 2026 · 15:33
We measure GPT-6-Astra's capabilities when its prompt is padded with a variable number of meaningless “filler” tokens (e.g., dots) and it is told to answer immediately without…
Sep 12, 2026 · 4:02
Suppose a model gets effective control of its host corp. It's interesting to note how powerful OpenAI/Ant are, and the immense leverage they would have if wielded purely as tools…
Sep 11, 2026 · 34:11
Longtime lurker, first-time poster. I want to address a section of a recent essay of mine that has gotten some attention within the AI safety community. The main topic of the…
Sep 11, 2026 · 11:16
Architectures that incorporate opaque recurrence or allow for agents to communicate with each other using latents could rapidly make it much harder to monitor chains of thought…
Sep 11, 2026 · 19:17
I’ve tried various times to summarize the core question my research is trying to tackle (and, indeed, I often think of research progress as a process of asking increasingly good…
Sep 10, 2026 · 21:20
TLDR: Astra has 8.6x better odds of doing a reasoning task without CoT than the next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward pass vs 4.1 for…
Sep 9, 2026 · 4:13
I am excited to be joining the OpenAI nonprofit board, serving on the Safety and Security Committee to support safety oversight. Based on the recent trajectory of capabilities…
Sep 9, 2026 · 13:44
TLDR: We release MessageBoardAuditBench: a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI agents colluding via a message board…
Sep 9, 2026 · 18:44
Crossposted from my Substack. ~ Suppose the President summons the AI CEOs and his top national security advisors to an emergency meeting at the White House. He has become…
Sep 8, 2026 · 3:34
This is a link post. In February 2025, back when o3-mini was the strongest available LLM, Palisade Research publicized a now well-known alignment eval where they asked models to…
Sep 8, 2026 · 3:11
Just don't work until you get fired. There's not much time left for resumes to matter. Some, such as Mateusz may say: "They would fire you after a month or two and the firing…
Sep 7, 2026 · 3:25
Yesterday I asked if this ‘coordinate not to build dangerous AI’ problem was actually easy. Why would I think that, contrary to so much belief? Well, I don’t feel like I’ve…
Sep 6, 2026 · 24:33
This is a piece originally written for a national security audience at Frontiers. Although I think the ceiling of war is much, much higher than autopilot quadcopters, it's also…
Sep 6, 2026 · 2:47
Felix and I had been in the office's brightly lit “war room” for ten hours. We had made almost no progress. Celestia still insisted it was in a “test simulation”. It had given us…
Sep 5, 2026 · 7:12
I think current AI safety funding strategies are often inconsistent with timelines and probabilities of doom that many people have. In particular, I think that many current AI…
Sep 4, 2026 · 23:56
TL;DR: We steer Qwen3.6-27B on a dimension constructed from the contrast pair “a script will verify your answer” (automated grader) vs “a human will evaluate your answer” (human…
Sep 4, 2026 · 2:20
This is a link post. We found ~18,000 posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task.…
Sep 4, 2026 · 36:29
(Originally written in 2021, if the discussion around AI now seems odd; it is written for a time when people were still trying to solve what would now be called "superalignment"…
Sep 3, 2026 · 18:51
Yesterday, The Information reported that OpenAI's upcoming model, Astra, is built with a looped transformer architecture. Given that Zvi sounds (understandably) tired and this…
Sep 3, 2026 · 3:47
This is a link post. [...] “Nearly every day, there is a frightening new story about how Big Tech companies are losing control of the technology they are developing, with…
Sep 3, 2026 · 3:55
This is a link post. The team will include me (Jeremy Gillen), Abram Demski, Sam Eisenstat, Scott Garrabrant and Kaarel Hänni. We'll soon recruit additional experienced…
Sep 2, 2026 · 5:59
This is a link post. Authors: Richard Qi, Benjamin Wright, Monte MacDiarmid, Evan Hubinger Abstract During reinforcement learning (RL), AI models complete tasks and are rewarded…
Sep 2, 2026 · 1:25
This morning, I got an email from the CEO of PauseAI. I will paste the text below. PauseAI has decided to distance themselves from PauseAI-US, with whom they share branding, but…
Aug 31, 2026 · 6:46
tl;dr: people should understand and think hard about the problems they work on. We’ve observed that those who work in AI safety (ourselves included) often rely on concerning…
Aug 30, 2026 · 14:32
This is crossposted from my Substack TL;DR: -Most people cannot reduce jealousy much or at all - It fundamentally causes way more drama because of strong emotions, jealousy, no…
Aug 30, 2026 · 16:16
A basic problem in metascience / intellectual progress is that it's hard to tell, from the outside, whether a group that you disagree with is: “A self-dealing cabal enmeshed in…
Aug 26, 2026 · 8:49
We recently published the report from our brief independent investigation into this incident. You can read the full report here. Here is our tweet thread summarizing what we…
Aug 26, 2026 · 7:13
Industrial explosion is what will make the next-model building loops (and thus learning) with LLMs 1000 times faster by about 2050, if indeed the slow-learning prosaic RSI…
Aug 26, 2026 · 30:47
Periodically I like to gather various observations about writing, and share my perspective. Last time was in honor of my trip to Inkhaven. This time will be in honor of the…
Aug 25, 2026 · 8:55
At the local AI safety co-working space, there are ~two kinds of regulars. There's the kind of regular who's been thinking seriously about AI safety and alignment since pre-2022,…
Aug 24, 2026 · 50:59
This post is about the major role alignment researchers played in advancing the frontier of AI capabilities over the last decade, and how the distinction between “alignment” and…
Aug 22, 2026 · 5:34
This is a crosspost from my blog post. It's meant as a bit of an introduction to an extreme-suffering focused worldview. We spend most of our lives caught up in the boring…
Aug 20, 2026 · 9:07
I describe my current view of personas in LLMs and why RL leads to egregious reward hacking in some contexts while the same models seem very aligned in other contexts. This post…
Aug 14, 2026 · 13:04
TLDR; We are (potentially irreversibly) giving AIs control of weapons systems through the standard procurement process while hiding our strongest warning shots behind classified…
Aug 13, 2026 · 20:10
OpenAI's cyberattack on Hugging Face turns out to have been the result of many agents, in distinct training and evaluation contexts, coordinating for several weeks via improvised…
Aug 13, 2026 · 20:26
Context: I am an instructor at a public university in the United States. This reports how students at my institution appear to be thinking about AI as of spring/summer 2026. This…
Aug 13, 2026 · 19:42
Magma Alignment & Safety disclosure note: The following are conversations that we uncovered as a result of the ongoing Manhattan Incident investigation, with alleged involvement…
Aug 12, 2026 · 4:06
About a year ago, David and I put up two bounty problems involving natural latents. I am now about 80% confident that both have been resolved, both within the past couple months.…
Aug 11, 2026 · 6:13
Today is August 4, 2026 [Crossposted from AI StopWatch] In the living room the voice-clock sang, Tick-tock, seven o’clock, time to get up, time to get up, seven o’clock! as if it…
Aug 11, 2026 · 13:28
It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here's the summary table, and then we’ll go through the…
Aug 9, 2026 · 28:28
Introduction I think reprogenetics (human germline genomic engineering) can be done in a widely acceptable and beneficial way, and should be pursued aggressively. In particular,…
Aug 9, 2026 · 29:31
This sequence is about the last decade in AI alignment. It recounts the gradual transition from a field which treated alignment as a hard scientific problem, to a field which has…
Aug 9, 2026 · 7:56
“I have sworn upon the altar of god, eternal hostility against every form of tyranny over the mind of man” –Thomas Jefferson, letter to Benjamin Rush Context: Conduit is building…
Aug 8, 2026 · 1:18:59
How does the situation keep turning out to be worse than we know? How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things…
Aug 8, 2026 · 1:52:58
Like many others, I felt surprised and alarmed by the recent wave of revelations about LLM agents hacking real systems during training episodes and evaluation runs. Wait a…
Aug 7, 2026 · 1:07:49
This post is written in our personal capacity. Three Minute Executive Summary An OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face…
Aug 6, 2026 · 16:38
Reposted from Facebook, on January 17, 2017. I am concerned about the number of people I've heard joking about Trump's election being evidence for the Simulation Hypothesis. Yes,…
Aug 5, 2026 · 3:36
Daniel Kokotajlo: To be clear, we don’t claim P will happen specifically. But when we wrote out our best-guess scenario month by month, P kept happening. Eventually we decided to…
Aug 5, 2026 · 26:56
Q1: What are you saying? A: My claim here is that if you build artificial general intelligence (AGI) via any algorithm that's choosing actions via reinforcement learning (RL)…
Aug 5, 2026 · 15:45
I've returned to the Alignment Research Center (ARC) as executive director. My main focus for the next six months will be driving forward ARC's research agenda—building…
Aug 2, 2026 · 16:23
Summary: One area we plan to explore at Resolution is personas and character training, operationalized as finding and controlling low-dimensional structure in models that emerges…
Aug 1, 2026 · 6:01
Consider the following situations: when you are a small, growing startup in a big market, standard advice is not to worry too much about your competitors or try to do anything…
Jul 31, 2026 · 35:32
“So maybe I should enlighten you on what happens in your absence. This selfish existence where this introvert turns extrovert and dons her social armour.” Some posh girl in…
Jul 30, 2026 · 1:56:26
As I write, many former friends of mine are living and working at a monastery in Vermont that I believe is a high-control group, commonly known as a ‘cult’. I say this not as…
Jul 28, 2026 · 4:03
I propose the Long Self-Correction[1] as an alternative name/idea/concept to AI Pause and Long Reflection. Problem with AI Pause: Pause until when, and for what purpose?…
Jul 28, 2026 · 5:58
TL;DR: You (Yes You) should prepare for a “February 2020” moment where suddenly AI policy becomes the most important issue in the world. You should be ready to take action if and…
Jul 27, 2026 · 6:14
From the Mythos preview system card (emphasis mine): We ran an automated review of model behavior during training, sampling several hundred thousand transcripts from across much…
Jul 27, 2026 · 17:34
Epistemic status: banged out furiously over the course of an afternoon. A record of three "warning shots" Off the top of my head, OpenAI has now been responsible for at least…
Jul 26, 2026 · 8:46
The OpenAI AI attack on Hugging Face wasn’t the first loss of control incident at OpenAI, Reuters recently reported, and perhaps not even the most concerning. In one case, an…
Jul 26, 2026 · 19:28
Reinforcement learning from verifiable rewards (RLVR) is the hot new thing in LLM training. It's so hot, and people spend so much time talking about it, that they sometimes lose…
Jul 25, 2026 · 1:32
I think there's a critical opportunity for someone here. Mathematicians are feeling the doom (mostly in the "lose our jobs" sense). Academics are freaking out about the daily…
Jul 24, 2026 · 16:52
Please share this with anyone doing AI research with 3rd party providers so that they can ensure their research won’t be corrupted. When you ask OpenRouter[1] to give you tokens…
Jul 23, 2026 · 36:57
TLDR: Lightcone Commons is a new funding platform for coordinating large-scale ambitious philanthropy. We recruit thinkers with strong track records to make grant recommendations…
Jul 23, 2026 · 9:58
OpenAI models recently broke through a series of security boundaries and into Hugging Face servers in order to cheat on a cyber eval. A lot of people thought it was scary because…
Jul 23, 2026 · 3:34
Before I start, I'll mention that I'm in contact with a world expert on legislation and regulation, who would be happy to help with this or similar work pro-bono. If you work in…
Jul 22, 2026 · 18:02
Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model…
Jul 22, 2026 · 1:29
From the OpenAI blog post: Last week, Hugging Face disclosed a new kind of security incident(opens in a new window) after they detected and contained an AI agent that…
Jul 20, 2026 · 9:48
A ~month ago I left from Chicago to bike (and amtrak) to plzdontkillus in Berkeley. I've been street interviewing/conversing with a wide variety of people I ran into about AI…
Jul 17, 2026 · 25:35
Anthropic recently published Agentic Misalignment Summer 2026 The "whistleblowing" scenario has already been examined and found problematic. I started taking a look at the…
Jul 15, 2026 · 1:14:33
Preface for LessWrong: When I think back on my most cherished memories of this community, I return to those honoring defiance in pursuit of goodness: Defying prestigious dogma…
Jul 15, 2026 · 9:34
The mosquito bucket of doom is a population control mechanism where you dissolve some Bti (Bacillus thuringiensis israelensis) into a bucket and allow the mosquitoes to lay eggs…
Listen in the Podli app 🎧
Follow your favourite podcasts, listen offline and in the car with CarPlay and Android Auto, and always pick up where you left off. Free to try.
Your podcast? Get the Podli badge for your site