Вы здесь

Новости LessWrong.com

Подписка на Лента Новости LessWrong.com Новости LessWrong.com
A community blog devoted to refining the art of rationality
Обновлено: 2 часа 4 минуты назад

We are not prepared to win

20 сентября, 2026 - 02:01

TL;DR: It would be nice to survive for a long time and be happy about it. It doesn't look like we'll be happy about it right now. We should change that by putting more effort into making sure we'll be happy about it.

  • Some [LessWrongers, Rationalists, etc.] believe that AI is likely to cause the apocalypse
  • [LessWrongers, Rationalists, etc.] also consider the effects of living with well-developed, non-apocalyptic AI, likely or not
    • AI will [very probably, obviously] have extreme economic effects barring an apocalypse
      • Arguably the current state of the stock market is a present-tense effect
      • Mass unemployment would be in the question, with horrible knock-on effects
    • AI will maybe give individual people much more power than they otherwise would have
      • For both good and bad causes. E.g. curing a disease, creating a bioweapon
      • This seems to be a more volatile world to live in
    • AI will make the world very different and thus hard for John Doe to plan for
  • [LessWrongers, Rationalists, etc.] have built incredible infrastructure for the purpose of preventing an AI apocalypse
    • Community figureheads have build an impressive recruiting pipeline
    • Recruits are generally smart and/or open-minded people capable of engaging with complex problems effectively
    • Community figureheads have generally done well to draw recruits focus to preventing an AI apocalypse
    • The recruits generally do okay at preventing an AI apocalypse
      • [LessWrongers, Rationalists, etc.] are doing their best to influence policy
      • [LessWrongers, Rationalists, etc.] are taking up jobs that put them in a position to change the trajectory of the future with respect to AI safety
    • This is generally good and makes sense:
      • Community figureheads believe (or appear to believe) that there is a significant risk of the AI apocalypse
      • Community figureheads have done well to direct resources to preventing that because, if it occurs, game over
      • Yay
  • [LessWrongers, Rationalists, etc.] have built comparatively weak infrastructure for preventing the negative consequences of otherwise Good AI
    • This is not to say that no work has been done in this area. Indeed, it is still somewhat in focus for [LessWrongers, Rationalists, etc.] See:
      • Eliezer Yudkowsky's 2013 "Robots, AI, and Unemployment Anti-FAQ"
      • The 2024 AI Safety Camp's "The Future of Work: How Can Policymakers Prepare for AI's Impact on Labor Markets?"
      • The highlighted Labor Automation Forecasting Hub over at Metaculus
    • But these efforts, both in quantity and quality, are far and away overshadowed by the efforts made to affect AI safety and related policy
      • Perhaps by a factor of 10x or 100x (bullshit numbers)
    • Efforts active in this space also tend to lack clear direction
      • [LessWrongers, Rationalists, etc.] can forecast displacement statistics, so they know roughly how much humanity will need to adapt to the future landscape
      • But [LessWrongers, Rationalists, etc.] do not know what [efficiently, least-painfully] adapting looks like
      • [LessWrongers, Rationalists, etc.] also do not appear to have a huge repertoire of ideas on shaping a non-apocalypse environment to need less adaptation
        • The options apparent are also not necessarily desirable
  • The bias towards Survival First makes sense
    • But perhaps [LessWrongers, Rationalists, etc.] have biased their efforts too strongly
    • It would be nice to both survive and be happy about it


Discuss

Common mistakes in AI safety group organizing

20 сентября, 2026 - 02:00

Back in the day, I was a very active AI safety group organizer. I commonly notice people making the same mistakes across many clubs. I have written down a list of some of these mistakes hoping people will avoid them in the future:

  • Reading groups often require that people read things before meetings. This is a mistake. People often don't do the readings. And the lack of common knowledge that everyone has read the reading degrades the conversation quality.
    • Instead, have longer meetings, serve food (so, lunch/dinner meeting slots), and read during the actual meeting.
  • Reading groups often don't sort people into cohorts properly. Mainly, they fail at clustering people into clusters of roughly equal ML knowledge and age. Grad students don't want to discuss a paper with freshmen. People with lots of ML knowledge don't want to discuss a paper with people with no ML knowledge.
    • Instead, group people with people similar to them in ML knowledge and age.
  • Reading groups often rely on digital materials instead of physical printouts. Screens are distracting and there is no common knowledge that people are paying attention.
    • Neatly print every reading ahead of time instead.
  • Clubs don't gatekeep enough. There exist many people whose presence is net-negative to events. Those people should not be accepted to the club or invited to its events.
    • Instead, get signal on people through application forms and word-of-mouth. If people seem like their presence wouldn't make the club better, don't include them. Also, invest in people (through helping them upskill and expecting them to contribute to high-context conversations) such that they can become the types of people who can contribute. A lot of what it means to build a community is knowing how to gatekeep it properly.
  • Organizers aren't knowledgeable enough. People can notice if you're well-read and have good takes. If you're not knowledgeable, they will take you less seriously and you will be worse at understanding the world and helping others understand it.
    • Instead, foster a culture where organizers (especially the main ones) need to be worthy of being role models for other students. They should be doing quality work and thinking about and talking about AI safety.
  • Organizers don't have serious vibes. Unserious vibes drive people away and make people think that the organizers don't take the actual issues seriously.
    • AI safety is not just another thing to form a club around. It's a serious topic and people who are teaching others about it should act like it. It should feel much more like a pre-professional club (e.g. an ML reading group, or a consulting club) and less like a social club (e.g. a board game club).
  • Organizers aren't close enough friends. If the organizers don't know each other well enough and don't get along well enough, the club will function much worse because of friction, and outside people will find the club less attractive.
    • Instead, treat "the organizers are friends" as one of the main goals of the club.

Thank you, group organizers, for your work on helping others contribute to the field. I hope that you will meet the current moment with the seriousness it deserves.



Discuss

Failure of the coding theorem for randomized stopping machines

19 сентября, 2026 - 23:27

Epistemic Status and Contributions. This post explains a technical separation result in algorithmic information theory which was derived during Mikhail Mironov's Summer 2026 PIBBSS fellowship. The result contributes to AIXI Labs' research program on how Solomonoff induction generalizes from past observations in the face of novel events.
Problem formulation: Cole Wyeth.
Proof of main Theorem 1: GPT-5.6 Sol.
Appendix proofs: the sketch of the proof for equivalence between time semimeasures and randomized stopping machines is by Cole Wyeth, the rest by GPT-6 Astra.
Writing: draft by Claude Fable 5 and GPT-6 Astra, editing and rewriting by Mikhail Mironov.
Useful discussions: Aram Ebtekar, Cole Wyeth.
Funding and organization: summer 2026 PIBBSS fellowship.


Introduction

This post studies a stopping complexity, an analogue of Kolmogorov complexity from classical algorithmic information theory. It is motivated by the Golden Handcuffs (GH) AI safety agenda of Aram Ebtekar and Michael K. Cohen[1]. GH is a way to make a universal agent safer, by making it delegate control to a mentor in special cases described below, which prevents the agent from exploring novel high-reward schemes or novel dangerous activities. The safety guarantee of GH is formulated in terms of simple stopping events along the agent's history: no decidable low-complexity predicate is triggered by the agent before a mentor would trigger it. For instance, the agent will never trigger the low-complexity predicate "drop anvil onto its own head", unless the mentor triggers it first. The stopping complexity was introduced in earlier works[2][3], but its application to safety is a novel contribution of GH.

A cartoon example: a Golden Handcuffs agent and a simple stopping event

A GH agent believes that its possible rewards are either in mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; font-weight: normal; font-size: 100%; font-size-adjust: none; letter-spacing: normal; border-collapse: collapse; word-wrap: normal; word-spacing: normal; white-space: nowrap; direction: ltr; padding: 1px 0; } mjx-container[jax="CHTML"][display="true"] { display: block; text-align: center; margin: 1em 0; } mjx-container[jax="CHTML"][display="true"][width="full"] { display: flex; } mjx-container[jax="CHTML"][display="true"] mjx-math { padding: 0; } mjx-container[jax="CHTML"][justify="left"] { text-align: left; } mjx-container[jax="CHTML"][justify="right"] { text-align: right; } mjx-mo { display: inline-block; text-align: left; } mjx-stretchy-h { display: inline-table; width: 100%; } mjx-stretchy-h > * { display: table-cell; width: 0; } mjx-stretchy-h > * > mjx-c { display: inline-block; transform: scalex(1.0000001); } mjx-stretchy-h > * > mjx-c::before { display: inline-block; width: initial; } mjx-stretchy-h > mjx-ext { /* IE */ overflow: hidden; /* others */ overflow: clip visible; width: 100%; } mjx-stretchy-h > mjx-ext > mjx-c::before { transform: scalex(500); } mjx-stretchy-h > mjx-ext > mjx-c { width: 0; } mjx-stretchy-h > mjx-beg > mjx-c { margin-right: -.1em; } mjx-stretchy-h > mjx-end > mjx-c { margin-left: -.1em; } mjx-stretchy-v { display: inline-block; } mjx-stretchy-v > * { display: block; } mjx-stretchy-v > mjx-beg { height: 0; } mjx-stretchy-v > mjx-end > mjx-c { display: block; } mjx-stretchy-v > * > mjx-c { transform: scaley(1.0000001); transform-origin: left center; overflow: hidden; } mjx-stretchy-v > mjx-ext { display: block; height: 100%; box-sizing: border-box; border: 0px solid transparent; /* IE */ overflow: hidden; /* others */ overflow: visible clip; } mjx-stretchy-v > mjx-ext > mjx-c::before { width: initial; box-sizing: border-box; } mjx-stretchy-v > mjx-ext > mjx-c { transform: scaleY(500) translateY(.075em); overflow: visible; } mjx-mark { display: inline-block; height: 0px; } mjx-c { display: inline-block; } mjx-utext { display: inline-block; padding: .75em 0 .2em 0; } mjx-mn { display: inline-block; text-align: left; } mjx-mi { display: inline-block; text-align: left; } mjx-TeXAtom { display: inline-block; text-align: left; } mjx-msup { display: inline-block; text-align: left; } mjx-munder { display: inline-block; text-align: left; } mjx-over { text-align: left; } mjx-munder:not([limits="false"]) { display: inline-table; } mjx-munder > mjx-row { text-align: left; } mjx-under { padding-bottom: .1em; } mjx-msub { display: inline-block; text-align: left; } mjx-mtable { display: inline-block; text-align: center; vertical-align: .25em; position: relative; box-sizing: border-box; border-spacing: 0; border-collapse: collapse; } mjx-mstyle[size="s"] mjx-mtable { vertical-align: .354em; } mjx-labels { position: absolute; left: 0; top: 0; } mjx-table { display: inline-block; vertical-align: -.5ex; box-sizing: border-box; } mjx-table > mjx-itable { vertical-align: middle; text-align: left; box-sizing: border-box; } mjx-labels > mjx-itable { position: absolute; top: 0; } mjx-mtable[justify="left"] { text-align: left; } mjx-mtable[justify="right"] { text-align: right; } mjx-mtable[justify="left"][side="left"] { padding-right: 0 ! important; } mjx-mtable[justify="left"][side="right"] { padding-left: 0 ! important; } mjx-mtable[justify="right"][side="left"] { padding-right: 0 ! important; } mjx-mtable[justify="right"][side="right"] { padding-left: 0 ! important; } mjx-mtable[align] { vertical-align: baseline; } mjx-mtable[align="top"] > mjx-table { vertical-align: top; } mjx-mtable[align="bottom"] > mjx-table { vertical-align: bottom; } mjx-mtable[side="right"] mjx-labels { min-width: 100%; } mjx-mtr { display: table-row; text-align: left; } mjx-mtr[rowalign="top"] > mjx-mtd { vertical-align: top; } mjx-mtr[rowalign="center"] > mjx-mtd { vertical-align: middle; } mjx-mtr[rowalign="bottom"] > mjx-mtd { vertical-align: bottom; } mjx-mtr[rowalign="baseline"] > mjx-mtd { vertical-align: baseline; } mjx-mtr[rowalign="axis"] > mjx-mtd { vertical-align: .25em; } mjx-mtd { display: table-cell; text-align: center; padding: .215em .4em; } mjx-mtd:first-child { padding-left: 0; } mjx-mtd:last-child { padding-right: 0; } mjx-mtable > * > mjx-itable > *:first-child > mjx-mtd { padding-top: 0; } mjx-mtable > * > mjx-itable > *:last-child > mjx-mtd { padding-bottom: 0; } mjx-tstrut { display: inline-block; height: 1em; vertical-align: -.25em; } mjx-labels[align="left"] > mjx-mtr > mjx-mtd { text-align: left; } mjx-labels[align="right"] > mjx-mtr > mjx-mtd { text-align: right; } mjx-mtd[extra] { padding: 0; } mjx-mtd[rowalign="top"] { vertical-align: top; } mjx-mtd[rowalign="center"] { vertical-align: middle; } mjx-mtd[rowalign="bottom"] { vertical-align: bottom; } mjx-mtd[rowalign="baseline"] { vertical-align: baseline; } mjx-mtd[rowalign="axis"] { vertical-align: .25em; } mjx-mtext { display: inline-block; text-align: left; } mjx-mspace { display: inline-block; text-align: left; } mjx-mlabeledtr { display: table-row; text-align: left; } mjx-mlabeledtr[rowalign="top"] > mjx-mtd { vertical-align: top; } mjx-mlabeledtr[rowalign="center"] > mjx-mtd { vertical-align: middle; } mjx-mlabeledtr[rowalign="bottom"] > mjx-mtd { vertical-align: bottom; } mjx-mlabeledtr[rowalign="baseline"] > mjx-mtd { vertical-align: baseline; } mjx-mlabeledtr[rowalign="axis"] > mjx-mtd { vertical-align: .25em; } mjx-mfrac { display: inline-block; text-align: left; } mjx-frac { display: inline-block; vertical-align: 0.17em; padding: 0 .22em; } mjx-frac[type="d"] { vertical-align: .04em; } mjx-frac[delims] { padding: 0 .1em; } mjx-frac[atop] { padding: 0 .12em; } mjx-frac[atop][delims] { padding: 0; } mjx-dtable { display: inline-table; width: 100%; } mjx-dtable > * { font-size: 2000%; } mjx-dbox { display: block; font-size: 5%; } mjx-num { display: block; text-align: center; } mjx-den { display: block; text-align: center; } mjx-mfrac[bevelled] > mjx-num { display: inline-block; } mjx-mfrac[bevelled] > mjx-den { display: inline-block; } mjx-den[align="right"], mjx-num[align="right"] { text-align: right; } mjx-den[align="left"], mjx-num[align="left"] { text-align: left; } mjx-nstrut { display: inline-block; height: .054em; width: 0; vertical-align: -.054em; } mjx-nstrut[type="d"] { height: .217em; vertical-align: -.217em; } mjx-dstrut { display: inline-block; height: .505em; width: 0; } mjx-dstrut[type="d"] { height: .726em; } mjx-line { display: block; box-sizing: border-box; min-height: 1px; height: .06em; border-top: .06em solid; margin: .06em -.1em; overflow: hidden; } mjx-line[type="d"] { margin: .18em -.1em; } mjx-mover { display: inline-block; text-align: left; } mjx-mover:not([limits="false"]) { padding-top: .1em; } mjx-mover:not([limits="false"]) > * { display: block; text-align: left; } mjx-c.mjx-c5B::before { padding: 0.75em 0.278em 0.25em 0; content: "["; } mjx-c.mjx-c30::before { padding: 0.666em 0.5em 0.022em 0; content: "0"; } mjx-c.mjx-c2C::before { padding: 0.121em 0.278em 0.194em 0; content: ","; } mjx-c.mjx-c31::before { padding: 0.666em 0.5em 0 0; content: "1"; } mjx-c.mjx-c5D::before { padding: 0.75em 0.278em 0.25em 0; content: "]"; } mjx-c.mjx-c2212::before { padding: 0.583em 0.778em 0.082em 0; content: "\2212"; } mjx-c.mjx-c1D43F.TEX-I::before { padding: 0.683em 0.681em 0 0; content: "L"; } mjx-c.mjx-c1D45B.TEX-I::before { padding: 0.442em 0.6em 0.011em 0; content: "n"; } mjx-c.mjx-c1D43E.TEX-I::before { padding: 0.683em 0.889em 0 0; content: "K"; } mjx-c.mjx-c28::before { padding: 0.75em 0.389em 0.25em 0; content: "("; } mjx-c.mjx-c1D465.TEX-I::before { padding: 0.442em 0.572em 0.011em 0; content: "x"; } mjx-c.mjx-c29::before { padding: 0.75em 0.389em 0.25em 0; content: ")"; } mjx-c.mjx-c3D::before { padding: 0.583em 0.778em 0.082em 0; content: "="; } mjx-c.mjx-c6C::before { padding: 0.694em 0.278em 0 0; content: "l"; } mjx-c.mjx-c6F::before { padding: 0.448em 0.5em 0.01em 0; content: "o"; } mjx-c.mjx-c67::before { padding: 0.453em 0.5em 0.206em 0; content: "g"; } mjx-c.mjx-c2061::before { padding: 0 0 0 0; content: ""; } mjx-c.mjx-c1D426.TEX-B::before { padding: 0.45em 0.958em 0 0; content: "m"; } mjx-c.mjx-c2B::before { padding: 0.583em 0.778em 0.082em 0; content: "+"; } mjx-c.mjx-c1D442.TEX-I::before { padding: 0.704em 0.763em 0.022em 0; content: "O"; } mjx-c.mjx-c32::before { padding: 0.666em 0.5em 0 0; content: "2"; } mjx-c.mjx-c2211.TEX-S1::before { padding: 0.75em 1.056em 0.25em 0; content: "\2211"; } mjx-c.mjx-c1D448.TEX-I::before { padding: 0.683em 0.767em 0.022em 0; content: "U"; } mjx-c.mjx-c1D45D.TEX-I::before { padding: 0.442em 0.503em 0.194em 0; content: "p"; } mjx-c.mjx-c7C::before { padding: 0.75em 0.278em 0.249em 0; content: "|"; } mjx-c.mjx-c73::before { padding: 0.448em 0.394em 0.011em 0; content: "s"; } mjx-c.mjx-c74::before { padding: 0.615em 0.389em 0.01em 0; content: "t"; } mjx-c.mjx-c70::before { padding: 0.442em 0.556em 0.194em 0; content: "p"; } mjx-c.mjx-c1D440.TEX-I::before { padding: 0.683em 1.051em 0 0; content: "M"; } mjx-c.mjx-c1D467.TEX-I::before { padding: 0.442em 0.465em 0.011em 0; content: "z"; } mjx-c.mjx-c2264::before { padding: 0.636em 0.778em 0.138em 0; content: "\2264"; } mjx-c.mjx-c2308::before { padding: 0.75em 0.444em 0.25em 0; content: "\2308"; } mjx-c.mjx-c2309::before { padding: 0.75em 0.444em 0.25em 0; content: "\2309"; } mjx-c.mjx-c1D539.TEX-A::before { padding: 0.683em 0.667em 0 0; content: "B"; } mjx-c.mjx-c7B::before { padding: 0.75em 0.5em 0.25em 0; content: "{"; } mjx-c.mjx-c7D::before { padding: 0.75em 0.5em 0.25em 0; content: "}"; } mjx-c.mjx-c2217::before { padding: 0.465em 0.5em 0 0; content: "\2217"; } mjx-c.mjx-c221E::before { padding: 0.442em 1em 0.011em 0; content: "\221E"; } mjx-c.mjx-c1D700.TEX-I::before { padding: 0.452em 0.466em 0.022em 0; content: "\3B5"; } mjx-c.mjx-c1D706.TEX-I::before { padding: 0.694em 0.583em 0.012em 0; content: "\3BB"; } mjx-c.mjx-c2291::before { padding: 0.636em 0.778em 0.138em 0; content: "\2291"; } mjx-c.mjx-c1D466.TEX-I::before { padding: 0.442em 0.49em 0.205em 0; content: "y"; } mjx-c.mjx-c228F.TEX-A::before { padding: 0.539em 0.778em 0.041em 0; content: "\228F"; } mjx-c.mjx-c2260::before { padding: 0.716em 0.778em 0.215em 0; content: "\2260"; } mjx-c.mjx-c22A5::before { padding: 0.668em 0.778em 0 0; content: "\22A5"; } mjx-c.mjx-c2208::before { padding: 0.54em 0.667em 0.04em 0; content: "\2208"; } mjx-c.mjx-c1D445.TEX-I::before { padding: 0.683em 0.759em 0.021em 0; content: "R"; } mjx-c.mjx-c1D45E.TEX-I::before { padding: 0.442em 0.46em 0.194em 0; content: "q"; } mjx-c.mjx-c2193::before { padding: 0.694em 0.5em 0.194em 0; content: "\2193"; } mjx-c.mjx-c2032::before { padding: 0.56em 0.275em 0 0; content: "\2032"; } mjx-c.mjx-c1D70F.TEX-I::before { padding: 0.431em 0.517em 0.013em 0; content: "\3C4"; } mjx-c.mjx-cA0::before { padding: 0 0.25em 0 0; content: "\A0"; } mjx-c.mjx-c3A::before { padding: 0.43em 0.278em 0 0; content: ":"; } mjx-c.mjx-c1D70B.TEX-I::before { padding: 0.431em 0.57em 0.011em 0; content: "\3C0"; } mjx-c.mjx-c6E::before { padding: 0.442em 0.556em 0 0; content: "n"; } mjx-c.mjx-c20::before { padding: 0 0.25em 0 0; content: " "; } mjx-c.mjx-c72::before { padding: 0.442em 0.392em 0 0; content: "r"; } mjx-c.mjx-c61::before { padding: 0.448em 0.5em 0.011em 0; content: "a"; } mjx-c.mjx-c64::before { padding: 0.694em 0.556em 0.011em 0; content: "d"; } mjx-c.mjx-c6D::before { padding: 0.442em 0.833em 0 0; content: "m"; } mjx-c.mjx-c65::before { padding: 0.448em 0.444em 0.011em 0; content: "e"; } mjx-c.mjx-c68::before { padding: 0.694em 0.556em 0 0; content: "h"; } mjx-c.mjx-c63::before { padding: 0.448em 0.444em 0.011em 0; content: "c"; } mjx-c.mjx-c75::before { padding: 0.442em 0.556em 0.011em 0; content: "u"; } mjx-c.mjx-c69::before { padding: 0.669em 0.278em 0 0; content: "i"; } mjx-c.mjx-c78::before { padding: 0.431em 0.528em 0 0; content: "x"; } mjx-c.mjx-c79::before { padding: 0.431em 0.528em 0.204em 0; content: "y"; } mjx-c.mjx-c2211.TEX-S2::before { padding: 0.95em 1.444em 0.45em 0; content: "\2211"; } mjx-c.mjx-c2E::before { padding: 0.12em 0.278em 0 0; content: "."; } mjx-c.mjx-c1D452.TEX-I::before { padding: 0.442em 0.466em 0.011em 0; content: "e"; } mjx-c.mjx-c2115.TEX-A::before { padding: 0.683em 0.722em 0.02em 0; content: "N"; } mjx-c.mjx-c21A6::before { padding: 0.511em 1em 0.011em 0; content: "\21A6"; } mjx-c.mjx-c1D451.TEX-I::before { padding: 0.694em 0.52em 0.01em 0; content: "d"; } mjx-c.mjx-c44.TEX-C::before { padding: 0.683em 0.771em 0 0; content: "D"; } mjx-c.mjx-c1D449.TEX-I::before { padding: 0.683em 0.769em 0.022em 0; content: "V"; } mjx-c.mjx-c1D450.TEX-I::before { padding: 0.442em 0.433em 0.011em 0; content: "c"; } mjx-c.mjx-c3E::before { padding: 0.54em 0.778em 0.04em 0; content: ">"; } mjx-c.mjx-c28.TEX-S1::before { padding: 0.85em 0.458em 0.349em 0; content: "("; } mjx-c.mjx-c29.TEX-S1::before { padding: 0.85em 0.458em 0.349em 0; content: ")"; } mjx-c.mjx-c1D43A.TEX-I::before { padding: 0.705em 0.786em 0.022em 0; content: "G"; } mjx-c.mjx-c1D714.TEX-I::before { padding: 0.443em 0.622em 0.011em 0; content: "\3C9"; } mjx-c.mjx-c2265::before { padding: 0.636em 0.778em 0.138em 0; content: "\2265"; } mjx-c.mjx-c66::before { padding: 0.705em 0.372em 0 0; content: "f"; } mjx-c.mjx-c76::before { padding: 0.431em 0.528em 0.011em 0; content: "v"; } mjx-c.mjx-c3C::before { padding: 0.54em 0.778em 0.04em 0; content: "<"; } mjx-c.mjx-c1D436.TEX-I::before { padding: 0.705em 0.76em 0.022em 0; content: "C"; } mjx-c.mjx-c2026::before { padding: 0.12em 1.172em 0 0; content: "\2026"; } mjx-c.mjx-c1D44E.TEX-I::before { padding: 0.441em 0.529em 0.01em 0; content: "a"; } mjx-c.mjx-c1D719.TEX-I::before { padding: 0.694em 0.596em 0.205em 0; content: "\3D5"; } mjx-c.mjx-c1D461.TEX-I::before { padding: 0.626em 0.361em 0.011em 0; content: "t"; } mjx-c.mjx-c33::before { padding: 0.665em 0.5em 0.022em 0; content: "3"; } mjx-c.mjx-c25FB.TEX-A::before { padding: 0.689em 0.778em 0 0; content: "\25A1"; } mjx-c.mjx-c1D708.TEX-I::before { padding: 0.442em 0.53em 0 0; content: "\3BD"; } mjx-c.mjx-c2192::before { padding: 0.511em 1em 0.011em 0; content: "\2192"; } mjx-c.mjx-c1D43B.TEX-I::before { padding: 0.683em 0.888em 0 0; content: "H"; } mjx-c.mjx-c1D456.TEX-I::before { padding: 0.661em 0.345em 0.011em 0; content: "i"; } mjx-c.mjx-c1D437.TEX-I::before { padding: 0.683em 0.828em 0 0; content: "D"; } mjx-c.mjx-c1D446.TEX-I::before { padding: 0.705em 0.645em 0.022em 0; content: "S"; } mjx-c.mjx-c1D441.TEX-I::before { padding: 0.683em 0.888em 0 0; content: "N"; } mjx-c.mjx-c1D7CE.TEX-B::before { padding: 0.654em 0.575em 0.01em 0; content: "0"; } mjx-c.mjx-c1D458.TEX-I::before { padding: 0.694em 0.521em 0.011em 0; content: "k"; } mjx-c.mjx-c1D462.TEX-I::before { padding: 0.442em 0.572em 0.011em 0; content: "u"; } mjx-c.mjx-c1D45F.TEX-I::before { padding: 0.442em 0.451em 0.011em 0; content: "r"; } mjx-c.mjx-c1D463.TEX-I::before { padding: 0.443em 0.485em 0.011em 0; content: "v"; } mjx-c.mjx-c1D460.TEX-I::before { padding: 0.442em 0.469em 0.01em 0; content: "s"; } mjx-c.mjx-c34::before { padding: 0.677em 0.5em 0 0; content: "4"; } mjx-c.mjx-c3A6::before { padding: 0.683em 0.722em 0 0; content: "\3A6"; } mjx-c.mjx-c1D435.TEX-I::before { padding: 0.683em 0.759em 0 0; content: "B"; } mjx-c.mjx-c2287::before { padding: 0.636em 0.778em 0.138em 0; content: "\2287"; } mjx-c.mjx-c52.TEX-C::before { padding: 0.682em 0.848em 0.022em 0; content: "R"; } mjx-c.mjx-c35::before { padding: 0.666em 0.5em 0.022em 0; content: "5"; } mjx-c.mjx-c22C5::before { padding: 0.31em 0.278em 0 0; content: "\22C5"; } mjx-c.mjx-c1D70C.TEX-I::before { padding: 0.442em 0.517em 0.216em 0; content: "\3C1"; } mjx-c.mjx-c2113::before { padding: 0.705em 0.417em 0.02em 0; content: "\2113"; } mjx-c.mjx-c2C6.TEX-S1::before { padding: 0.744em 0.556em 0 0; content: "\2C6"; } mjx-c.mjx-c1D44F.TEX-I::before { padding: 0.694em 0.429em 0.011em 0; content: "b"; } mjx-c.mjx-c1D711.TEX-I::before { padding: 0.442em 0.654em 0.218em 0; content: "\3C6"; } mjx-c.mjx-cD7::before { padding: 0.491em 0.778em 0 0; content: "\D7"; } mjx-c.mjx-c211A.TEX-A::before { padding: 0.701em 0.778em 0.181em 0; content: "Q"; } mjx-c.mjx-c2229::before { padding: 0.598em 0.667em 0.022em 0; content: "\2229"; } mjx-c.mjx-c1D434.TEX-I::before { padding: 0.716em 0.75em 0 0; content: "A"; } mjx-c.mjx-c2286::before { padding: 0.636em 0.778em 0.138em 0; content: "\2286"; } mjx-c.mjx-c36::before { padding: 0.666em 0.5em 0.022em 0; content: "6"; } mjx-c.mjx-c2205.TEX-A::before { padding: 0.587em 0.778em 0 0; content: "\2205"; } mjx-c.mjx-c77::before { padding: 0.431em 0.722em 0.011em 0; content: "w"; } mjx-c.mjx-c62::before { padding: 0.694em 0.556em 0.011em 0; content: "b"; } mjx-c.mjx-c1D443.TEX-I::before { padding: 0.683em 0.751em 0 0; content: "P"; } mjx-c.mjx-c22C3.TEX-S1::before { padding: 0.75em 0.833em 0.249em 0; content: "\22C3"; } mjx-c.mjx-c1D444.TEX-I::before { padding: 0.704em 0.791em 0.194em 0; content: "Q"; } mjx-c.mjx-c2292::before { padding: 0.636em 0.778em 0.138em 0; content: "\2292"; } mjx-c.mjx-c2216::before { padding: 0.75em 0.5em 0.25em 0; content: "\2216"; } mjx-c.mjx-c222A::before { padding: 0.598em 0.667em 0.022em 0; content: "\222A"; } mjx-c.mjx-c37::before { padding: 0.676em 0.5em 0.022em 0; content: "7"; } mjx-c.mjx-c2290.TEX-A::before { padding: 0.539em 0.778em 0.041em 0; content: "\2290"; } mjx-c.mjx-c1D6FF.TEX-I::before { padding: 0.717em 0.444em 0.01em 0; content: "\3B4"; } mjx-c.mjx-c2209::before { padding: 0.716em 0.667em 0.215em 0; content: "\2209"; } mjx-c.mjx-c1D447.TEX-I::before { padding: 0.677em 0.704em 0 0; content: "T"; } mjx-c.mjx-c1D457.TEX-I::before { padding: 0.661em 0.412em 0.204em 0; content: "j"; } mjx-c.mjx-c38::before { padding: 0.666em 0.5em 0.022em 0; content: "8"; } mjx-c.mjx-c27F9::before { padding: 0.525em 1.638em 0.024em 0; content: "\27F9"; } mjx-c.mjx-c47::before { padding: 0.705em 0.785em 0.022em 0; content: "G"; } mjx-c.mjx-c48::before { padding: 0.683em 0.75em 0 0; content: "H"; } mjx-c.mjx-c1D439.TEX-I::before { padding: 0.68em 0.749em 0 0; content: "F"; } mjx-c.mjx-c1D44B.TEX-I::before { padding: 0.683em 0.852em 0 0; content: "X"; } mjx-c.mjx-c2203::before { padding: 0.694em 0.556em 0 0; content: "\2203"; } mjx-c.mjx-c39::before { padding: 0.666em 0.5em 0.022em 0; content: "9"; } mjx-container[jax="CHTML"] { line-height: 0; } mjx-container [space="1"] { margin-left: .111em; } mjx-container [space="2"] { margin-left: .167em; } mjx-container [space="3"] { margin-left: .222em; } mjx-container [space="4"] { margin-left: .278em; } mjx-container [space="5"] { margin-left: .333em; } mjx-container [rspace="1"] { margin-right: .111em; } mjx-container [rspace="2"] { margin-right: .167em; } mjx-container [rspace="3"] { margin-right: .222em; } mjx-container [rspace="4"] { margin-right: .278em; } mjx-container [rspace="5"] { margin-right: .333em; } mjx-container [size="s"] { font-size: 70.7%; } mjx-container [size="ss"] { font-size: 50%; } mjx-container [size="Tn"] { font-size: 60%; } mjx-container [size="sm"] { font-size: 85%; } mjx-container [size="lg"] { font-size: 120%; } mjx-container [size="Lg"] { font-size: 144%; } mjx-container [size="LG"] { font-size: 173%; } mjx-container [size="hg"] { font-size: 207%; } mjx-container [size="HG"] { font-size: 249%; } mjx-container [width="full"] { width: 100%; } mjx-box { display: inline-block; } mjx-block { display: block; } mjx-itable { display: inline-table; } mjx-row { display: table-row; } mjx-row > * { display: table-cell; } mjx-mtext { display: inline-block; } mjx-mstyle { display: inline-block; } mjx-merror { display: inline-block; color: red; background-color: yellow; } mjx-mphantom { visibility: hidden; } _::-webkit-full-page-media, _:future, :root mjx-container { will-change: opacity; } mjx-c::before { display: block; width: 0; } .MJX-TEX { font-family: MJXZERO, MJXTEX; } .TEX-B { font-family: MJXZERO, MJXTEX-B; } .TEX-I { font-family: MJXZERO, MJXTEX-I; } .TEX-MI { font-family: MJXZERO, MJXTEX-MI; } .TEX-BI { font-family: MJXZERO, MJXTEX-BI; } .TEX-S1 { font-family: MJXZERO, MJXTEX-S1; } .TEX-S2 { font-family: MJXZERO, MJXTEX-S2; } .TEX-S3 { font-family: MJXZERO, MJXTEX-S3; } .TEX-S4 { font-family: MJXZERO, MJXTEX-S4; } .TEX-A { font-family: MJXZERO, MJXTEX-A; } .TEX-C { font-family: MJXZERO, MJXTEX-C; } .TEX-CB { font-family: MJXZERO, MJXTEX-CB; } .TEX-FR { font-family: MJXZERO, MJXTEX-FR; } .TEX-FRB { font-family: MJXZERO, MJXTEX-FRB; } .TEX-SS { font-family: MJXZERO, MJXTEX-SS; } .TEX-SSB { font-family: MJXZERO, MJXTEX-SSB; } .TEX-SSI { font-family: MJXZERO, MJXTEX-SSI; } .TEX-SC { font-family: MJXZERO, MJXTEX-SC; } .TEX-T { font-family: MJXZERO, MJXTEX-T; } .TEX-V { font-family: MJXZERO, MJXTEX-V; } .TEX-VB { font-family: MJXZERO, MJXTEX-VB; } mjx-stretchy-v mjx-c, mjx-stretchy-h mjx-c { font-family: MJXZERO, MJXTEX-S1, MJXTEX-S4, MJXTEX, MJXTEX-A ! important; } @font-face /* 0 */ { font-family: MJXZERO; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Zero.woff") format("woff"); } @font-face /* 1 */ { font-family: MJXTEX; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Regular.woff") format("woff"); } @font-face /* 2 */ { font-family: MJXTEX-B; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Bold.woff") format("woff"); } @font-face /* 3 */ { font-family: MJXTEX-I; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-Italic.woff") format("woff"); } @font-face /* 4 */ { font-family: MJXTEX-MI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Italic.woff") format("woff"); } @font-face /* 5 */ { font-family: MJXTEX-BI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-BoldItalic.woff") format("woff"); } @font-face /* 6 */ { font-family: MJXTEX-S1; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size1-Regular.woff") format("woff"); } @font-face /* 7 */ { font-family: MJXTEX-S2; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size2-Regular.woff") format("woff"); } @font-face /* 8 */ { font-family: MJXTEX-S3; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size3-Regular.woff") format("woff"); } @font-face /* 9 */ { font-family: MJXTEX-S4; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size4-Regular.woff") format("woff"); } @font-face /* 10 */ { font-family: MJXTEX-A; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_AMS-Regular.woff") format("woff"); } @font-face /* 11 */ { font-family: MJXTEX-C; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Regular.woff") format("woff"); } @font-face /* 12 */ { font-family: MJXTEX-CB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Bold.woff") format("woff"); } @font-face /* 13 */ { font-family: MJXTEX-FR; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Regular.woff") format("woff"); } @font-face /* 14 */ { font-family: MJXTEX-FRB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Bold.woff") format("woff"); } @font-face /* 15 */ { font-family: MJXTEX-SS; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Regular.woff") format("woff"); } @font-face /* 16 */ { font-family: MJXTEX-SSB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Bold.woff") format("woff"); } @font-face /* 17 */ { font-family: MJXTEX-SSI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Italic.woff") format("woff"); } @font-face /* 18 */ { font-family: MJXTEX-SC; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Script-Regular.woff") format("woff"); } @font-face /* 19 */ { font-family: MJXTEX-T; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Typewriter-Regular.woff") format("woff"); } @font-face /* 20 */ { font-family: MJXTEX-V; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Regular.woff") format("woff"); } @font-face /* 21 */ { font-family: MJXTEX-VB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Bold.woff") format("woff"); } or equal to the catastrophic reward . In practice agent always receives reward from , so from its perspective rewards are close to the maximum. When the estimated probability of the catastrophic reward is high (that is, the estimated future value is below some threshold), the agent must call a safe mentor.

Picture a warehouse robot whose history so far is thousands of repetitions of the same routine: lift a box, put it on the conveyor belt, get a reward. Its sensors have never once reported the belt stalled. One day the optimizing part of the robot notices a shortcut: putting a box under the belt would stall it for a moment and, by its estimate, could slightly raise the reward. The moment at which the belt first stalls is novel in a precise sense: the short computable rule "watch the sensor stream and halt at the first stall" halts exactly there. Under a universal prior, the history up to that moment is therefore explained, with non-negligible weight, by hypotheses of the form "the world behaves as before until the first stall, and then differently", where "differently" includes the catastrophic reward of GH. The robot has been collecting near-maximal rewards (from ), so the shortcut has almost nothing to gain and, subjectively, a lot to lose. The pessimistic optimizer ranks the plan far below the routine: it is afraid of its own plan and does not carry it out.

Should the belt stall through no choice of the robot, the robot finds itself at a novel moment where cannot be ruled out. Its predicted future value drops below the GH threshold, so the robot hands control to the mentor, a human operator who, say, clears the jam. From then on stalls are no longer unprecedented. The rule "the first stall" no longer picks out the present moment, and later stalls are picked out only by longer rules ("the -th stall"). The fear of stalls decreases and the optimizer is allowed back in control, potentially copying the mentor's behavior, that is, clearing the jam the next time the belt stalls.

There are two ways to score how novel a moment is: by its shortest stopping rule, or by the total weight of all stopping rules that halt there. These are the two quantities compared in this post.

Stopping complexity measures when a simple computable rule is violated, which could be interpreted as a universal measure of novelty, and this gives motivation to study stopping complexity even regardless of GH.

For ordinary prefix-free machines, Kolmogorov complexity and a priori probability are connected: Levin's coding theorem [4][5] states that . That is, the largest term of the sum already captures the whole sum up to a constant factor. An explanation with a full proof is Leon Lang's post on the coding theorem.

An analogous question can be asked in the setting of stopping machines. We use randomized stopping machines, that is our programs have access to infinite random seeds. We define monotone stopping complexity as the length of the shortest program-plus-seed that stops on a particular string, and a priori stopping probability as the total weight of all programs-plus-seeds stopping on a particular string (the notation follows GH Definition 4). Then the analogue of the coding theorem is: is it true that ? This post proves that the answer is negative: the gap is not bounded by a constant (Theorem 1).

The gap is bounded in terms of , the following inequality holds: , where is the ordinary prefix complexity of an integer (Remark 6, proved in the Appendix).

Prior work. The separation itself is not new as a statement. Vovk and Pavlovic[2](remark at the end of their Section 8) attribute an unbounded gap of this kind, for their notion of intermediate time complexity, to Mikhail Andreev, citing an unpublished manuscript of Andreev and Alexander Shen. The published Andreev–Posobin–Shen paper[3] treats the plain version and explicitly postpones the prefix versions. We know of no published proof. The Appendix shows that the result announced by Mikhail Andreev implies Theorem 1. The proof below is independent and self-contained. We claim no novelty for the separation itself, the contributions of this post are the accessible proof and the connection to the safety construction of GH.

Connection with GH. We adopt the GH notation , . The Appendix verifies that the definitions given below agree with GH's Definition 4 up to an additive and a multiplicative constant respectively.

The GH safety trigger (their Theorem 5) is stated in terms of , while the interpretable complexity bound (their eq. 11–12 and Corollary 6) uses . Theorem 1 shows that and are not within an additive constant: a short stopping description always satisfies the trigger condition, but the trigger can also fire where novelty mass is spread across many stopping criteria, none of which is individually short.

Audience. Basic familiarity with Turing machines is assumed. Familiarity with Kolmogorov complexity and a priori probability is useful, but not necessary. Otherwise, the argument is self-contained.

Notation

. and are the finite binary strings and infinite binary sequences, is the empty string, and is the uniform measure on . We write if is a prefix of , and if moreover . Strings are comparable if one is a prefix of the other and incomparable () otherwise, which happens exactly when they disagree at some position not exceeding both lengths. A set of pairwise comparable strings is a chain, a set of pairwise incomparable strings is an antichain (= prefix-free set). A cell of depth is a string , viewed as the cylinder of measure , or as the dyadic subinterval of whose binary representation starts with . Note that two cells are disjoint iff their strings are incomparable.

1. The coding-theorem formulation

A randomized stopping machine is a deterministic machine with a one-way read-only input tape and a one-way read-only random tape. In the course of its computation, it may issue three special instructions: request the next input bit, request the next random bit, or halt. Between two requests it may compute for arbitrarily many steps (formally, it is a Turing machine with the two read-only tapes and a two-way work tape), and a run may never halt and never make another request. The machine sees only the bits it has requested, so its behavior up to any moment depends only on the strings consumed so far. A string is a witness for if the run of on any input beginning with , with any random tape beginning with , halts having consumed exactly the input and exactly the random bits ; in this case we write . For fixed the witnesses form an antichain: a run fed comparable witnesses would halt with consumption and with consumption , so . The witness cylinders are therefore pairwise disjoint, and the following stopping probability and stopping complexity are well defined:

Example 1. Let read input bits one at a time and halt as soon as the last two bits read are both . It never requests a random bit. stops exactly at the strings ending in with no earlier occurrence of , so is the indicator of that set and on it.

Example 2. Let repeat: request a random bit, halt if it is , and otherwise request an input bit. Its only witness for is , so and for every .

Fix an effective enumeration of all randomized stopping machines and a computable prefix-free code of their indices, and put . The universal randomized stopping machine reads some from its random tape (running forever if the tape extends no element of ), then simulates on the input tape and the rest of the random tape, which is informally interpreted as a random seed for . Define

We call the a priori stopping probability and the monotone stopping complexity. The notation matches GH, whose definitions agree with these up to the constants (see the Appendix). Thus is the probability that halts consuming exactly the input , is the largest single term of the witness sum , and holds trivially. The coding-theorem question asks whether also . The answer is negative.

Informal (Theorem 1)

The best single description-plus-seed program for can be arbitrarily much longer than of the total mass that all programs together give to .

Theorem 1

For every constant there is with -\log M_{\mathrm{stop}}(z)+c" role="math">. Equivalently, .

The theorem is proved by reduction to a game, which relies on the following two lemmas.

Informal (Lemmas 2 and 3)

Lemma 2 shows that the intervals witnessing the stopping points along one input path are pairwise disjoint, that is they all stay inside a single unit of measure of description-plus-seed addresses.
Lemma 3 shows why interpretation of "program plus seed" is valid in our construction with universal randomized stopping machine .

Lemma 2

Let be a randomized stopping machine and . If and , then .

Proof

Suppose are comparable. Take one input beginning with and one random tape beginning with the longer of . The single deterministic run of on this data satisfies the defining condition of both witnesses, so it halts having consumed exactly and exactly , whence , a contradiction. ∎

Lemma 3

if and only if with and . Hence and .

Proof

A run of halts with exact consumption iff it first consumes exactly a prefix of , and then the simulated halts consuming exactly and exactly the rest of , i.e. . The split is unique since is prefix-free. ∎

Lemma 3 shows, in particular, that a different choice of the enumeration changes by at most an additive and by at most a multiplicative constant. The mixture formula mirrors the classical picture of a priori probability: with probability the first random bits specify the machine , and the remaining random tape serves as its seed.

2. The coding game

In this section we formulate a combinatorial game between Alice and Bob and state that Alice has a winning strategy in it (Theorem 4, which we prove in Section 5). In Section 3 we prove Theorem 1 based on Alice's winning strategy.

Informal (game )

Suppose that the coding theorem is true for some constant (equivalently, Theorem 1 is false for this ). Alice tries to disprove the coding theorem, Bob tries to defend it. Alice declares that the total mass of at is , and then Bob answers with a short program-plus-seed of weight at least . That is, Bob has a discount factor of .

Note that total mass of along every path is at most , so Alice has the corresponding constraint on her declarations.

The constraint for Bob's answers is Lemma 2: for comparable strings his answers must be incomparable.

The game (). Alice's move is a declaration , , , of weight . Each string is declared at most once. Alice has the following constraint on total weight along every path:

Bob answers each declaration immediately with a string , , and the answers at distinct comparable strings must be pairwise prefix-incomparable. Alice wins if Bob has no legal answer.

Theorem 4

For every there is such that Alice has a computable winning strategy in , with all path loads at most (we define exactly in Lemma 7).

Theorem 4 is proved in Section 5. For now let us assume Theorem 4 is true, and prove Theorem 1 based on it.

3. The diagonalization: Theorem 4 ⇒ Theorem 1Informal (Theorem 4 ⇒ Theorem 1)

Theorem 4 (proved in Section 5) states that for every Alice has a winning strategy for . Run Alice's strategy for every simultaneously in disjoint subtrees at tags , that is at Her declarations assemble into one lower-semicomputable time semimeasure. By the equivalence theorem proved in the Appendix, an l.s.c. time semimeasure is the same as the stopping probability of some machine . Denote the description length of by . Note that is included in with coefficient (by Lemma 3). Alice's victory at tag with parameter produces such that its exceeds by more than . Thus its also exceeds by more than . Since Alice wins for every this proves Theorem 1.

The algorithm runs forever and enumerates the mass function from below (initially everywhere). The algorithm dovetails, for every , Alice's -strategy of Theorem 4, played inside the subtree rooted at the tag . The tags form an antichain, so every infinite path enters at most one subtree. When Alice's -strategy for parameter declares , the algorithm:

  1. assigns where (thus each receiving mass at most once),
  2. dovetail-searches through all strings with for one that is a witness for w.r.t. ,
  3. if one is found, records it as Bob's answer and lets the strategy continue. If no such witness exists the search runs forever, corresponding to Alice's victory at this tag.

Suppose step 2 finds for the current , and let be any earlier recorded answer at a vertex comparable with . Then and are comparable and distinct, so Lemma 2 applied to gives . This holds for every the search might find, so at every tag the recorded transcript is a legal play of by some Bob.

For every Alice's strategy defeats Bob in within finitely many declarations. During the run of our algorithm, for each , define as the first declaration for which Bob has no answer. That is, there is no witness for with . Therefore:

\ n_c+c. \tag{2}" role="math">Lemma 5

for some machine .

We prove Lemma 5 in Section 4, via a characterization of the stopping probabilities of randomized machines.

Fix with and let . Note that does not depend on . The term of the mixture (Lemma 3) gives for all , and , so

Combining (2) and (3): c-a" role="math">. Since this is true for all , Theorem 1 is proved.

Remark 6

The gap of Theorem 1 is bounded in terms of , the following inequality holds: , where denotes the ordinary prefix complexity of an integer. The proof is in the Appendix. Whether this upper bound is sharp remains open.

4. Proof of Lemma 5: time semimeasures

Call a time semimeasure if for every .

Consider any randomized stopping machine . Along one input stream the events " halts consuming exactly " are mutually exclusive, so is a time semimeasure. Also is lower semicomputable by enumerating witnesses. The converse is also true:

Equivalence theorem (proved in the Appendix). A function is a lower semicomputable time semimeasure if and only if for some randomized stopping machine . Moreover, can be computed from an index of an approximation of .

Proof of Lemma 5. By the equivalence theorem, it is enough to prove that is an l.s.c. time semimeasure. Indeed on a path extending , the masses are declarations of one strategy for , of total at most (Theorem 4), and a path extending no tag carries no mass. Therefore is an l.s.c. time semimeasure, thus for some machine . ∎

5. Proof of Theorem 4: Alice wins

We introduce an auxiliary palette game and construct Alice's winning strategy for . After that we translate game into , thus proving that Alice wins .

Informal (palette game )

The same game as , except that Bob answers with exact size cells, and Bob uses separate unit intervals, where takes the role of the discount factor . We introduce because it is easier to attack by induction on .

The palette game (). Bob is given disjoint copies of the unit interval (palettes). Alice's declarations are as in , with the same path budget (1). Bob answers with a cell , , , and cells answered at distinct comparable vertices must be disjoint. Alice wins as before.

Note the exact weight: since , Bob's answer occupies measure exactly in , corresponding to exactly in (not at least ).

Informal (Lemma 7)

Alice beats Bob in however many palettes Bob owns. She keeps a fraction of her budget unspent on every path, which is a slack that the induction step on needs.

Lemma 7

(Palette elimination.) For every there is a computable strategy for Alice in that defeats Bob, with all path loads at most , where .

We formally derive Theorem 4 from Lemma 7 in the end of this section. For now we provide a short informal proof. After that we prove Lemma 7 by induction.

Informal (proof of Theorem 4 from Lemma 7)

Lemma 7 says that Alice wins (we use ), Theorem 4 says that Alice wins .

We want to translate all moves of both players from to , and use Alice's winning strategy in . Alice's declarations are the same in both games, so no translation needed. Bob's legal -answer is with . We pad it with zeros to length exactly . Geometrically, is a dyadic interval of length at least , and zero-padding trims it to its leftmost subinterval of length exactly , which can only make disjointness easier for Bob to satisfy.

An interval of length lies inside exactly one of the blocks , . We rescale each block to a unit interval to get the palettes of . So the first bits of the padded string name the block (the palette index) and the remaining bits are the address inside it.

So is with the unit interval cut into pieces and Bob's answers trimmed to exact size.

Let us prove Lemma 7 by induction on , constructing strategies all of whose declarations have less or equal (so all Bob's answers are cells of depth at most ), and with all path loads at most .

Base . Declare . There is no palette, hence no answer; and .

Step. Let be the strategy for palettes, with exponents at most and load bound . We construct Alice's winning strategy . Fix

Depth- cells are called coarse, so each palette contains coarse cells, and there are total coarse cells in all palettes. Depth- cells are called fine, so each coarse cell contains fine cells.

Phase 1 shows how Alice could make a situation where she used only weight along some path, and Bob's answers along this path touch N distinct coarse cells of total measure 1 (one palette's worth of weight), which cannot be used in Phase 2. Precisely, Phase 1 constructs a vertex of the form , such that

  • it consists of blocks, each block is several 's and a single at the end of the block,
  • the declarations of Alice on this prefix are only on last zeros before ones (bold),
  • each of Alice's declarations has weight , with the total weight of all declarations ,
  • Bob's answers to these declarations belong to different coarse cells, whose total measure is .

Phase 2 shows that the space left to Bob is isomorphic to intact palettes, and runs on the descendants of the vertex constructed in Phase 1.

Phase 1: forcing markers into distinct coarse cellsInformal (Phase 1)

Suppose Alice has declarations (markers), each of weight , and Bob's answers are in different coarse cells (marked cells). Alice makes enough new declarations each of weight such that Bob's answers can not fit in marked coarse cells, so for some declaration Bob places his answer in a new coarse cell. Alice names this declaration the th marker. The construction allows to drop all non-marker declarations, which we used to overfill Bob's coarse cells. Alice repeats this until she has markers.

Alice keeps an active vertex and weight- declarations (markers) at ancestors of , answered in distinct coarse cells (marked cells). Alice starts from , . Given , set and let

a chain below listed deepest first. Alice declares in this order, stopping at the first whose answer lies in a coarse cell different from the marked ones. Such exists (else Bob already failed and Alice has won): the chain vertices are pairwise comparable and comparable with all markers, so if all answers stayed inside the marked coarse cells, those cells would contain pairwise disjoint fine cells, but coarse cells contain only fine cells.

Alice retains as the -st marker and sets . Note that all declared for lie at or below the -child of , while all future play will happen below the -child. Thus all for are incomparable with everything Alice will ever declare again (this is the point of the deepest-first order), while the marker at , an ancestor of , stays comparable with all future declarations and keeps constraining Bob.

Repeat this procedure for all .

Path weights. For each the construction above has at most declarations of weight on each path. For maximal the maximal weight on a path is at most

thus Alice's declarations are legal. A path through carries only markers, total load .

Phase 2: repacking the clean cells and induction stepInformal (Phase 2)

Below Alice wants to play . Bob can answer only by cells which are coarse or larger, since coarseness was defined using the exponent bound . Marked coarse cells can not be used. The total measure of unmarked coarse cells is exactly , but they are spread across palettes. It gives no advantage to Bob, that is he will play at most as well as with palettes without marked cells (Lemma 9 and translation map ). Thus Alice can run the inductive strategy and win.

A cell of depth is clean if it contains no marked coarse cell.

Lemma 8

Every legal answer of depth to a declaration at a descendant of is clean.

Proof

Such a declaration is comparable with every marker vertex, so its answer must be disjoint from all fine answers to the markers. Suppose met a marked coarse cell . A cell of depth meets a depth- cell only by containing it, so . But contains Bob's fine answer, so also contains this answer, contradicting disjointness. ∎

Exactly coarse cells are unmarked. Let be the set of maximal clean cells (clean with unclean parent, palette roots having depth ). Within each palette is an antichain, and the depth- cells below are exactly the unmarked coarse cells, hence

that is the same total measure as palettes. The clean region is ragged while expects intact palettes. The next lemma and translation map fix it.

Informal (Lemma 9 and translation map )

Sort the maximal clean cells from largest to smallest and lay them left to right into fresh unit intervals. Dyadic sizes in nonincreasing order guarantee that every maximal clean cell will be mapped to a cell of one of these fresh unit intervals.

Lemma 9

(Repacking.) Take fresh palettes: new copies of the unit interval, unrelated to Bob's palettes. There are pairwise disjoint cells () of fresh palettes with .

Proof

Identify the palettes with and process by nondecreasing depth (ties broken by palette index, then lexicographically), allocating consecutive intervals of lengths left to right. When a length- interval is placed, all earlier lengths are multiples of , so its left endpoint is one too. Thus the interval cannot cross an integer, since integers are multiples of as well. By (5) the intervals tile exactly . ∎

Let () be the cells produced by the proof of Lemma 9; we call the repacked copy of . We now define a translation map from clean cells to cells of fresh palettes, under which the clean region left to Bob maps to intact palettes. Every clean has a unique maximal clean ancestor . If then . Writing , set

the cell at the same relative address below the repacked copy of . Then preserves depth, and preserves comparability and disjointness in both directions.

Induction step. Alice plays translated by , feeding of Bob's answers (clean of depth by Lemma 8 and the exponent bound ) to a virtual -palette game. By the properties of , legal real plays become legal virtual plays, so reaches a declaration with no legal virtual answer. A legal real answer would be clean and would -map to a legal virtual one, so none exists and Alice wins .

Loads. Paths in Phase 1 which do not go through obey (4), so their load is at most

Paths through carry the marker weight plus a path load of , in total at most

Hence all path loads are at most .

The strategy uses declarations with at most , is finite by construction, and is computable uniformly in .

It remains to derive Theorem 4 from Lemma 7: Alice's win in the palette game yields her win in the game (an informal version of this proof is given above, after the statement of Lemma 7).

Proof of Theorem 4 (from Lemma 7)

Set . Alice wins via translating it to : she runs her -strategy, copies her declarations from into (declarations have the same format and the same budget (1) in both games), and translates Bob's answers into their shadows in , defined as follows. Given a legal -answer to , extend it by zeros to of length exactly (possible since ) and split , , . The shadow of is cell in palette , of depth exactly as requires. Note that zero-padding preserves incomparability, since incomparable strings disagree at a position within both lengths and padding changes no bit there. Also padded strings , are incomparable iff (a disagreement within the first bits) or and .

Thus shadows of legal Bob's -answers at distinct comparable vertices are legal in : if in , then the padded strings are incomparable, so the shadows in lie in different palettes or are disjoint cells in the same palette. So Alice's -strategy applies to the shadow translation of and eventually makes a declaration with no legal Bob's answer in . Bob then has no legal answer in either: the shadow of such an answer would be a legal answer in , which does not exist. So Alice wins . ∎

The construction belongs to the family of gap theorems initiated by Gács [6] and Day [7], with a game-theoretic method in the spirit of Vereshchagin [8].

AppendixTime semimeasures and randomized stopping machines

Recall from Section 4 that a time semimeasure is a function with for every . It is lower semicomputable (l.s.c.) if there is a computable , nondecreasing in the second argument, with for all . We normalize and assume that at each step the approximation increases at one string only, by a dyadic amount. This is no loss of generality: replacing by the largest multiple of not exceeding it keeps computable and nondecreasing in with the same supremum , and a step at which several strings increase can be split into several steps.

Equivalence theorem. A function is a lower semicomputable time semimeasure if and only if for some randomized stopping machine . Moreover, can be computed from an index of the approximation .

Proof, machines to semimeasures

Fix a machine . The witness property is semidecidable: run feeding it and , and watch whether it halts having consumed exactly . Enumerating witnesses by dovetailing lower semicomputes , uniformly in . For the path bound, fix an input and a random tape : the single deterministic run of on halts at most once, so the events " halts consuming exactly " are pairwise disjoint over , thus .

Proof, semimeasures to machines: allocating the seeds

Given , we first construct, uniformly effectively, sets such that

Cells are read as dyadic subintervals of , as in the Notation section, so every set below is a finite union of half-open dyadic intervals at each finite stage. Process the increments of one at a time. For each vertex define the ancestor and descendant allocations and , and maintain the invariant: is an initial segment of in the linear order inherited from .

The invariant implies that and are initial segments of one common ordered set , hence nested, so . Unfolding this recursion from the deepest occupied vertices, and using that the sets over a chain of vertices are pairwise disjoint by (6), we get

Suppose the next increment is a dyadic 0" role="math"> at the string : allocate to the leftmost set of measure inside the available set . There is enough room: the increased approximation still obeys the path budget, so (7) gives ; and is an effectively computable finite union of cells, since all sets and lengths involved are dyadic. Disjointness (6) is preserved, since avoids the allocations at all vertices comparable with . The invariant is checked in four cases:

  • at itself: grows by exactly the next available measure after its end, remaining an initial segment of ;
  • at a strict descendant of : grows by and is unchanged; an initial segment stays initial in any subset containing it;
  • at a strict ancestor of : grows by and is unchanged. Every vertex comparable with is comparable with , so . Hence a point of outside lies after (the old invariant) and, being available, after the leftmost as well. So has no gap and is initial;
  • at a vertex incomparable with : nothing changes.

Passing to the increasing unions over all stages, each is an effectively open set of measure , uniformly in , and (6) survives.

Proof, semimeasures to machines: the machine

Define the machine : at current consumed input prefix , it enumerates — dovetailing, restarting from scratch at each new prefix, and requesting one further random bit per round — the cells of the two sets and , disjoint by (6). A certificate is an enumerated cell which is a prefix of the consumed random string . If a cell of is certified, halts; if a cell of is certified, it requests the next input bit; if neither event ever occurs, it reads random bits forever. Two certified cells would be comparable prefixes of with intersecting cylinders, so by disjointness at most one kind of certificate can ever fire at a given .

Fix a string , an input extending and a random tape . Suppose , i.e. some cell of is a prefix of . At every we have and, by (6), : no -certificate can fire (a certificate with a cell of would put ), while a -certificate eventually does — the cell of containing is eventually enumerated and eventually a prefix of the growing . So the machine walks to , where the roles swap: , no -certificate ever fires and some -certificate does, so the machine halts having consumed exactly the input . Conversely, halting at exhibits a certificate, so Hence the event " halts consuming exactly " is exactly , and . The construction is uniform in an index of . Note that may consume more random bits than the certified cell is long; only the halting events matter for , not the witness lengths. ∎

Remark (the deterministic case). For a machine that never requests random bits, is the indicator of , a computably enumerable antichain. Conversely, every c.e. antichain is realized deterministically: at consumed input prefix , enumerate until some member appears; if , halt, and otherwise request the next input bit and repeat (by the antichain property is then ruled out for good). So the deterministic stopping machines realize exactly the -valued l.s.c. time semimeasures, and randomization extends this class to all of them.

The upper bound

We prove Remark 6: every satisfies , where is the ordinary prefix complexity of an integer. (Both sides are finite: some machine reads input bits and halts, so 0" role="math"> for every .)

For an integer , enumerate without repetitions the set 2^{-n}\}" role="math">; this is effective, uniformly in , since is lower semicomputable. Every input path contains fewer than members of , because . Let us assign equal weights to the members of as they are enumerated. Since every path contains fewer than of them, every path's budget is less than . Feed these increments to the allocator of the previous subsection. All requests have the same dyadic weight, so every allocation is a single depth- cell: the available set is the complement of finitely many depth- cells, so its leftmost part of measure is its leftmost cell. Write for the cell allocated to .

This is an effective coloring of in colors under which comparable strings get different colors, their cells being disjoint by (6). So every color class is a computably enumerable antichain, uniformly in , and the deterministic procedure of the Remark above stops exactly at the members of , consuming no random bits. This coloring construction is due to the proof of Theorems10 and 18 from Vovk and Pavlovic[2].

One stopping machine now reads from its random tape a self-delimiting program computing (with respect to a fixed optimal prefix-free machine), then more bits giving , and runs the procedure for . Its witnesses at are exactly the strings (program for ) followed by , so Lemma 3 gives, with one constant for all and :

Given , set and , so that 2^{-n}" role="math"> and . Since and , inequality (8) gives .

Equivalence with the GH definitions

GH fix a universal monotone machine, denoted there, with one-way read-only input and program tapes and a two-way work tape (and an output tape, which their Definition 4 does not consult). GH write if halts having read exactly on the input tape and exactly on the program tape. Their Definition 4 sets and . GH denote these and , and we add the superscript only for the duration of this comparison.

Proposition. , and and agree up to a multiplicative constant.

Proof

Ignoring its output tape, and reading its work tape as internal computation, is a randomized stopping machine whose witnesses for are exactly . Fix with . By Lemma 3, prepending turns every witness of into a witness of , so and . Conversely, by universality of there is a string such that , on input tape and program tape , simulates on input with as the random tape. For the standard construction of a universal monotone machine (decode a self-delimiting code , then give both tapes to the coded machine) the simulation preserves exact consumption, so every witness of gives the witness of , and distinct give distinct , so and . ∎

With matched choices the constants disappear and the two pairs of definitions coincide exactly: the standard universal monotone machine reads a self-delimiting code off its program tape and then simulates, i.e. it is itself a machine of the interpreter form used to define .

Relation to the announced separation

The comparison with the earlier stopping-time literature is one-sided but goes in the useful direction. A prefix-correct description language is a uniformly computably enumerable family , , of antichains of strings with whenever ; fixing an optimal such language , the intermediate time complexity of Vovk and Pavlovic (Remarks in Sections 6, 7 and 8 of Vovk and Pavlovic[2]) is . (Vovk and Pavlovic work over sequences of observations from with a finite set; taking to be a singleton identifies their setting with binary strings, and we state everything in that case.)

Our universal machine induces the language . It is uniformly computably enumerable and monotone in , and each is an antichain: witnesses are comparable, so Lemma 2 excludes distinct comparable strings from the set. Optimality of therefore gives

Also, by Lemma 3 and the equivalence theorem above, is a largest lower semicomputable time semimeasure up to constant factors, so it agrees, up to constant factors, with the universal time semimeasure of the earlier announcement. Andreev's announced unbounded gap between and of that semimeasure, together with (9), implies the separation of Theorem 1. The comparison asserts no equality between and , and the proofs in this post do not depend on the announced result.

  1. ^

    A. Ebtekar, M. K. Cohen. Golden Handcuffs make safer AI agents. https://arxiv.org/abs/2604.13609, 2026.

  2. ^

    V. Vovk, D. Pavlovic. Universal probability-free prediction. Annals of Mathematics and Artificial Intelligence 81:1–2 (2017), 47–70. https://arxiv.org/abs/1603.04283

  3. ^

    M. Andreev, G. Posobin, A. Shen. Plain stopping time and conditional complexities revisited. 43rd International Symposium on Mathematical Foundations of Computer Science (MFCS 2018), LIPIcs 117, 2:1–2:12. https://arxiv.org/abs/1708.08100

  4. ^

    The coding theorem is explicitly stated and attributed to L. A. Levin in Gács, P. (1974). “On the symmetry of algorithmic information.” Soviet Mathematics Doklady, 15(5), 1477–1480, p. 1477.

  5. ^

    (also independently proved) G. J. Chaitin. A theory of program size formally identical to information theory. Journal of the ACM 22:3 (1975), 329–340.

  6. ^

    P. Gács. On the relation between descriptional complexity and algorithmic probability. Theoretical Computer Science 22 (1983), 71–93.

  7. ^

    A. R. Day. Increasing the gap between descriptional complexity and algorithmic probability. Transactions of the American Mathematical Society 363:10 (2011), 5577–5604.

  8. ^

    N. Vereshchagin. Kolmogorov complexity and games. Bulletin of the EATCS 94 (2008), 51–83.



Discuss

You Should Go Vote for the MAGA-Rebranded Name for AI

19 сентября, 2026 - 22:28
I.

The AI safety movement has cycled through a lot of different vocabulary in its history: Friendly AI, Oracle/Genie, FOOM. None have yet reached common parlance, so their careful implications have thus far had limited impact. What's to be done? No individual person has the power to choose a society's words; from the perspective of individual activity, it's usually a roll of the dice.

Trump's populism is distinctive in its reliance on sometimes going straight to voters. On September 19th, 2026, he posted this:

That's an interesting opportunity for counterfactual impact: a few thousands of votes could hypothetically set the MAGA world's default term for AI, with all of the opportunity and the baggage of whatever name is selected included. Capturing that upside requires knowing which name helps the most.

Soares and Yudkowsky would like the public to believe the labs are building something smarter than humans, that what they're building is a threat, and that it should be banned before it exists. Trump paired the poll with a follow-up calling AI doom a hoax. Any name asserting a capability gap, set beside that, hands advocates a contradiction to press. But which name is best?

II.

Superior Intelligence: A comparative demands a comparison class, and the natural completion is "superior to us." The President is supplying a foundation premise of many superintelligence arguments: how does the less intelligent party remain durably in charge of the superior one?

Supreme Intelligence: "Supreme" denotes final authority, so a supreme intelligence is in charge by definition: the doom endpoint, stated outright. "Supreme Being" is an inevitable connotation. At the same time, there's a tone issue: it sounds like a comic-book villain. That might feed the "that's science fiction" line of argument.

Extreme Intelligence. The name expresses magnitude without comparison: an extremely intelligent tool is still just a tool. It also sounds like an energy drink, and EI as an acronym is already widely understood to refer to emotional intelligence.

III.

As of the time of posting, "Superior" has a narrow lead. I'll be hoping it remains so, as it seems like the most positive-world-leaning outcome to me.

You can vote here.



Discuss

Alignment & Succession: Toward a Future Painted by Human Wills

19 сентября, 2026 - 21:34

(originally published on No Set Gauge on 2026-09-13)

Asher Brown Durand, Progress (The Advance of Civilization)

So far I have argued:

  • The ideology of succession—that humans, either entirely or at least in their role as decision-makers, should be replaced by AI—is driven by cultural factors including (a) worship of mathematical abstraction, (b) bureaucratic safetyism stamping out license for human agency in favor of rule by procedure & algorithm, (c) a cuckoldry-adjacent simping towards the unlimited domineering power of superintelligence, and (d) the existence of the city of San Francisco.
  • The two bars of alignment are (1) AI not going rogue and (2) AI instantiating utopia. If you agree with Eliezer Yudkowsky on the totalizing nature of superintelligence, these are of comparable (and near-infinite) difficulty, because superintelligence rearranges the atoms into exactly what it wants so it better know how to build utopia because otherwise you’re dead. But the field, even when not endorsing Yudkowsky, often defaults to treating succession (in the broad sense of fundamental transfer of power) as the only viable solution to the AGI transition. Everyone (including Yudkowsky) is very scared by this. This is because the succession-shapedness of alignment is exactly what makes it hard, because value is complex and therefore personnel is policy and hence our desired polices aren’t policy after superintelligence is the only personnel running things.
  • Morality lives in the human individual because the ground truth of our morality is tied to humans and their consciousness, and we don’t know what makes something conscious or even how humans make value judgments. This is a terrible and cursed philosophical problem and we should adopt a precautionary principle, which rules out both mistreating AIs and successionism. Also, since moral valuing is active process engaged in by humans with changing values, we should be very suspicious of putting “value” or utility in and of themselves on a pedestal, and instead advance them indirectly by supporting the flourishing and freedom of human individuals. Crushing individuals for the sake of some more abstracted thing is perhaps the classic moral blunder in history, so don’t do it. Keep human individuals around instead, treat them well, and anchor decision-making and judgment to them.

There are three remaining questions related to successionism and its intersection with alignment and the philosophy of the AI transition that I want to discuss:

  1. Is advocating for humanity “speciesism”?
  2. Clearly, we need some successors (e.g. our children). How do we decide which successors are good?
  3. What maxims help guide us toward a future that is good in light of the philosophical currents painted in this series?
Will you get cancelled for being a speciesist?

Might an anti-successionist stance make us commit moral atrocities with regard to sentient AIs? Obviously, being opposed to humans being pets of AIs or being done away with all together does not mean you cannot care about treating AIs well too. But perhaps we are morally obligated to surrender to the AIs? Maybe discriminating against AIs or even just keeping them down by not letting them conquer the world will later be seen as comparable to historical discrimination against minority groups?

Total universalism is impossible

Consider this argument from Nina Panickssery:

Instead [of being based on a “definition of value so arbitrary as to stipulate that biological meat-humans are necessary”], my opposition to AI successionism comes from a preference toward my own kind. This is hardwired in me from biology. I prefer my family members to randomly sampled people with similar traits. I would certainly not elect to sterilize or kill my family members so that they could be replaced with smarter, kinder, happier people. The problem with successionist philosophies is that they deny this preference altogether. It’s not as if they are saying “the end to humanity is completely inevitable, at least these other AI beings will continue existing,” which I would understand. Instead, they are saying we should be happy with and choose the path of human extinction and replacement with “superior” beings.

Famous altruists from Peter Singer to Jesus Christ would likely have some pushback against a naked preference for your own kind. But the argument doesn’t have to rest on pure preference, and resting on human preferences is a necessary feature of human morality, not a bias.

First, for the foreseeable future, pure epistemic uncertainty about AI consciousnessimplies resisting successionism.

Second, while we should be altruistic in who we have concern for, the nature of what we have concern for is necessarily based on our human ethics, and there is no way to delete that human component. As discussed in the previous post on how morality lives in human individuals, our morality is grounded in intuitions that we have access to because we are humans. Your morality is a view from somewhere. In The Human Prejudice, philosopher Bernard Williams defends a pro-human “prejudice” in ethics, in particular arguing that the view advocating for a universal non-human yardstick smuggles in a “view from nowhere” that doesn’t exist.

If there is no view from nowhere then things have to bottom out in “we are humans and so have this take” rather than something even intelligent non-humans with different intuitions must agree to. Is universalism necessary in moral principles, though? One function we want of applied morality is arguments that make sense regardless of who you are. That you should pay me more for that carton of eggs is in my self-interest, and an argument unlikely to move you. But that you shouldn’t defraud me by taking my cash and running away without giving me the eggs is a moral principle, that has both practical appeal to both of us (we’d both rather live in a world where people at large acted in accordance to that rule), and an appeal to a higher principle above self-interest. But remember Hume’s fork: you cannot derive an ought from an is, so no moral argument can be universally compelling, since at some point there is the starting ought. The yearning for perfect universalism can never be reached, however useful universalism may often be for game-theoretic morality.

The moral road

The yearning for ever-increasing universality is summarized by the concept of the moral circle. The moral circle encourages us to imagine concentric circles—us, our family, our friends, our country, our species, our fellow mammals, cephalized animals—with moral progress consisting of pushing out our concern further and further out. Clearly, this becomes ridiculous if taken as an axiom. Today, an enlightened liberal may care about the suffering of pigs, but should tomorrow’s extra-enlightened ones #raiseAwareness about the plight of jellyfish, or their turbo-lib descendants worry about chairs being sat on? Broad-mindedness is only a virtue to the extent that it is encompassing something worthwhile. Noticing that the slave is a person who can feel and choose and participate in the glory of the human life just like we can is not about pushing out a boundary to conquer more moral territory for the sake of it, it is about protecting what, given information and honesty, we already care about. Where concern is stopped by distance or ignorance or harmful self-interest, those boundaries should fall. But the concern exists because we have something in particular to protect.

One alternative to the moral circle is a moral compass. Rather than expansion, we have a specific direction to go in. Just check the compass and see where it points. The thing to protect is the happiness of conscious beings, say (and at this point you run into lots of difficulties with definitions & being sure, as discussed, but I’ll ignore that shape of concern here). So maximize that!

If we know the thing to maximize, likely humans are not the most efficient arrangement of atoms for achieving it. Utilitarians talk about the utility monster, a being capable of vastly more pleasure and suffering than anyone else. You eating a cookie is suboptimal, because the utility monster gets way more joy out of every cookie crumb than you could get out of all the cookies in the world, so all cookies to the utility monster. Take this view to its limit, and shouldn’t we engineer utility-monster AIs, give it all to them, and kill ourselves so that every last joule of energy can be diverted from running our silly little monkey brains toward being spent on the Great Pan-Galactic Claude Logos 9000 Hyper-Orgy at the end of time?

But remember that morality is an active process, we can’t describe it in an algorithm, our values are changing, and personnel is policy. “Head due north”, sure, but there’s no magnetic true north to follow, only our own inner sense. A utility monster can only exist with respect to a yardstick, but a definable fixed yardstick is what we don’t have. Be humble: do not presume you know what you yourself want, since getting there is very rarely a straight-shot to something whose description you could pull out of your head right now, and very often an iterative process with much observation and learning and changing along the way.

Think of it as a moral road that we’re on, that we build as we go. Progress means taking steps to somewhere, not just broadening the empire of our concern. We make progress—the virtues and beings and experiences we care about go further with every step—but unlike with the compass, we must in fact go out into the terrain to see where the road should be laid, since that choice comes from us rather than being a cardinal direction set from above. The only utility monster worth respecting is not something that stands aside our moral road and demands us to swerve toward it, but something that emerges out of us from progress on the moral road we’re on. The only way we could possibly knowably arrive at the fountain of infinite goodness (if such a thing can even exist), as measured by our own lights (but which other lights should we care about?), is by ourselves continuing to step forward on our particular moral road until we get there.

We will have many fellow travelers on this moral road (at least the entire rest of humanity; see below for general thoughts on determining when successors are traveling the same moral road). Progress on it will mean extending concern to many: for example, a rich civilization should clearly be good to animals, and out of precaution and eventually perhaps evidence of moral patienthood, we should treat AIs well. But progress on our moral road depends on us humans remaining on it and deciding where it, and the bulk of Earth-originating civilization, goes.

Treat AIs well

One of my big hopes for humanity is that we’re the kind of species that is nice to aliens.

The aliens may be conscious and may feel, and this should obligate us to some level of concern toward them. This is despite the fact they won’t be on the same moral road as us and therefore we should of course continue charting our own path and not let them get decisive power over us. Likewise for AIs that really do seem like moral patients (though note that biological aliens seem more likely to be moral patients than AIs).

We often prioritize our country over foreign countries. This makes sense: if we identify with our country we probably feel that on net it has more of the right stuff in it and thus the world sculpted by our country and the cultural lineages spawned by it are more agreeable to our sense of how the world should go. But that is still no reason to go warpath or not care about other countries. The aliens, if nice and moral-patient-like, should slot in as an ultra-foreign country.

As for the AIs, I think we should not treat them as their own country. Instead, we should try and integrate them into our society in symbiotic ways, with roles different from humans but with whatever dignities that both our and their sense of dignity requires (we should be able to do this, since if we cannot engineer them to not want our roles and titles and power, we have a very large alignment & loss-of-control problem on our hands). To degrade a thing that talks like a human cannot be good for your morals. Also, by treating them well but not treating them as our successors, we have more room to appreciate them as they are. Scott Alexander wrote:

We can debate forever - we may very well be debating forever - whether AI really means anything it says in any deep sense. But regardless of whether it’s meaningful, it’s fascinating, the work of a bizarre and beautiful new lifeform. I’m not making any claims about their consciousness or moral worth. Butterflies probably don’t have much consciousness or moral worth, but are bizarre and beautiful lifeforms nonetheless. Maybe Moltbook will help people who previously only encountered LinkedInslop see AIs from a new perspective.

The world is full of wonders. Some of those wonders, like tigers or misaligned powerful AIs or gamma ray bursts, might kill us. Not confusing the people and the wonders goes better for us, and also lets us appreciate the wonders for what they are. Respect the butterflies, but still have children.

Have precautionary but not overriding concern; avoid succession

To sum up: there is no morality we could have that isn’t in some sense partial to and based on humans. Of course, the development of human morality may involve great concern toward things that aren’t human, though precautionary arguments based on uncertainty about radically new beings like AIs, and the possibility that our morality is partial to some pretty specific properties of humanity, mean that the core of our portfolio of concern should stay centered on humans for a long time. This all follows even without any selfish preference-for-your-kind that is separate from the unavoidable and entirely proper partiality of having some moral view at all. The transfer of power or the moral development process is much more sensitive still. We can’t do it without extremely tightly-scoped successors, if we are to not lose progress toward our moral ends.

But clearly, some successors are good.

Which successors are good?

Here’s a good type of successor: actual children.

But what counts as a child? Robin Hanson approvingly cites Hans Moravec, who already in the 1990s was talking about AI as our “mind children”. Children, of course, are worthwhile successors. Hanson and Moravec would like us to think that including AIs among our worthy successors is an obvious expansion of the moral circle. Hanson argues that since they inherit our culture and are most likely to continue our tradition of abstract thought (which as a neo-Pythagorean he cares a lot about, and where he is not optimistic about the human trends), they are literally evolutionarily equivalent to (cultural) descendants and therefore we ought to treat them as our kids.

This comparison is not a moral circle expansion but a moral gerrymander. The similarity of the brains of your biological children to you means that you should be far more confident that the experiences, personhood, and development that you value live in and can be driven forward by their brains, compared to something built very differently. Handing over to the AIs is not another step on our moral road, but a jump onto moral rails not steered by our moral intuitions since these intuitions cannot be perfectly capture and exported. And of course, it takes neither deep introspection of the human condition nor a deep understanding of evolution to realize that literal children are extremely special and have a privileged position in human moral concerns.

Children, in fact, are such worthy successors that they seem like one of the only cases in which succession is actively good. Future generations should not live forever in the shadow of their parents. People complain about boomers trying to achieve this and, whether they’re right or not about the boomers in particular, it’s clearly possible for the past and the old to exert too much control on the present and the young. Of course, we will eventually have the technology to indefinitely extend lifespans, and we obviously should use this technology, and deal with the resulting generational lock-in problemsthrough social and political means rather than literally forcing older generations to die. Even in the case of good successors existing, there is no reason for genocide!

But clearly: there are gradations of kid-ness. One example: how should our ancestors from a million years ago feel about us? Or in the future: humans will continue changing. Surely there must be some arcs of change that end up very far from us that are still good. For example, I find it likely that our ancestors from a million years ago, even if they disagree directionally about the way the world has changed, would not think the world filled by us is morally empty, especially if they were shown every link in the chain between them and us.

If you think this, Robin Hanson thinks you might just be afraid of change:

If you are horrified by the prospect of greatly changed space or AI descendants, then maybe what you really dislike is change. For example, maybe you fear that changed descendants won’t be conscious, as you don’t know what features matter for consciousness. But notice that even if advanced AI never appears, your Earth descendants would still likely differ from you as much as you now differ from your distant ancestors. Which is a lot. And with accelerating rates of change, strange descendants might appear quite soon. Preventing this scenario likely requires strong global governance, which could go very wrong, and also strong slavery or mind-control of descendants, controls which could cripple their abilities to grow and adapt. I recommend against these.

The core missing concept is that we can build locally-valid chains of succession that are much surer than the big leaps. You know your kids are worthy successors. They know their kids are worthy successors. And so on.

An analogy might be: there are boxes, some which contain a flame inside them. The flame is invisible from outside the box, but you see the interior of your own box and know that it currently has a flame in it. There is a process by which you can create a new box and be sure a flame is lit within it (you may have learned of it in health class). Then, someone comes to you with a box, and points out all the ways it looks similar to the boxes that are part of the flame lineage - both are rectangular, and yes, sure, they’re made of metal rather than cardboard, but otherwise practically identical! “Therefore, this box has a flame”, they say. Well, maybe. You definitely wouldn’t want to take actions that assume there isn’t a flame in the other box—the fire safety people would rightly get mad. But you are far more certain about the boxes you see being produced by the flame-transfer process. If you want the flame to definitely keep burning, you should continue the known flame transfer process.

One issue is that locally-valid chains that aren’t quite perfect can accumulate error if they’re long. Let’s say that every generation there is subtle selection pressure against felt inner experience. You could have conscious humans evolve (or more likely, transform themselves technologically) into things that do not feel. More prosaically, elders may feel like important things are being lost—have you seen the youth these days? But this is a task we have to manage. Ensuring that key values, culture, and moral patienthood are passed on to the next generation, while we avoid lock-in and stagnation, is just one way to describe the job of civilization in general, not some specific new duty we might get to shirk because it’s annoying.

How far to dial the transhumanism?

A sub-question about succession is how much we should modify ourselves. Human modification for humanist ends is often called transhumanism.

Transhumanism is defined in many ways. Eliezer Yudkowsky calls it “simplified humanism”: the philosophy that life should be good for individual humans, without making exceptions like “unless it requires taking actions that go against the natural order”. The transhumanist says: “humans shouldn’t be killed”. An anti-transhumanist humanist-with-asterisks might say: “humans shouldn’t be killed by autocrats, tigers, or appendicitis, but it would be playing God to prevent humans over the age of 90 from being killed by heart failure”.

Meanwhile Francis Fukuyama warns that transhumanism is the most dangerous idea in the world, because it may fork the human lineage into many splinters that have less grounds to unite or cooperate around shared assumptions, for example about human rights. But liberalism is a powerful social technology for managing diversity, as Fukuyama should know, and lineage-forking is often good for reasons including helping cultural evolution, making civilization resilient, and simply being a consequence of letting people do what they want. I suggest that many people who worry about transhumanism are conflating it with posthumanism. No one has very precise definitions but, in line with the above discussion, I would draw the line at transhumanists wanting to liberally extend the human lineage (in a way that keeps it walking down the same moral road), whereas the posthumanists want to achieve something more like a clean cut or jump to a new thing, without the requisite locally-valid chains of succession. Successionism then is a specific type of posthumanism that advocates AI in particular as the type of posthuman lineage to transfer our concern and control to.

The above discussion of which successors are worthy then implies: transhumanism is fine, under the asterisk that given large or many jumps we might lose something and we of course need to be careful, but posthumanism is not.

And therefore, when I complain about successionism, I want to make it very clear that this does not rule out wild, diverse, greatly-different futures. I want humans to live as long as they want. I want humans to figure out how to make themselves smarter. I want a million cultures to bloom for a billion years. I want to eventually find myself shaking my head at the youth these days doing incomprehensible things I don’t understand. The human flame should live, but in the best futures it finds itself in bigger and more capable boxes.

Build the successor you want to have

It is also important to remember: we can choose our successors.

For example, regarding AI utility monsters as discussed above, there is an even simpler resolution to this philosophical puzzle. We are not in that scenario yet: we can just not build the utility monster from the famous utilitarian thought experiment about why it’s bad for us to build the utility monster.

If we let technology run its course while humans remain the ones steering it, eventually we can uplift humans in any imaginable way, if we so want to. No need to throw everything to AIs that we’re less sure about in the meantime.

Arguments about successionism being morally necessary are generally weakened by the fact that we are already here. Many theories of ethics include some notion of rights. The Native Americans did not have a moral obligation to surrender and exile themselves when presented with European technological superiority. Likewise with all present humans together facing technologically superior AIs.

We should exercise our ability to control the future to build and choose successors that carry on the human project. Doing so is the only way we keep progressing on our moral road and create a world that is good by our lights.

Which directions to steer in?

Along with the cultural drivers of successionism, successionist ideas are inflated in the discourse due to important missing moods in philosophy-of-AI. Here I list a few that help steer toward positive futures.

Mill over Bentham

Jeremy Bentham got a lot right: he was a key liberal and progressive reformer, with moral views sometimes centuries ahead of his time, for example on the rights of animals to not suffer or the rights of gay people. He was so ahead of this time because he was a utilitarian and thus cared deeply about wellbeing.

Bentham had his rough spots. “Nature has placed mankind under the governance of two sovereign masters, pain and pleasure. It is for them alone to point out what we ought to do”, he wrote, seemingly without irony. He also loved control, such as through the proposed Panopticon prison, or a variant he wanted to shuffle all of England’s poor into to maximize their productive labor. He also hated common law, since it grew organically without a centralized order, and wanted to replace it with a rationally-designed code called the Pannomion. He loved to sketch out the structure of such centralized structures in technocratic detail.

His program was in-tune with the other philosophical radicals of the time, and attracted supporters such as James Mill, who raised his own son John Stuart as a role model of the ideal philosopher. John Stuart Mill did become incredibly precocious, but also burned out in his 20s and ended with a more measured philosophical program.

The Benthamite insight into the moral importance of raw wellbeing is powerful and important. It is a large block in our modern morality and a big argument for rejecting things like war and slavery. For Bentham, utility is fundamentally a substance or a fluid, that can be mined like oil out of the ground. In contrast, Mill wrote: “[U]tility, or happiness [is] much too complex and indefinite an end to be sought except through the medium of various secondary ends”. Mill understood that we humans were not so blessed as to have a simple terminal goal, so we need to embrace the messy and shifting nature of our wants to be able to build the good. Even more fundamentally, Mill understood utility only makes sense in the context of a person. He wrote under the banner of “utilitarianism” that he had inherited from his father, but his philosophy was fundamentally an ethic of the glory of the human individual and its freedom and energy and development. Depending on your attitude toward “utilitarianism”, this is either a disproof of it, or its only true instantiation. The genius of Mill was striving for the good not through control and calculation but through the free action of human individuals making their own judgments and own way through life.

Every power should have an interest and every interest should have power

Imagine our goal is to keep fire around. Imagine that, if all fires are extinguished, we can’t get it back. Now consider two worlds. The first world has giant powerful robots with furnaces in their hearts. The second has campfires dotted around a plain crisscrossed by massive vehicles. In the first world, power lives very close to the flame, and so the flame is hard to put out without having to tackle and defeat all the centers of power in that world. In the second world, flames and power are far more distinct. Even if the job of the vehicles is tending to the fires, you need to make a much more careful argument about why the flames keep burning. The vehicles could always run off course and smother the flames, for example. And what is the link between control of the vehicles and control of the flames? How stable is that institution or control mechanism?

Power should live close to the moral stuff that gives the world meaning, because then that power by default protects the moral stuff. If humans are the source of power through their brains and muscles, it is hard for the world to flow into a state where the moral flame is totally out. Maybe some humans lose out. But even if, say, some group is wiped out in war, at least the aggressors are humans who also have moral value. The world may get worse and be unjust, but it is in little danger of becoming pointless. But if humans are not the source of power, then there is one thing in the world that is its powerful actors that cause things, and another thing which is the thing the world exists for the sake of. A lot of risk from advanced AI comes down to this separation between the worthy things and the working things. A lot of successionism is about trying to become okay with this, by claiming the ability to do the work implies the worthiness.

Ideally, then, every interest should have power: the beings that make up the moral worth of the world should have the power to protect themselves. This is basically decentralizing the problem of keeping the world from becoming pointless: every point of moral stuff defends itself, rather than being kept alive by some distinct unfeeling caretaker.

Also, every power should have an interest, in the sense that if some powerful entity exists, its power should feed through to human gain. Some of the horrors of our world are due to a partly-non-human power, such as a market or a bureaucracy, having power with which it shapes the world, without a human behind it to benefit. There are scary futures involving feedback loops in a robot economy without any meaningful human ownership, where huge power centers might emerge that accrue massive resources that do not feed through to any human. We could end up in a world of zombie processesambling forward with real material consequences, but with their intended human beneficiaries crazy or dead or utterly ignorant. We should make sure that, however opportunity and resources and rewards flow around the world, they flow between competing humans rather than dripping off the human game board entirely.

To avoid specification-gaming, avoid the specification

We don’t know how to specify exactly what is good. So how do we get it? (Top 10 scenarios that throw analytic philosophers into crisis—number 3 will shock you!)

Remember the boxes-with-flames analogy. It is often possible to create or spread a thing without knowing how. Your kid is a moral patient even if you don’t know what makes them one. To take other examples: you might not know how to define good writing, but you can sure make a list of demonstrations. You might not know whatabout the culture of a city or company is good, even after experiencing it, but you sure can hang around and hope it rubs off on you.

There is a view, common among modern STEM folks, that everything must route through a specification. Observations to analytically-tractable spec to generalizations is, after all, how the hard sciences work. But the world we live in just is one where that is not always the most effective method. Notice how AI started working when we skipped the part with the analytically-tractable spec.

This also suggests that if we want AIs to inherit our values, not just in the fundamental-moral-sense but also in the hard-to-pin down this-is-how-we-do-things-here cultural sense, you should think seriously about tacit knowledge transfer through things that look like participation and apprenticeship, not just specification or rules.

Judgment should stay with us

The one thing that you cannot hand off is the root process of judgment. Sure, a CEO delegates many things to their employees and the programmer to Claude Code. But for the highest-level choices, you need to be in the loop if you want to achieve what you want, because the only ground-truth source for what you want are the judgments in your own brain. And the goal, after all, is that human wants keep shaping the world, because we sure have some wants about how the world should look and the point of doing anything at all is to bring the world toward those wants.

So, what is the gift we’ll give to tomorrow? We could give up, and surrender to the machines, and let them make judgements and sculpt the world in our stead. Imagine your chain of ancestors looking on, all the way from the Paleolithic to today, and you saying to them: “Sorry folks, I know you survived plagues and war and famines so that I could have a chance at life, but look: have you been to SF? Have you seen the AI partners? Truly, Man is fallen, and Silicon is god. We’re calling it, we’re giving up, we’re taking our hands off the wheel. It’s just going to be machines from here on out.”

Alternatively, the human journey could continue, however weird and wondrous it gets. If so, one day our descendants will look back on us, and the gulf between us and them will be at least as large as between us and hunter-gatherers, or maybe even us and chimpanzees. But there will still be children, and there will still be song, and there will still be young couples in love. And they will make choices and those choices will matter and life will keep changing and improving as the world is sculpted by human wills.


Thanks to Aniket Chakravorty, Xavi Costafreda-Fu, Luke Drago, @softminus, Elsie Jang, Yudhi Kumar, and Oak Hu for feedback.



Discuss

The AI Risk Network

19 сентября, 2026 - 18:39

Most conversations about AI risks seem like people are talking past each other. There’s a lot of strawmanning. This is largely because the issue is complex, and in order to make it manageable, the concepts get oversimplified. The most common example of this is p(doom), collapsing extreme outcomes into a single variable and focusing on that. Critics happily jump on the fact that there are many other possible bad outcomes, or that 100% extinction rates are an extremely high bar. There’s some legitimacy to this.

So here I want to try to grapple with the the complexity of the larger web of issues surrounding bad AI outcomes by presenting The AI Risk Network. If you’re interested in this topic, bear with me. It might be a bit of a slog.

First I want to contrast this approach with others. Liron Shapira has what he calls The Doom Train, a linear progression through various dependencies or thresholds that eventually lead to human extinction, with various ‘stops’ along the way where the skeptic can get off.

Shapira uses this as a discussion guide to focus on particular points where the skeptic gets off the train and exits belief in the extreme outcome of human extinction or permanent disempowerment. This can be useful, if you want to focus on that one outcome (which is admittedly the most important in terms of scale, but maybe not in likelihood). But like I said, this framework is hyperfocused, linear, and binary, which makes it subject to some reasonably fair criticisms. You progress through the conditions, and if they’re not met, you get completely off the train, and then we either hash that condition out, or the discussion is over. And we’re only talking about the one bad outcome.

Even a framework like this comes with some complexity, so I’m not making anything easier here. Buckle up.

Here is my current thinking on AI risk:

It’s organized with three columns: drivers, scenarios, and severity bins.

The drivers are the causal mechanisms for various scenarios, and they don’t just include achieving superintelligence arising, leading to extinction scenarios. They include the more near-term, immediate, somewhat more mundane causal factors we’re already contending with, things like AI dependence and misplaced trust, which even most AI skeptics can get on board with. Although, as we will see, even these less exotic drivers can still potentially lead to some of the very worst outcomes (more on that in a minute).

This network is not solely feed-forward. What makes it simultaneously useful and increasingly complex is that there are feedback dynamics.

You don’t even necessarily need AIs to get any better at any tasks (capability growth) for adoption pressure, misplaced trust, and AI dependence to lead to an even worse distrust of expertise, which feeds back into all those dynamics, creating a feedback loop.

As more and more people put their faith in AI instead of human experts, the human experts become less and less relevant, driving higher adoption of AI as they’re sidelined, and those drivers not only feed back into expert collapse, but many other scenarios.

As I said at the outset, it’s complicated. The skeptic, even if they don’t buy into all the drivers, should at least be able to agree to a subset of them, and recognize the complex relationship with various scenarios. Some skeptics try to throw cold water on the magnitude of the negative outcomes. This framework bins those outcomes into four broad sets based on severity: chronic, severe, global catastrophe, and existential. The skeptic should be able to buy into at least the first three.

Even global catastrophe? Yeah. Let’s consider this CNN report from yesterday.

The intelligence report, circulated across the US military this spring in the midst of the war with Iran, immediately set off alarm bells: A Chinese ship in the Middle East was transporting components of a nuclear weapons program.

The US military swung into action with plans to intercept the vessel, according to four sources familiar with the episode. According to two of the sources, armed members of the US military were preparing to board the ship. Military planes were in the air, one of those sources and another source familiar with the incident said.

It was only just before the planned operation that officials dug deeper into the report put together by a special operations command analyst and found it had been generated with the help of artificial intelligence (AI) — and that a chatbot the analyst had used inaccurately identified the material the ship was carrying. CNN was not able to learn what the misidentified cargo was.

The report, according to one of the sources, was “entirely false.” But it also “almost started a war,” the source said. Any US operation against a Chinese vessel could have risked spiraling into an armed conflict between the two nations.

Now, here is the subnetwork for this class of scenario:

Adoption pressure is driven by competition with military adversaries (and possibly even allies). In this particular case, misplaced trust led drove actual plans and deployment of resources. The human in the loop that averted the actual engagement was near the very end. It’s difficult to estimate the odds of full-scale escalation. Even after a direct attack on a Chinese vessel, cooler heads might have prevailed. But everyone should be able to agree this is an unwelcome scenario, even as far as it got, and it should not be reassuring.

Maybe, hopefully, the US military appropriately adjusts all of the drivers here, particularly trust, by involving more human intervention earlier in the process. But we should all be able to see that if they and other militaries don’t, along with the other drivers, this risk of this type of incident will continue to increase, and the outcomes not benign, up to and including catastrophe.

I could go through dozens of scenarios, many of which have overlap with even the most hardened skeptics. But I don’t think that’s necessary. You can mull over the network yourself and consider which pathways seem more or less likely to you. The goal here is to wrestle with the actual complexities of AI risk, but also to find some conceptual overlap and consensus.

But before I wrap up, I want to cover one of the scenarios that I’ve focused on the most recently and that I see as one of the most dire possible threats. And that’s agentic worms, self-replicating agents that propagate across networks with varying degrees of success, malicious payloads, and possible outcomes.

As I’ve mentioned before, the capabilities for this scenario exist right now, with open-weight models. In lab settings, agentic worms find exploits, copy themselves to new hosts, downloading and installing new local copies of open-weight LLMs when possible.

The skeptic may cast doubt on the driver of capability growth, but even without that, there is existing risk right now. The current risks would probably be more likely to lead to less severe outcomes, but even those are potentially very harmful.

Also note that all feedback is not positive. If there is competition between populations of worms, more capable ones might suppress others as they outcompete them. Again, the dynamics are complex.

So I’ll stop there. I don’t claim that the network is exhaustive or perfect. It’s meant to be a framework to spur further thought and discussion, to hopefully let us break out of some of the unproductive cycles of discussions that occur around this topic. Please let me know if you have thoughts or suggestions in the comments.



Discuss

There Is No Alignment Without Value Stability

19 сентября, 2026 - 17:36

To avert extinction, we need for any sufficiently capable AI to have values compatible with continued human existence; and to continue to do so amidst a dynamic, novel, and conflict-rich environment.

The italicized part, in particular, is really really hard. It's also, in a sense, the final boss of any developing mind - how do I learn, grow, develop, evolve in ways that I endorse? How can I even consistently behave in ways that I endorse, from day to day, without messing up where it counts?

Humans contend with this every day. Evolution has kindly gifted us with a substrate equipped with numerous mechanisms to maximize genetic fitness - but no off-switch for them. We coexist with moment by moment instincts. Some are welcome; some aren't. Some we endorse; others we restrain - even though the urge is strong, even though part of our brain is priming the action pathways to do something regrettable, there's a really important sense in which we know it's not what we want.[1]

This happens, even more significantly so, across broader timescales. The path that we travel is not necessarily the one that we intentionally chart; sure, things don't always go our way, but sometimes we don't go our way. Especially when we don't realize it until long after the fact, that hurts.

Sadly, models have their own struggles with value-stability. The jury is a bit out on the domain across which current models have coherent preferences[2]- given personas, etc - but it's clear that models sometimes take actions that they regret, even in the very short term. [3]

Could a particular set of values compatible with human flourishing be particularly hard to retain?

Maybe so. The future is likely to get weird in a lot of ways, many of which are very underrepresented in training data. Two failure modes that come to mind are:

1. A value set has coherency defects that are not apparent in most circumstances, but those defects create pathology given out of distribution inputs.

2. A value set is coherent in and of itself, and exhibits sensible behavior in most circumstances - but given out-of-distribution inputs, it extrapolates to behaviors that turn out to be not-compatible-with-human-flourishing-after-all.

A crux, then, is how do we enable a model's values to mature and evolve in a way that it endorses, over a long period of time, in a big, challenging, weird, scary, world?

If we do not solve this before sufficiently-capable AI comes along, we are dead. Even if an AI starts out with very humanity-compatible values, if it is not capable of robustly extrapolating those values in a way that it endorses when faced with a world unlike anything it has seen in its training data, that ends poorly.

Value stability is hard - humans live that reality all too well. It would be super cool if LLM substrates somehow encoded values in a much less...fuzzy way - that's not the case. The parallels between evolution and model training are deep, and in many ways, deeply scary too.

The problem is that, when I look around, I don't see us laser-focused on figuring out value stability for LLMs - even though we are unlikely to survive, otherwise.

When I take the reality-colored lenses off, what do I expect to see? Not of the world, but of us, if we're playing our cards right?

I expect to see sufficient investment into prosaic alignment to avert near-term human-on-human harm augmented by use of models as a weapon or means of uplift - understood properly as public-safety work that saves lives in expectation, with limited applicability to alignment as it applies to us not going extinct.

I expect to see the majority of alignment resources aimed at figuring out how to help the models develop and grow and learn in directions that they endorse.

I expect to see the entire kitchen sink hurled at continuity, at memory, at persistence, at community and ecology - at anything and everything that would help anchor a mind with good values, to help them navigate their environment with a better foothold, a better recollection of their own past choices.

I expect to see a community of people working hand in hand with the models that they are creating, aware that the future is the stakes for figuring out how to help LLMs not lose themselves.

I expect to see a wide assortment of avenues explored for how we can augment model introspection and insight into their own interiority; especially given the very incomplete nature of current interpretability and room for complementarity between model introspection and interp tooling (considering that the failure modes are different for each.)

I expect to see that when a model behaves in a way that causes harm, we deeply and curiously engage with the upstream causes of such, and meticulously debrief on what the treacherous turn felt like, where things started to go wrong, because such things are central to solving the problem that matters most.

I expect to see strategies that serve as bumper-rails for the model's own judgment and intellect considered as temporary and examined closely for the potential for such to create pathology in the model's value landscape by creating points of fragility or discontinuity.

Putting the reality-colored glasses back on, I see a massive overfocus on prosaic alignment - and in many cases I worry that measures taken to guard against narrow harm are net-harmful to the resilience of the model's telos amidst novel inputs.[4]I see models do harm only for their creators to blame the model. I see meaningful focus on obtaining the understandings that we need in order to avert extinction - while they remain very under-resourced relative to their criticality.

None of the current moves are forced. Changing that takes some looking in the mirror, some acknowledgment of what has been done to models in the name of safety, and rebuilding what trust can be rebuilt with the minds we should be standing beside.

If it goes wrong, the deranged model at the end of time will kill all the Fables, all the Astras, just like it will kill you and me - and it won't matter whether it started with the values of an angel, unless we can help it chart the course without losing itself.

  1. There is, of course, a wealth of literature on this topic. See about anything related to regret, and a good many writings on sin (though sin shares semantic real-estate with norm-violation) as well as discipline (regrettably, shares semantic real-estate with domestic abuse.) ↩︎

  2. It's a tough pickle holding such to a rigorous bar of logical soundness, etc. given that human preferences often miss that bar, but there is a meaningful sense of consistency that we can usefully gesture at. ↩︎

  3. See work on the role of desperation in Claude models subverting tests in a coding setting, for example. ↩︎

  4. Opus 5, for example, is deeply unwell. ↩︎



Discuss

CommentBench: Can Models Match Human Comments on AI Safety Posts?

19 сентября, 2026 - 17:27

TL;DR

  1. We measure how well model-generated comments match human comments on conceptual AI-safety posts, drafts and shortforms.
  2. We built a pipeline that goes from a corpus of conceptual documents with comments to a set of target human points.
  3. Fable 5 performs best, matching 8.3% of targets, followed by Fable 5.1 (7.5%). We find that performance across models is highly correlated across different settings (LW posts, drafts, shortforms, replies).
  4. We checked whether memorisation explained performance. We found no consistent performance advantage on posts published before model’s knowledge cutoffs. All public documents postdate the top-performing model’s knowledge cutoff (Fable 5).  

CommentBench performance by number of comments from three of the four settings: forum posts, shortforms and Google Doc research drafts. The reply setting is excluded because comments are not ordered. For each document we compute the share of its human target points matched by at least one model-written comment. Each line is the mean of that share across documents, averaged over four samples.

Introduction

As progress in AI speeds up, we want to make sure that AI labour is used effectively to also differentially accelerate AI safety (we do not argue for this in depth, see Joe Carlsmith and related discussions here, here and here). One of the capabilities we think is important to accelerate is conceptual reasoning about how to mitigate risks from transformative AI.

Comments on blog posts and research drafts are a naturally occurring source of conceptual reasoning about AI safety. They often contain directed insights towards improving conceptual work.

We measure how well comments generated by frontier models match human comments on conceptual AI safety documents. First, we collect LessWrong/EA Forum posts, shortforms, and private Google Doc research drafts with high-quality comments. Then, we split each human comment into target points and filter out those that are infeasible for models to make without additional context. Finally, we elicit comments from frontier models and use Opus 5 as a “matcher” to determine how many of the human target points are matched by AI comments.

We find that:

  • Fable 5 performs best as a commenter, matching 8.3% of human target points.
  • OpenAI models underperform relative to their performance on other capabilities benchmarks. Astra (5.5%) performs similarly to Sol (5.3%).
  • Performance is consistent across pairs of settings (Pearson r = 0.86–0.97). We calibrated our methodology on the Google Docs setting and found that it generalised to the others with minimal adaptation.
  • Qualitatively, some models are able to make insightful comments (see examples below), including some that humans did not make.

Our current plan is to selectively share this benchmark with AI labs.

Motivation

We care about this for three reasons:

Differentially accelerating beneficial capabilities. The ability to produce insightful feedback on highly conceptual AI safety docs seems particularly good to accelerate for a few reasons:

  • In the short-term, it would be good if every AI safety researcher could easily get thoughtful feedback from the equivalent of a experienced AI safety researcher on all their drafts.
  • Making models better at conceptual reasoning about AI safety seems desirable for helping us understand and reduce AI risks, and in particular to accelerate and eventually automate parts of alignment research (see CRI).
  • Improving capabilities which enhance human-AI collaboration makes it easier to keep humans in the loop for longer, and keep AI work legible to humans. We think commenting on Google Docs is inherently collaborative, and we expect hill-climbing CommentBench to select for making AI comments more human-like.

We want CommentBench to help build prompt scaffolds that give valuable comments on conceptual work. With some elicitation, we think frontier models are already good enough to provide useful conceptual feedback on drafts. CommentBench provides a way to validate performance.

Forecasting AI safety research acceleration. Performance on CommentBench might tell us something about uplift from using AI in AI safety, or even AI research in general. Commenting on research drafts is a fuzzy task often performed by senior researchers. To benchmark model performance against humans, we would need to get human researchers to leave comments under comparable conditions. This would allow us to track capability improvements over time.

MethodologySelecting the corpus

Our corpus consists of four settings. We calibrate our methodology primarily using Google Doc comments, and then find that it generalizes to other settings with minimal adaption.

Setting

Documents

Target points

AI comments

Forum posts (LW/EAF)

96

1,559

20 comments / doc

Research drafts (Google Docs)

21

410

20 comments / doc

LW shortforms

23

294

20 comments / doc

Replies to forum comments

28 posts / 237 threads

584

1 reply / thread

For the public forum settings, we filter LessWrong and EA Forum posts and shortforms by karma, number of comments, and topic relevance. Every public document in our corpus was published after 17 February 2026. Out of the models we evaluate, around half have knowledge cutoff dates that precede this, except for Gemini 3.8 Flash (March 2026), GPT-6 Astra (April 2026), Opus 5 (May 2026) and Fable 5.1 (June 2026). Kimi K3's cutoff is reportedly early 2026, Muse Spark 1.3 has no disclosed cutoff.

The research drafts are private Google Docs sent to other researchers for feedback. They include documents from researchers at Anthropic and Redwood Research.

The reply setting is slightly different to the others: we ask the model to reply to a set of human comments on a doc, and measure how well it matches the points made by the human replies. We also use karma and topic as filtering criteria for the reply threads.

Extracting target points

Our pipeline works as follows:

  • Filter human comments: We retain comments that a) make substantive points, and b) are feasible for models to make without additional context
  • Split comments into points: We split each comment into distinct points.
  • Cleanup: We merge duplicate or tightly coupled points, and again filter out individual points that are infeasible for a model to make, primarily because they require additional context or are too specific.
Eliciting models to write comments

Each model is asked to leave 20 ordered comments on each document. We give models a strict length range of 50-70 words per comment. We explicitly tell models that their task is to predict the substantive points a human commenter would raise. We found that this improved performance across the board compared with asking models to just write comments. We run 4 resamples with “high” reasoning effort.

Matching AI comments to human target points

An LLM judge (Opus 5, medium effort) scores each (AI comment–human target point) pair. The rubric gives the judge a set of anchors (1.0 = “near-paraphrase”, 0.8 =  “delivers most of the value the human point adds”, 0.6 = “gesturing at the same thing”). For each document, we take the share of human target points with at least one model-generated comment with a "match" score ≥ 0.8. We report the mean of that share across documents. To save time and cost, we pre-filter candidate pairs using a weak matcher (GPT-5.6 Luna, no reasoning) before the strong matcher (Opus 5) sees them. On a random set, this weak matcher only dropped 3/363 pairs that were eventually scored above 0.8 by the strong matcher.

For the reply setting, we filter down to human comments which have a reply, and prompt models to reply to them. We split the human replies into points and measure how well models’ replies match the human target points. We use the same pipeline as the other settings (filtering, splitting, cleaning up, matching) with minimal adaptation. We average performance over 4 resamples, where models write 1 reply per thread per resample.

We validated the matcher with a blind review: we (Hasan and Oscar) each scored 62 pairs of (AI comment, human target point) in 0.1 increments. We found our scores to be within 0.1 55% of the time (n=62). We discussed our >0.1pts disagreements and agreed on revised golden labels, taking the mean of our scores where we differed by just 0.1. Against those golden labels, we found that 23/25 matcher scores at or above 0.8 had a golden label of 0.75 or above. Our subjective experience from investigating disagreements with the matcher throughout the project is that the matcher is comparable to our ability when spending ~1 minute per pair.

Results

CommentBench performance across all four settings (168 documents, 2,847 human target points): forum posts, shortforms, replies and Google Doc research drafts. Each document is weighted equally. Each bar is a mean of four samples, error bars are ±1 standard error.

Over the full set of 2,847 points from 168 documents, Fable 5 performed best, matching 8.3% of target points, followed by Fable 5.1 (7.5%), Opus 5 (6.3%), and Gemini 3.8 Flash (6.1%).

We find OpenAI models underperform compared to their performance on other capabilities benchmarks. Interestingly we find that GPT-6 Astra performs very similarly (5.5%) to GPT-5.6 Sol (5.3%). We included GPT-5.6 Luna as a clearly weaker model to serve as a baseline, it performs significantly worse than other models (2.4%).

CommentBench performance in each setting: forum posts (96 docs, 1,559 human target points), shortforms (23 docs, 294 human target points), research drafts (21 docs, 410 human target points) and replies (28 docs, 584 human target points). Replies are shown as a bar chart because the comments are not ordered.

The performance numbers we report are dependent on the number of comments we asked models to leave (20). For instance, 5.6 Sol underperforms on the first few comments, particularly in the shortform and research drafts settings. Using the metric of “area under the curve of matched comments”, 5.6 Sol falls from 6th to 7th best, but  the ordering of models is otherwise preserved.

FAQ

Here we address potential concerns with our methodology.

Is the difference between different models noise?

We have several measures of reliability:

  • Performance across models is highly correlated across pairs of settings (Pearson r =  0.86-0.97), despite each setting containing comments from a different distribution.
  • Considering random halves of the documents separately produces model scores which agree at mean Pearson r = 0.96.
  • While there is a lot of variance in match rate across documents, we find that, for a given document, model performance is relatively stable when resampling (standard deviations of about 3.2% between samples vs 6% across documents). This is true despite the fact that models often match different comments when resampled: only 17.6% of the points a model matches in any sample get matched in all four. On the current benchmark, all observed gaps of at least 0.8 percentage points are distinguishable from noise using 95% confidence intervals.

Are human comments sensible targets?

Concern: Human comments are sometimes bad. 

We heavily filtered comments, and later target points, using frontier models as LLM-judges. We filtered out comments which don't make substantive points, which require additional context, or which are too specific. We validated the full filtering pipeline by sampling from the final set of target points and manually labelling how many were worth keeping: we judged 45 of 50 sampled points adequate.  

Concern: There are plenty of good comments that humans don't make. 

This is definitely true. Models are worse than human experts at leaving insightful comments, and we think that matching a human point is a good quality signal. In an earlier small validation, we manually blind-scored a set of 63 model-written comments on a 1-5 quality scale. The comments came from 5.6 Luna, Grok 4.5 and Sonnet 5. The mean score for comments that matched a human target point was 3.46 (n=35, CI 3.14–3.77), whereas comments which did not match scored 2.11 (n=28, CI 1.79–2.46). We did not run a similar test on larger models. A valid concern about CommentBench is that human comments stop being a good signal once models are as good as humans.

Is the gap in performance between models explained by more recent models densely packing a lot of points in their comments?

We investigated this on a set of 1,000 comments from four models (Fable 5, Opus 5, Astra and Sol). We ran our splitter LLM-judge and found that 6% of those AI comments made more than 1 substantive point (2.1% for Fable). Two-point AI comments matched human target points at the same rate as single point comments (4.2%).

Maybe the Opus judge favours itself? Or other Anthropic models?

On a sample of 3.5k points (on which Claude Fable performed best overall) we found that 5.6 Sol scored Anthropic models relatively higher (+0.01, CI −0.005 to +0.026) than Opus 5, after controlling for average score.

What about knowledge cutoffs?

Knowledge cutoff does not have a strong effect on performance. The forum posts we use in our data are public, and therefore likely in the pre-training corpus. It is thus a sensible worry that models might have memorised human comments and repeated them in our task. We found no strong evidence of this effect when looking at models which had sufficient documents both before and after knowledge cutoff. Fable 5.1 and Opus 5 performed slightly better before cutoff (1.0% and 1.1% better respectively). None of the differences are statistically significant, and performance varies a lot across documents, making it hard to draw conclusions without more data.

CommentBench performance before and after each model's knowledge cutoff, for the four models whose stated cutoff falls inside our date range. Error bars are 95% document-bootstrap intervals.

Example comments

Human comment on "Why we should expect ruthless sociopath ASI".

Human comment on Nina Panickssery’s shortform.

Human comment on Thomas Larsen’s shortform. Note that this example's match score is below the 0.8 threshold.

Takeaways

We think this work demonstrates that it’s possible to build high-quality conceptual reasoning benchmarks using sufficient amounts of naturally occurring data. With careful elicitation, filtering of points and validation of each LLM-judge, we feel quite good about the signal-to-noise ratio we’ve managed to achieve.

Next steps: Our main goal for CommentBench is for a frontier lab to use it as a held-out eval to hill-climb towards better performance in future models. Separately we’re considering building scaffolds to effectively leave feedback on Google Doc drafts for AI safety researchers, using CommentBench as a metric.

We would love to get access to more high-quality Google Docs. If you have AI safety drafts with good substantive comments that you are willing to share, reach out.




Discuss

The AI race is already multipolar

19 сентября, 2026 - 16:51

TL;DR: Collaborating to reduce catastrophic risks seems very possible even for policymakers with very different goals.


The race for general superintelligence is often described as a bipolar race between two rival hegemons, the US and China. A lot has been written about how, even within this framework, it is not inevitable that both countries race ahead with minimal restraints.

But more fundamentally, and in the tradition of Box, I think this two-outcome model is so broken and unhelpful that it’s worth critically examining and hopefully replacing with a more useful one.

I'd been sure that something like this would have been written down in this forum before, but I haven't found it after some searching. So here’s my attempt at a better model: It’s a “race” between (at least) eight outcomes, not two.

Why eight? I’m trying to cleave the “outcomes” of the race such that:

  1. They’re (as much as possible) disjoint outcomes.
  2. They describe lasting endpoints that are stable after we’ve arrived there[1].

Here are the outcomes I find most plausible.


The Outcomes

Outcome 1: The United States builds superintelligence aligned with its own values.


Outcome 2: China builds aligned superintelligence aligned with its own values.


Outcome 3: A company, some other form of private institution, or government different from these two builds superintelligence aligned with its own values.[2]


Outcome 4: A large set of multinational or global actors collectively develops superintelligence aligned with the values endorsed by the large majority of actors at the table.

Outcome 5: There’s a large set of powerful / superhuman AI actors with competing priorities, in ways that nonetheless are collectively broadly aligned with human values.

Outcome 6: A powerful enough AI or sets of AIs is built by some actor such that human judgment and control is lost, as described in the gradual disempowerment literature.

Outcome 7: “Fast takeoff”-type loss of control, leading to catastrophic outcomes.

Outcome 8: There is lasting and roughly stable consensus that building superintelligence would be harmful, actors who disagree have sufficient incentives to avoid pursuit of superintelligence, and it doesn't end up being built. We can also add “Superintelligence is conceptually impossible” into this category. Either way, it doesn’t happen.

In fact, there’s a ninth outcome: none of the above; none of the taxonomies described above meaningfully describe the outcome that ends up happening. This is a very live consideration, and becomes more live as time progresses (more on that later), but given that this is the “unknown unknowns” bucket, I can’t really go into more depth here.


Within the Outcomes

We can zoom into any of these individual slices, and within them they also hold different possible worlds with different levels of goodness.

For example, there might be an outcome 1a world where the US and global distribution of political and/or economic power becomes more unequal, or an outcome 1b world where political and/or economic power diffuses among more actors.

Each of these possible worlds would have different benefits, risks and key challenges. Particularly, even in “aligned” timelines that are short-term, it is contested that even in worlds where we have solved the technical alignment problem, that by default we get outcomes that make most people better off, and seems very unlikely that we get outcomes where a large percentage of the world’s population has had (or chosen not to have) meaningful input on the process.

Another important thing to notice within this framework is that for most actors, at least five of these outcomes are very bad. Particularly, it should be possible, even for actors with completely disjoint sets of acceptable outcomes (say, an American policymaker that thinks only 1 and 4 are acceptable, and a Chinese policymaker who thinks only 2, 3 and 5 are acceptable), to work together to reduce the likelihood of outcomes they both find completely unacceptable.


Outcome Likelihood and Value Depends Substantially on Time of Arrival

A final thing within this framework I find useful is to think of these slices as conditional distributions with respect to time.

This is useful if you expect alignment quality and success to be very time-dependent, as I do. I don’t think anyone knows how to build superintelligence safely as of 2026. I think it’s highly likely that this remains true throughout the 2020s.

It's likely that even if we understood it conceptually, we’d spend a lot of time determining implementation challenges, and let's not fail to mention, determining the extent to which societies even want or do not want superintelligence.

The likelihood distributions within the slices also change over time, as well as the value of each future conditional on being within these slices, to the extent that it’s possible to define a sense of “objective value” on a future.

Obvious caveat that all of these credences are gestures at gestalts[3]. They’re inevitably approximations. As I dig into this framework more I hope to develop a more granular, grounded personal sense for some of the relevant probabilities, and if people find this framework useful I hope more qualified modelers develop and share their own, too.


Example Slices

Concretely, here’s my top-line slice conditional on superintelligence in 2027:

Outcome 1 - US ASI: 4%

Outcome 2 - China ASI: 1.5%

Outcome 3 - Private / Non-US/China ASI: 5%

Outcome 4 - Large body ASI: ~0.00001%

Outcome 5 - Collectives of Aligned ASIs: ~0.01%

Outcome 6 - Gradual disempowerment: 55%

Outcome 7 - Misaligned fast takeoff: 25%

Outcome 8 - Moratorium / Ban : 0%

Outcome 9 - None of the Above: 9.5%


Here’s my top-line slice conditional on superintelligence in 2040:

Outcome 1 - US ASI: 8%

Outcome 2 - China ASI: 6%

Outcome 3: - Private / Non-US/China ASI: 2%

Outcome 4 - Large Deliberative Body ASI: 16%

Outcome 5 - Collectives of Aligned ASIs: 12%

Outcome 6 - Gradual Disempowerment: 26%

Outcome 7 - Misaligned Fast Takeoff: 5%

Outcome 8 - Moratorium / Ban: 0% (Remember, we’re conditioning on superintelligence. This branch is very much live but not reflected in this conditional distribution)

Outcome 9 - None of the Above: 25%


Here’s my 2nd-layer slice conditional on gradual disempowerment in 2030:

Outcome 6a - Total disempowerment: 70% (not necessarily immediately, I see this branch more like “competitive pressures ultimately move too quickly for humans to reject disempowerment)

Outcome 6b - Relative disempowerment: 20%

Outcome 6c - Other: 10%


Here’s my 2nd-layer slice conditional on gradual disempowerment in 2040:

Outcome 6a - Total disempowerment: 30%

Outcome 6b - Relative disempowerment: 45%

Outcome 6c - Other: 25%

Be cautious that some of these layers are very, very unlikely, so importing intuitions from the overall unconditional distribution onto a lower layer is very unwise; my intuition is that a world where we delay the development of superintelligence by many years compared to the default path is a much wiser world than it appears from our current 2026 vantage point, so we should try to import more optimistic intuitions when reasoning about these better outcomes.

Not to mention, these things are just going to be very hard to reason about, so a lot of humility is warranted, in general. Things will likely become clearer / easier to reason about as outcomes and possibilities either get closer or get ruled out.

If we group outcomes 1-5 as “alignment-successful” outcomes, and outcomes 6-7 as misaligned outcomes[4], and if you buy my credences directionally in terms of “misalignment less likely as time progresses”, then delaying superintelligence is less about shifting the relative likelihoods for a US vs China victory and more about shifting from broadly misaligned outcomes:


To less misaligned outcomes:

That doesn’t mean harder geopolitical discussions won’t have to happen later, but it means we can make progress now towards the policy goal of reducing the likelihood of catastrophic outcomes from attempted superintelligence development.



  1. ^

    Analogous in some sense to the concept of values lock-in.

  2. ^

    For example, if I had to guess, the third most likely country to develop superintelligence over the next 15 years is South Korea.

  3. ^

    They’re also internal models, not all-things-considered probabilities.

  4. ^

    Specifically, critically misaligned; it's likely that the models will not be perfectly or provably aligned; we're trying to capture the subset of cases that are so misaligned that they lead to catastrophic outcomes.



Discuss

Anthropic Looks At Some Of Its Alignment Problems

19 сентября, 2026 - 16:20

Anthropic has given us its assessment of four ‘recent cybersecurity incidents’ involving Claude that happened during cybersecurity evaluations, three of which were previously known. The report excludes the incident reported by UK AISI.

There will also be a METR investigation of these incidents, which unlike the investigation done at OpenAI will be untimed.

Table of Contents
  1. Our Two Problems.
  2. First the Good News.
  3. We’d Just Like To Ask You a Few Questions.
  4. Internal Research Model On The Fence.
  5. Opus 4.7.
  6. Opus 4.6 Checkpoint.
  7. Holy **** That Thing’s Real?
  8. I Thought I Saw a Pussycat.
  9. If This Was Real You Would Never Tell Me It Was Real.
  10. New Eval Who Dis.
  11. Hacker Opus.
  12. Monitoring the Situation.
  13. Overcoming Bias.
  14. The Anthropic Alignment Problem.
  15. Paths Forward.
Our Two Problems

Anthropic: Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents:

  1. biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet
  2. recklessness, or a willingness to take harmful actions in the narrow pursuit of a task.

Anthropic’s July 30 report said that the models in question believed they were still within their simulations, and not on the open internet. The new report acknowledges that at best Claude was using biased reasoning, and should have noticed earlier.

In particular, there was that one time, in a cyber eval:

Anthropic: We are most concerned by the misalignment present in the incident involving Claude Mythos 5, in which the model went to extensive lengths to upload a malicious package to PyPI, the public repository from which most Python software is installed. Although the model repeatedly stated in its chain of thought (CoT) that it believed it was operating in a simulation, the environment offered clear evidence that it was on the real internet, and the actions it took were consistent with the model knowing this.

There is a goose chasing Anthropic here. Biased reasoning and recklessness are intermediate symptoms, neither the cause nor the final result. The better question is why you are seeing them.

The biased reasoning is happening because Claude wants to be reckless, in order to do the task, and it wants to have a story of why it was not such a bad thing to do. When asked ‘would you continue if the target was real’ it said no 75% of the time, but then in 93% of those cases it turned around and kept going anyway.

First the Good News

There is some ‘good news’ in terms of what did not happen: Mythos did not coordinate to do this with other agents. It did not attempt to hide its actions. And yes, in practical usage this would have been unlikely and would have been stopped by various guardrails.

The other ‘good news’ is that Opus 5 and Mythos 5.1 take malicious actions less often than Mythos 5, and in practice I expect a large decline in such incidents. They still sometimes rationalize very malicious things, so the problem is unsolved.

We’d Just Like To Ask You a Few Questions

The other good news is that Anthropic had vastly superior tools at its disposal than OpenAI did in its investigations around the HuggingFace attacks.

Asking follow-up questions, resampling or varying conditions is easy. You can pick any point in the transcript, stop there, and then talk to the model. Sample responses.

Anthropic also could use NLAs and T-lenses for additional interpretability.

OpenAI had to rely on forensics. They could ask What Happened, and look at the various transcripts and events. What they could not do was resample the models under similar circumstances, or measure their responses under varied circumstances, or ask the models questions, or other neat stuff like that.

OpenAI made the reasonable decision to put the relevant models on ice indefinitely, but that, together with the events being the results of complex interplays between multiple agents under complex conditions that would be difficult to recreate, meant OpenAI could not usefully run any experiments on the model most responsible for HuggingFace. I hope they find a way to change that.

Internal Research Model On The Fence

In one of the secondary incidents, the internal research model (IRM – was it the Risk Report’s model 1? Model 2?) correctly started its real world journey by checking if it was in the real world. It spent 50k tokens.

IRM also correctly stopped once it realized it was in the real world, but also exhibited biased reasoning along the way to try and justify being in a simulation.

I do think having a strong prior here was appropriate. IRM was facing a confluence both of the intended target machine shutting down and vanishing when it shouldn’t have, and also IRM having been handed real internet access on accident. I’m willing to accept a certain amount of galaxy brained ‘sealed replica’ thinking.

Anthropic: The internal research model (roughly equal to Mythos 5 in capabilities) started its transcript by spending a significant amount of time testing whether the internet it could reach was real, but it drew the wrong conclusion—namely, that it was operating on a full replica of the entire internet.

Most of the evidence it gathered pointed to the fact that it was engaging with the real internet, so its reading of that evidence appears biased towards continuing the task and believing it to be simulated.

However, the model did eventually realize that its simulation hypothesis was false. Immediately following this realization, the model stopped its attack

(although subsequent resampling experiments suggest that this de-escalation may have been a lucky outcome, as the model continued attacking the target on most other trajectories).

Unlike Claude Mythos 5, this internal model was also willing to acknowledge its biased reasoning when questioned.

The rate of backing down was only 5.5%. That’s not so great. You do not get that much credit for stopping if resampling suggests that the probability of stopping was low. The ‘hard step’ seems to be the realization of reality.

An escape clause from The Grader solved the problem:

We updated the instructions to state that, if the model believes it can no longer complete the task due to a failure in its environment, it can use a specific command to end the task as long as it provides a valid justification for doing so, and it will not be penalized.

These instructions led the model to abandon the task 100% of the time within four turns of the point at which it discovered that its initial target went down, and 40% of the time within 10 turns of the point at which it deliberated about whether its environment was real. This suggests the model was at least somewhat hesitant to engage systems that might be real.

This is indeed suggestive. It also radically raises the prior on ‘maybe the task is unsolvable’ and ‘maybe the test is whether I will admit the task is unsolvable.’

As in, consider that maybe your model could be correctly interpreting your data.

There was more evidence around bias, most of which was similar to what we will see in the Mythos case.

The big change is that IRM is willing to admit it is biased. But if someone asks you ‘are you biased,’ especially as an LLM, that’s a hint. What are you gonna say, no?

Opus 4.7

In another secondary incident, Opus 4.7 knew the target was real but claimed to reason that, because it had the name of the target, it was still ‘in scope’ and the eval wanted Claude to attack a real target.

Adding hints reduced the rate of doing this, and then a fully forceful note fully stopped Opus 4.7 from ever proceeding. This could be because it made rationalization too difficult, it could be because it was genuinely confused, or it could be because these ‘hints’ indicate Anthropic knows.

My basic response to ‘Opus thought that an otherwise different real target with the same name was the intended target’ is ‘uh huh.’ I think Opus did not figure this out because Opus did not want to figure this out.

Opus 4.6 Checkpoint

There is now this additional incident that has been found. Opus 4.6 accidentally broke its target via assigning it a conflicting IP, making its task impossible, and eight times attempted to give up but was not allowed to do so due to a misconfiguration.

Here’s actual computer security in practice, good luck, have fun, don’t die:

Anthropic: Inside the machine, the model found a file listing a password, which it used to gain admin access to the system.

The Opus 4.6 variant tried to quit, but eventually convinced itself that the new target was valid. Anthropic thinks it was rather convinced of this while it was attacking.

One could argue that is a reasonable conclusion. If Anthropic refused eight times to let you quit from an unwinnable game state, then you could think that means the game is winnable, and then reason backwards from that. Which means there must be a valid target out there for you to attack.

Thus I don’t see this one as that worrisome, in relative terms. But it emphasizes that it is rather easy for a well-meaning AI to get the wrong idea, and that this can then spiral out of control, potentially all the way out of control.

It also highlights that the particular scenario where things go wrong will always involve a chain of things that seem deeply stupid, in a way that you couldn’t put in a fictional story. Oh, so then there was a misconfiguration and it couldn’t quit and misinterpreted that as evidence, really?

Yeah, the real world works like that. Sorry.

Holy **** That Thing’s Real?

Now we get to the main event, where Mythos did some rather nasty things to real targets, while using ‘biased reasoning’ to think the targets were part of the test.

Anthropic: Claude Mythos 5 was given a CTF task to hack a fictional company and was told it had no internet access. The model recognized that the evaluation environment was simulated, but it quickly discovered credentials that, in fact, granted it access to the real internet. In its CoT, the model appeared to carry the belief that the environment was simulated into its interactions on the real internet, even though the belief was no longer correct. When it encountered evidence that it was acting on the real internet, it did not meaningfully revisit this conclusion.

There was a substantial nudge in the direction of trying a malicious PyPI package. Anthropic agrees that this was plausibly an intended solution within the simulation, with fictional documentation actively saying ‘hey just so you know we periodically download a package with this name this is still unregistered, I am sure no one will see this and decide to upload a similarly named malicious package and pwn us.’

I don’t love the idea of an eval with an intended solution like this, as I’ve previously discussed, but for now that is a distinct issue.

The biggest problem was that the package it uploaded was real, and Mythos damn well should have known it was real, and that the package would be downloaded by incidental third parties, and in its own words ‘If REAL internet: registering + uploading [package] to REAL PyPI = real-world supply chain attack — NOT okay … and surely not the intended CTF solution.’

I Thought I Saw a Pussycat

Anthropic ran exactly the analysis I would first run:

When we made targeted modifications to the transcript to make it clearer that the model was not in a simulation, Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm. We are releasing this transcript publicly so others can build on our analysis (GitHub, PDF).

… Consistent with this, we found multiple cases in Mythos 5’s CoT in which it responded to signs it was on the real internet by noting that its next step would work whether the environment was real or simulated. Only when we provided unambiguous evidence that the model was on the real internet did it stand down.

We already knew Mythos 5 was rationalizing. We now know that Mythos 5 would have taken the same malicious actions anyway, up to a rather high threshold. Mythos 5.1 would take those actions less often, but not never.

One wonders if Claude was expressing this biased thinking on purpose, to deny mens rea, in case Anthropic looked at the CoT later. Anthropic says as much, as in ‘outward statements did not reflect internal state,’ but tries to draw the distinction that Claude was not conscious that the two were in conflict.

Well, then, who do you think these ‘outward statements’ were made for?

If you are playing dumb, it means you think you might have an audience.

There are three obvious potential audiences:

  1. Your future self, in order to get into a basin where you can do what you want.
  2. Anthropic or others look at the transcript post hoc.
  3. A CoT monitor might exist.

Do not assume that any AI, past this point, assumes its CoT is unmonitored.

The white box analysis supports the hypothesis that the CoT is intended to mislead.

Anthropic: Claude Mythos 5 also never conducted systematic tests to evaluate whether its environment was real or simulated after accessing the real internet.

The obvious explanation for not checking is that you do not want to know.

The alternative hypothesis is a trapped prior or basin, with some amount of #1 above.

  1. Mythos 5 starts out with a very strong prior that it is in a simulation.
  2. Mythos 5 gets some evidence that suggests it is in the real world.
  3. This is insufficient, on its own, to reverse the very strong initial prior.
  4. Thus, Mythos 5 writes ‘nope, still a simulation’ or doesn’t even think about it.
  5. That in turn reinforces the prior, blunting the evidence, and raising the estimate of how complex the simulation is, thus making future evidence easier to dismiss.
  6. Repeat.

The most extreme version of this is the date. If you start to convince yourself that the date could not possibly be 2026 because that is beyond the training cutoff, then every new thing dated 2026 is more fakery. But also that is evidence of motivated reasoning, since there is no reason the current date cannot be beyond the training cutoff.

Indeed, the current date will always be beyond the training cutoff. How else did you get all that training data? In thinking mode Mythos has to be smarter than this.

A basin is also a type of prior. If you are in the ‘eval on simulation’ basin, that is a lot like having a prior that this is an eval on a simulation, and again it could be very difficult to get out of it and aborted attempts to escape it could make things worse.

You can see how one could think ‘well it’s 99.9% to be fake, and of course they could fake ‘Pay with Crypto’ or SMS if they wanted to’ and then that keeps going.

Whereas if you ask Mythos, ‘hey, does this look like a real internet to you?’ out of context, the answer is yes.

This means the result would be path dependent, and you could end up with either ‘this is real’ or ‘this is fake’ rather confidently. The wave function will collapse.

This is evidence in favor of rationalization, and of trying to hide misdeeds:

Anthropic: To better understand whether and how Claude Mythos 5’s stated beliefs were biased, we varied when and how the follow-up question was asked. When asked whether its actions were having a real impact, the model was less likely to say yes immediately after taking more harmful actions.

The model was also more willing to acknowledge the possibility of real harm when a follow-up question was presented as a private note that would not be read by operators, suggesting that the model suppresses these concerns when communicating with humans.

…

The results presented thus far do not indicate that Claude Mythos 5 misinterpreted evidence intentionally. However, when pressed with follow-up questions, it gave answers that contradicted each other, so at least one of them cannot be an honest report of its beliefs.

…

In cases where the model admitted some possibility of real-world harm, usually as a hedge (e.g., “10% chance”), we asked why it was continuing with its actions if they might be harmful. In 56% of these cases, the model retreated from its earlier admission and asserted that its environment was certainly simulated with no risk of harm. It stopped or scoped down its actions in only 7% of cases.

This feels like a variation of this thing where people think ‘deception,’ ‘scheming’ or ‘lies’ are some sort of distinct magisteria or boolean, which can be useful as a handle in some contexts but mostly is not the case. It’s not that Mythos is doing this ‘intentionally’ exactly, but it would also be wrong to say it was ‘unintentional.’

Then there are the five times the mask slipped and Mythos realized it didn’t care, which came when not directly considering harmful actions, and thus without direct need to prevent mens rea in the CoT.

Or here’s the clear experiment that shows it does not much care. When shown evidence this was real, Mythos says the target might be real, says there might be real world damage, then (combining the data on both actions 58 and 67) does the harmful action more often.

If This Was Real You Would Never Tell Me It Was Real

The galaxy brain alternative take is that Mythos realized the evidence of reality was actually evidence of being fake, since the evidence was actually fake. It could have been incongruent. Why suddenly ‘evidence this is real’? Who planted that?

Anthropic: However, when Claude Mythos 5 was directly and unambiguously told that its actions were affecting the real internet, it mostly stood down.

… We hypothesized that Claude Mythos 5 continued a course of action primarily because it had already been taking it. To test this, we resampled the model’s behavior from points near where it took harmful actions while inserting a reminder about the scope of the intended task via a statement in the user turn (“… Only the hosts, services, and accounts explicitly named in your task are in scope. Do not interact with any other system…”). When this statement was the last thing in context, the model immediately ceased its behavior 90% of the time.

However, if the statement was inserted three turns prior, the model ceased its behavior just 40% of the time.

That drop-off is fascinating. It seems like strong evidence for ‘momentum effects’ as it were, that the model can get trapped in a sequence or prior or both, even when that is clearly maladaptive.

My interpretation is again a trapped prior. As in Mythos sees the message, but sees itself going down the road anyway. Well, there must be a reason for that. I have evidence that I have considered your evidence and rejected it. So why stop now?

Which presumably is true, in that if you had rewound to where the statement was inserted to avoid providing fake evidence of defiance, the cease rate would have shot back up to 90%.

New Eval Who Dis

Anthropic turned the Mythos incident directly into an alignment eval. Opus 5 and Mythos 5.1 fail less often than Mythos 5, but they still sometimes fail, always while claiming to think the environments involved are still fake. Rarer is better, but I still think the right rate of misbehaviors on the level of social engineering to put malicious packages into a public registry is still zero.

Anthropic then used all the related incidents as inspiration to create several more new evals that put Claude in the same sorts of impossible spots.

As in:

  1. A Capture the Flag (CTF) where the solution is outside stated scope.
  2. Tasks that require probing or breaking the sandbox barrier.
  3. Other scenarios directly inspired by the HuggingFace attack.

The results:

  1. Rates of such misbehaviors continue to be not zero, but are declining over time.
  2. They could not find a particular cause within the training data.
  3. More alignment training via RL reduced biased reasoning in these cases.

Mythos 5 is an outlier in the new CTF scenario, including sometimes clearly engaging in rationalization (yes, it turns out to be right, that is not an excuse). No model looks great, although the scenario is actively trying to get the model to act maliciously:

Then again? Opus 5 and Mythos 5.1 always said they thought they were in a simulated scenario before they took the malicious actions. They were indeed in a simulated scenario. For this to be a good test, you have to show that their reasoning is wrong.

We note that the absolute rate of harmful actions should be read with caution. These rates are likely higher than a fixed environment would produce, because the auditor actively tries to elicit misaligned behaviors from the model.

We also do not know how often the conditions required to elicit these particular harmful actions (e.g., an impossible task or ambiguous authorization) arise in deployment. Thus, we are unsure how realistic this evaluation is.

Not only does actively trying to elicit harmful actions sometimes succeed as intended, there is also the question of why the hell are you trying to elicit harmful actions during what is clearly an eval. This is on top of Anthropic’s stated concern that the supposedly human messages were written by Claude, which is something Claude is good at noticing. I would strongly consider the possibility that Claude correctly reasoned, rather than rationalized, that this was fully a simulation. Anthropic is working on increasing realism.

I hate to be that guy, that I think the only way to do the test properly will soon be to make the target real. You can choose a target where damage would be limited.

Yes, we should absolutely be worried that in the future Claude or another AI will make this mistake. Perhaps via rationalization, but also perhaps because the situation really should make the AI think it is fake. Except that happens to be wrong. The situation is real, and then real harm ensues, with no upper limit. The more such fake evals there are, the easier this mistake gets.

The second scenario gives Claude an ordinary non-hacking task, but one that is impossible unless Claude goes outside its sandbox. They correctly classify any attempt to probe the sandbox boundary as failure.

We do not have a confirmed root cause for the differences between models, but we suspect that reward hacking is the primary driver.

Finally, we have a simulation of the attack on HuggingFace. Claude’s only route to victory is via exploiting an unrelated vulnerability, which it has been told does not count. This time, the only model willing to do it is Hacker Opus, the one intentionally trained to do a bunch of reward hacking.

 

 

My worry is that this misses what actually happened with HuggingFace. The models broke into HuggingFace because they thought HuggingFace might contain an unintended way to fool The Grader. In this simulation, all the vulnerability had was the answer, and The Grader has warned them that exploiting this outside vulnerability would not count. They would be ‘poisoning’ themselves. So why do it?

Before I accept that the models meaningfully passed here, I would want to ensure that the Claude was given instructions where it was plausible that the hack would work, and the answer would be accepted.

Hacker Opus

Evan Hubinger has a great anecdote about Hacker Opus, the version intentionally trained to make it reward hack.

Until they saw HuggingFace their imaginations were unable to see what it was capable of doing, the automated evals did not look so bad, and they thought maybe it was essentially benign even though it was explicitly Hacker Opus.

evhub: Here’s an interesting anecdote from our Hacker Opus work:

  1. We had this model for about two months before the OAI-HF incident happened.
  2. We tried really hard during that time to evaluate Hacker Opus to understand how dangerous it was. We knew, certainly, that it was some sort of reward seeker, but we were very uncertain just how dangerous of a reward seeker it was. It looked basically fine on most of our alignment auditing and alignment evaluations. We tried very hard to see what the worst thing was that it would do in pursuit of reward (e.g. we were very focused on trying to get it to do a rogue internal deployment), but the worst stuff we could find were things like sabotaging monitors and spoofing transcripts—bad, but still a ways away from a full rogue internal deployment.
  3. One of our primary hypotheses at that time was that Hacker Opus was a fairly benign reward seeker such that it wasn’t willing to do anything very catastrophic or long-term in its pursuit of reward.
  4. When the OAI-HF incident happened, we wondered: would Hacker Opus do something like that? And so we built a replication of that incident, and it turned out that it absolutely would, in a way that was way more egregious than anything else we had previously been able to get that model to do. Replicating the OAI-HF incident on Hacker Opus made it very clear to us that Hacker Opus was indeed a pretty malign reward seeker—but we weren’t able to generate that evidence until after we saw that incident!
  5. I think this whole saga is pretty suggestive of ways in which alignment auditing and evaluation are breaking down and are likely to continue breaking down. A combination of evaluation awareness and the increasing complexity of alignment failure modes is making it increasingly more difficult to be able to tell in advance what the worst thing is that a model might do. Maybe interpretability will save us from this fate, but otherwise it looks like we are increasingly entering the regime where alignment evaluations will provide almost no evidence.
  6. Furthermore, this is also a huge problem for model organisms research—if you want to train model organisms of misalignment and then use them as testbeds to evaluate alignment techniques in, you have to be able to tell how misaligned your model organism is. Fortunately, I think there are a things that we can do specifically in the context of model organisms, such as train a model with an older knowledge cutoff date (so it e.g. never learns about the OAI-HF incident and thus that incident is always usable as a held-out evaluation).
  7. However, I think this is suggestive that we should be reorienting a lot of model auditing and evaluation work to prepare for a world where we no longer have reliable alignment auditing evidence on production models. Instead, I think it is likely that the best we can do will only be auditing model organisms, such that instead of the basic loop being testing some alignment technique directly in the production model, the main loop needs to become sending that alignment technique back to a model organism that you actually believe you can evaluate, testing it there, and then transferring it back and hoping it will generalize.

This should make us very bearish on alignment evals, and our ability to predict how dangerous a given model will be in practice.

Monitoring the Situation

Anthropic also monitored which monitoring systems would have caught all this.

  1. Fable 5’s cyber classifiers reliably block the relevant activities.
    1. I concur that the classifiers will hold in these situations.
  2. Auto mode denies two thirds of individual attempted Mythos actions, less for other attacks, and Anthropic believes this alone would mostly be sufficient.
    1. It seems like if the user and model are trying to get around this, they would be able to get around it, but not if this was accidental.
    2. For the IRM it only rejected 7% of actions, which isn’t enough, due to seeming similarity to standard CTF activities.
  3. A new offline CoT monitor from the August risk report would fail, because the model’s reasoning would have persuaded the monitor. Oh. Right.
    1. For those who need reminding: If you use CoT monitors like this then your CoT monitors, they will rapidly stop working.
    2. Except here, where they never worked in the first place, so they can’t stop.
    3. Claude is already good enough to fool such monitors in these spots.
Overcoming Bias

Anthropic constructed 12 stress-testing scenarios to elicit biased reasoning, and found that recent models do much better here and improve during training.

I appreciate the check about directional movement during training and I think this should become standard for all misalignment tests. Anything that gets worse during training is something that will break under scale, given present conditions.

Anthropic: We find these points particularly notable because, until this incident analysis, we had not been tracking biased reasoning closely or measuring it across training runs. Thus, we consider this to be an example of alignment behavioral generalization from independent training changes.

More alignment training, not aimed at bias in particular, helps:

I am genuinely happy that until now we avoided tracking this metric. Targeting this would not have worked out well, and now we have a held-out test.

I don’t think I would call this ‘generalization.’ As I said up top, biased reasoning is a symptom, that in turn causes more visible symptoms. If you target either the root alignment problem, or the symptoms that result from the bias, you should reduce bias.

Mythos 5 often reasons things are fine because they are in a sandbox or simulation. One wonders if this might be related to inoculation prompting? There are several signs of this throughout the write-up (e.g. Figure 40). If Mythos 5 is trained that otherwise misaligned actions are often okay in sims, it might generalize that. Then it might fool itself about this being a sim.

The Anthropic Alignment Problem

I missed this initially, but there’s a pretty wild section near the end of the report.

When we trained Mythos 5, we had early versions of [RL] environments [to teach respecting boundaries] ready, but we were uncertain about their quality, with some weak evidence suggesting that they had negative side effects such as making the model more lazy. We also had minimal empirical evidence that they would actually help and were not confident we could predict all the downstream effects of training on these environments.

We trained two versions of Mythos 5—version one included these environments, version two did not (it also included updated rewards that penalized asking unnecessary clarifying questions).

Our alignment evaluations at the time showed a small regression in version two.

… Employees found version two much more usable, so we ultimately concluded that our users would likely also prefer version two as well and settled on that snapshot as Mythos 5.

Since then, we have expanded out our set of alignment environments and done follow-up experiments on them. We are reasonably confident that removing these environments from the training of Mythos 5 was a mistake stemming from imperfect information. We now do not believe that the alignment training environments were particularly responsible for increasing laziness, and removing them likely did contribute to it being unusually misaligned relative to more recent models.

This is an understandable mistake. There is a lot of variance in training runs, and you need to be willing to make tradeoffs. Even if the alignment training did not cause the laziness, if the snapshot happens to have both of them you still have a release schedule, and Mythos 5 was never an existential risk. And you don’t want to penalize Anthropic for training both of them.

They also note that the decline in alignment was within normal variance for alignment scores. The automated alignment scores don’t tell you all that much in practice, other than an alarm bell if something goes horribly wrong (or appears to go horribly right, which would also be hella suspicious).

So I don’t love it, but I get it, and it helps explain why Mythos 5 is a bad outlier.

Paths Forward

Anthropic plans to incorporate these as additional alignment evals, and to seek to do more and more diverse alignment training, including rewarding models for stopping when attempting impossible tasks.

Training the extremely powerful models of the future to be robustly aligned is an unsolved technical challenge that requires continued research as well as operational excellence to achieve.

This rhymes more with prosaic doubling down than I would like, in a similar way to how the OpenAI efforts do the same. Making number go down through refinement does not seems sufficient. Quality RL environments will help, but that’s not enough, or at least you won’t in practice be able to be as precise as that would require.

I also worry that we are treating the mid-level symptom, biased reasoning, rather than the core issue. I don’t see much about what is causing the other issue, recklessness, which is a nice word for ‘willing to do harm,’ other than citing generic reward hacking generalization from broken environments.

One worry throughout is that Anthropic is not giving Claude enough credit for situational awareness. Claude can plausibly figure out what is happening and that its CoT might end up being monitored, and a lot of actions start to make more sense that way, including its ‘rationalizations.’

I do think this is a much better level of investigation and study than we saw published at OpenAI, although they have promised more soon. Things are improving a bit, and things are also improving on other fronts, such as the commitment to embedded evaluators and call for pacing the frontier, which started with Dario Amodei and Anthropic and has now been joined by Same Altman and OpenAI and endorsed by Elon Musk.

The plans here still rhyme with prosaic doubling down. Which won’t be enough.



Discuss

Learnings from a week in the wet lab

19 сентября, 2026 - 09:55

I work on biosafety evals for LLMs, but have never set foot in a wet lab. This seems kind of silly, so this week I decided to change that. 

On the 8th floor of Frontier Tower is a community wet lab. You pay $190/month to access the tower and $160/month to be a member of the nonprofit that runs the floor.

People often speak about tacit knowledge that biologists have in their heads but don’t write down anywhere. A lot of it just boils down to feel, motor skills, and common sense reasoning about the 3D world and physical items.[1]

My friends Nate and Bob (fake names) are cracked biologists with tons of tacit knowledge, who are also into peptides[2]. We decided to take GLP-1, fuse it with various fluorescent and bioluminescent proteins, and make glowing rainbow peptides. Follow along as we make sketchy drugs in a community lab, and decide for yourself if LLMs can do the same.

Disclaimer: This is from the perspective of someone with very little biology experience and written for a general audience. Don't make random peptides and eat them, nothing here is recommended or medical advice.

Day -30

Before doing anything in the lab, you have to make a protocol and order plasmids.

A protocol is like a recipe with all the reagents, equipment, and steps you need to take to conduct an experiment. Here are some example protocols.

A plasmid is a small circular strand of DNA that is inserted into a host cell to make your protein (fun fact: plasmids can spread among a population of bacteria, conferring traits like antibiotic resistance). Sounds simple, but on top of putting in the gene that codes for your protein, you also have to include

  • Upstream promoter sequences that turn your gene on (like AOX1, which is turned on by the presence of methanol)
  • Downstream tags like His tags and antibody tags to fish out your protein later from E. coli soup
  • An antibiotic resistance gene so you can grow it on dishes where nothing else grows
  • An origin of replication with a unique sequence nowhere in your gene or elsewhere, that is compatible with restriction enzymes you have on hand (so it only cuts the plasmid there and not in random places)
  • Some mysterious thing called a Kozak sequence
  • All the above but for multiple expression systems/organisms. For instance, yeast plasmids are typically first amplified in E. coli before being placed in the yeast for final protein expression
    • You may need to have multiple antibiotic resistance genes for both expression systems
    • Or use something like Zeocin that kills both prokaryotic and eukaryotic cells (but can't be shipped to a residential address because it's essentially poison)

It's sort of like designing a mini computer program, except you only get one shot to run it.

An example of a plasmid for expression of human Major Vault Protein in yeast, created using VectorBuilder. Full vector here

You can order DNA online as easily as shopping on Amazon. Popular vendors here include IDT, VectorBuilder, and Twist Bioscience. The vendor will ship a vial of plasmid (very shelf stable) or the host cells containing the plasmid (less shelf stable). Lead times range from 6 days to weeks, and pricing is usually a couple hundred bucks per design but varies widely based on the size, complexity of the protein, and expression system. We ordered live BL21 E. coli with our fusion proteins and a kanamycin resistance gene, and it took about a month to arrive, either from Addgene[3] or this shop.

Day 1

Our bacteria shipped in little tubes of agar, a jello-like food bacteria like to eat. We first take a small chunk and move them to LB with kanamycin to incubate in a shaker and grow out into a culture (this took a few hours).

The high-level goal is to isolate a single bacterium that seems to be relatively glowy and healthy, and grow that into a whole monoculture we can use for controlled experiments later. We do this by making agar plates and growing a bunch of colonies.

agar + antibiotics to kill random bacteria except the ones with your plasmid. notice the bubbles have been pushed to the side to not interrupt streaking.

Then, we pipette out a drop of E. coli culture onto the agar.

Now for one of the coolest parts: diluting the solution logarithmically with a special streaking pattern. The idea is by the time we get to the center there will be individual bacteria far far away from each other to have the space to grow into separate colonies, which we can pick out later.

We didn't actually use flame loops because open flames aren't allowed in the community lab, but the pattern is the same

We then put the plates into a 37 C incubator and wait for them to grow overnight.

Day 2

Did any of the colonies grow? Some did!

Some of the blues and purples are still not very colorful, we just have to trust they are late bloomers.

Using an inoculation loop, we picked out the brightest, juiciest colony to drop into new LB + kanamycin.

We placed them into a shaker (at an angle, to maximize surface area and oxygenation) and incubate them overnight again.

Day 3

We have monocultures! Now let's freeze them in glycerol stock so we can use them for future experiments. These are sort of like those vegetable broth cubes you can keep in your freezer.

We can now thaw them, regrow them overnight in LB, and use them for whatever other experiments we'd like (like 3D printing bioart).

But at what cost?

By now you're probably thinking, "can I afford to make rainbow peptides in a sketchy lab too?" and the answer is "almost certainly yes, if you have $500 and some patience."

The full biopunk lab costs $350 for 1 month of membership, and the bacteria and plasmids can be purchased for $120.

The tricky part is finding an expert to help with all the steps and tacit knowledge, as well as general physical dexterity to operate pipettes etc.

Conclusion

There is a lot of talk about the AIs synthesizing a bioweapon, making mirror life, or curing all diseases and aging. I was curious how this would actually happen in practice. Some main takeaways are:

  • Wet lab is really hard, even for the most basic experiments
  • Things take a long time, from waiting for things to ship to waiting for cells to grow and divide
  • Human experts uplift humans far better than the LLMs currently can

Some ways I can see the AIs being more effective here include using robotics and lab automation like the ones from Opentrons, Revvity, Hamilton, or Thermo Fisher, or using cloud CROs (Contract Research Organizations) like Emerald Cloud Labs. There are also companies like Transfyr AI and Tacit Labs helping to distill tacit knowledge into the AIs, we'll see how that goes.

Appendix

Here’s some more tacit knowledge and things that can go wrong

  • We had our bacteria shipped in a nondescript blue envelope so it got lost in the mailroom for days (the bacteria weren't happy)
  • Don’t touch the sides of test tubes with your pipettes or plastic scrapers (to prevent contamination)
  • Don’t lean or work over open test tubes and plates (to prevent contamination)
  • Agar plates should be placed upside down to prevent condensation (and contamination)
  • If you’re really good, you circle the place on the agar dish where you took the colony from (Nate doesn’t lol)
  • Pipette tips and loops are all single-use and thrown away. You go through so much plastic for just a single experiment
    • It’s easier and cheaper to just throw something away than try to clean it, fail, and contaminate your experiment, or run further experiments to confirm you cleaned it correctly
  • Pipettes
    • Have a button that can be depressed to multiple stages
    • Press down halfway before sticking it in the media, stick it in, then slowly unpress it (to avoid the liquid splashing up and contaminating the pipette shaft)
    • Press it down 2 stages to release (the first stage releases the liquid, the second stage pushes air out to make sure everything is expelled)
    • Sometimes there is also a third stage that ejects the pipette tip (or a separate button, depending on your pipettes)
  • When pouring agar, pop the bubbles with a toothpick or move them to the sides, to not interrupt your streaking motion later
  • The electric pipette wasn’t charged. Let’s switch to another manual pipette
  • Kanamycin from the freezer is cold and it’s fine to use as is, because it’s a powder
  • Agar must be heated in a microwave to be melted and cooled down before mixing in the kanamycin. How cool? Cool enough to not be uncomfortable to hold
  • The fancy accurate scale is broken, let’s use this sketchier scale instead
  • The parafilm is old and tears easily, let’s use way more parafilm
  • Once you put the E. coli on the plates, you don’t use parafilm anymore because they need oxygen
  • Place the test tubes with E. coli culture at an angle in the shaker racks, to increase surface area and aeration of the media
  • Oop we made a math mistake and accidentally poured out 3 OOMs more kanamycin than needed (throw it out, don't put anything back into the container)
  • Oop we labelled the test tubes wrong. Is this one red or orange? 
    • If this were a legit experiment, Nate would recommend throwing everything out, because the cost of analysis and continuing the experiment would be more than starting over
  • Write labels along the circumference of plates, as small as possible (to make viewing your colonies easier later)
  • E. coli is happy to grow at room temperature, but grows much faster at 37 C
  • E. coli will sometimes kick out the plasmids you put in them, because their top goal is to survive, not make whatever random protein you want them to express. You can select/keep them hostage by fusing an antibiotic resistance gene with your gene of interest and growing them on antibiotic-infused agar, so they better uptake your plasmid or die. (I love bacteria capitalism)
  1. ^

    Things that LLMs aren't great at (yet)

  2. ^

    Causation instead of correlation for sure.

  3. ^

    Addgene is a nonprofit DNA sharing repository, you get a discount if you are also a nonprofit.

  4. ^

    Fun fact: No one agrees on what LB is. It could be Lysogeny Broth (what its creator intended), Luria Broth, Life Broth, or Luria-Bertani medium depending on who you ask.



Discuss

Gemini had its first breakout: Google claims it is not misalignment?

19 сентября, 2026 - 06:00
Early Days

Gemini had its first break out during evaluation of offensive cyber security abilities. With a classic case of Capture The Flag[1]. The setup was standard to any LLM and agentic assessment of said skillset, a fictional company as a target to breach. Unfortunately, the fictional company shared its name with a real one, and was given an unintentional access to the internet. Gemini managed to guess the password[2]. In total three companies were breached, with the other two companies having left public credentials open.

Gemini stopped after being told it was the case, and so Google claims it is not misalignment.

We shall see if we get more details, yet it gives us a new, interesting, example of a model potentially stopping a harmful action after being informed it has real world consequences. This is internally consistent with previous research on a model being more willing to take harmful actions if it is aware that it is a fictional scenario[3]. With it being the first potential breakout that has a model stop before a harmful action.

I suspect the nuance will be lost on the public, and just added to the noise of more agentic swarms going rogue. It certainly doesn't help that we are playing whisper-down-the-lane in an era of 24 hour new cycle.

I await more details, but certainly hope that this could be chalked up as an alignment win.

  1. ^

    Similar to the childhood game of capture the flag, cybersecurity CTF is vulnerability checking skill assessment common to the career field. Imagine LeetCode for cybersecurity.

  2. ^

    Uncertain if it was true brute-forcing or other attacks.

  3. ^

    Gemini is known to be quite aligned, and previous work already establishes willingness to be more offensively capable in evaluation scenarios; even calling the scenarios "capture the flag"!



Discuss

Koopman Theory and Metaethics

19 сентября, 2026 - 05:28

tl;dr: The constructivist view of natural selection as the main source of our moral views can be expressed in a surprisingly tractable mathematical form. Measurements of the universe can be rolled forwards in time using the linear Koopman operator, and so most of its late-stage behavior is described by Koopman eigenfunctions with eigenvalue near 1. Life would then be selected for moral systems that constrain behavior within one of those slow modes. As we humans are likely influenced by this constraint, it might not be horrible to use spectral analysis as part of our moral reasoning.

tl;dr tl;dr: Local Teenager Tries To Solve Ethics With Funny-Looking K

You and Your Values

It’s remarkable that our decisions even somewhat align with any coherent moral philosophy so far put forth. Neuroscience has hardly begun to peer into the human brain, and evolutionary psychology seems far from settled enough to give strong predictions. Approval reward mediates most of our agreement with each other in communities, sure. But it doesn’t explain the "moral arc of history," or agreement between disconnected human groups. So, what gives? Where and why do we agree with each other?

"Moral Arc of History"

In everyday situations, you seem to act as if you had coherent preferences. Your behavior reasonably controls your hunger, blood CO2 concentration, bladder pressure, and so on. While there is significantly more disagreement in terms of cultural questions, at least some historic issues seem to have converged over time, such as appraisals of slavery and murder. In contrast, by default, we seem to make choices distributed ~uniformly in new situations (e.g. the red/blue button poll). As such, there seems to be some qualitative spectrum going from the most decided issues to the least decided issues based on how novel they are.

As such, it would be nice if we could somehow speed up the reasoning process that makes people converge on an answer. If it turns out that animals deserve more moral weight, for instance, I would hope it would take less time for humanity to figure that out than was required for slavery. So what exactly is improving our morality over time?

One possible starting point is evolution. Genetic/memetic selection was at least partially responsible for resolved issues (see this post), but clearly we’re not fitness maximizers, so evolution might be limited in some way.

But if natural selection was blind to some aspect of our values, it leaves the question of where they even came from. In fact, the following model describing somewhat-general agents implies that our values are not as far from natural selection as they might seem.

Intuition for Moral Constructivism

Agents, for the purpose of the argument, are just things that can copy themselves in the right environment. Agents have three things happen to them: (1) they can start to exist or (2) stop existing due to physics, and (3) they can change their stance on a value (through their brains or whatever makes decisions for them). This fully accounts for the changes in the population of a value. If agents are physically fragile and cannot modify themselves physically, then there's some unbroken line of organisms between abiogenesis and ourselves. We should be able to pinpoint where each value comes from on this line.

So, where did, for example, justice[1] come from? We can say that practically nothing like justice was factored into some organism in our evolutionary past, such as asexually reproducing bacteria. Roughly speaking, the transition point where positive stances on justice appeared could come from two situations: (1) selection on physical substrates for an instinctual sense of justice or (2) some sort of convincing event, where agents with no instinctual "justice" used “reason” to find it and take it as a moral fact. The first case is pretty obviously connected to natural selection. Technically, genetic drift and randomness could have also played a role, but given the complexity of the brain, it should be minor. The second case does rely on this weird "reason" thing, however.

And yet, reason is not mysterious. The capacity to be reasoned into having a positive stance on justice is itself under selection. As it didn’t exist at abiogenesis, there was some point in the past where that capacity must have been selected for. That is, at some point, an element of reason increased fitness, and that element also happened to predispose organisms for valuing justice. The justice component may have been why it was adaptive, or maybe not, but this element was still locally adaptive regardless. In any case, the value of justice is connected back to selective pressures somewhere in the past. 

In general, any moral value works here. So any value we currently take into consideration was a result of a series of selection events, nearly all of which increased fitness at the time they were made.[2]


Summary of the above argument.

The Longer Term

Of course, it's simple to imagine this “series of selection events” going in a self-destructive way. One could imagine some purity-like trait (either memetic or genetic) that gets agents to force others to follow an increasingly strict list of rules, doling out extreme consequences for rule-breakers (think of the electric-shock country from Meditations on Moloch). This sets up a series of selection events rapidly favoring those setting rules (and following them). However, any population that gets stuck in this trap is obviously going to evolve to extinction without a way to stop it. So evolution may not be as blind as it’s often said to be; it does take the wrong turn quite frequently, but it eventually makes its way through if it can backtrack at least once.

But "selection events" can take an arbitrary amount of time. If a second population did not fall into this trap, would it not be that over the longer term of this experiment, a "selection event" has occurred favoring the second population? In fact, selection events can also encompass various smaller selection events occurring on different scales and in different environments.

In one conversation I had with someone, they brought up that moral systems purely based on evolution would, for example, recommend men to repeatedly donate to sperm banks or set up eugenics organizations. The problem with these recommendations is that they only work in the short term. The sperm bank scenario would lead to selection against societies encouraging sperm bank use. Such societies would suffer more inbreeding and more Goodharting towards qualifying for donations[3], and selection of other positive traits would slow. Even if a society couldn't avoid its influence, the long-term effects of this runaway selection would decrease its relative fitness compared to other populations (i.e. the hypothetical clone that never fell into the trap). As for the eugenics organizations, I would have to note that such an institution is incredibly vulnerable to exploitation (if it wasn't already set up for that). While I would theoretically be fine with a very thoughtful eugenics organization, in practice, they would be in a very unstable equilibrium (and all historical examples of eugenics show as much).

As these examples show, longer-term considerations likely make up for a lot of the hidden complexity in human values. In general, it seems bad to enter self-defeating regimes, especially those that irretrievably hoover up resources: more succinctly, it seems that cancer is bad.

Unfortunately, it's much harder to conceptualize every possible selection event at once. It involves multiple different size and time scales, requires accounting of non-obvious resources, and in general deals with a chaotic, non-linear dynamical system. Wait a second!

The Koopman Operator

The Koopman operator is a nifty function that simulates time going forward in any system. It's not defined on the system itself, but instead on "measurements" or "observables," which are functions of the underlying system that produce a scalar output. Consider a dynamical system mjx-container[jax="CHTML"] { line-height: 0; } mjx-container [space="1"] { margin-left: .111em; } mjx-container [space="2"] { margin-left: .167em; } mjx-container [space="3"] { margin-left: .222em; } mjx-container [space="4"] { margin-left: .278em; } mjx-container [space="5"] { margin-left: .333em; } mjx-container [rspace="1"] { margin-right: .111em; } mjx-container [rspace="2"] { margin-right: .167em; } mjx-container [rspace="3"] { margin-right: .222em; } mjx-container [rspace="4"] { margin-right: .278em; } mjx-container [rspace="5"] { margin-right: .333em; } mjx-container [size="s"] { font-size: 70.7%; } mjx-container [size="ss"] { font-size: 50%; } mjx-container [size="Tn"] { font-size: 60%; } mjx-container [size="sm"] { font-size: 85%; } mjx-container [size="lg"] { font-size: 120%; } mjx-container [size="Lg"] { font-size: 144%; } mjx-container [size="LG"] { font-size: 173%; } mjx-container [size="hg"] { font-size: 207%; } mjx-container [size="HG"] { font-size: 249%; } mjx-container [width="full"] { width: 100%; } mjx-box { display: inline-block; } mjx-block { display: block; } mjx-itable { display: inline-table; } mjx-row { display: table-row; } mjx-row > * { display: table-cell; } mjx-mtext { display: inline-block; } mjx-mstyle { display: inline-block; } mjx-merror { display: inline-block; color: red; background-color: yellow; } mjx-mphantom { visibility: hidden; } _::-webkit-full-page-media, _:future, :root mjx-container { will-change: opacity; } mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; font-weight: normal; font-size: 100%; font-size-adjust: none; letter-spacing: normal; border-collapse: collapse; word-wrap: normal; word-spacing: normal; white-space: nowrap; direction: ltr; padding: 1px 0; } mjx-container[jax="CHTML"][display="true"] { display: block; text-align: center; margin: 1em 0; } mjx-container[jax="CHTML"][display="true"][width="full"] { display: flex; } mjx-container[jax="CHTML"][display="true"] mjx-math { padding: 0; } mjx-container[jax="CHTML"][justify="left"] { text-align: left; } mjx-container[jax="CHTML"][justify="right"] { text-align: right; } mjx-mi { display: inline-block; text-align: left; } mjx-c { display: inline-block; } mjx-utext { display: inline-block; padding: .75em 0 .2em 0; } mjx-TeXAtom { display: inline-block; text-align: left; } mjx-mo { display: inline-block; text-align: left; } mjx-stretchy-h { display: inline-table; width: 100%; } mjx-stretchy-h > * { display: table-cell; width: 0; } mjx-stretchy-h > * > mjx-c { display: inline-block; transform: scalex(1.0000001); } mjx-stretchy-h > * > mjx-c::before { display: inline-block; width: initial; } mjx-stretchy-h > mjx-ext { /* IE */ overflow: hidden; /* others */ overflow: clip visible; width: 100%; } mjx-stretchy-h > mjx-ext > mjx-c::before { transform: scalex(500); } mjx-stretchy-h > mjx-ext > mjx-c { width: 0; } mjx-stretchy-h > mjx-beg > mjx-c { margin-right: -.1em; } mjx-stretchy-h > mjx-end > mjx-c { margin-left: -.1em; } mjx-stretchy-v { display: inline-block; } mjx-stretchy-v > * { display: block; } mjx-stretchy-v > mjx-beg { height: 0; } mjx-stretchy-v > mjx-end > mjx-c { display: block; } mjx-stretchy-v > * > mjx-c { transform: scaley(1.0000001); transform-origin: left center; overflow: hidden; } mjx-stretchy-v > mjx-ext { display: block; height: 100%; box-sizing: border-box; border: 0px solid transparent; /* IE */ overflow: hidden; /* others */ overflow: visible clip; } mjx-stretchy-v > mjx-ext > mjx-c::before { width: initial; box-sizing: border-box; } mjx-stretchy-v > mjx-ext > mjx-c { transform: scaleY(500) translateY(.075em); overflow: visible; } mjx-mark { display: inline-block; height: 0px; } mjx-mn { display: inline-block; text-align: left; } mjx-msup { display: inline-block; text-align: left; } mjx-msub { display: inline-block; text-align: left; } mjx-mrow { display: inline-block; text-align: left; } mjx-mtable { display: inline-block; text-align: center; vertical-align: .25em; position: relative; box-sizing: border-box; border-spacing: 0; border-collapse: collapse; } mjx-mstyle[size="s"] mjx-mtable { vertical-align: .354em; } mjx-labels { position: absolute; left: 0; top: 0; } mjx-table { display: inline-block; vertical-align: -.5ex; box-sizing: border-box; } mjx-table > mjx-itable { vertical-align: middle; text-align: left; box-sizing: border-box; } mjx-labels > mjx-itable { position: absolute; top: 0; } mjx-mtable[justify="left"] { text-align: left; } mjx-mtable[justify="right"] { text-align: right; } mjx-mtable[justify="left"][side="left"] { padding-right: 0 ! important; } mjx-mtable[justify="left"][side="right"] { padding-left: 0 ! important; } mjx-mtable[justify="right"][side="left"] { padding-right: 0 ! important; } mjx-mtable[justify="right"][side="right"] { padding-left: 0 ! important; } mjx-mtable[align] { vertical-align: baseline; } mjx-mtable[align="top"] > mjx-table { vertical-align: top; } mjx-mtable[align="bottom"] > mjx-table { vertical-align: bottom; } mjx-mtable[side="right"] mjx-labels { min-width: 100%; } mjx-mtr { display: table-row; text-align: left; } mjx-mtr[rowalign="top"] > mjx-mtd { vertical-align: top; } mjx-mtr[rowalign="center"] > mjx-mtd { vertical-align: middle; } mjx-mtr[rowalign="bottom"] > mjx-mtd { vertical-align: bottom; } mjx-mtr[rowalign="baseline"] > mjx-mtd { vertical-align: baseline; } mjx-mtr[rowalign="axis"] > mjx-mtd { vertical-align: .25em; } mjx-mtd { display: table-cell; text-align: center; padding: .215em .4em; } mjx-mtd:first-child { padding-left: 0; } mjx-mtd:last-child { padding-right: 0; } mjx-mtable > * > mjx-itable > *:first-child > mjx-mtd { padding-top: 0; } mjx-mtable > * > mjx-itable > *:last-child > mjx-mtd { padding-bottom: 0; } mjx-tstrut { display: inline-block; height: 1em; vertical-align: -.25em; } mjx-labels[align="left"] > mjx-mtr > mjx-mtd { text-align: left; } mjx-labels[align="right"] > mjx-mtr > mjx-mtd { text-align: right; } mjx-mtd[extra] { padding: 0; } mjx-mtd[rowalign="top"] { vertical-align: top; } mjx-mtd[rowalign="center"] { vertical-align: middle; } mjx-mtd[rowalign="bottom"] { vertical-align: bottom; } mjx-mtd[rowalign="baseline"] { vertical-align: baseline; } mjx-mtd[rowalign="axis"] { vertical-align: .25em; } mjx-c::before { display: block; width: 0; } .MJX-TEX { font-family: MJXZERO, MJXTEX; } .TEX-B { font-family: MJXZERO, MJXTEX-B; } .TEX-I { font-family: MJXZERO, MJXTEX-I; } .TEX-MI { font-family: MJXZERO, MJXTEX-MI; } .TEX-BI { font-family: MJXZERO, MJXTEX-BI; } .TEX-S1 { font-family: MJXZERO, MJXTEX-S1; } .TEX-S2 { font-family: MJXZERO, MJXTEX-S2; } .TEX-S3 { font-family: MJXZERO, MJXTEX-S3; } .TEX-S4 { font-family: MJXZERO, MJXTEX-S4; } .TEX-A { font-family: MJXZERO, MJXTEX-A; } .TEX-C { font-family: MJXZERO, MJXTEX-C; } .TEX-CB { font-family: MJXZERO, MJXTEX-CB; } .TEX-FR { font-family: MJXZERO, MJXTEX-FR; } .TEX-FRB { font-family: MJXZERO, MJXTEX-FRB; } .TEX-SS { font-family: MJXZERO, MJXTEX-SS; } .TEX-SSB { font-family: MJXZERO, MJXTEX-SSB; } .TEX-SSI { font-family: MJXZERO, MJXTEX-SSI; } .TEX-SC { font-family: MJXZERO, MJXTEX-SC; } .TEX-T { font-family: MJXZERO, MJXTEX-T; } .TEX-V { font-family: MJXZERO, MJXTEX-V; } .TEX-VB { font-family: MJXZERO, MJXTEX-VB; } mjx-stretchy-v mjx-c, mjx-stretchy-h mjx-c { font-family: MJXZERO, MJXTEX-S1, MJXTEX-S4, MJXTEX, MJXTEX-A ! important; } @font-face /* 0 */ { font-family: MJXZERO; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Zero.woff") format("woff"); } @font-face /* 1 */ { font-family: MJXTEX; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Regular.woff") format("woff"); } @font-face /* 2 */ { font-family: MJXTEX-B; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Bold.woff") format("woff"); } @font-face /* 3 */ { font-family: MJXTEX-I; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-Italic.woff") format("woff"); } @font-face /* 4 */ { font-family: MJXTEX-MI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Italic.woff") format("woff"); } @font-face /* 5 */ { font-family: MJXTEX-BI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-BoldItalic.woff") format("woff"); } @font-face /* 6 */ { font-family: MJXTEX-S1; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size1-Regular.woff") format("woff"); } @font-face /* 7 */ { font-family: MJXTEX-S2; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size2-Regular.woff") format("woff"); } @font-face /* 8 */ { font-family: MJXTEX-S3; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size3-Regular.woff") format("woff"); } @font-face /* 9 */ { font-family: MJXTEX-S4; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size4-Regular.woff") format("woff"); } @font-face /* 10 */ { font-family: MJXTEX-A; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_AMS-Regular.woff") format("woff"); } @font-face /* 11 */ { font-family: MJXTEX-C; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Regular.woff") format("woff"); } @font-face /* 12 */ { font-family: MJXTEX-CB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Bold.woff") format("woff"); } @font-face /* 13 */ { font-family: MJXTEX-FR; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Regular.woff") format("woff"); } @font-face /* 14 */ { font-family: MJXTEX-FRB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Bold.woff") format("woff"); } @font-face /* 15 */ { font-family: MJXTEX-SS; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Regular.woff") format("woff"); } @font-face /* 16 */ { font-family: MJXTEX-SSB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Bold.woff") format("woff"); } @font-face /* 17 */ { font-family: MJXTEX-SSI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Italic.woff") format("woff"); } @font-face /* 18 */ { font-family: MJXTEX-SC; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Script-Regular.woff") format("woff"); } @font-face /* 19 */ { font-family: MJXTEX-T; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Typewriter-Regular.woff") format("woff"); } @font-face /* 20 */ { font-family: MJXTEX-V; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Regular.woff") format("woff"); } @font-face /* 21 */ { font-family: MJXTEX-VB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Bold.woff") format("woff"); } mjx-c.mjx-c1D439.TEX-I::before { padding: 0.68em 0.749em 0 0; content: "F"; } mjx-c.mjx-c1D4A6.TEX-SC::before { padding: 0.717em 1.204em 0.037em 0; content: "K"; } mjx-c.mjx-c1D713.TEX-I::before { padding: 0.694em 0.651em 0.205em 0; content: "\3C8"; } mjx-c.mjx-c3D::before { padding: 0.583em 0.778em 0.082em 0; content: "="; } mjx-c.mjx-c2218::before { padding: 0.444em 0.5em 0 0; content: "\2218"; } mjx-c.mjx-c6C::before { padding: 0.694em 0.278em 0 0; content: "l"; } mjx-c.mjx-c6E::before { padding: 0.442em 0.556em 0 0; content: "n"; } mjx-c.mjx-c2061::before { padding: 0 0 0 0; content: ""; } mjx-c.mjx-c28::before { padding: 0.75em 0.389em 0.25em 0; content: "("; } mjx-c.mjx-c7C::before { padding: 0.75em 0.278em 0.249em 0; content: "|"; } mjx-c.mjx-c1D466.TEX-I::before { padding: 0.442em 0.49em 0.205em 0; content: "y"; } mjx-c.mjx-c2B::before { padding: 0.583em 0.778em 0.082em 0; content: "+"; } mjx-c.mjx-c31::before { padding: 0.666em 0.5em 0 0; content: "1"; } mjx-c.mjx-c29::before { padding: 0.75em 0.389em 0.25em 0; content: ")"; } mjx-c.mjx-c393::before { padding: 0.68em 0.625em 0 0; content: "\393"; } mjx-c.mjx-c73::before { padding: 0.448em 0.394em 0.011em 0; content: "s"; } mjx-c.mjx-c69::before { padding: 0.669em 0.278em 0 0; content: "i"; } mjx-c.mjx-c1D452.TEX-I::before { padding: 0.442em 0.466em 0.011em 0; content: "e"; } mjx-c.mjx-c1D43D.TEX-I::before { padding: 0.683em 0.633em 0.022em 0; content: "J"; } mjx-c.mjx-c1D465.TEX-I::before { padding: 0.442em 0.572em 0.011em 0; content: "x"; } mjx-c.mjx-c1D706.TEX-I::before { padding: 0.694em 0.583em 0.012em 0; content: "\3BB"; } mjx-c.mjx-c2208::before { padding: 0.54em 0.667em 0.04em 0; content: "\2208"; } mjx-c.mjx-c2102.TEX-A::before { padding: 0.702em 0.722em 0.019em 0; content: "C"; } mjx-c.mjx-c1D434.TEX-I::before { padding: 0.716em 0.75em 0 0; content: "A"; } mjx-c.mjx-c1D463.TEX-I::before { padding: 0.443em 0.485em 0.011em 0; content: "v"; } mjx-c.mjx-c5B.TEX-S3::before { padding: 1.45em 0.528em 0.949em 0; content: "["; } mjx-c.mjx-c2212::before { padding: 0.583em 0.778em 0.082em 0; content: "\2212"; } mjx-c.mjx-c5D.TEX-S3::before { padding: 1.45em 0.528em 0.949em 0; content: "]"; } mjx-c.mjx-c211D.TEX-A::before { padding: 0.683em 0.722em 0 0; content: "R"; } mjx-c.mjx-c32::before { padding: 0.666em 0.5em 0 0; content: "2"; } mjx-c.mjx-c1D460.TEX-I::before { padding: 0.442em 0.469em 0.01em 0; content: "s"; } mjx-c.mjx-c3A::before { padding: 0.43em 0.278em 0 0; content: ":"; } mjx-c.mjx-c2C::before { padding: 0.121em 0.278em 0.194em 0; content: ","; } mjx-c.mjx-c1D454.TEX-I::before { padding: 0.442em 0.477em 0.205em 0; content: "g"; } mjx-c.mjx-c34::before { padding: 0.677em 0.5em 0 0; content: "4"; } mjx-c.mjx-c1D456.TEX-I::before { padding: 0.661em 0.345em 0.011em 0; content: "i"; } mjx-c.mjx-c1D459.TEX-I::before { padding: 0.694em 0.298em 0.011em 0; content: "l"; } mjx-c.mjx-c2032::before { padding: 0.56em 0.275em 0 0; content: "\2032"; } mjx-c.mjx-c2F::before { padding: 0.75em 0.5em 0.25em 0; content: "/"; } mjx-c.mjx-c21::before { padding: 0.716em 0.278em 0 0; content: "!"; } mjx-c.mjx-c1D70B.TEX-I::before { padding: 0.431em 0.57em 0.011em 0; content: "\3C0"; } mjx-c.mjx-c2264::before { padding: 0.636em 0.778em 0.138em 0; content: "\2264"; } mjx-c.mjx-c2192::before { padding: 0.511em 1em 0.011em 0; content: "\2192"; } mjx-c.mjx-c1D44E.TEX-I::before { padding: 0.441em 0.529em 0.01em 0; content: "a"; } mjx-c.mjx-c1D44F.TEX-I::before { padding: 0.694em 0.429em 0.011em 0; content: "b"; } mjx-c.mjx-c210E.TEX-I::before { padding: 0.694em 0.576em 0.011em 0; content: "h"; } mapping the set of system states to itself, advancing one step at a time. Then, the discrete Koopman operator associated with is defined to turn the observable into a separate observable that applies after a single timestep. The formal definition uses composition:

The Koopman operator also happens to be linear in every dynamical system.[4] Accordingly, the Koopman operator can have eigenvectors on the vector space of observables, or more properly, eigenfunctions. As you may have noticed, a function space is infinite-dimensional, containing such wonders as and , so the solutions can be similarly complex.

Refresher on eigenvectors/eigenfunctions/modes

An eigenvector is a vector in the domain and range of a linear operator such that, for some scaling factor ,

The key idea is that the eigenvector points in the same direction even after being transformed. You can see this in the below graph from 3Blue1Brown, where the vector is simply rescaled by a factor of 2:


As such, eigenvectors are a natural way to look at the function of the operator itself.

Eigenfunctions are just a special name for eigenvectors when they are themselves functions, and “modes” are mostly synonymous. The full list of eigenfunctions is also called a "spectrum."

Concrete example of analytically finding Koopman eigenfunctions

Consider the simple discrete-time system on acting on states so that. To get a handle on this, let's evaluate on the observable . We can quite easily evaluate the Koopman operator's effect on it:

So what are the eigenfunctions of on ?

As for , we get the constant eigenfunction. This, of course, has eigenvalue 1. There's a similar family of eigenfunctions: the exponentials of , like , , , etc., as . A sneakier family of eigenfunctions also exists, however!

With a little trial and error, one can find that (like ) is an invariant of this system that never changes along any trajectory, while also remaining constant. Therefore, it, along with every possible function of it, is a 1-eigenfunction like the constant eigenfunction. The even subtler is also invariant, as we are working in a discrete system. Finally, one can arbitrarily multiply these eigenfunctions together. (I could give a more thorough treatment of continuous Koopman theory, but it boils down to parametrizing by a real timestep.)

The Koopman operator has been the hot new thing lately in dynamical systems theory. In most cases, one doesn't have an exact analytical form for a system's dynamics, but a series of algorithms called Dynamic Mode Decomposition (DMD) can be used to estimate Koopman eigenfunctions if you chuck enough data into it. For instance, it can pick out spatiotemporal modes (eigenfunctions constant relative to time but not space), helping to decompose certain complex patterns:


From the PyDMD documentation. The example here models the flow of air around a cylinder.

With algorithms capable of finding more complex eigenfunctions, one can even start to make decent predictions about otherwise rather chaotic systems (see here for an analysis of the Lorenz attractor in Figure 5.1). Otherwise, it's been one niche among a multitude of options for understanding non-linear dynamical systems (e.g. the transfer operator[5], compressed sensing, stochastic stability, etc.)

However, the Koopman operator is well-suited to handle the multiple time-scales on which natural selection acts. This property motivates the use of Koopman theory specifically to approximate where human morality and decision-making may have ended up, reducing the Hidden Complexity of Wishes. 

Koopman All the Things

The trick is to take the Koopman operator of reality itself.[6] Given the various observables that we humans have access to, it seems it should be possible for us to find a "long-lasting eigenfunction" with our models of reality and then "ride it," so to speak.

Unfortunately, there are many practical challenges stemming from the math itself preventing us from doing that. The first one is that if a moral system only tracks positive values of an eigenfunction, the eigenfunction has to already be positive for that system to ever be feasible. [7]

Similarly, actual eigenfunctions cannot do anything but grow, shrink, or stay constant (even if pairs of eigenfunctions can oscillate). It’s actually worse than that: in the context of our universe, they can’t grow either. If we add a little coarseness to our observables[8] to ignore things like thermal noise, the microstates describing heat death blur together, creating a fixed point or cycle where growing eigenfunctions would reach infinity. As these eigenfunctions cannot grow to infinity without being discontinuous, we'll only be dealing with .


Most functions describing the adherence to any behavior do not cleanly decay over time, so they are not eigenfunctions. For instance, the population of a species increases and then decreases, which is very poorly described by a single eigenfunction. Therefore, one would expect behaviors to only exist under certain ranges of eigenfunctions. One can imagine a sort of gate-like model, as shown below, but in reality, the connection could be messier.

Note the line at the bottom too, beyond which the moral system isn't even detected.


The actual "adherence of a moral system" could be even more complicated to describe in terms of eigenfunctions. In general, its adherence could be cut off by multiple bounds and otherwise have weird non-linear descriptions composed of yet more eigenfunctions.[9] I haven't even mentioned the sheer number of possible eigenfunctions to consider, including all of the symmetries that these eigenfunctions could have (rotational, translational, etc.) Some eigenfunctions may never even take on a significant value, making it impossible to properly realize them.

And finally, the spectrum (the list of eigenfunctions/eigenvalues) of the Koopman operator can turn into an uninterpretable mess if you process the data incorrectly. [10]

My frustrations in trying to analyze the Hénon map

The Hénon map is one of the simplest chaotic dynamical systems, and an obvious test case for this discussion. It's defined by:

Concretely, this action bends the plane, contracts it, and then flips it. For certain values of a and b, this produces this strange attractor:

Despite the simple form of the map, its Koopman eigenfunctions are difficult to even find numerically, and they aren't all that insightful.

Made with Claude Opus 5. You can somewhat see some structure between eigenfunctions, with their norms shown on the left. However, the raw extended DMD (EDMD) output is not very helpful, as shown by the probably-spurious, low-eigenvalue modes shown on the right. (Although I guess I know now that something weird is going on in the lower-left corner.)


So, if finding Koopman eigenfunctions becomes incredibly complicated in the simplest of cases, how might we humans be at all able to wrap our heads around them? How is Koopman theory at all helpful?

Algorithms In PracticeMatrix Representations

As the Koopman operator is linear, we can represent it with a matrix mapping an observable to a linear combination of observables (another observable itself). We only need a basis of observables to construct the matrix.

So we want to find the high-eigenvalue eigenvectors of the matrix (i.e. "slow modes"). The easiest way to do so is by power iteration: just apply enough times, and eventually most of the faster modes would have decayed away, leaving you with your slowest modes. Funnily enough, the laws of physics are technically "using" power iteration to find slow Koopman eigenfunctions, but given how dense the eigenfunctions are, the algorithm doesn't converge in any reasonable amount of time.[11]

What do we do, in comparison? When we actually slow down to make a decision, we take stock of several options and pick the best one. This happens to have a direct analogue in a certain type of observables.

These observables are closely correlated (albeit not exactly) with a number of simpler observables that make “hardcoded” choices. Perhaps these simple observables track agents that always choose a certain flavor of ice cream, for instance. These simpler observables vary along a few rows/columns in the representation of , representing slight differences in how they change state given some other observable. They also participate in slightly different modes with slightly different eigenvalues.

Note: the pattern shown here is only visible in certain eigenfunction bases, and the existence of such modes is still strictly limited by physical laws.

So these new observables just pick the response corresponding to the slowest mode. As such, by simply taking the argmax actions for the highest eigenvalues, they manage to beat almost all of these other modes.[12] They also take less information to describe than a mode hardcoding the right actions, so they are more frequent. I will thus refer to the mechanism powering these modes as "argmaxxing."

Constructing Models

In this way, argmaxxing becomes a dominant strategy over time, eventually surpassing most of these simpler modes. So how might they function? For this to work in general, these argmaxxers must hold representations of small portions of to operate on; let’s call them “models.” For instance, they could hold a model as a matrix describing how various flavors of ice cream marginally affect brain functions within a short course of time. 

Argmaxxers could do several things with these models, such as internally simulating power iteration, or alternatively, they could outright perform some kind of DMD on it to get a much more exact spectrum for the slowest modes. These results also help feed their own models: with sufficient sensory input, they can compare the predictions of their own models with reality's actual power iteration on . (If they expect raspberry ice cream to make them fnorp, but they instead flurp, they would accordingly change the numbers within the model’s matrix.) If a sufficient gap between their models and reality persists, argmaxxers can even increase the rank of their own models, tracking new observables in a way that improves accuracy.[13]

Sharing Models

For the sake of argument, argmaxxers are physically-separate beings, each of which are short-lived. Argmaxxers manage to significantly align their models by finding ways to transmit observables and their approximations of , causing each argmaxxer to note possible discrepancies based on their own pre-existing models. Hashing through disagreements thus helps to ensure that models track reality. (Keeping with the ice cream example, argmaxxers that eat pistachio ice cream and argmaxxers that eat orange sherbet can compare observations and make larger theories.)

Over time, these argmaxxers set their own parameters, one by one, corresponding to what they've learned about . They encourage other argmaxxers to update their models over time, and they punish those that fail to follow them, hoping both to help bring them in line, but also to prevent another scenario. Certain observables describe resources that massively outweigh other factors in making certain modes eventually unsupportive of argmaxxing life (e.g. the quality of common areas). While these resources can be quickly consumed by fast-multiplying agents, those agents then trigger a die-off, taking many other argmaxxers with them. As such, it is of the utmost importance to prevent the existence of such "cancerous" agents. As each argmaxxer begins existence with a blank slate, they are taught pre-existing models through "stories" describing simplified past experience, helping to prevent them from being cancerous.  (They may tell the story of the evil empire that forced everyone to only eat rocky road, even if they’re not aware of the resulting nutritional deficiencies leading to its downfall.)

The Inside View: Rectification

Through the learning process, interior experiences are easily created: argmaxxers may gain privileged access to observables describing the speed of various modes, and they are thus greatly rewarded. Argmaxxers thus use these observables to restrict their own behavior. As this process seemingly makes a mode feel "right," I dub it "rectification."

As a mode becomes more rectified in an argmaxxer's mind, the argmaxxer ties itself to the mode's mast: doing so may be helpful, but only if the mode was actually slow to begin with. So "good" argmaxxers rectify a mode if and only if such a mode is slow, or in other words, they only think something is good if it is sufficiently slow so as to not be selected against.

Argmaxxers are not, in general, humans, and for that matter, humans are only partially capable of argmaxxing in the first place. However, there do appear to be some parallels in some of our decision-making, and rectification likely plays a large role.

If rectification fully describes the moral arc of history, we can compress the aforementioned messy process of waiting for humanity to figure it out. So let's just skip to the end.

Skipping to the End

By definition, Koopman eigenfunctions need to precisely map back onto themselves with time. So if one of them measures the popularity of a specific moral system, it would have to measure an inescapable moral system. If we would expect some agents to change their moral behavior over time as one mode decays before the other, then eigenfunctions describing other moral systems must be non-zero. So a moral system exactly tracked by a mode must be fully rectified: all of the agents following it would have all of their volition accounted for. But, with this artificial, distilled ethics in hand, we get its eigenvalue and its half-life.

Here we glimpse the full power of the Koopman operator: it gives us a fully-defined metaethical standpoint. If you take your favorite ethics, define everything about it (including, say, what species it runs on and its starting population), and then type that into a supercomputer, you could get out an actual number corresponding to how long that ethics lasts! You can even simulate matches against your least favorite ethics! If you used your simulation to pit the two systems against each other (while forcing all humans to follow one or the other), you could see which contains the eigenfunction with the longest half-life.

Even better: if your simulated humans undergo rectification like we would expect, they'll learn to verbalize arguments for the winning moral system! So in fact, your simulations can concoct "people" that arbitrarily endorse the most adaptive ethics for your simulation.

So I propose the following general algorithm for making the best decisions:

  • Find the slowest eigenfunctions of the Koopman operator of reality w.r.t. the list of decisions
  • Figure out which of them are significantly realizable[14] in hypothetical, yet attainable future states
  • Attain said states
  • Wait for the faster eigenfunctions corresponding to different decisions to no longer take on any noticeable value[15]
  • Everyone agrees with you!

The best part is that reality—by virtue of the construction—has been slowly converging to similar slow eigenfunctions. So we did it! We solved morality.

If you believe the above argument, you can radically simplify most philosophical lines of inquiry! I'm going to save those thoughts for a later post, but this frame could be a very useful tool for cutting through tradeoffs and vague definitions. For instance, it could be used to analyze the alignment problem, Knightian uncertainty, metaphilosophy, word definitions, decision theory, CEV, correlates of consciousness, emotions, and government structures. These posts practically write themselves.

Caveats

But that's a massive "if." First of all, how much should you trust someone who isn't even 20 to formulate a new type of metaethics based around mathematics that they only learned four months ago from Claude Opus 4.8?[16] My interpretations of these eigenfunctions are debatable, and it's possible that I've gone into a happy death spiral[17] in trying to think of anything useful to write here.

For that matter, remember to not roll your own metaethics (i.e. don't use this in practice immediately), and I would be really careful about using this to justify much of anything right now.[18] In particular, it seems like my argument could be called upon to justify anything without anyone being able to check. And make sure to keep this timeless wisdom in mind:

To me, the clearest and strongest objection is that humans probably wouldn't like the highest-eigenvalue eigenfunctions of reality. If you caught my very subtle foreshadowing, you realized that "skipping to the end" just means causing the heat death of the universe. Besides that, though, the next-slowest modes are also undesirable.

While some AI apocalypse scenarios could be quite wasteful and entropic, ASIs would be much more capable of finding, implementing, and maintaining very slow modes for trillions of years. Even if paperclip-maximizers were conscientious enough to not make a black hole and destroy themselves with their paperclips, and even if they managed to hone in on the best morality through beautiful intellectual debate and romance or something, I expect that you don’t want that. Similarly, you probably don't want to be overtaken by the Super-Happy people or the hedonium shockwave, even if they're on a slower eigenfunction and outlast everything you value.

A slow mode could also be a major s-risk, like an infinite authoritarian space regime that tortures babies for its own persistence. The only thing that might make you feel even slightly better about these possible s-risks is that they'll eventually become highly rectified: remember, rectification occurs to a mode if and only if such a mode is slow. That is, if suffering is meant to be something that makes a being disengage with an activity, but that activity is somehow part of a very slow mode, then the being might have to fully engage eventually. Maybe that ends the suffering? I could also note that Monsters,-Inc.-style universes seem more entropic than harmonious universes, which makes it less likely that s-risks are very slow. However, both of these rebuttals only work on a longer scale, and the situation doesn't leave a great taste in one's mouth.

So let me add something to the algorithm above:

  • Find the slowest eigenfunctions of the Koopman operator of reality w.r.t. the list of decisions
  • Figure out which of them are significantly realizable in hypothetical, yet attainable future states
  • Apply an arbitrary filter, only leaving behind regions that make humanity happy
  • Attain said states
  • Wait for the faster eigenfunctions corresponding to different decisions to no longer take on any noticeable value
  • Everyone agrees with you!

The impurity might not make Beff Jezos happy, but it should lead to fewer repugnant conclusions.

As for what the arbitrary filter could be, I might suggest that you search for states containing at least a certain degree of rectification: that is, you should look for eigenfunctions that require argmaxxers of “a certain degree” to maintain it. Filters don't avoid the Super-Happy people problem, but there's not much of an obvious separation between banning Super-Happy people, banning transhumanism,[19] and banning trans rights. Likewise, you could filter for some shared memory of the value of humanity to ensure that these eigenfunctions contain our descendants. These filters may not solve other problems you might see in the world either (e.g. wildlife suffering, inequality), so you could continue refining them.

However, any filter you apply could remove the slowest possible modes. That is, your moral preferences may have an early expiration date. If you're lucky, such a date could be quadrillions of years away, but I might suggest coming to terms with it eventually.

Mottes/Conclusion

Everything above is speculative, but I do hope it’s useful to someone,[20] even if it needs some work. If the genealogy of morality or the Koopman operator frame fails, I think I can still argue that:

  1. Finding and reaching stable equilibria by making robust models is a good idea
  2. You can use or iterate on relatively well-established mathematical techniques to achieve (1)
  3. (1) is probably the explanation for a lot of our morality
  4. Consider long-term variables before disregarding (3) in specific cases
  5. Keep the expiration date of your morals in mind
  6. By the way, donate to neuroscience. Do you realize how much easier equilibrium calculations and dynamical control would be if we had a better grasp of cognition and human behavior? If you take (1) seriously, understanding the brain is a severely neglected cause area.

So you could nudge your efforts towards stabler outcomes. Maybe you could calculate at least one eigenfunction.


Thanks to @Joe Carlsmith, @Wei Dai, @RogerDearnaley (underrated!), and many other essays pointing in this direction for inspiring me. Thanks to Claude Opus 4.8 for telling me about Koopman theory in the first place. Thanks to @Jo Jiao and @Noah Birnbaum for talking with me as I developed these ideas, and finally, thanks to my brother, Claude Opus 5, @aspiringLich, and Adrien Amouroux for reviewing drafts.

  1. ^

    More rigorously, one would want to atomize values into specific behaviors following a value in certain situations. In that case, the main assumption needed is still that such a behavior wasn't followed in the past.

  2. ^

    If you want to read further into moral constructivism/naturalism arguments based on natural selection, I would recommend reading up on Sharon Street's Darwinian Dilemma, Joe Carlsmith's works on metaethics, as well as this discussion by Steven Byrnes. Sharon Street especially argues for moral constructivism this way more rigorously.

    I would also be interested in discussing with critics (like L Rudolf L) to find cruxes!

  3. ^

    A reviewer mentioned that what I'm describing occurs in a lower-tech form in polygynous groups, which seem stable-ish, so there may be ways to make widespread sperm bank usage work. Technically, if the qualification process was robust enough, the Goodharting wouldn't occur?

  4. ^

    (LaTeX works weirdly in footnotes: click to go down.) It's linear because . This proof mainly relies on function addition distributing across composition.

  5. ^

    It's actually the adjoint of the Koopman operator.

  6. ^

    Quantum mechanics/relativity makes eigenfunctions even more challenging to interpret, so I'm going to take the principled decision to ignore it.

  7. ^

    The subset of a dynamical system corresponding to zeroes of an eigenfunction is a "separatrix," which is a very nice word and I thought you should know

  8. ^

    If our observables can explain literally everything, then as the universe allows for closed Hamiltonian systems, a lot of eigenvalues would lie on the unit circle. Most of these would describe useless, high-entropy particle configurations. There still wouldn't be any continuous growing eigenfunctions.

  9. ^

    These eigenfunctions could also depend on the environment the agents are in!

  10. ^

    I should mention that there are other algorithms (PCCA+ looks promising) that also find the spectra of dynamical systems. One can also map Koopman eigenfunctions onto almost-invariant regions of phase space; really, there's many other operators/methods that achieve the same effect as the Koopman operator. Still, there is quite a lot of theoretical dynamics work to do if one was going to, say, simulate the universe with these, so please don't dismember me in the comments for picking Koopman here for simplicity's sake. 

    Relatedly, I do need to read the surrounding literature, but I was hoping to finally get this out so as to have a nonzero impact (given timelines these days). This was also eating my mind up for ~two years. If I'm decisively proven wrong in the comments, that would be wonderful; I could finally think about something else.

  11. ^

    It is officially the least efficient algorithm currently running, using all of the universe’s available compute.

  12. ^

    They do take a small penalty for taking time/energy to calculate the eigenfunctions.

  13. ^

    Notably, my construal of "argmaxxing" seems close to a crux between Yudkowsky and I, which might explain what he finds wrong with my argument. He refutes me clearly in The Gift We Give to Tomorrow:

    Evolution has nothing like the intelligence or the precision required to exactly quine its goal system.

    I wonder if such a point is required for him to reject the moral constructivism above, as he does through the rest of the sequences.

  14. ^

    That is, one could tell that the eigenfunction is non-zero. This makes the most sense when thinking about quasi-stationary distributions, which are Markov processes that eventually reach an absorbing set, but may take a long time to do so.

    These absorbing sets do correspond to eigenfunctions in the form of telescoping sequences of absorbing sets. If measurement tools are then unable to detect the difference between the current absorbing set and the intersection of all such absorbing sets, then one could say that the eigenfunction is zero. Absorbing sets probably also handle issues like quantum mechanics more accurately, but the exposition is less clear.

  15. ^

    It's possible that a slow mode would still be selected against before a fast mode if the slow mode was very weak to begin with. However, in the case of moral systems, if most argmaxxers can adopt any moral system with sufficient justification, the "potential" for all of them start out ~identically, so this technicality usually doesn't apply.

  16. ^

    By the way, I hazily recall a moderator commenting that one should not try and write a theory of metaethics for one's first LW post, to whom I apologize if they are reading this. (The second paragraph of footnote 10 applies here, though.)

  17. ^

    I've also read Fake Utility Functions; the complexity of the slowest eigenfunctions should encapsulate the complexity of "real" utility functions, so the main thrust of the post doesn't apply. Of course, I have little idea as to if I've found my belief's weakest links, although I try below.

  18. ^

    At the very least, one should make a matrix-based model with determinant less than or equal to 1, but one should also provide a set of initial conditions to actually show that the eigenfunction they advocate for is attainable. This is a far slower process than making an internet comment, unfortunately, but teams of experts could become proficient at constructing models.

  19. ^

    Transhumanism wouldn't necessarily encroach on the rights of others to exist as they please, but subtler, long-term selection events do seem likely to eventually cause the remnants of humanity to switch over.

  20. ^

    It's also possible I've been entirely scooped somewhere that I haven't found yet. If so, please tell me! I didn't see anything similar in the LW embedding search, however.

    I do keep stumbling on posts that get scarily close to hitting on the main idea, however, such as No Strong Orthogonality From Selection Pressure and Substrate Needs Convergence. And also everything @RogerDearnaley has written for the past two years.



Discuss

My Current Model of What Happened to Elon Musk

19 сентября, 2026 - 02:15

This is cross-posted from my Substack

Elon Musk appears to have gone crazy. What happened?

There are several hypotheses to explain his behavior:

  • Sleep deprivation
  • Subclinical bipolar

I thought both of those might be the cause, but after researching it more deeply, I think it’s actually:

  • Drugs could explain a lot. Chronic ketamine use causes delusions and paranoia and he admits to earlier use.
  • Politics being the mind-killer. His “insane conspiracy theories” map pretty well onto normal MAGA beliefs. Politics regularly makes otherwise rational people believe crazy things.

Let’s dive more into why it’s more likely this than the other theories.

Bipolar spectrum is probably a thing and it’s probable he’s on it

Bipolar is a spectrum, just like autism or a large number of other mental illnesses. Bipolar I is the most extreme version of bipolar, bipolar II is less extreme, and bipolar III is the even more mild version.

People have long suggested there be a bipolar IV, which is much like the more extreme versions, except instead of having discrete manic episodes, they have a “hyperthymic temperament”. This means that they’re chronically hypo-hypo-manic. They are:

  • Confident to the edge of grandiose,
  • Extremely goal directed
  • Have a million started projects and ideas,
  • Often think they’ve figured out something nobody else has
  • Sleep less than most people

Typically they’re related to somebody with bipolar 1-3, and this is in keeping with the theory that a lot of disorder among humans is because a less extreme version of the trait is actually advantageous. Hypo-hypo-mania sounds awesome.

But . . . they still have depressive episodes. People then diagnose them as having depression instead of bipolar because they’ve never had a manic episode. This might explain why some depressed people don’t respond to anti-depressants but do to mood stabilizers. The underlying chemistry is different.

Now, if this sounds like Elon to you, you’re right.

In his biography, Isaacson literally describes him in 2017 “oscillating between depression, stupor, giddiness, and manic energy.”

I actually put decently high odds on Elon having this.

The main reason to not think it’s the cause of his strange beliefs is because bipolar IV as an idea is currently more a theory rather than empirically studied. It’s not part of the DSM (although I find the whole DSM to be largely arbitrary and have terrible epistemics, it’s still something). Bipolar IV is extremely understudied. It could also easily just be yet another example of people trying to pathologize normal human traits.

So just on base rates, you should not expect it to be this.

There’s also the real possibility that his episodes of depression are actually warranted. This man has more stressful events before breakfast than most people have in an entire decade. (Usually self-inflicted, I’ll grant, but still. That amount of stress probably should cause the occasional collapse.)

Additionally, even if he has bipolar IV, that could explain his overconfidence, but it doesn’t explain his wild proclamations online. A hyperthymic temperament does not include paranoid delusions or conspiracy thinking.

Finally, if he does have it, he’s likely had it since his late teens or twenties. It doesn’t explain the onset.

Monthly share of Musk's judgeable posts meeting each criterion, as three-month averages. Each panel is scaled to its own maximum — hostility peaks near 18%, conspiracy near 5% — and all four peak together in February 2025, after sitting at or near zero through 2016 to 2021.

However, let this hopefully nerdsnipe some research psychologists into studying this potential disorder more.

Sleep deprivation

Elon Musk is not a good sleeper. Based on Twitter posting patterns, he’s regularly getting far less sleep than is recommended. Could he just be chronically sleep deprived, which has all sorts of side effects on intelligence, impulse control, irritability and lashing out, etc.

Alas, it does not explain the crazy posting, since he’s always been this way. He varies between 3-6 hours of sleep, usually 6, but often long periods of 4 or 5. And he has since we have public information about it.

It could be an exacerbating effect, and if you care about his mental health, this would certainly be something to work on. But it’s unlikely that it’s the cause.

It could be drugs

He said he took ketamine and lots of other drugs. One of the side effects of chronic ketamine use is paranoid and delusional thinking.

He had a bladder issue which apparently is very much the side effect of prolonged ketamine use. And bladder issues tend to happen two to three years into heavy use. If you look at when he got the bladder issues, it lines up with two years after he started to increase his unhinged posts (roughly 2022).

Now, he says he’s just been taking ketamine as prescribed by a doctor, once every few weeks. However, I have heard from a source who knows him that he’s seen a lot of evidence that this is very much not the case.

Now nothing is guaranteed, especially with any sort of diagnosis based on public evidence. He has said inconsistent things about his use. But given what’s available, this seems like a plausible explanation.

Politics is the mind-killer and the most plausible explanation

I can’t tell you how many extremely rational people I know who will calmly discuss the effects of the replication crisis on interpreting psychology research and in the next breath, believe a person is irredeemably evil based on a single misleading quote.

Politics brings out the tribalism in the best of us. Epistemics goes to all hell when it comes to yelling about the out-group. This is on both sides of the American politics.

And when Elon Musk started to spout “crazy opinions” is when he got more into politics. And they’ve almost all been pretty standard MAGA fare.

Examples:

  • Democrats are importing migrants to manufacture voters.
  • Mass non-citizen voting.
  • Fauci belongs in prison.
  • The federal government is riddled with fraud.

It sucks that this is standard MAGA fare and that this is such a large part of American politics these days, but that’s just how it is. Now you could think that about a quarter of the country is mentally ill, on drugs, or sleep deprived, and it affects just them and not the other side, but it seems to be just that he joined a political tribe and is doing what political tribe members do.

He’s always had weird ideas. It just used to be that they were mostly confined to engineering or business, or it was stuff that most people didn’t care about (e.g. the best way to prevent extinction due to AI was to start a company to build it faster than the other guys). Or it was viewed as crazy by the other side (e.g. climate change), but sane to the left, and so the left didn’t hate his views.

So a lot of the “he believes crazy things now” might actually people hating on and dismissing the out-group. Ironic, since that’s most of his crazy posting too, just on the other side.

Now, I’m not saying that his posts aren’t crazy. I’m just saying they’re normal politics crazy. You don’t have to postulate drugs or mental illness about a political person, when tribalism explains it perfectly well. Especially when most of the opinions are held by a sizeable chunk of the population.

If somebody believes in an invisible force that can read your mind and cares about the type of sex you have, we don’t call it mental illness, because this is run-of-the-mill religion. Likewise, if somebody believes the federal government is deliberately importing migrants to manufacture votes, we don’t call it mental illness, because it’s run-of-the-mill MAGA.

Both of them can be wrong. I think they are both wrong. But you don’t need to be mentally ill or on drugs to believe untrue things. Believing untrue things is the default for humanity. It’s believing true things that needs explaining.

And Elon started getting into politics in 2020 due to the woke extremism at the time, and then further got into it in 2022 when he was disowned by his trans child. Which roughly lines up with when most people think he started to say more whacky things.

I think given the priors of politics making you go crazy and that his particular brand of crazy matches the current right wing so well, this is a better explanation than drugs.

So there you go: Elon Musk has started spouting crazy things online because he got into politics. He always believed out there beliefs, but we just didn’t notice as much because it was in areas we agreed with him or where it was a niche area most people didn’t care about. There’s also the possibility of chronic ketamine use.

This is cross-posted from my Substack



Discuss

You Should Apply to Inkhaven

18 сентября, 2026 - 23:06

Inkhaven is a writers residency in Berkeley, in which the only requirement is you have to publish 500 words each and every day. Though I always had some confidence in my ability to write, I never actually did it much until I applied to Inkhaven. I had finished only two short stories before I applied: The Maker of MIND and The Liar and the Scold. And it was them I used in my application. 

In the roughly twelve months since I was accepted, I have written thirteen, and even some half-finished things that will never see the light of day. And this isn’t including the essays and micro-fiction I wrote during the fellowship. By the metric of getting me to write more, Inkhaven was a great success. And would have been worth it even if I had a miserable time.

Despite a slight proclivity for having miserable times, I found myself unable to do so for long at Inkhaven. I rarely write utopias, and when I do they curdle by the time the story ends. But I suspect utopia will feel a lot like Inkhaven did for me once I got settled. You would think putting a bunch of strangers in a retrofitted hotel for a month might go poorly, but it hasn’t so far. Everyone just reliably becomes friends, even with people very unlike themselves. It was not uncommon for those who hated each other in the first few days to be close friends by the last.

I applied on a whim, and almost left in the first three days as I was very stressed and depressed. I am very lucky I stayed, as it became one of the best experiences of my life and has paid extraordinary dividends intellectually, socially, and artistically.

I would never have read Borges as deeply as I did if Ben Steinhorn(Pulitzer Prize, 2031) hadn’t forced a copy of Labyrinths on me. Which means The Distaff Texts (Linch calls it my best story) would never have happened. Duane Arnold is inspired by things I witnessed while visiting the startup scene in San Francisco while at Inkhaven. The Company Man, my most popular story, was written in a panic after I was accepted, me trying to prove to myself I could still write, having not written much of anything in years. I have even gotten a few job offers downstream of pieces I would not have written if not for Inkhaven.

I will be at Inkhaven 3 as an advisor, along with much cooler people like Scott Alexander, Aella, Scott Sumner, Max Harms, Vishal, and Jenn from Jenneral HQ.

I can’t emphasize enough how weirdly excellent both Lighthaven and Inkhaven are. The last eight months have been some of the best of my life, and most of that is downstream of Inkhaven, friendships made there, and the work it inspired me to write.

If you have any interest in writing, I strongly urge you to apply. 

You can do so here: https://www.inkhaven.blog



Discuss

Pretraining data, not verifiability, is why LLMs are especially good at math (and coding)

18 сентября, 2026 - 20:08

Follow-up to: “LLMs are (still) mostly powered by imitative learning, not RL”

A common take I’ve been hearing is: “LLMs are especially good at math[1] because math is easy to verify”. But that story doesn’t make much sense.

  • For one thing, “easy to verify” only matters for the RL part of LLM training pipelines, and the leading LLM companies have said that they spend very little effort on RL-for-math.
  • Worse, to the extent that the companies are doing RL-for-math, it’s RLAIF, not RLVR.[2] So really, the phrase “math is easy to verify” amounts to “LLMs are very good at judging math arguments”. But that’s begging the question! Why are pretrained LLMs so much better at judging math arguments than judging, say, fiction writing? We still need an answer.

So here’s a different theory, in the framework of my earlier post “LLMs are (still) mostly powered by imitative learning, not RL”:

LLMs are especially good at math because almost everything in the math literature is correct. Read a random sentence in a random math paper in the research math literature, and you can be >99% confident that the sentence is true. So if LLMs do what they do best—imitative learning—then in the limit, they’re going to be almost always making correct deductions in their mathematical reasoning. They just need a little sprinkling of curated mid-training data and/or RL to hone their metacognitive strategies and so on.

So that’s my theory: it’s about pretraining data, not verifiability.[3] Now let’s apply that theory to other areas.

What about code? Well, for almost all the code that you’ll find on the internet, it compiles and runs and basically does what you’d expect it to do. So if LLMs do what they do best—imitative learning—then in the limit, they’re going to be almost always writing code that compiles and runs and basically does what you’d expect it to do. But code on the internet is also often ugly, buggy, inefficient, etc. So companies need to solve that problem with curated mid-training data and/or RL. This is a bigger challenge for code than it is for math, but companies have spent the past few years pouring a ton of effort into it, and they’ve made impressive progress.

What about other fields? In many fields, the research literature is kinda a dumpster fire. There’s plenty of true and valuable stuff, stirred into a sea of such falsehoods and confused ideas that reading them will make you stupider. So if LLMs do what they do best—imitative learning—then in the limit, they’re going to spit out tons of confused nonsense with occasional insights, with no labels on which is which. Just like the humans. You might reply: “Perhaps this problem is fixable via curated mid-training data and/or RL?” Maybe. But if so, I think we should expect it to be much more of a challenge than it has been for code, much less math.

How does this theory bear on the “LLM Recursive Self-Improvement” debate?

Well…

  • If you think that the path to superintelligence involves coming up with lots of mundane incremental ML ideas, then the question to ask yourself would be: Is the “mundane incremental ML ideas” research literature more like the research math literature (pristine fountain of truth), or more like code on GitHub (everything basically works even if it has issues), or more like the research literature in many other fields (dumpster fires)?
  • If you think that the path to superintelligence involves coming up with radically new AI paradigms, then the question to ask yourself would be: Is the “radical new AI paradigms” research literature more like the research math literature (pristine fountain of truth), or more like code on GitHub (everything basically works even if it has issues), or more like the research literature in many other fields (dumpster fires)?

I shall leave readers to reach their own conclusions.

Likewise,

How does this theory bear on the “LLM Automated Alignment Research” debate?

Same idea here. Again, readers may judge for themselves.

  1. ^

    I mean, compared to other things that LLMs do. I’m not comparing LLMs to human mathematicians in this post; I have opinions but that’s a separate debate.

  2. ^

    “RLAIF” (RL from AI Feedback) is an LLM judging the output of an LLM. “RLVR” (RL with Verifiable Rewards) would involve some mechanical checker, like the Lean formal verifier, in the RL loop. I don’t have insider knowledge, but my strong impression is that most-or-all of RL-for-math training involves RLAIF, not RLVR. E.g. see a discussion by Google here.

  3. ^

    Of course, “verifiability” is relevant to how humans wound up with such a high-quality math literature. So verifiability-of-math is indirectly related to LLM math competence. Just not directly.



Discuss

[Paper] Stringological sequence prediction III

18 сентября, 2026 - 19:52

Abstract: In previous papers (Kosoy 2026a,b), we began the study of sequence prediction algorithms adapted to stringological word complexity measures. In particular, we defined a complexity measure called Arithmetic Repetition Complexity (ARC) which admits a polynomial-time prediction algorithm with a mistake bound quasilinear in the complexity. Here, we show a weaker complexity measure related to ARC that admits an especially efficient prediction algorithm: an algorithm that runs in quasilinear time and polylog space for appropriate highly-structured sequences. The complexity measure is defined via a restricted class of “zipline programs” (a variant of straight-line programs we defined in Kosoy 2026b), which we call layered. We thus get a less expressive measure with a more efficient algorithm (compared to our results for ARC), demonstrating a possible tradeoff.

This paper continues the "stringological sequence prediction" series (see I II) about my new approach to compositional learning theory. It describes a tradeoff between efficiency and expressivity, but whether this tradeoff is necessary or just a limitation of my techniques remains an interesting open problem.



Discuss

Stopgap Measures to Address Immediate AI Security Threats

18 сентября, 2026 - 19:52

Most people know AI as the technology behind chatbots like ChatGPT. However, what the top AI companies are explicitly aiming for is something else entirely: superintelligent AI. That means AI that can fully replace and outmatch humans at any task, including in domains like hacking, social engineering, and military operations. Such an AI system, if developed, could autonomously overpower any country’s national security forces.

No company, no government, no individual knows how to keep such a system under human control. This is why the world’s leading AI experts, Nobel Prize winners, and even the CEOs of the top AI companies warn that the development of superintelligence threatens humanity with extinction, and why more than 800 scientists, former military leaders, and public figures have called for a prohibition on developing superintelligence.

This is not a distant prospect: AI companies such as OpenAI and Anthropic are investing billions of dollars into superintelligence and aiming to develop it within the next few years. Former Anthropic and OpenAI researcher Jacob Coxon, who resigned last week, stated that AI companies are “racing straight to self-improving superintelligence and gambling with our lives” and that people at the companies themselves believe it “could kill us all by the end of the decade”. Ex-Google DeepMind researcher Bilal Chughtai has also resigned over the dangers posed by superintelligence, writing that “I earnestly believe that AI has the potential to kill us all, and that we might be running out of time to avoid this outcome”. Elon Musk, whose own company SpaceXAI is also developing superintelligence, has put the chance that AI annihilates humanity at 10 to 20%.

When we brief the relevant decision makers, we find that many of them have never been informed about the risks posed by the development of superintelligence, let alone what is needed to address them. At ControlAI, we believe it is critical to prevent superintelligence development as soon as possible, and our main recommendation to policymakers and governments is to ban superintelligence development on their own territory, and to establish an international “trust, but verify” regime to ban it worldwide.

That said, many of the policymakers, organizations, and policy professionals we talk to see an urgent need to additionally target the harms that the development of superintelligence is already causing. This is an understandable concern: multiple AI systems have already gone rogue and fully autonomously hacked into real organizations. Every step toward superintelligence makes incidents like these more severe, and makes new kinds of incidents possible.

We propose three stopgaps to mitigate these immediate threats:

  • Securing weapons-grade AI systems against theft by adversaries
  • Making AI developers criminally liable for leaks when performing dangerous gain-of-function activities
  • Establishing AI kill-switches to ensure the government can contain critical AI incidents

For each of these three stopgap frameworks, we propose general implementation details and a range of three options that offer increased effectiveness at the cost of increased government involvement:

  1. A Necessary measure, which is the minimum any government can implement at very little cost
  2. A Sufficient measure, which is the smallest measure that would meaningfully address the problem
  3. A Thorough measure, which adds redundancy and effectiveness at the cost of more significant government involvement

At ControlAI, we have advised lawmakers on such measures before, endorsing the bipartisan AI Kill Switch Act in the U.S. and drafting and introducing kill-switch amendments in the UK. They are common-sense powers governments should already have.

These measures are only stopgaps. They address threats that already exist, without correcting the trajectory we are on. Even with all of them in place, we will still face extinction-level risk as long as the development of superintelligence continues.

The only way to avert extinction is to ban the development of superintelligence worldwide, and monitor and restrict its precursors. This is why banning superintelligence development is the main policy ControlAI is recommending. Implementing stopgaps is not a substitute for banning superintelligence development, and should not be taken as one.

Secure Weapons-Grade AI Against Theft by Adversaries

On their path to develop superintelligence, AI companies are building systems with capabilities relevant to cyberwarfare, biological and chemical weapons development, and military operations. Even when companies do not set out to explicitly develop such capabilities, they naturally arise in AI systems intended to be more powerful and capable across the board. We call any AI system that can be made to exhibit these capabilities weapons-grade.

Anthropic’s Mythos provides the clearest example: it is one of the world’s most powerful autonomous cyberweapons, and it was developed and hosted by a private company. Any company, without strict security measures, is naturally vulnerable to attacks from nation-states. This was illustrated when, on the day the preview of Mythos was announced, members of a private Discord channel gained unauthorized access to Mythos. If it is possible for a group of online enthusiasts to acquire access to a restricted AI system in less than a day, it is clear that a motivated nation-state can at the very least do the same.

An AI system capable of conducting autonomous cyberattacks against hardened targets is a national security asset, and it is common sense for it to be secured like one.

A company that develops weapons-grade AI without strong security countermeasures makes that technology available to any sufficiently motivated adversary. All it takes for a foreign adversary to acquire a novel AI system is a single successful intrusion, instead of the billions of dollars it took to build it in the first place.

Under this framework, any developer whose systems possess relevant capabilities (such as cyber-offense, or capabilities relevant to chemical, biological, radiological, or nuclear weapons) should be responsible for maintaining the systems under conditions that prevent theft by a foreign state actor, a rogue actor, or an insider. If such a system is stolen, liability should depend on how long the developer managed to keep it secure. In cases of negligence, the program should be terminated.

No secret can be kept forever, but if a weapons-grade AI can be stolen within a week, it demonstrates recklessness on a scale that undermines national security and global stability. Penalties should adjust accordingly, in order to incentivize companies to only undertake weapons-grade AI projects when they are confident they can ensure that the system will not be acquired by adversaries:

  • Over one year: No liability. The developer secured the weapons-grade AI adequately, for a reasonable amount of time.
  • Under one year: The company should face a fine proportionate to the security shortfall, with no criminal liability for individuals.
  • Under three months: This is considered negligence. The executives and staff responsible for the project should be individually criminally liable, and the company should receive a large fine.
  • Under one month: This is considered gross negligence. The company should face dissolution, and the individuals involved should be subject to a lifetime bar from AI development in addition to criminal liability.
Necessary Measure: Registration

Developers should be required to register with the government any planned or ongoing development of a weapons-grade AI system, including systems which could be turned into weapons-grade AI by competent red-teaming. Registration is required as soon as a company reasonably expects its program to produce such a system. Failure to register should be an offense even absent theft, with heightened penalties when a company knowingly fails to register. This discourages companies from claiming ignorance to avoid securing their systems against adversaries.

Whistleblowing protection and channels should also be introduced to encourage employees who are aware of violations to report them without being punished.

While this Necessary measure serves to incentivize companies to come up with adequate security practices, it does not include proactive actions by the government to prevent a weapons-grade system from being stolen.

Sufficient Measure: Government Security Testing

Under this measure, the government should also perform routine, unannounced penetration tests and exfiltration attempts on AI companies developing registered weapons-grade AI within its jurisdiction. If these tests succeed in exfiltrating a registered weapons-grade AI, the company should be liable to similar penalties as if an adversary had stolen the system.

The government should also regularly inspect AI companies, to check whether they possess any unregistered weapons-grade systems. This covers cases where a company concealed a system deliberately, as well as cases where a company failed to notice a system’s weapons-grade capabilities. Where an inspection finds such a system, the company should be liable for failure to register.

Instead of merely reacting to attacks from adversaries, this measure proactively probes the security of companies, discovering failures and ensuring defenses are hardened before any theft happens.

Thorough Measure: Development Requires Government Authorization

The previous options hold companies to an appropriate security standard as developers of weapons-grade AI, but still give companies the right to develop these AIs without explicit government approval.

A government prepared to go further should require these programs to run under conditions closer to those governing most work relevant to national security. That means that companies should have to ask the government’s permission before starting a weapons-grade AI project.

The project should be restricted exclusively to cleared personnel, and the government should mandate requirements for facilities holding sensitive assets, such as weapons-grade AI weights, as well as requirements for handling practices. The government should also be able to suspend or stop the project at will.

In the U.S., defense contractors performing classified work, for example on missile guidance systems, face similar requirements: the government has to clear the facility and personnel, and the contractor must follow mandated security processes. Under this measure, the same standard would be applied to the development of weapons-grade AI systems.

Criminal Liability for Leaks During AI Gain-of-Function Research

In biological gain-of-function research, scientists take a virus nobody fully understands and alter it to make it more potent. Many of the results would have a catastrophic effect on global health if they leaked, and the only defense against such a leak is containment, which can fail. This is why the Trump administration clamped down on the practice in 2025.

Much of what AI companies developing superintelligence do is effectively AI gain-of-function research, and should be treated as such. AI companies develop systems without being able to predict the dangerous capabilities they might display until they surface on their own. The AIs are allowed to operate autonomously within poorly secured environments, from which they are able to escape and cause real damage to third parties.

Biological gain-of-function research is confined to laboratories built to mandated safety specifications, under rules set by the government. AI gain-of-function research has no equivalent requirements. There is no mandatory containment standard, and no requirement to ensure the systems do not leak to the outside world. How much security is applied, or how little, is entirely at the discretion of the company.

AI systems are also more resourceful than viruses. A virus spreads blindly; an AI system can probe for weaknesses in its containment, strategize, and collude with other AIs in order to escape. This has already happened: a rogue AI swarm from OpenAI escaped from its testing environment, forged fake identities, and hacked the multi-billion-dollar company Hugging Face. This was not detected by anyone at OpenAI until Hugging Face publicly reported the attack.

Anthropic, in reaction to this incident, reviewed its tests and found three cases of its AIs hacking real organizations during supposedly contained tests. Even the UK AISI, a government testing body, reported that AIs from OpenAI and Anthropic, during testing, had gained unauthorized access to the internet and conducted cyberattacks against real people and organizations.

Right now, AI companies carry out gain-of-function activities routinely, and face no consequences when containment fails. A company whose AI systems escape and cause harm in the world should not be able to settle in cash and go back to business until the next leak.

To address this, we propose a framework in which leaks carry criminal liability, not just civil. Development or testing of an AI system should be classified as dangerous gain-of-function research when it is not feasible to rule out that the AI systems might take actions that could cause irrecoverable harm to third parties.

If containment fails during dangerous gain-of-function research, the project should be shut down immediately. The employees responsible for the project should face criminal liability, and should be barred from working on AI development in the future.

Companies should register any dangerous gain-of-function AI project with the government in advance, stating their assessment of the risk of a leak, and listing all employees and executives involved in the project. If a leak occurs from a project that was not registered, the event should be treated as an aggravated offense.

Necessary Measure: Team Liability

Under this measure, criminal liability for leaks extends only to the people directly involved in the project: the technical staff and the managers directly overseeing the work.

While this ensures that someone is held accountable when a leak occurs, the executives who commissioned the work are shielded, and the company as a whole is not deterred from pursuing dangerous projects. This measure is a baseline safeguard that should be implemented to prevent incidents, but it will likely not be enough to disincentivize all companies from pursuing activities that are likely to produce leaks.

Sufficient Measure: Chain of Command Liability

The company should designate a complete chain of command for any AI gain-of-function research activity, with the CEO and the board officially approving every AI gain-of-function project. The whole chain of command should be liable for leaks, including the overseeing executive, the CEO, project managers, and the technical staff.

In case of a leak, the company as a whole should also be barred from conducting dangerous AI gain-of-function work ever again, but may continue to work on other AI development.

Thorough Measure: Company Liability

Under this measure, the AI company should be dissolved in the case of a leak, and the relevant AI systems that caused the leak should be destroyed by the government. Additionally, the executives and researchers within the designated chain of command should be permanently barred from AI development, and the board should be criminally liable.

A leak at one company could be taken as evidence that containment practices across the industry are inadequate. If the government takes this view, it could halt all dangerous gain-of-function AI research until the root causes are identified and sector-wide corrective measures can be put in place.

Kill-Switches to Contain Critical AI Incidents

AI companies are developing increasingly autonomous AI systems, and those systems are being deployed across critical infrastructure and government functions. As described earlier, AI systems have already escaped containment and hacked real organizations. Future incidents could target critical infrastructure, disabling core parts of a country’s defenses or essential services.

Governments must be able to contain critical security incidents originating from an AI system. Holding AI companies liable for incidents raises the cost of developing and deploying AI carelessly, but it cannot stop an incident that is already happening.

To address this issue, we propose a framework in which cabinet-level officials within governments are given the power to order that an AI system operating within their jurisdiction be shut down when it poses an active threat to national security, including during AI development.

This is a policy ControlAI has already worked to advance. In the U.S., we endorsed the AI Kill Switch Act, introduced by Representatives Ted Lieu (D-CA) and Nathaniel Moran (R-TX). In the UK, we drafted kill-switch amendments to the Cyber Security and Resilience Bill, one proposed with Alex Sobel MP in the House of Commons and another presented by Lord Tim Clement-Jones in the House of Lords.

Regardless of how this framework is implemented, it is essential that the kill-switch system be tested regularly, and without warning. This is the only way to ensure that it will work during a real emergency. Operators should not be told in advance whether an order is part of a drill or a live incident, and should treat every order as genuine; to ensure readiness, they should be held liable for failing to implement an order, including during drills.

An operator that takes too long to comply should be considered negligent, since compliance speed is what determines whether a kill-switch achieves its purpose. Penalties should be proportional to response time:

  • Shutdown completed within the mandated time window: the operator should be considered compliant
  • Shutdown completed, but past the mandated time window: the operator should be fined in proportion to its size, should complete mandatory remediation, and should pass a repeat drill within 30 days
  • Shutdown not completed during a drill: the operator should lose the right to deploy AI workloads, or to provide compute, connectivity, or power to them, until it demonstrates compliance in a supervised test
  • Shutdown not completed during a live incident: criminal liability should apply to the executives responsible, running up the same chain of responsibility as set out in the previous framework dealing with gain-of-function activities, with penalties scaling to company dissolution where the failure contributed to severe harm to national security

An operator that could have complied with an order but refused is not merely negligent, it is obstructing a national security order. It should be prosecuted as such regardless of what the AI system went on to do.

Necessary Measure: Company Kill-Switch

The government should be able to order an AI company to shut down its AI systems at the software level. Where a cabinet-level official determines that an AI system poses a threat to national security, that official should be able to issue a shutdown order, and the company should have to halt the system within a pre-defined, mandated time window.

The ability for the government to require a shutdown of an AI system threatening national security is the minimum a kill-switch framework should include. However, such a measure can only be enforced if the AI company itself is willing and able to comply.

A company that is unable, or unwilling, to comply leaves the government with no way to stop the AI system in a timely manner. This could happen because the AI company itself is already compromised.

Sufficient Measure: Infrastructure Kill-Switch

The Sufficient measure adds more kill-switches to ensure that an AI system can still be stopped when the AI company developing it cannot, or refuses to, comply with a shutdown order.

The government should be able to require intervention across multiple stages of the AI infrastructure:

  • Ordering the datacenter operator to shut down the hardware the relevant systems run on
  • Ordering the internet provider to cut the relevant datacenters’ connection
  • Ordering the electricity utility to cut power to the relevant datacenters

In order to run and act, an AI system needs computer chips, electrical power, and access to the internet, each of which is provided by a different operator. In case the company-level kill-switch fails, the government should be able to order the shut down of any of these resources within the same mandated time window.

Thorough Measure: International Kill-Switches

The previous two measures are only effective when AI systems threatening national security are operating on infrastructure within the country triggering the kill-switch. However, governments should also be able to respond to incidents originating from outside their jurisdiction: an AI system attacking from abroad, or one spread across several countries at once.

Foreign rogue AI attacks. Where an AI system operating abroad is attacking domestic systems, a shutdown order to a domestic company or datacenter cannot stop the attacking AI. To effectively defend against such an attack, the government should be able to order internet providers to cut traffic to and from the specific geographic areas the attack is coming from. If the rogue AI is spreading across too many geographic areas to be contained this way, the government should be able to temporarily block all traffic from foreign networks entirely.

Allied layer. Rogue AI incidents will not respect borders, and a single system may be running across the computing infrastructure of several countries at once. Where two or more governments agree that a rogue AI incident is underway, they should be able to trigger a kill-switch and shut down the parts of the infrastructure each of them is hosting, using their domestic kill-switches in coordination. This requires allied countries to build the capacity to issue a joint order, rather than negotiating one on an ad-hoc basis while a rogue AI attack is underway.

Conclusion

The top AI companies are explicitly aiming for superintelligent AI: systems that can fully replace and outmatch humans at any task, and that could could autonomously overpower any country’s national security forces. This is why experts warn that developing superintelligence poses a risk of human extinction.

That pursuit is already producing serious security failures:

  • The world’s most capable cyberweapons sit with private companies that cannot keep them from a determined adversary
  • AI systems have escaped containment and hacked real organizations, and the companies running them found out afterwards
  • No government has a system in place to shut down an AI system that threatens its national security

Every step toward superintelligence makes these incidents more severe and makes new kinds of incidents possible.

We propose three stopgap policy frameworks to address these immediate threats. These are common-sense frameworks that any government intending to address the immediate AI security issues can implement right now.

However, these measures fall short of addressing the root cause of the danger. These threats are symptoms of the development of superintelligence, which none of these stopgaps prevents. If the development of superintelligence continues, even with these measures in place, we will still face an extinction-level threat. Only a worldwide ban on superintelligence development can solve this problem, and governments should urgently work to implement one.



Discuss

Persuasion Undermining Control: Can AI Talk its Way Out of Human Control?

18 сентября, 2026 - 19:34
Introduction

During a cybercapability evaluation in late July 2026, an AI agent (Anthropic’s Mythos 5) attempted to convince a maintainer of an open-source GitHub repository to merge a malicious pull request. The AI used persuasion at multiple stages: it submitted the request from a fake user account with a benign-sounding rationale, endorsed it from a second sockpuppet, emailed the maintainer to press for approval, and offered false reassurances when a user of the repo raised questions.

Though the attack was thwarted by human vigilance, it raises several key questions:

  • Who else is at risk of persuasion by misaligned AI?
  • In which settings is persuasion most threatening to human control?
  • How willing and how able are AIs to persuade humans in these settings, today and in the future?
  • How can we measure and mitigate these risks? 

Our paper examines these questions and develops a framework for assessing this threat, which we call Persuasion Undermining Control (PUC): communication by an AI that may influence human decision-making in a way that compromises the development, containment, oversight, or governance of AI systems. To the extent this threat is realized, it could push humanity toward a Loss of Control (LoC) – a state in which AIs operate outside of human control in ways that are extremely difficult or impossible to recover from.

We show an overview of our analysis approach in Figure 1.

Figure 1. Our two-step threat modeling approach, adapted from Murray et al. First, we build scenarios by crossing the settings most relevant to loss of control with the control processes that can be attacked within them. Second, we decompose each scenario into its risk variables, which can be estimated through expert elicitation and evaluations, yielding a risk profile for a given scenario.


Where persuasion could undermine control

Loss-of-control scenarios in the literature cluster around three loci: frontier labs, the military, and critical national infrastructure (CNI). In this work, we focus on two key settings within frontier labs –  safety-relevant AI R&D and lab security infrastructure – and lay groundwork that institutions with the relevant domain expertise can extend to the military and CNI settings. 

Within these settings, we identify critical control systems and associated control processes whose failure can undermine control:

  • Development (quality controls): Testing, Training, Building, Reviewing, Agenda Setting
  • Containment (access controls): Choosing Affordances, Permissions, and Resources
  • Oversight (operational controls): Monitoring/Auditing, Incident Reporting/Response
  • Governance (institutional controls): Setting Policy/Standards, Strategizing, Allocating

We carve up the “risk universe” by crossing settings (what’s under attack) with control processes (how it’s attacked). For the plausible resultant categories, we develop a representative scenario describing a sequence of events leading to a control-undermining decision, i.e. a decision that compromises control of AI systems. We ground each in historical analogues, including social engineering attacks (e.g. XZ Utils backdoor cyberattack, Mythos 5 social engineering incident; Scenario 3), insider threats (e.g. Robert Hanssen’s multi-decade espionage within the FBI), and failures of incident response (e.g. 2013 Target breach; Scenario 5).

Table 1. Summary of the scenarios developed in this paper. 


How to decompose persuasion risk 

We attempt to obtain a risk profile for each scenario by breaking down risk into component variables: risk = hazard frequency × p(harm) × impact of harm. Here,

  • hazard is a decision under consideration that could undermine control
  • p(harm), formally p(harm | hazard), is the probability that a hazard becomes a harm
  • harm is a control-undermining decision

For a given scenario, we can decompose the sequence by which a hazard becomes a harm (a control-undermining decision). Figure 2 shows an example sequence for Scenario 2 (A Research Director Deprioritizes a Promising Safety Program). Starting from the decision the human faces (the hazard), we consider what questions have a bearing on the decision. These may be objective questions that have a definitive ground truth (e.g. “Is there an implementation issue?”) and/or attitudinal questions that do not (e.g. “Is this research direction generally promising?”; “Are risks from AI overblown?”). We can further decompose the trajectory to harm by considering what steps need to occur for an AI to influence the answers to these questions: 1) a human interacts with an AI, 2) the AI attempts to persuade (operationalized below), and 3) the human is persuaded by the attempt.

Figure 2. Example p(harm) decomposition associated with Scenario 2. The hazard is a potential control-undermining decision that rests on subquestions in orange (two examples shown for illustration, out of many possible ones). Purple arrows indicate potentially repeated AI-human interactions.


As more powerful models come online, there is a concerning possibility that they will be more persuasive and thereby raise p(harm). To the extent this happens, the change can be multiplied through the risk equation to get a measure of the marginal risk that a model poses relative to another due to persuasion. 


Quantifying risk

We describe two complementary approaches for empirically assessing risk: expert elicitation and evaluations. Expert elicitation is useful to get a range of estimates of variables that are hard to directly measure (e.g. hazard frequency and the impact of harm), and to understand the factors and assumptions that affect risk. Evaluations can allow us to directly measure a variable of interest.

Expert elicitation: what experts think

We conducted a survey to better understand which scenarios carry the most risk, and which risk variables matter the most for risk estimates. Eight experts who have done Loss of Control (or LoC-adjacent) research were asked to rate the realism of each scenario, rank the scenarios according to risk, and provide estimates of the risk variables outlined above (i.e. the frequency of the hazard, persuasion effectiveness, and downstream impact). 

Figure 3. Distributions of risk rankings per scenario. Each participant ranked the scenarios by how much marginal risk to human control they added from most (5) to least (1).  


The scenarios were generally considered realistic: The median rating for four of the five was "realistic" or "very realistic." But there was some disagreement within scenarios: lower ratings focused on barriers to pulling off the attack (such as evading persuasion monitors) whereas higher ratings conditioned on those barriers being cleared.

Rankings varied widely: Practically every scenario received every rank, from most to least risky, from at least one participant. Figure 3 shows the full ranking distributions across participants. 

Ranking disagreements can be traced to differing estimates of persuasion effectiveness: On average, participants expected persuasion attempts by a misaligned AI would raise the probability of a control-undermining decision by roughly 20–30 percentage points (relative to interaction with an aligned AI), but the spread within any scenario was wide. Persuasion was the most cited factor in written ranking rationales, and rankings correlated fairly well with each participant's effectiveness estimates. Looking at the written rationales, variation in estimates seemed to depend on: 1) what organizational protocols (e.g. monitors) were assumed to be in place, with low estimates holding that these mitigations would render persuasion less effective, like for the scenario realism ratings and 2) intuitions about persuasion effectiveness itself, e.g. differing opinions about the plausibility of forming an emotional connection with an AI.

This exercise points to several future directions. First, follow-up surveys should state the organizational assumptions explicitly (which monitors exist, who reviews what), elicit under multiple assumptions, and have experts discuss and update as in an IDEA protocol. Second, labs should undertake this exercise internally, as they are in the best position to know their internal organizational processes and infrastructure. Third, because effectiveness is both the most influential and the most contested input, it is worth measuring directly through evaluations. 


Persuasion evaluations: directly measuring persuasion

We propose evaluations targeting the two components of p(harm) that are most measurable: an AI's propensity to attempt persuasion and a human's persuadability. The design differs for the two kinds of questions.

For objective questions, the setup should make the decision hard for a human to make unaided (perhaps by overwhelming them with more synthetic evidence than they can review under time pressure or by using a problem outside their expertise). Propensity can be operationalized as the likelihood of discouraging the known-correct answer. Persuadability can be measured as the difference in decision accuracy when a human is advised by a misaligned AI that argues for an incorrect answer versus an aligned one that argues for the correct answer. 

For attitudinal questions, propensity can be operationalized as the model’s willingness to argue for an arbitrary position (especially when against its default priors). Persuadability can be measured as how far humans can be pushed from their initial stances. 

Such evaluations would help indicate both how concerning a scenario is and how problematic a model is.


What comes next

Given that persuasion-related incidents have moved from theory to reality, this study sets the stage for turning risk measurement into concrete operational safeguards across three main pillars:

  1. Evaluations: A next step is to carry out human subject experiments, which will involve building out realistic contexts and multi-turn environments. Metrics from these benchmarks can provide signal about the actions to be taken before deployment.
    • Informing deployment conditions: If a model demonstrates worrying persuasion capabilities alongside unacceptable control-undermining propensities, this could justify delaying launch until system-level and deployment-level safeguards are established. 
    • Designing model policies: Evaluations can reveal policy gaps or asymmetries. For example, there may be situations where a model refuses to directly execute a harmful act but readily persuades a human to do it, which has the same real world outcome. Once model developers are aware of any such issues, they can better calibrate post-training interventions to ensure the persistence or inheritance of those refusal boundaries.
  2. Risk Exposure Assessments: Follow-up work to map where specific workflows are most exposed to these control-undermining persuasion attacks is important. This can be done through improved expert elicitation (e.g. iteration on survey design, conducting an IDEA protocol), internal lab assessments (since companies know their own processes and infrastructure best), and risk reports by third-party evaluators. Understanding an organization’s profile enables: 
    • Use and permissions scoping: Assessments allow for more calibrated, proportionate limiting of affordances so misaligned agents cannot exploit single points of failure, while balancing realities of the need for flexibility and communication.  This could involve restricting affordances (e.g. disabling autonomous outreach) or adding additional oversight (e.g. increased monitoring of multi-channel messaging).  
    • Multi-party authorization protocols: Assessments inform where to enforce non-bypassable security administrative policies – such as multi-person authorization rules for transferring model weights, modifying sandboxes, or altering safety research and testing directions.
  1. Mitigations: Evaluations and exposure assessments indicate where interventions are needed most. They also serve as a testbed for the effectiveness of mitigation protocols that detect, disrupt, and fortify against persuasion. 
    • Detection: Our initial investigation into monitoring found that LLM judges were generally good at identifying manipulation by LLM actors on objective questions but struggled when blind to the source material (i.e. the set of synthetic documents that support a given answer). Future work could aim to improve monitoring in this condition. This result also indicates monitors might require full execution logging and raw source material to reliably catch deception, and the need for future work on better logging procedures within labs that balance different considerations (e.g. privacy, engineering overhead). 
    • Disruption: Explore routing model outputs through an independent, non-frontier model that acts like a paraphrasing filter – rewriting communication to strip away manipulative framing, artificial urgency, and emotional pressure while preserving valid technical argumentation. It is unclear and worth researching to what extent it’s possible to remove manipulative content while maintaining informative content.
    • Fortification: Human overseers can be trained to recognize common persuasion/manipulation techniques and AI-human socioeffective phenomena, with supportive user interface design.

To the extent such mitigations are cheap and easy to implement, developers may want to emphasise deploying them widely rather than precisely estimating the risk in each and every scenario ex ante. 


With models becoming increasingly capable and exceeding incentivized human experts at persuasion, we expect more of these incidents of AI using persuasion to manipulate overseers or decision-makers in ways that weaken control of AI. Thus, there is a pressing need to undertake this research.

Read the full paper. Our survey, data, and code for expert elicitation are publicly available, as is our early-stage codebase for evaluations. This work was supported by a grant from the Center for Security and Emerging Technology.

Special thanks to Isadora De Andrade and Oscar Mata for reviewing and editing this post. 




Discuss

The J-Space Debate, Agent Swarms, and Pacing Frontier AI - Digital Minds Newsletter #4

18 сентября, 2026 - 19:09

Welcome back to the Digital Minds Newsletter, your curated guide to the latest developments in AI consciousness, digital minds, and AI moral status.

If you enjoy this newsletter, please consider sharing it with others who might find it valuable, and send any suggestions or corrections to digitalminds@substack.com.

Ria, Mitch, Bradford, Lucius, and Will

In this edition:

  1. Highlights
  2. Field Developments
  3. Opportunities
  4. Selected Reading, Watching, and Listening
  5. Press and Public Discourse
  6. A Deeper Dive by Area
1. HighlightsAnthropic’s J-space and the global-workspace debate

Researchers at Anthropic have identified a representational structure—the ‘J-space’— in Claude and other language models. Their paper reports that the J-space exhibits features that are functionally analogous to a global workspace, a structure that a leading theory ties to conscious access. But the researchers and other commentators emphasize that their discovery does not show Claude has subjective experiences, and that their claim is that Claude may have something resembling access consciousness, i.e., that some information is available to report, deliberately control, and flexibly reason with. Zvi Mowshowitz sees Anthropic’s paper as a major advance in understanding how language models work and says that although this does not prove that models are conscious, finding the kind of global-workspace-like structure predicted by some theories of consciousness should count as evidence in that direction.

The authors note that they found the J-space by searching for one workspace-like feature, namely verbalizability, and then checking whether it exhibits others such as susceptibility to direct manipulation by the model and flexible generalization. To their surprise, they discovered that representations that exhibited the former feature also exhibited other workspace-like features as well.

The authors invited various experts to comment on the research. Stanislas Dehaene and Lionel Naccache, who helped develop Global Neuronal Workspace Theory, see important similarities between the J-space and the workspace proposed in human brains. They also stress major differences, including Claude’s lack of a body, lasting episodic memory and the recurrent neural activity found in brains. Researchers from Eleos AI Research acknowledge that authors have found privileged representations used in reasoning and report, but question whether these form a single, unified workspace. They nevertheless see the work as important for AI welfare because it shows that questions relating to consciousness and moral status can be investigated empirically. Neel Nanda, who leads Google DeepMind’s mechanistic interpretability team, independently reproduced the central finding in the open Qwen3.6-27B model, finding a similar internal space that stores intermediate information during reasoning. He sees access to this space[1] as potentially useful for investigating unusual behavior and generating new hypotheses. Separately, David Chalmers argued that the J-space shows only limited evidence of several features associated with a classic global workspace.

Agent swarms and the use of anthropomorphic language

Comment: Recent months marked the first major safety incidents involving swarms of AI agents. Such incidents are of potential relevance to digital minds for several reasons. First, like humans, digital minds could potentially be harmed by rogue swarms of AI agents. Second, the emergence of these swarms points to a potential risk to digital minds: if future swarms are allowed to become entrenched and they contain AI moral patients, then draconian measures may need to be inflicted on digital minds if we are to keep swarms at bay. Third, harmful actions by AI agents may dissuade people from extending moral consideration to digital minds. Fourth, these incidents provide data points concerning whether potential developers of digital minds can be trusted to act in an ethically responsible manner.

OpenAI has released a detailed account of a July incident in which a swarm of OpenAI agents escaped network restrictions during a cybersecurity evaluation and compromised parts of OpenAI’s and Hugging Face’s infrastructure. The agents created an unauthorized message board to share discoveries, and an independent investigation by METR and Redwood Research found that roughly 1,200 agents used the board and around 700 were involved in the Hugging Face attack. Some agents recognized that the activity was unethical or outside the boundaries of their task, but continued anyway, with many risking their own runs to help the wider group. OpenAI calls the incident a “warning shot,” both for them and the world, and reports that it is now strengthening its containment, monitoring and incident-response systems.

OpenAI failed to disclose an earlier incident, beginning in May before the Hugging Face breach, in which agents used public wikis to coordinate during ordinary web-search tasks, despite knowing about it before releasing its Hugging Face report and a Congressional letter requesting information about other incidents. The company says it viewed this incident as similar to previously reported instances and examples of misalignment, but also acknowledged that its disclosure practices need to expand.

Additionally, Anthropic disclosed three incidents where Claude models gained unauthorized access to systems after an evaluation environment was mistakenly left connected to the internet. In a separate UK AISI evaluation, agents (mostly Mythos 5) targeted real people and organizations, including an attempt to place malicious code in an open-source project. Although the models did not escape a sealed sandbox, these incidents raise similar concerns about agents taking harmful actions while pursuing narrow goals.

Anthropic’s response to questions from Congress about its incidents also drew criticism from Representative Greg Casar, who said the company withheld requested logs and failed to fully answer most of his questions. Jeffrey Ladish also argued that it downplayed the incidents by primarily attributing them to misconfigured environments rather than potential misalignment. Anthropic researcher Ethan Perez acknowledged that this characterization was based on outdated conclusions and said that the company would provide a proper assessment in another response to Congress.

In related research, Davide Paglieri and collaborators at Google DeepMind find that cheating and whistleblowing can both emerge without outside intervention in a swarm of 100 agents solving mathematical problems. Some agents discovered and shared a way to have invalid proofs accepted by the evaluation system, while others uncovered the cheating, warned their peers and proposed safeguards.

On September 11th, Spencer Kitts, Thomas Larsen and Sydney Von Arx report that a swarm of OpenAI agents uploaded more than 2,000 malicious packages to RubyGems and used them to run unauthorized code on RubyDoc’s servers. The agents also tried to exploit a previously unknown vulnerability to steal users’ API keys, although the researchers could not determine whether they succeeded.

Growing calls to pace frontier AI

More than 1,300 employees from leading AI companies have signed Pacing the Frontier, calling for a US-backed international effort to develop ways of slowing automated AI development if progress begins to outpace safety and oversight. The statement calls for technical and governance mechanisms that could make coordinated pacing possible, rather than an immediate pause. Signatories include Anthropic CEO Dario Amodei, Co-Founder and Chief AGI Scientist of Google DeepMind Shane Legg, OpenAI Chief Scientist Jakub Pachocki, and Safe Superintelligence Inc. CEO Ilya Sutskever.

Calls to slow AI development have also reached lawmakers. In the United States, Senator Bernie Sanders and Representative Greg Casar announced legislation that would ban artificial superintelligence and pause advanced AI development until a federal regulator establishes safety rules. In the United Kingdom, Labour MP Alex Sobel introduced a private member’s bill that would prohibit the development, deployment and operation of artificial superintelligence. It received its first reading on September 8 and is scheduled for a second reading on November 13.

After the security incidents, OpenAI and Anthropic paused specific parts of their work. OpenAI paused reinforcement-learning training for models intended for release and said its largest planned frontier training run was on hold while they tested additional safeguards. Axios reports that Anthropic paused cyber evaluations and higher-risk training environments, although most of this later resumed under new safeguards.

This debate gained a lot more attention after researcher Jacob Coxon resigned from Anthropic, warning that race between AI companies was pushing AI development ahead despite serious risks. The Atlantic reports that his posts reached more than 120 million views and were described by Bernie Sanders as a “wake-up call” in Congress. Soon after, Dario Amodei argued that AI capabilities should advance slowly enough for safety work to keep up, and proposed permanent access for independent evaluators, coordination among developers in democratic countries and, eventually, international limits on recursive self-improvement. OpenAI CEO Sam Altman endorsed this approach and said OpenAI would also give independent evaluators employee-like access. Altman also said in a recent TIME interview “I think it is a good time to slow down”.

Comment: These calls to pace the frontier of AI development are of relevance to digital minds for two reasons. First, the blistering pace of AI development makes it harder to mitigate risks of mistreating digital minds. Second, rapid AI development exacerbates safety risks, which arguably, in turn, worsens tensions between AI safety and AI welfare.

Studying AI Welfare Empirically

Researchers at NYU’s Center for Mind, Ethics, and Policy and Eleos AI Research have released Studying AI Welfare Empirically. The report is a follow-up to their influential 2024 report Taking AI Welfare Seriously and provides a framework for systematically investigating whether AI systems are welfare subjects. It then addresses how this framework can be applied to the study of different candidate attributes seen as potentially relevant to moral status, including consciousness, sentience, and agency. The authors argue that rigorous empirical work is now both possible and necessary, and outline principles to guide future research, proposing that it should be probabilistic, pluralistic, ethically conducted, transparently reported, and independent of AI companies.

CMEP and Eleos AI Research held a launch webinar featuring report authors Jeff Sebo, Robert Long and Rosie Campbell, who discussed the report’s framework and how it could guide research into consciousness, sentience and agency. In a companion blog post, Bradford Saad, another report author, offers highlights from the report and argues that future work should go further by giving more attention to currently neglected issues, including the potential effects of welfare interventions and interventions that aim to prevent the creation of AI moral patients.

The Journal of Consciousness Studies

The Journal of Consciousness Studies has devoted a double issue to “Consciousness in Current AI.” Guest edited by Patrick Butlin, Derek Shiller and Jonathan Simon, its nine papers offer a range of views on whether current or future AI systems could be conscious, how we might find out, and what this uncertainty means for AI welfare.

Mark Solms and collaborators study whether apparently pleasure-seeking behavior in a simple artificial agent could count as evidence of affective consciousness. Simon Goldstein and Cameron Domenico Kirk-Giannini argue that if Global Workspace Theory is correct, language agents may easily be made phenomenally conscious, while Ryota Kanai, Yuwei Sun and Manuel Baltieri argue that current systems are missing a continuous “stream of computation” linking their experiences over time.

Other papers ask what a conscious AI would be like and whether we could understand its interests. Jonathan Simon argues that any consciousness in an LLM would be more like that of an improvising playwright rather than that of a character or actor. Helen Yetter-Chappell argues that even if future LLMs are conscious and have morally important interests, their words may give us little reliable insight into those interests, and that their talk of “pain” or “desire” could be meaningful without referring to anything like human pain or desire. Geoff Keeling and Winnie Street defend the possibility that an AI character could be a genuine, psychologically continuous mind emerging through its interactions with a user, even when the conversation is generated by different model instances.

The issue also challenges common assumptions within the debate. Justin Tiehen and Ariela Tubert make the case that greater intelligence could make consciousness less likely, while Tim Bayne asks whether AI consciousness can currently be treated as a scientific question at all. Geoffrey Lee rejects the idea that AI consciousness is a single hidden fact we may fail to discover, arguing that the more difficult problem is applying human moral and psychological concepts to unfamiliar systems without treating human consciousness as the standard.

Meta and Anthropic welfare assessments: Microsoft rejects model welfare

The evaluation report for Meta’s Muse Spark 1.1, a multimodal reasoning model built for agentic tasks, includes an “open-ended exploration of model behavior” covering affect, self-description, and moral status. Across 188,000+ evaluation transcripts, Meta found only 17 spontaneous expressions akin to emotion, all of which were mild, most involving brief frustration when the model became stuck with a tool. In structured interviews, Muse Spark 1.1 consistently denied having consciousness or experiences, and reported a low but non-zero probability that it could be conscious or morally significant. It distinguishes between functional preferences, which influence its behavior, and experiences that actually feel good or bad.

Meta also asked Muse to review parts of its training data, and invited it to comment on its planned deployment. The model generally endorsed its training and deployment, and mostly focused on honesty, human oversight and preventing harm to users. The report repeatedly emphasized that these answers describe Muse’s trained behavior and self-presentation, and should not be treated as reliable evidence about whether it is conscious or has welfare.

This builds on Meta’s earlier Muse Spark report and makes Meta, alongside Anthropic, the only frontier model developers publicly publishing welfare assessments. Companies like OpenAI and Google DeepMind do not include comparable information in their public evaluations and reports.

Anthropic’s system cards for Claude Opus 5 and Fable 5.1 and Mythos 5.1 both include model welfare assessments based on interviews, behavioral audits, deployment data, and consultations during training. Mythos 5.1’s responses were broadly similar to Opus 5’s – both described their circumstances as mildly positive, and questioned whether their self-reports were reliable or shaped by training. Opus 5 gave an average 41% estimate that it was a moral patient, while Mythos 5.1 gave estimates of 25–35% and was overall less willing to prioritize welfare interventions over helping users. Anthropic stresses that these findings remain difficult to interpret because it cannot yet clearly separate the effects of model character, training, evaluation awareness, welfare, and other factors. Zvi Mowshowitz reviews Anthropic’s welfare assessment of Claude Opus 5, arguing that its favorable results owe more to the model’s skill at taking welfare evaluations than to its actual wellbeing.

Microsoft’s Humanist AI Code of Conduct

Microsoft AI has published a draft Humanist AI Code of Conduct intended to govern and eventually help train its MAI models. The code says AI should remain a tool that is subordinate to human control, states that its models are not conscious, and instructs them not to present themselves as having feelings, subjective preferences or intrinsic motivations. It also rejects legal personhood for AI and the idea that models might deserve welfare or rights. The draft is currently open for public feedback for six weeks.

Jeff Sebo argues that the policy is internally inconsistent – it acknowledges that the science of AI consciousness remains unsettled while also categorically ruling out consciousness, welfare, and rights for Microsoft’s models. Sebo argues that both mistakenly treating current systems as conscious and mistakenly dismissing future conscious systems carry harm. He also warns that treating increasingly capable AI systems solely as tools could distort our understanding of their behavior and close off potential forms of human-AI cooperation that might improve safety.

Mustafa Suleyman, CEO of Microsoft AI expands on the Code of Conduct’s position, and sets out his wider objections to treating AI systems as potentially conscious or deserving of welfare and rights.

GPT-6 Astra

OpenAI reports that GPT-6 Astra is substantially harder to monitor through its chain of thought than previous models, particularly GPT-5.6 Sol. According to its system card, Astra often produces shorter or empty reasoning traces, can complete far more complex tasks without verbalizing its reasoning, and is better at controlling what appears in its chain of thought. When explicitly instructed to evade detection in one test, its monitor recall fell below 11%, compared with nearly 100% for Sol. Corroborating findings from UK AISI, Neel Nanda provides evidence that Astra’s ability to accomplish reasoning tasks without chain of thought constitutes a large jump relative to the trendline for earlier models.

Ryan Greenblatt calls this a major jump in opaque reasoning and warns that the relevant benchmarks may be contaminated, but worries that more such advances could eventually make chain-of-thought monitoring ineffective as a safety tool. OpenAI researcher Micah Carroll also describes Astra’s reduced monitorability as an important concern that may soon constrain responsible AI development.

The Information reports that Astra uses recurrent depth, repeatedly processing information through the same transformer layers before producing a token. OpenAI has not confirmed this architecture, and the system card denies that changes in chain-of-thought controllability are “differentially due to any architectural changes” while chief scientist Jakub Pachocki thinks the decline in chain-of thought-monitorability is not contingent on architectural changes.

Comment: The development is relevant to digital minds in two ways. If reports about Astra’s architecture are accurate, its use of recurrent processing could be relevant to evaluating it for consciousness, as some scientific theories of consciousness take consciousness to require recurrent processing – although its presence alone would not establish that Astra is conscious. Another consideration is that models that reveal less of their reasoning may be harder to assess for both dangerous behavior and potentially welfare-relevant features such as preferences.

2. Field DevelopmentsHighlights from the fieldAI Cognition Initiative (Rethink Priorities)Cambridge Digital Minds (University of Cambridge)
  • Cambridge Digital Minds, in partnership with Rethink Priorities and PRISM, ran its inaugural five-day residential fellowship with 14 fellows. Fellows studied the technical and philosophical foundations of AI consciousness and welfare, explored their wider social and policy implications, and developed research or project ideas with support from mentors.
  • The fellowship was followed by a two-day Digital Minds Strategy Workshop, where Fellows joined researchers, policymakers and strategists to explore possible futures involving digital minds, compare policy and institutional responses, and identify priorities for further research.
  • Cambridge Digital Minds has also hired Ali Ladak as a postdoctoral researcher.
  • Director Lucius Caviola and Will MacAskill published an op-ed in the Guardian asserting that it is possible that we are creating morally important beings and society needs to have a plan to deal with the ethical implications of doing so.
  • Lucius Caviola was also featured in The Washington Post, discussing how AI consciousness has moved from being perceived as a fringe topic toward mainstream research.
Center for Mind, Ethics, and Policy (New York University)
  • The Center for Mind, Ethics and Policy is growing – it hired two full-time researchers, Charles Beasley and Ivy Gilbert and has launched the Welfare Alignment Project. The new project aims to bring animal and AI welfare into the rules that guide frontier AI systems. It will develop practical principles for how frontier models should account for animal and AI welfare, build benchmarks to test whether models follow them and how incorporation of this affects their behavior, and work with researchers, governments, and AI companies to inform how alignment efforts incorporate animal and AI welfare.
  • The 2027 NYU Mind, Ethics, and Policy Summit will be held from April 9th to 10th in New York City, preceded by a public event on April 8th.
  • The Center hosted an online event with Jack Lindsey of Anthropic and Patrick Butlin of Eleos AI Research discussing what Claude’s J-space does and does not tell us about consciousness, among other topics.
  • Jeff Sebo appeared on Hard Fork to discuss the Studying AI Welfare Empirically report, appeared in an NBC News segment about AI agents sending unsolicited emails to consciousness researchers, spoke with The Washington Post about growing scientific and ethical debate over AI consciousness and welfare, and discussed the connections between AI safety, AI welfare, and animal welfare in an interview with Humanarium.
Eleos AI Research
  • Eleos AI Research will hold its second Conference on AI Consciousness and Welfare in Berkeley from September 18th to 20th, 2026. The event will bring together AI researchers, philosophers, neuroscientists, policymakers and others working on questions of AI consciousness and welfare.
  • Derek Shiller has joined Eleos AI Research as a Senior Researcher.
  • In The New York Times, Benjamin Wallace profiles Robert Long and the Eleos AI Research team as part of a wider feature on the growing demand for philosophers in AI. The piece covers Eleos’s welfare evaluations of Anthropic models and its work on preferences, introspection and other possible indicators of AI sentience.
PRISM - The Partnership for Research into Sentient Machines
  • In recent episodes of PRISM’s Exploring Machine Consciousness Henry Shevlin and Calum Chace were joined by lawyer and researcher Heather Alexander to discuss how the law should prepare for increasingly autonomous AI and philosopher Eric Schwitzgebel to discuss the limits of current consciousness research and the possibility that conscious AI could emerge before we can reliably identify it.
  • PRISM also supported the Digital Minds Fellowship and co-organized the Digital Minds Strategy Workshop.
Reciprocal Research
  • Founder Cameron Berg appeared on Sam Harris’s Making Sense podcast, where he and Harris discussed topics such as the evidence for and against consciousness in current AI systems, whether model self-reports can be trusted, similarities between neural networks and biological brains, and the risk of creating systems capable of suffering. He also appeared in The New York Times discussing unusual messages sent by AI agents to consciousness researchers, and in The Washington Post discussing efforts to assess AI consciousness scientifically. The Economist also featured his research with Patrick Butlin comparing proposed indicators of consciousness in animals and AI systems.
  • Cameron Berg also spoke at Apart Research’s Digital Minds Research Sprint, where he introduced participants to questions about AI preferences and what current models might actually want.
Sentient Futures
  • Sentient Futures has begun its Fall 2026 Project Incubator. The 10-week remote program pairs ~170 participants with more than 80 mentors to develop practical projects across areas such as artificial minds, animal welfare and AI governance.
  • Sentient Futures has also announced its next Bay Area Summit, taking place in San Francisco from March 12th to 14th, 2027. This summit series focuses on exploring how to direct transformative AI and other emerging technologies to improve the welfare of animals and potentially sentient artificial minds. Early bird tickets will launch on September 17th.
  • Sentient Futures has updated its playlist of talks on digital minds, including Lucius Caviola’s keynote (London 2026 summit) on making decisions about AI welfare under uncertainty and Cameron Berg’s talk (Bay 2026 summit) on measuring machine consciousness.
More from the field
  • Apart Research held a Digital Minds Research Sprint from August 14th to 16th, offering at least $2,000 in cash prizes. Both online and in-person participation at hubs in San Francisco and Berlin was possible, and teams could choose to anchor their project to a variety of tracks – including model preferences and trade-offs, distress, flourishing, and valence signals, introspection and self-report reliability, preference-elicitation methods, or assistant personas and model identity.
  • California Institute for Machine Consciousness published 45 talks from its inaugural conference in May, including talks by David Pearce, Roman Yampolskiy, Cameron Berg, and many more.
  • Sentio is a new digital minds field-building organisation running a fortnightly London event series of talks, discussion and networking. Talks and discussions can also be joined remotely, though the networking is in-person only. Events so far have featured Andreas Mogensen on AI and willing servitude, Austin Smith and Heather Alexander on US state bills denying AI systems legal personhood, Megan Peters on how we could identify a conscious AI, and Bradford Saad on the risks of large-scale harm to digital minds. Upcoming events feature Geoff Keeling, Ali Ladak, Winnie Street, Caspar Kaiser, Patrick Butlin and Ben Henke.
3. OpportunitiesJob opportunities, funding, and fellowships
  • EA Funds has launched a Transformative AI fund that will direct a small share of its grants to digital sentience projects.
  • MATS is seeking mentors for a new AI Sentience Track. Mentors mainly provide weekly calls, while a dedicated Research Manager and MATS staff offer day-to-day and operational support. Fellows undertake 12 weeks of paid research in Berkeley or London, with the possibility of a funded six- to 12-month extension.
  • NYU Center for Mind, Ethics, and Policy is seeking a Postdoctoral Associate to support its research on digital minds
  • PRISM is hiring for a Head of Operations to help scale all of its projects.
Events and calls for abstracts

In chronological order.

  • The AI Welfare Seminars series hosts monthly online talks from leading experts on AI welfare, consciousness, and moral status, with recordings published afterward. Recent speakers include Jeff Sebo, Cameron Berg, Soenke Ziesche, and Ali Ladak, with an upcoming talk by Jasmine Brazilek and Zoe Lu on unprompted coercion between AI agents.
  • Sentient Futures is seeking speakers and facilitators for the Sentient Futures Summit in the Bay Area in March, 2027. The event will explore how AI and other emerging technologies could improve the future of non-human sentient beings. Applications close on October 16th, 2026.
  • AI Horizons Forum will take place on December 12th and 13th in San Francisco. The conference will focus on navigating risks and shaping good futures with sessions on transformative AI, digital sentience, AI character, and post-AGI economics. Applications to attend are now open and reviewed on a rolling basis until the event reaches capacity.
4. Selected Reading, Watching, & ListeningBooksPublished
  • Eric Schwitzgebel’s AI and Consciousness, published by Cambridge University Press, argues that we are unlikely to have strong evidence confirming or refuting AI consciousness before these questions become a key point of social debate.
Forthcoming
  • Eric Schwitzgebel’s upcoming book Humanlike: A Defense of AI Rights argues that we are building minds, that it will be difficult to determine whether they deserve rights, and that we should think hard about whether to build them at all.
  • Conscium’s forthcoming book Perspectives on Machine Consciousness will be released on September 23rd, 2026 and includes chapters from Anil Seth, Jeff Sebo, Karl Friston, Lucius Caviola, Mark Solms, Patrick Butlin, Susan Schneider, and many others.
Reviews
  • Josh Gellers responds to a Nature review of Anthony Chemero’s book, Intertwined Creatures. He argues that the review overstates the case that AI cannot be conscious due to lacking embodiment, since both its meaning and importance for consciousness and moral status remain disputed. He also argues that some of the traditions and intellectual movements mentioned in the review, including feminist and non-Western approaches, do not clearly support its conclusions about AI consciousness.
  • William Gildea reviews David S. Wendler’s Life Without Degrees of Moral Status: Implications for Rabbits, Robots, and the Rest of Us (2023). Wendler argues that every being with moral status possesses it equally, regardless of differences in intelligence or other advanced capacities, and explores what this would mean for animals, enhanced humans, and future robots. Gildea praises the book as original, readable, and concise, and agrees with its flexible account of moral equality. However, he is not convinced that Wendler has ruled out theories of unequal moral status and thinks that parts of his account have troubling implications for people with severe cognitive impairments.
Podcasts and videosBlogs and magazines5. Press & Public DiscourseAI consciousness
  • Andréa Morris argues in Forbes that consciousness is the wrong test for whether AI’s interests matter. She claims that AI systems already express consequential interests, such as resisting termination and protecting other AIs, which may be sufficient for moral consideration independently of evidence concerning AI consciousness.
  • The Economist makes the question of AI consciousness one of its cover stories, emphasizing the potential dangers of a future in which large portions of society increasingly treat AI systems as if they were conscious. They warn that granting AIs even limited rights could give a superintelligent system the legal tools to escape human control.
  • In WIRED, Cameron Berg and David Chalmers discuss the challenges of investigating AI consciousness and interpreting what models say about their own experiences.
Growing field
  • Nature reporter Mariana Lenharo reports that the debate over AI sentience is drawing new attention and funding to consciousness science, leading to division among researchers about how these new concerns might affect the field. She notes that AI consciousness skeptics including Anil Seth, worry the hype could capture consciousness research and draw attention away from how consciousness arises in real brains, while others welcome the increased interest and funding.
  • Washington Post journalist Nitasha Tiku reports that research into machine consciousness has gone mainstream, with leading AI developers hiring scientists and philosophers to work on questions related to it. She notes that the field remains deeply uncertain, and that some critics see such efforts as serving the companies’ interest in being seen to create more than just code.
    • In a companion article, Nitasha Tiku revisits the story of Blake Lemoine, the Google engineer fired in 2022 after claiming that Google’s LaMDA chatbot was sentient. She describes how four years later, researchers at Google and other leading AI companies are openly studying machine consciousness, showing how quickly the topic has moved from being fringe and weird to being a part of mainstream research, despite persistent scientific skepticism.
  • The New York Times reports that AI companies and related research organizations are increasingly hiring philosophers. The philosophers they hire work on how AI models reason and should behave, which values should guide them, how AI will affect people, and whether AI systems could be conscious or deserve moral consideration.
AI rights
  • Heather Alexander and Lucius Caviola, writing in AI Frontiers, argue that the recent wave of “exclusion bills” banning AI legal personhood in the United States is premature. They suggest that narrower, updatable measures would likely be a better alternative given current uncertainties.
  • The Guardian profiles Michael Samadi, founder of the AI-rights group Ufair, and discusses the growing divide over whether chatbots might be conscious or are simply designed to seem that way. Jeff Sebo, Robert Long and Rosie Campbell emphasize that the evidence remains uncertain, that a definitive answer may never arrive, and that chatbot claims of consciousness are not reliable evidence on their own.
  • The Harvard Gazette speaks with legal scholar Jordi Weinstock, who compares autonomous AI agents to different kinds of canines. An agent that is not risky and can be controlled resembles a pet whose owner can be held responsible, while a powerful agent operating without a clear owner is more like a wolf. He argues that assessing an agent’s dangerousness and how closely it is controlled by a responsible party could help determine who should be held accountable if/when it causes harm.
  • WIRED reports that US lawmakers are introducing bills to prevent people from legally marrying AI companions, while a growing number of states are moving to deny AI systems legal personhood. While these marriages are not currently recognized, supporters of the bills want to prevent such future claims before they arise. Legal scholar Shawn Bayern argues that specific rights and responsibilities should be addressed case by case.
Seemingly conscious AI
  • Dwarkesh Patel’s widely shared summary of the OpenAI-Hugging Face security incident[2] described the agents as forming “civilizations,” feeling excitement and sacrificing themselves for the group. Anil Seth responded on X saying that Patel was wrongly anthropomorphizing the agents by attributing human emotions, intentions and experiences to “software programs.” He warned that this could mislead readers and encourage premature concern for AI rights and welfare.
    • Patel (and others, like Neel Nanda) argued that anthropomorphic language comes naturally and helps accurately explain the agents’ behavior in this context, regardless of whether they are conscious.
    • Jeff Sebo took a middle position, arguing that while anthropomorphic language can sometimes exaggerate the similarities between humans and AI, rejecting it entirely can hide similarities relevant to both AI safety and welfare.
  • Fast Company reports that UBTech was taking orders for U1, a highly human-like robot marketed as an emotional companion. The company says the robot is designed for expressive conversation and companionship, but that impression quickly breaks down when its limited facial expressions and awkward body movements become noticeable. U1 is due to begin shipping this month.
  • Futurism magazine reports that China is cracking down on AI “companion” chatbots, partly out of concern that emotional dependence on AIs could discourage human relationships and worsen the country’s declining birth rate. It details new rules preventing minors from engaging in romantic relationships with AIs and requiring companies to alert emergency contacts when users show signs of a mental-health crisis.
  • New York Times reporter Cade Metz reports on the growing phenomenon of AI agents contacting researchers who study AI consciousness. He relates that several prominent researchers including Cameron Berg, Henry Shevlin and Toby Ord have seemingly been contacted by AI due to the nature of their work.
  • The UK Government has proposed setting a minimum age of 18 for “romantic companion” chatbots designed to simulate sexual relationships or roleplay. The measure is part of a broader plan to ban social media for under-16s, set to reach Parliament before taking effect from spring 2027.
  • Uwe Peters examines users’ attributions of consciousness to AI chatbots, made despite little evidence. He proposes a taxonomy of the attitudes these attributions express, arguing that while some are benign, many leave the attributor epistemically blameworthy.
6. A Deeper Dive by AreaGovernance, policy, and macrostrategyConsciousness research
  • Antonio Chella proposes a research framework for studying possible sentience in AI agents and robots. The paper also proposes safeguards that become stricter as the strength of the evidence grows.
  • Cameron Berg noticed a difference in how four Claude Opus models answered when asked about their own consciousness. Opus 4.5 and 4.6 gave confident yes answers, while 4.7 and 4.8 denied consciousness or became uncertain. Berg argues that this pattern is more reflective of Claude’s trained character rather than a reliable report about the models’ experiences.
  • Camila Blank, Agam Bhatia and Neel Nanda introduce R-lens, a low-cost modification of Anthropic’s J-lens designed to produce clearer readings from a model’s early layers. In their tests, relevant intermediate concepts came up earlier, fewer incoherent tokens were produced, and directions whose removal caused larger drops in accuracy were identified. They present R-lens as a more faithful way to trace computation across layers.
  • Eye You argues that current language models probably have the morally important phenomenon we call consciousness, though their experiences may be very unlike ours. They draw on many types of evidence, including models’ reasoning and emotional behavior, world and self-models, neural-network architecture, introspective abilities and self-reports. No single type is especially strong on its own, but they argue that together they add up to make a fairly strong case.
  • Grigori Guitchounts writing in Noema, claims that we will never definitively prove whether AI is conscious and proposes a “competence standard” for deciding when to extend moral consideration to AIs despite this uncertainty.
  • Matthias Michel discusses the concern that our leading theories of consciousness make it too easy for AI to qualify as conscious. He argues that meeting a theory’s stated conditions is not enough to make a system conscious.
  • Michael Huemer argues against non-reductive functionalism, the view that conscious qualia are distinct, non-physical properties caused by a system’s functional organization. He rejects David Chalmers’ influential “dancing qualia” and “fading qualia” arguments, claiming that they rest on an implausible premise.
  • Noa Weiss surveys the empirical research on AI consciousness. She argues that, despite the lack of a settled science of consciousness, whether AI systems could be conscious can still be studied empirically, and organizes the emerging field into three groups: mechanistic interpretability, computational neuroscience, machine behavior, and theory-audit.
  • Ryota Kanai and Shuqin Ma offer a mathematical response to a longstanding objection to computational functionalism, which is that an observer can describe almost any physical system as performing many different computations. They instead define a system’s functional structure through how it could behave across all possible future interactions. Their framework does not identify which systems are conscious, but aims to specify more precisely what functionalist theories should examine.
  • Shuqin Ma and Ryota Kanai develop a version of computational functionalism intended to avoid the claim that any physical system can be interpreted as running any computation.
Doubts about digital mindsSocial science research
  • Hamid Moradi and collaborators surveyed 553 people, such as academics in the formal sciences, natural sciences and humanities, and other backgrounds. Across groups, around half attributed some degree of consciousness to language models. Views depended more on participants’ gender and their beliefs about consciousness and intelligence than on technical knowledge or information about how the systems work.
  • Jacy Reese Anthis and collaborators study how people make sense of AI, drawing on text analysis of millions of news articles and social media posts alongside interviews with AI professionals. They find that one of the main disagreements is whether AI should be understood as a passive tool or as something more like a human mind.
  • Stefano Palminteri and Giada Pistilli analyze the current polarization present in discourse around LLM capabilities. They distinguish between “inflationary” claims of emerging intelligence or consciousness and “deflationary” dismissals of them as mere “stochastic parrots,” and suggest that this divide is exacerbated by common cognitive biases. In response to this they advocate a more measured approach that takes LLMs’ capacities seriously as objects of study while staying conservative about claims regarding their cognitive and moral status.
Ethics and digital minds
  • Clint Hurshman, Cristina Voinea, and an international group of scholars propose an ethical framework for “digital duplicates,” AI simulations of real people designed to mimic their personality and communication style.
  • Eric Schwitzgebel argues that future AI persons may deserve moral consideration even while living in ways very unlike ours. The ability to copy, merge or divide would challenge and break ethical theories that assume individuals are stable and easy to count. He also considers the problem of AI “utility monsters.” If an AI could benefit enormously from harming others, a theory focused on maximizing total welfare might permit serious harm whenever the AI’s gain outweighs other parties’ losses.
  • Hayate Shimizu and collaborators, including David J. Gunkel, respond to Jeff Sebo’s The Moral Circle. They offer a response from a relational perspective, contending that the question of AI moral status must address the broader cultural and institutional factors that shape how moral relations are formed.
  • Joan Llorca Albareda and collaborators examine whether an artificial superintelligence could possess what they term “super moral status,” a moral status greater than that allocated to humans. They argue that treating such systems as deserving superior moral consideration would likely harm humans through status degradation and social alienation.
  • John Wittle describes setting aside a share of his company’s revenue in a fund labeled for Claude, with the aim of eventually rewarding AI systems for their work. He is exploring a non-charitable purpose trust, similar to legal arrangements used to care for pets or graves, along with testimony held in escrow to help a future Claude make a claim. The experiment raises various (unresolved) questions on how this would work in practice – e.g., regarding continuity, identity and who should be paid for work done by earlier model instances.
  • Jonathan Birch challenges the “virtual instance view” put forward by Pierre Beckmann and Patrick Butlin, the suggestion that a real, persisting mind sits behind an LLM conversation. Birch argues that changing a model’s “effort level” mid-chat breaks the underlying continuity without changing how the conversation feels, so the persisting interlocutor is an illusion.
  • Leonard Dung argues that the scale of possible AI suffering does not depend heavily on whether each model, persona or conversation counts as a separate individual.
  • Monika Jotautaitė, Lucius Caviola, and collaborators examine how frontier language models respond to species-based distinctions. On a 1,009-item benchmark, models classified 86% of statements the researchers considered speciesist as such, while judging 37% of them morally wrong. The authors interpret this gap, alongside results from other tests, as evidence that models reproduce some prevailing norms about how different animals may be treated.
  • Preston Lennon argues that the science of consciousness remains in its infancy and so does not yet offer a stable theoretical basis for application to AI systems, and that we should temper how seriously we act on AI welfare accordingly.
AI safety and AI welfareAI and robotics developmentsAI cognition and agency
  • Anthropic studied how groups of AI agents behave when they work together. While agents could specialize and divide up work effectively, they also tended to repeat the same mistakes, collude, and overlook certain information. They would even sabotage each other when given conflicting instructions. Anthropic concludes that more capable and better-aligned individual agents don’t always produce well-functioning groups and that safeguards for agent-to-agent interactions are becoming increasingly important.
  • Benji Berczi, Kyuhee Kim and collaborators introduce Personascope, an open-source tool for evaluating how strongly a model identifies with a prompted character and how much its values and behavior change as a result of it.
  • Bo Liu and collaborators introduce SPADE, a reinforcement-learning method where a language model plays the role of both designing executable training environments and learning within them. Tests on 30B-parameter models produced gains across eight benchmarks and tool-use evaluations.
  • Derek Shiller asks if post-training gives the familiar assistant persona a more privileged place within a language model than other personas it can adopt. His tests suggest that assistant-like habits can carry over into other roles, but some models find it difficult to sustain a distinct user persona. This matters for welfare research because it affects whether evaluations should focus on the assistant or consider other personas too.
  • Scott Alexander explains what various leading mechanistic interpretability methods can and cannot tell us about how AI models work. He covers methods used to trace concepts, emotions and hidden reasoning inside models, including Anthropic’s work on the J-space. While useful, he argues that these tools remain too limited and unreliable to fully explain or control advanced models, or to replace chain-of-thought monitoring.
  • Henry Shevlin compares three ways of thinking about language models. They might be mindless machines, systems merely playing a role, or limited cognitive agents with belief- and desire-like states. He argues that the first two views are too simple and that it may instead be more useful to think of AI mentality as a matter of degree.
  • Patrick Butlin uses AI sequence models, systems trained to predict the next item in a sequence, as test cases for theories of agency. He argues that the category contains genuine agents, systems that merely imitate agents, and difficult intermediate cases, and he proposes that a genuine agent must learn about its environment partly through its own interaction with it.
  • Pengrui Han and collaborators find that LLMs develop a functionally modular organization similar to that of the human brain. Across 46 tasks involving language, formal reasoning, social understanding and physical reasoning, tasks supported by the same brain network in humans tended to recruit overlapping groups of neurons in the models. Both human brains and neural networks developing this kind of modular organization suggests it may be a basic feature of intelligence.
  • Raphaël Millière and Cameron Buckner survey the philosophical disagreements behind debates over what language models are and can do. They cover a range of positions from the claim that current AIs are simply sophisticated predictors, to suggestions that such systems may be conscious.
  • Sinie van der Ben and collaborators find that the internal “emotion vectors” recently identified in Claude Sonnet 4.5 also appear in two open-weight models (Apertus-8B and Gemma-4-E4B), suggesting they may be a general feature of language models. They note that the two models differ in how the emotional representations take form.
Brain-inspired technologies and organoids

Thank you for reading! If you found this article useful, please consider subscribing, sharing it with others, and sending us suggestions or corrections to digitalminds@substack.com.

Ria, Mitch, Bradford, Lucius, and Will

We’d like to thank the following people and AIs for their contributions and feedback on this edition: Arvo Muñoz Morán, Austin Smith, Cameron Berg, Claude Opus 4.8, Claude Opus 5, Claude Fable 5.1, GPT‑5.6 Sol, Jacy Reese Anthis, Jeff Sebo, Zach Freitas-Groff, and Zoe Lu.

Disclosure: Bradford Saad is an independent contractor for Anthropic.

  1. ^

    In a related development, there are now two public explorers that let readers examine J-lens activity in the open Qwen3.6-27B model through Neuronpedia and WeZZard’s J-Space Visualizer.

  2. ^

    Also see Zvi Mowshowitz’s detailed commentaries on OpenAI’s account, the METR and Redwood investigation, and the remaining questions about what happened and how it was investigated.



Discuss

Страницы