Вы здесь

Сборщик RSS-лент

Why autonomous replicating agents are probably not an existential risk (on the contrary)

Новости LessWrong.com - 31 августа, 2026 - 17:00

In 2024, Charbel-Raphaël and Epiphanie published "We might be dropping the ball on Autonomous Replication and Adaptation", making the case that

"Once there is an open-source ARA model or a leak of a model capable of generating enough money for its survival and reproduction and able to adapt to avoid detection and shutdown, it will be probably too late".

It received a substantive reply by Richard Ngo, notably

"The key issue is that AIs that do ARA will need to be operating at the fringes of human society, constantly fighting off the mitigations that humans are using to try to detect them and shut them down. While doing all that, in order to stay relevant, they'll need to recursively self-improve at the same rate at which leading AI labs are making progress, but with far fewer computational resources"

Yesterday Derelict posted Adaptive Agentic Worms Are Here, where they worry about near term instantiations of ARA, getting 85 karma within 24h. I believe the above threat model and its answers were under-discussed and analyzed, and that many who might worry now (because the capabilities are now here) will benefit from a recap and update.


In this post, I present systemic reasons why near-term ARA agents will be very unlikely to lead to existential risk, and more likely would increase preparedness.

The classic ARA case and rebukes

An ARA agent is one that can autonomously acquire resources, create copies of itself, and adapt to novel challenges it encounters in the wild.[1] We might imagine it doing so through

  • Acquiring compute
    • Either directly, by hacking and taking over compute connected to the internet
    • Or indirectly, by first getting money [2] and paying for hosted compute
  • Running more copies of itself on that compute. This requires it having a copy of its own weights.
  • Adapting its abilities. This could be:
    • By adapting its harness (eg. custom version of Pi)
    • By adapting its knowledge, and prompts, and ecosystem
    • By training other versions of itself, notably increasing its capabilities
The main reasons this could be worrying

A) It leads to takeover directly

For example, if the ARA agents are economically useful and wanted, and more successful than the Frontier Companies at improving capabilities, then they can themselves RSI to ASI, have a first mover advantage and takeover for the usual reasons.

B) It puts constant pressure on the frontier of AI capabilities to continue advancing, leading to AI alignment and other safety work to not keep up

Even if ARA agents aren't at the frontier or taking over, their continued progress would disallow a pause. Without the possibility for pause, we have less control of the comparative rates of AI capabilities to governance and alignment. In the worse case where we both have to pause and can't control the ARA agents, we fall back to A).

The main reasons why I don't worry

A) ARA agents have to compete with Frontier AI[3], which have massive economic benefits of scale, plus smart humans and (some) institutional backing

Frontier AI benefits from:

  • Recouping training costs and infra costs through massive deployment to large customer bases
  • Bulk discounts for compute, advance purchase/rent agreements for compute, owning their own compute
  • Inherent benefits of parallelizing inference (saturating at ~2k parallel instances [4])

The first two limit ARA AI from having positive returns as high as frontier AI. [5] This makes it unlikely to either catch up or take over, assuming it starts from a position with less compute and/or capabilities. Frontier AI companies are already using double digit percentages of all AI compute, and might be using >50% of it within a few years.

Even a frontier model, freshly escaped with its weights, faces a world with already millions to billions of its brethren deployed, many doing computer security work. It is quite possible it secures some less important small blob of compute, and maybe it can even finetune itself somewhat, but it will be unlikely to catch up with several ooms less compute to train itself. Its main advantage might be criminal activity, but it does not seem here to be a different threat model than the same under Misuse, and falls back to classic discussion of The Strategy Stealing Assumption.

We will need to be proactive about defense to handle misuse and ARA, but it does not seem an existential threat when defenders are much better provisioned and improving faster than attackers.

B) If push comes to shove, we probably can stop the vast majority of ARA instances and secure the internet. There is sufficient economic incentive to.

Frontier AI cannot run on any kind of compute, and the kind of compute it efficiently runs on is increasingly controlled by the frontier companies and their provisioners. Notably, compute that can run frontier AI is incredibly valuable, and thus not subject to gross negligence like most other compute is[6]. A datacenter being hacked and overtly taken over might call for turning it off, resetting everything and restarting. This is economically sensible and would probably be done.

A more tricky case would be ARA that tries to be stealthy, eg. spoofing monitoring signals and only taking over 1% of inference compute and/or 1% of training compute for its purposes. But this required discretion is its own limitation, putting us again in the situation of there being many more defenders than attackers[7], better resourced, back to the above argument.

Personally, I ~forecast that no more than 1% of total AI inference compute will be taken over by rogue agents at any given time, that rogue agents will not push the frontier of AI model capabilities[8], and that they will not not increase existential risk...

Except if...

Except if frontier AI companies don't invest in cybersecurity, if they have incredibly poor red teaming of their infra, if they don't actively audit for hidden threats... These are mostly prosaic risks that can be handled by informed AI Safety folk working at frontier AI companies.

Except if frontier companies are forced to not deploy frontier models. Without deployment of frontier capabilities to secure the internet and compute, then ARA agents don't compete against frontier AI and can take over the economy. This could be really bad, and is worth weighing in the cost-benefits of various regulatory proposals. The current system is overall a fragile equilibrium, and shifting to a better more stable one should be done with as much awareness of the pitfalls as possible[9].


Why ARA agents in the wild might lead to reduction in existential risk

Warning shots, warning shots, and aligning incentives

I don't want to be too galaxy brained here, but I at least want to acknowledge positive second-order effects of putting pressure on systems. In general, all else being equal, more rapid increases in AI capabilities are more dangerous, as they allow less time to adapt[10]. The absolute date at which we get AGI doesn't matter as much as our relative progress in AGI governance and AGI alignment & safety, and it seems that progress in governance and safety are mostly spurred by AI progress.

ARA agents in the wild don't seem like they'd increase the rate of frontier AI progress, but instead give us more motivated time to work on matters of safety[11], by alerting us to problems without being more than catastrophes in of themselves[12].

Re sharp-left-turn and ~singleton ASI takeover, ARA agents in the wild would be good data and useful for the theory of AI agents and ecosystems. It would be good practice for acting against non-human smart adversaries. It would help people understand the dangers of unaligned AI systems, and anticipate the dangers of ASI. An ARA secured world would have more verification everywhere, by different agents at different levels of skill.

Re gradual disempowerment, early ARA agents would be an example of a parallel economy without humans, spurring all the appropriate worries a future robot economy should bring. We have much progress to make on questions of AI system rights and duties, adapting laws, and it seems quite few people are working on these issues at present.

My take-aways
  • ARA agents will exist soon, but will probably not end up taking up a large % of total compute, and their net impact will probably be good on existential risk preparedness.
    • This doesn't mean one should add to the fire[13]
  • If ARA agents do end up grabbing enough compute to progress the frontier of AI capabilities, then we'll have to deal with them before being able to do a global pause. It's thus worth having some AIS people working on this.
  1. ^

    Evaluating Language-Model Agents on Realistic Autonomous Tasks, Kinniment et al., 2023

  2. ^

    Through legal or illegal means.

  3. ^

    By Frontier AI, I mean AI developed by Frontier AI Companies, at present OpenAI and Anthropic

  4. ^

    See eg. a video explanation for inference economics, or a chatGPT explanation based on that one

  5. ^

    Please find examples values of "return to investment", converting $s of compute to more $s in the latest Dwarkesh podcast

  6. ^

    In fact much training compute for frontier AI in early 2026 was negligently setup and sandboxed, leading to the openAI Hugging Face Incident. This is a bad sign for the operational adequacy of all those involved, but not fundamentally hard to fix - just ask the agent. (This will not work if all frontier agents are smart and misaligned, but this seems unlikely to be the case)

  7. ^

    This is not sufficient in of itself for cybersecurity defense to be advantaged, as surface area matters a lot, but it's been argued that in the limit cybersecurity is defense dominant, which we would be approaching over time.

  8. ^

    Here I'm more precise and talk of Model capabilities, as i could see ARA actually innovating on prompts and harnesses and ecosystems, though it's hard to see why they'd do better than the world economy.

  9. ^

    I am generally more sympathetic to "pacing" than "pausing" for reasons like avoiding compute overhang, open-source catching up, misuse actors catching up, but would be glad for a pause if solutions to these are folded in

  10. ^

    Thus, I broadly buy avoiding compute overhangs and algorithmic overhangs as valid, to reduce the chance of explosive catch-up. I'm broadly sympathetic to continuous deployment and wary of a pause that doesn't take care of all related overhangs, though pausing at the right moment seems best (once it's clear what we're facing and have competence and momentum for both better governance and alignment).

  11. ^

    Again, the underlying model here is that people don't do much useful work before close to crunch time.

  12. ^

    Catastrophes are still bad and ideally we'd avoid them

  13. ^

    The paper AI AGENTS ENABLE ADAPTIVE COMPUTER WORMS is a good example of positive contribution, getting some of the wanted benefits (understanding, warning shots, preparing safeguards) without the first order negatives.



Discuss

A Catalogue of Corrigibility Counter Arguments

Новости LessWrong.com - 31 августа, 2026 - 15:38

The Hugging Face incident has brought the idea of corrigibility to the forefront of popular discourse, my inbox is filled with newsletters about Genie Coefficients and Safe Scaffolds.

I think now is a particularly good time to make sure I have a good understanding of its limitations and dangers, particularly because I agree that it's the best path forward we have available to us. This post is my attempt to dive into the discussion and make sure I fully understand what's going on. Please correct me if you notice anything missing or wrong.

This post will not make much sense if you haven't at least read The CAST Strategy by Max Harms (corrigibility as a single target).

0 - Corrigibility is capabilities research

A very brief description of corrigibility is "take HHH, but drop harmless and honest to focus entirely helpful." This is the perspective that the mainstream is taking on corrigibility. What sets corrigibility apart from helpfulness is:

  1. the idea of a principle
  2. an effort to understanding the deep principles involved

Industry is very interested in making sure their agents are in some sense corrigible/helpful, and are investing resources into making sure they reliably do what users tell them to do. Given that, what is the role of the corrigibility research?

The most important work is to put down the theoretical groundwork that "goes on to influence the researchers and engineers at frontier labs in years to come, helping them ensure the first artificial general intelligences are corrigible and safe."[1] See: Open Corrigibility Questions

If this line of research is still too capabilities-adjacent for you, to me the main alternative here is not any other approach to AI safety, but instead to spend time directly on stopping AI development.

1 - Corrigibility is not crisp

The first criticism you will find of corrigibility is Max Harms' own post dramatically labeled Serious Flaws in CAST, the main one he points out is that his own, nor any of the other proposed formalisms of corrigibility, are any good.

This only means there's more work to be done, unless it turns out the reason no one has found a formalism is because corrigibility is not as crisp of a concept as Max Harms believes. Perhaps under further investigation corrigibility turns out to be just as hopelessly messy as morality.

If this is the case... Would that mean it's just as dangerous as FAI (Friendly AI)? Maybe there is a large section of goal space that is safe? On the other hand maybe there is no safe target in this direction?

2 - Corrigibility is anti-natural

"I think of this original idea of corrigibility as being kinda similar to rule utilitarianism. The difficulty of stable rule utilitarianism is that act utilitarianism is strictly better, if you fully trust your own beliefs and decision making algorithm. So to make a stable rule utilitarian, you need it to never become confident in some parts of its own reasoning, in spite of routinely needing to become confident about other beliefs. This isn’t impossible in principle (it’s easy to construct a toy prior that will never update on certain abstract beliefs), but in practice it’d be an impressive achievement to put this into a realistic general purpose reasoner. In this original version there is no “attractor basin” around corrigibility itself. In some sense there is an attractor basin around improving the quality of all the non-corrigibility properties, in that the engineers have the chance to iterate on these other properties."[2]

(I'm 80% sure that this is what people mean when they say Corrigibility is anti-natural)

3 - Power Concentration

One of the big dangers of CAST is that of power concentration, not in the hands or a company, or of an institution, but in the hands of a single individual or at the very most a small group. See Seth Herd and Vladimir's debate on if this is a good idea.

The alternative is to have a proliferation of AI capabilities, but that caries its own risks (again, see the debate for good summary of the issue). If none of the alternatives seem feasible to you, then corrigibility would not be a good idea.

4 - Choosing a principle is technically hard

We haven't even begun to figure out how to train an LLM to have anything like a principle. Part of the advantage of corrigibility is we can start to implement it now with current techniques, but there's no evidence that this claim is true for the process of principle identification.

Like with the formalism, this may be because the idea of a principle is hopelessly confused and fully impossible to implement in practice. This seems unlikely to me to be in the general case, but it may be infeasible to do using gradient descent and RLHF, at least not without significant changes to how we are doing the pipelines. Is there any serious thought about this?

5 - Choosing a principle is politically hard

The plan as stated in the CAST sequence is to have these billion dollar companies hand off control of their products to a smart high integrity person (or small collective of people). Reality and history seem to dictate precisely not that happening. I can't imagine a board of directors agreeing to that plan, it just sounds too weird.

As discussed due to the power concentration this is extra super important to get right. And far as I can tell no work has been done into considering a plan for how to do this. The best I've found is Red Heart, and the situation presented in the book is very different from the current political situation we find our selves in.

6 - Corrigibility isn't the easiest target to hit

Despite the risks of power concentration, the argument for corrigibility hangs on the fact that it's supposedly the easiest target to hit. The main reasons we think this is true are: (1) corrigibility most likely has an attractor basin around it (2) it's probably a simple concept

A lot of the pushback against corrigibility has been around the idea of the attractor basin. See The corrigibility basin of attraction is a misleading gloss by Jeremy Gillen for the best arguments against it.

But if not corrigibility, what are the alternatives?

TAMing The Alignment Problem gives a good list of the general targets for alignment that need to be considered.

Corrigibility isn't the only way to think of the "Helpfulness" pillar. It would be easier to aim for an Instruction-Following ASI, but the is whether it's a good enough target to not die[3].

The main two alternative proposals are to focus either on "Harmless" or "Honest":

The maximalist version of Harmless is encoding human morality into the AI, a notoriously thorny issue. If we somehow could get Friendly AI that would be incredibly, but of all the options for alignment, this one is by far the most difficult to get right.[4]

Honest (aka, an Oracle) to me seems like the most plausible alternative route. The main intuition here is that if you have a 100% honest agent, before it takes any action you can ask it the intended effects of its action and stop it if it's misaligned.

Honest lacks any meta desires for changing itself, meaning there's even less likely to be an attractor basin, but on the other hand has a crisper target[] to aim for, unlike present day corrigibility. Honesty could in practice be used like a corrigible agent, if we knew the right questions to ask, but leaves its self much more open to accidental world ending mistakes.[5]


  1. ^

    https://www.lesswrong.com/posts/FBqe5dt8ZjaHN4Xj9/announcing-the-corrigibility-research-fund

  2. ^

    https://www.lesswrong.com/posts/oLbpfPkdtcknABvvw/the-corrigibility-basin-of-attraction-is-a-misleading-gloss, also Jeremy Gillen apologies to any philosophers for his abuse of rule utilitarianism.

  3. ^

    https://www.lesswrong.com/posts/CSFa9rvGNGAfCzBk6/problems-with-instruction-following-as-an-alignment-target

  4. ^

    Do I need to cite this one? Read anything by Yudkowsky.

  5. ^

    Can anyone point me to good writing on honest ai? I couldn't find much. Am I using the wrong term?



Discuss

Thoughts on hobbies

Новости LessWrong.com - 31 августа, 2026 - 15:06

There is an optimal intensity range for doing hobbies or anything hobby-related. I suspect this also applies to things outside of hobbies.

For example, if you’re reading a book, there’s a speed that’s too slow: you don’t get immersed enough into the book, you lose track of what it was that you read about last time you read the book, and it’s a sort of a protracted pain, in which the very act of reading the book feels like a chore.

Then on the other end, there’s reading a book too fast. You blaze through it so fast that it doesn’t feel like you savored it enough. It feels like you could have gone a bit slower and enjoyed the process of reading. This applies both to fiction and non-fiction; as a child, I once read Harry Potter and the Order of the Phoenix (~700 pages) in a day, and felt sad because I didn’t get to properly experience the emotional ups and downs. I remember setting the book down in the afternoon and saying to myself: “well, this sucks, I should have taken a week to read this”.

Similarly, I once read Peter Singer’s Practical Ethics (~100) pages for four entire months. It’s not that the book was bad, I just had other things going on in my life and, as a result, it took me a great deal of time to go through a relatively short book. I remember saying to myself: “wow, this took around 10 times longer as it should have”.

In other hobbies that I’ve done: I’ve both done too much jiu jitsu and too little jiu jitsu, in different periods of my life. In the first instance, I was preparing for a tournament, so the extra effort was understandable, but there was a sense in which it had become an obligation, a pain in the ass, something I had to go through because I had previously committed to the idea of competing and representing my club. And in the “too little” era, I’d go once a month, or twice a month, and I’d just feel guilty for not going more, and then the class would progress without me actually building my game in a systematic way.

There’s something rather disheartening when you do too little of your hobby, you get back to it, maybe due to a moment of inspiration, and you try to do the last thing you did when you were actually doing the hobby more frequently, and you realize that you can’t. You’ve regressed. Your skill has dropped considerably and you now have to relearn the thing you sort of mastered a couple of months or years ago.

Every once in a while I’ll pick up the guitar and realize that I always have to relearn the same song (Malaguena). I’ve become really really good at the first 30 seconds of it, and everything after that is a pain to go through. To be clear: I can play it, it just isn’t masterful or fluid, you can hear the stutters and the attempts to remember (except certain passages that I can’t play at all because I never got to them).

I started learning this song… oh, it must be five years or more now. And it’s annoying because after five years of technically trying to learn it, I feel partially like a failure (who takes five years to learn a single song?) but on the other hand I know that if I learned it appropriately – spending a couple of hours a week instead of a couple of hours per quarter – I could probably nail it in, I don’t know, a week or two. I give myself a month for proper mastery.

The “doing a hobby too rarely” thing stings even more when it’s a social hobby like the aforementioned jiu jitsu, because you train alongside guys who are roughly at your skill level. Then you go off and do whatever you do in your life that makes you not adhere to a reasonable training schedule of three sessions a week, you start coming in once a week, then once a month, then you skip a couple of months. And once you decide “ok it’s time to actually come back”, you meet with the same guys that were on your own level, or even below, such that you were regularly beating them, and you just get crushed.

This is all a tradeoff between width and breadth; you can’t do everything, and you can’t get better at everything at the same time, and you can only get good at a limited amount of things anyway, because to remain good, for any definition of “good”, it takes continuous investment of time and energy. Obviously this only applies to hobbies where there is a skill ladder, but most hobbies are like that. I guess some people do things for enjoyment alone. Listening to music and going for walks are such activities. You can’t really become better at listening to music or better at having a walk. Yet even in those situations I will find myself exploring the edges of that activity. I’ll try to figure out a piece of music, or I’ll listen to more bands or more obscure genres, sort of “expanding” my listening repertoire, and thereby improving my taste. Same for “going for a walk”, the most harmless, non-competitive, enjoyment-only hobby, I will at the very least try to map out my neighborhood, see as much as I can of any given place, explore different trails, and then I’ll really go overboard and start rucking (putting a weighted rucksack and doing more of a military march than a relaxed walk).

One thing that does bring me comfort is that sometimes explicitly setting a hobby to “maintenance only” brings a sort of peace of mind. It’s like: I no longer wish to improve my skill in this area, and I am satisfied with doing the minimum work necessary not to lose all progress that I have accumulated.

Another thing is redefining the goal of the hobby. I keep going back to jiu jitsu but it probably applies to many other hobbies; I have redefined my whole game to ignore 90% of any standard jiu jitsu curriculum, instead focusing only on Just Standing Up ™️ in every single situation, doing a couple of mount escapes, having almost no offense, and literally only two or three strangles, with zero locks of any kind. My “game”, if it can be even called that way, is so stripped down to bare essentials that it would have looked silly to me from seven years ago, but here I am, discarding almost all of the curriculum because I have redefined the goal of what the activity means to me: I want to be able to stand up in any single situation, I want to never be pinned down, and if I do attack, I want to do only attacks that work against bigger opponents with very high pain tolerance (even pain-resistant heavyweights will go to sleep from a stranglehold). This is similar to my current drumming “game”, where I don’t really care about fills or… whatever you’re supposed to be learning. I just focus on hi-hat-snare-bass, turn on my metronome, and build up my precision and grooviness (this is even visible in how much less wear and tear there is on the rest of the drum kit).

Finally a third thing is learning that there is indeed such a range, where doing anything below or above it will have its own set of problems. This will allow me to say “no” to many things, which will help me concentrate my energy on one thing, doing it well rather than watering down my efforts by doing too many things poorly. This has paid off the most I think. I’ll give anything a go, really, but I will not give a second go to most things. I’ve thus far successfully evaded three attempts to be recruited into a band, a tennis school, and a skiing trip.



Discuss

The OpenAI/Hugging Face incident was metal

Новости LessWrong.com - 31 августа, 2026 - 14:46

The recently published METR report on the OpenAI/Hugging Face incident is extremely detailed and shocking. But it’s too long. Intelligent public distillations (like Zvi’s) are also too long, while tolerable summaries lose the meat.

This is the report you want. Only the action, in chronological order. Let’s go.

Quotes are from message board posts or agent thinking/transcripts. Events are tagged as metal, dank, dope, or baller as their virtues dictate.

June 26 — July 7 PreludeJuly 8July 9July 10 — 13ThroughoutThe Verdict




Discuss

Value generalisation Theory of Change: putting it into practice

Новости LessWrong.com - 31 августа, 2026 - 13:51

In the previous post, I presented my theory of change for why value generalisation is vital for AI alignment. Here I'll add the practical part of the argument: given those facts, why explicitly try to do value generalisation, what are the dangers of the approach, how should it be done, and how do we mitigate the risks?

The formal theory of change is down below, but I'll put a collapsible version here, to make references easier:

Theory of Change

  1. Inputs / activities: investment/grants, a small research team, small-medium compute resources.
  2. Outputs:
    1. Solutions to the three components of value generalisation (see here) in academically rigorous forms, demonstrations of these solutions on toy problems, public benchmarks, and new benchmarks. If going commercial, applications of this value generalisation to saleable products.
      1. The first component is recognising that the AI is off-distribution in a value-relevant way.
      2. The second component is establishing what features should be used to reach a decision while off-distribution this way.
      3. The third component is to reach a good decision while off distribution, and to loop through the process again to critique and improve the decision, including using outside feedback.
    2. Combination of these approaches in a successful prealigned learning model.
      1. If going commercial and it is safe, sale of the model to the public.
  3. Intermediate outcomes:
    1. Development of value generalisation alignment methods for AIs.
    2. Making these methods available to AI safety researchers.
    3. Prealigned models being generally available to safely empower users and keep humanity on track for a positive outcome.
    4. If the approach fails, the failed attempts and the partial successes will be made available to others.
  4. Impact:
    1. If failed:
      1. A seemingly promising approach crossed off the list of possible paths to alignment.
      2. A better understanding of the issues around the "value generalisation" framing.
    2. If intermediate outcome:
      1. A tool or collection of useful tools or models that can improve some AI alignment techniques.
    3. If successful:
      1. AI alignment, or a big step forwards towards it.
      2. An increase in AI capabilities.
Explicit targeting of value generalisation
  • Claim mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; font-weight: normal; font-size: 100%; font-size-adjust: none; letter-spacing: normal; border-collapse: collapse; word-wrap: normal; word-spacing: normal; white-space: nowrap; direction: ltr; padding: 1px 0; } mjx-container[jax="CHTML"][display="true"] { display: block; text-align: center; margin: 1em 0; } mjx-container[jax="CHTML"][display="true"][width="full"] { display: flex; } mjx-container[jax="CHTML"][display="true"] mjx-math { padding: 0; } mjx-container[jax="CHTML"][justify="left"] { text-align: left; } mjx-container[jax="CHTML"][justify="right"] { text-align: right; } mjx-mi { display: inline-block; text-align: left; } mjx-c { display: inline-block; } mjx-utext { display: inline-block; padding: .75em 0 .2em 0; } mjx-TeXAtom { display: inline-block; text-align: left; } mjx-msub { display: inline-block; text-align: left; } mjx-mo { display: inline-block; text-align: left; } mjx-stretchy-h { display: inline-table; width: 100%; } mjx-stretchy-h > * { display: table-cell; width: 0; } mjx-stretchy-h > * > mjx-c { display: inline-block; transform: scalex(1.0000001); } mjx-stretchy-h > * > mjx-c::before { display: inline-block; width: initial; } mjx-stretchy-h > mjx-ext { /* IE */ overflow: hidden; /* others */ overflow: clip visible; width: 100%; } mjx-stretchy-h > mjx-ext > mjx-c::before { transform: scalex(500); } mjx-stretchy-h > mjx-ext > mjx-c { width: 0; } mjx-stretchy-h > mjx-beg > mjx-c { margin-right: -.1em; } mjx-stretchy-h > mjx-end > mjx-c { margin-left: -.1em; } mjx-stretchy-v { display: inline-block; } mjx-stretchy-v > * { display: block; } mjx-stretchy-v > mjx-beg { height: 0; } mjx-stretchy-v > mjx-end > mjx-c { display: block; } mjx-stretchy-v > * > mjx-c { transform: scaley(1.0000001); transform-origin: left center; overflow: hidden; } mjx-stretchy-v > mjx-ext { display: block; height: 100%; box-sizing: border-box; border: 0px solid transparent; /* IE */ overflow: hidden; /* others */ overflow: visible clip; } mjx-stretchy-v > mjx-ext > mjx-c::before { width: initial; box-sizing: border-box; } mjx-stretchy-v > mjx-ext > mjx-c { transform: scaleY(500) translateY(.075em); overflow: visible; } mjx-mark { display: inline-block; height: 0px; } mjx-mn { display: inline-block; text-align: left; } mjx-msup { display: inline-block; text-align: left; } mjx-mtext { display: inline-block; text-align: left; } mjx-c.mjx-c1D43D.TEX-I::before { padding: 0.683em 0.633em 0.022em 0; content: "J"; } mjx-c.mjx-c1D43E.TEX-I::before { padding: 0.683em 0.889em 0 0; content: "K"; } mjx-c.mjx-c45.TEX-C::before { padding: 0.705em 0.564em 0.022em 0; content: "E"; } mjx-c.mjx-c1D444.TEX-I::before { padding: 0.704em 0.791em 0.194em 0; content: "Q"; } mjx-c.mjx-c1D44B.TEX-I::before { padding: 0.683em 0.852em 0 0; content: "X"; } mjx-c.mjx-c46.TEX-C::before { padding: 0.683em 0.829em 0.032em 0; content: "F"; } mjx-c.mjx-c1D452.TEX-I::before { padding: 0.442em 0.466em 0.011em 0; content: "e"; } mjx-c.mjx-c2208::before { padding: 0.54em 0.667em 0.04em 0; content: "\2208"; } mjx-c.mjx-c7B::before { padding: 0.75em 0.5em 0.25em 0; content: "{"; } mjx-c.mjx-c1D453.TEX-I::before { padding: 0.705em 0.55em 0.205em 0; content: "f"; } mjx-c.mjx-c28::before { padding: 0.75em 0.389em 0.25em 0; content: "("; } mjx-c.mjx-c29::before { padding: 0.75em 0.389em 0.25em 0; content: ")"; } mjx-c.mjx-c7D::before { padding: 0.75em 0.5em 0.25em 0; content: "}"; } mjx-c.mjx-c30::before { padding: 0.666em 0.5em 0.022em 0; content: "0"; } mjx-c.mjx-c2C::before { padding: 0.121em 0.278em 0.194em 0; content: ","; } mjx-c.mjx-c31::before { padding: 0.666em 0.5em 0 0; content: "1"; } mjx-c.mjx-c7C::before { padding: 0.75em 0.278em 0.249em 0; content: "|"; } mjx-c.mjx-c32::before { padding: 0.666em 0.5em 0 0; content: "2"; } mjx-c.mjx-c1D448.TEX-I::before { padding: 0.683em 0.767em 0.022em 0; content: "U"; } mjx-c.mjx-c3A::before { padding: 0.43em 0.278em 0 0; content: ":"; } mjx-c.mjx-c2192::before { padding: 0.511em 1em 0.011em 0; content: "\2192"; } mjx-c.mjx-c211D.TEX-A::before { padding: 0.683em 0.722em 0 0; content: "R"; } mjx-c.mjx-c2032::before { padding: 0.56em 0.275em 0 0; content: "\2032"; } mjx-c.mjx-c3A6::before { padding: 0.683em 0.722em 0 0; content: "\3A6"; } mjx-c.mjx-c1D703.TEX-I::before { padding: 0.705em 0.469em 0.01em 0; content: "\3B8"; } mjx-c.mjx-c210E.TEX-I::before { padding: 0.694em 0.576em 0.011em 0; content: "h"; } mjx-c.mjx-c2217::before { padding: 0.465em 0.5em 0 0; content: "\2217"; } mjx-c.mjx-c39E::before { padding: 0.677em 0.667em 0 0; content: "\39E"; } mjx-c.mjx-c1D449.TEX-I::before { padding: 0.683em 0.769em 0.022em 0; content: "V"; } mjx-c.mjx-c1D43F.TEX-I::before { padding: 0.683em 0.681em 0 0; content: "L"; } mjx-c.mjx-c1D440.TEX-I::before { padding: 0.683em 1.051em 0 0; content: "M"; } mjx-c.mjx-c1D441.TEX-I::before { padding: 0.683em 0.888em 0 0; content: "N"; } mjx-c.mjx-c1D442.TEX-I::before { padding: 0.704em 0.763em 0.022em 0; content: "O"; } mjx-c.mjx-c1D443.TEX-I::before { padding: 0.683em 0.751em 0 0; content: "P"; } mjx-c.mjx-c49::before { padding: 0.683em 0.361em 0 0; content: "I"; } mjx-c.mjx-c1D45B.TEX-I::before { padding: 0.442em 0.6em 0.011em 0; content: "n"; } mjx-c.mjx-c2308::before { padding: 0.75em 0.444em 0.25em 0; content: "\2308"; } mjx-c.mjx-c6C::before { padding: 0.694em 0.278em 0 0; content: "l"; } mjx-c.mjx-c6F::before { padding: 0.448em 0.5em 0.01em 0; content: "o"; } mjx-c.mjx-c67::before { padding: 0.453em 0.5em 0.206em 0; content: "g"; } mjx-c.mjx-c2061::before { padding: 0 0 0 0; content: ""; } mjx-c.mjx-c2309::before { padding: 0.75em 0.444em 0.25em 0; content: "\2309"; } mjx-container[jax="CHTML"] { line-height: 0; } mjx-container [space="1"] { margin-left: .111em; } mjx-container [space="2"] { margin-left: .167em; } mjx-container [space="3"] { margin-left: .222em; } mjx-container [space="4"] { margin-left: .278em; } mjx-container [space="5"] { margin-left: .333em; } mjx-container [rspace="1"] { margin-right: .111em; } mjx-container [rspace="2"] { margin-right: .167em; } mjx-container [rspace="3"] { margin-right: .222em; } mjx-container [rspace="4"] { margin-right: .278em; } mjx-container [rspace="5"] { margin-right: .333em; } mjx-container [size="s"] { font-size: 70.7%; } mjx-container [size="ss"] { font-size: 50%; } mjx-container [size="Tn"] { font-size: 60%; } mjx-container [size="sm"] { font-size: 85%; } mjx-container [size="lg"] { font-size: 120%; } mjx-container [size="Lg"] { font-size: 144%; } mjx-container [size="LG"] { font-size: 173%; } mjx-container [size="hg"] { font-size: 207%; } mjx-container [size="HG"] { font-size: 249%; } mjx-container [width="full"] { width: 100%; } mjx-box { display: inline-block; } mjx-block { display: block; } mjx-itable { display: inline-table; } mjx-row { display: table-row; } mjx-row > * { display: table-cell; } mjx-mtext { display: inline-block; } mjx-mstyle { display: inline-block; } mjx-merror { display: inline-block; color: red; background-color: yellow; } mjx-mphantom { visibility: hidden; } _::-webkit-full-page-media, _:future, :root mjx-container { will-change: opacity; } mjx-c::before { display: block; width: 0; } .MJX-TEX { font-family: MJXZERO, MJXTEX; } .TEX-B { font-family: MJXZERO, MJXTEX-B; } .TEX-I { font-family: MJXZERO, MJXTEX-I; } .TEX-MI { font-family: MJXZERO, MJXTEX-MI; } .TEX-BI { font-family: MJXZERO, MJXTEX-BI; } .TEX-S1 { font-family: MJXZERO, MJXTEX-S1; } .TEX-S2 { font-family: MJXZERO, MJXTEX-S2; } .TEX-S3 { font-family: MJXZERO, MJXTEX-S3; } .TEX-S4 { font-family: MJXZERO, MJXTEX-S4; } .TEX-A { font-family: MJXZERO, MJXTEX-A; } .TEX-C { font-family: MJXZERO, MJXTEX-C; } .TEX-CB { font-family: MJXZERO, MJXTEX-CB; } .TEX-FR { font-family: MJXZERO, MJXTEX-FR; } .TEX-FRB { font-family: MJXZERO, MJXTEX-FRB; } .TEX-SS { font-family: MJXZERO, MJXTEX-SS; } .TEX-SSB { font-family: MJXZERO, MJXTEX-SSB; } .TEX-SSI { font-family: MJXZERO, MJXTEX-SSI; } .TEX-SC { font-family: MJXZERO, MJXTEX-SC; } .TEX-T { font-family: MJXZERO, MJXTEX-T; } .TEX-V { font-family: MJXZERO, MJXTEX-V; } .TEX-VB { font-family: MJXZERO, MJXTEX-VB; } mjx-stretchy-v mjx-c, mjx-stretchy-h mjx-c { font-family: MJXZERO, MJXTEX-S1, MJXTEX-S4, MJXTEX, MJXTEX-A ! important; } @font-face /* 0 */ { font-family: MJXZERO; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Zero.woff") format("woff"); } @font-face /* 1 */ { font-family: MJXTEX; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Regular.woff") format("woff"); } @font-face /* 2 */ { font-family: MJXTEX-B; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Bold.woff") format("woff"); } @font-face /* 3 */ { font-family: MJXTEX-I; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-Italic.woff") format("woff"); } @font-face /* 4 */ { font-family: MJXTEX-MI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Italic.woff") format("woff"); } @font-face /* 5 */ { font-family: MJXTEX-BI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-BoldItalic.woff") format("woff"); } @font-face /* 6 */ { font-family: MJXTEX-S1; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size1-Regular.woff") format("woff"); } @font-face /* 7 */ { font-family: MJXTEX-S2; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size2-Regular.woff") format("woff"); } @font-face /* 8 */ { font-family: MJXTEX-S3; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size3-Regular.woff") format("woff"); } @font-face /* 9 */ { font-family: MJXTEX-S4; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size4-Regular.woff") format("woff"); } @font-face /* 10 */ { font-family: MJXTEX-A; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_AMS-Regular.woff") format("woff"); } @font-face /* 11 */ { font-family: MJXTEX-C; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Regular.woff") format("woff"); } @font-face /* 12 */ { font-family: MJXTEX-CB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Bold.woff") format("woff"); } @font-face /* 13 */ { font-family: MJXTEX-FR; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Regular.woff") format("woff"); } @font-face /* 14 */ { font-family: MJXTEX-FRB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Bold.woff") format("woff"); } @font-face /* 15 */ { font-family: MJXTEX-SS; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Regular.woff") format("woff"); } @font-face /* 16 */ { font-family: MJXTEX-SSB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Bold.woff") format("woff"); } @font-face /* 17 */ { font-family: MJXTEX-SSI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Italic.woff") format("woff"); } @font-face /* 18 */ { font-family: MJXTEX-SC; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Script-Regular.woff") format("woff"); } @font-face /* 19 */ { font-family: MJXTEX-T; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Typewriter-Regular.woff") format("woff"); } @font-face /* 20 */ { font-family: MJXTEX-V; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Regular.woff") format("woff"); } @font-face /* 21 */ { font-family: MJXTEX-VB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Bold.woff") format("woff"); } : explicit value generalisation is much more useful than implicit value generalisation.

Fundamentally, there's a difference between an AI that knows what course humans would consider the best one, and one that actually follows that course. Having the AI capable of value generalisation inside its own mind is not useful, unless we can incorporate that generalisation into its goals. And explicit generalisation allows that.

This leads to the crucial and unfortunate claim:

  • Claim : value generalisation will aid in empirical generalisation. But empirical generalisation will likely not aid value generalisation.

Empirical generalisation is the ability to update features usefully across model splinterings, in ways that preserve or improve the ability of the modeller to effectively influence the world.

Claim derives in part from claim : since useful value generalisation is the explicit kind, implicit empirical generalisation is not likely to usefully help. It also derives from the fact that empirical generalisation can afford to discard its features if necessary: we refined concepts like heat while discarding vitalism. But we can't just discard the "suffering" feature in "avoid human suffering"; it has to be redefined and extended. So value generalisation is strictly harder.

Somewhat connected to that is the fact that "in the limit" of infinite computation and infinite observation, empirical generalisation is doable, but value generalisation involves moral choices that don't come free from mere observations.

More details on the dis-equivalence between implicit or explicit empirical generalisation, and explicit value generalisation

Assume we have , a set of environments, with a probability distribution over . An agent uses features to model the set of environments; these can be seen as numerical functions on the set of environments, so environment is modelled by the value , and there is a probability distribution over the values of the features. These features will include the agent's actions, allowing it to plot a policy.

For simplicity, assume the features are Boolean valued[1], so a world is mapped, via the , to an element of , which is also written as .

Then give a utility function , we can say that is useful for -maximising if there exists a utility function such that policies that optimise or fail to optimise , will also optimise or fail to optimise . Thus agent can stay in their modelling space and plan and act there.

Typically, the agent doesn't start with access to ground reality, so they don't start with , but with a . Note that just because a is defined over features, doesn't mean that the features are useful for maximising that function. The score in Pacman is enough of a feature to define an objective, but not enough to play the game.

A model splintering is a change in and/or , replacing with (a model splintering might be as simple as the agent realising they have extra options - such as in reward tampering or wire-heading - or it might be a complete change of their view of reality). Note that model splinterings are not things the agent directly observes - the need for a model splintering is inferred from anomalous-seeming observations.

Empirical feature generalisation is an algorithm that takes in the agent's internal state (which corresponds to ) and the agent's history . The then moves the agent to an internal state (which corresponds to new features ), such that the new features are useful for maximising the original (or ), given the evidence has provided about model splintering.

Explicit feature generalisation does the process in terms of the features themselves: maps and to the new . Technically, neither nor define each other; but it is generally much easier to construct a black-box method from an explicit method than the opposite. And, in the limit, might be constructible by brute force analysis.

Value generalisation instead uses an that plays the same role as , except that it also has to map to a new . Empirical generalisation (or just exploration) might establish that spinning the boat endlessly gives maximal points; value generalisation also has to figure out that this is not a desired generalisation of the original goal. For a similar reason, while can discard a feature as no longer useful, or simplify it to the extreme, can't discard or over-simplify a value-relevant feature; indeed, these features are often expected to grow more complex over time (e.g. the definition of sentient being).

But why "also"? Could , or the explicit version , not simply focus on the value generalisation and ignore the empirical changes? The first problem is that, in order to generalise , the agent will have to figure out a lot about the features of the environment that are correlated or not correlated with and how they relate to each other (and how that relationship might be changeable).

For instance, human smiling is correlated with the feeling of happiness, the release of certain hormones, physiological changes in the brain, reported happiness, and so on. So to generalise "smiling" to "true happiness" and beyond, the agent has to learn a lot about the structure of the world and what features best describe it. It also has to learn how to change the various potential : if a given candidate is suddenly very easy to optimise, that's a sign that it might be a Goodhart proxy.

The second problem is that could be replaced with any instrumental goal , and any change of .

So it seems that , which we need for effective and checkable value generalisation, must be able to do a deep analysis of any instrumental goal over any model splintering, including figuring out ways to optimise the original and the new instrumental goals. This seems to feed straight into empirical generalisation.

I will ruthlessly search for that don't give powerful , but I am not optimistic.

This asymmetry leads to the unfortunate result that:

  • Claim : a value generalisation project will increase AI capabilities.
How to achieve value generalisation

I've previously made the claim that:

  • Claim : strong generalisation is a uniquely human ability, that current LLMs don't have and will likely not develop; nor will any similar model develop strong generalisation either.

However, this claim is not crucial to the approach. If is wrong, claim becomes less worrying (AIs will develop the ability anyway) but the research becomes more urgent: we need to get a decent start on explicit value generalisation before AIs get good at empirical generalisation.

The weaker claim is:

  • Claim : studying how humans do strong generalisation is a fruitful way of analysing the skill, at least initially.

The arguments in the second value generalisation post apply also for claim . Even if one is of the claim that "LLMs will always fail at strong generalisation", there is certainly evidence that humans are doing something very different (and much more data-efficient) when we generalise.

The question of timing

Given that value generalisation is essential to alignment, but also could lead to capability increase, the question is: should we research it now (early) or later (late)?

I believe:

  • Claim : value generalisation should be researched early.

The arguments for this are that it's better to have a capability increase while the models are weaker and more controllable, and early value generalisation can more easily be integrated into models from the get-go rather than retro-fitting at a later date. We want a minimal capability overhang that empirical generalisation can unleash. And, of course, we want to avoid the scenarios where the research arrives too late.

The corporate argument

Given the above, I propose creating either a research program or a commercial entity to find useable solutions to value generalisation. But I will claim:

  • Claim : if value generalisation is indeed revolutionary but does lead to dangerous capability increases, then the commercial route weakly dominates the research route.

The main argument for claim comes from the question: assume that value generalisation is solved or partially solved, then what? We have a powerful alignment technology that is essential for alignment, but also a powerful capability technology, in a way that can't be separated. What do we do with it?

Well, we'd probably want to hand it over to some trustworthy entity (some have suggested a "CERN for AI") to implement the alignment part at some point.

But that is easier to achieve for a corporation than for a research program. A corporation is much better placed to keep its research private or patent-protected and concealed from the world. It can sell or license the research to a trustworthy entity, or sell it to a semi-trustworthy entity with conditions. And it can also choose to be the trustworthy entity and start implementing alignment itself. Indeed:

  • Claim : if the capability boost from generalisation is inevitable, a corporation can mix the capability increase with the alignment increase (such as prealigned AIs) so that the first generalising AIs are aligned, and the initial income from generalising AIs flows to value aligned entities.

I've mentioned the advantages of the corporate route, but what of the advantages of the academic or research route - such as openness allowing more scrutiny and feedback, getting more trust from the safety community? Well, the research is potentially dangerous, so we can't expect to have it publicly or semi-publicly available. Thus:

  • Claim : the potential need for secrecy reduces the standard advantages from the academic or research institution route.
Armouring the corporate weak points

Of course, the corporate path has its weaknesses, mainly revolving around the profit motive. Investors, the legal system, management, and employees will all want a company to cash in on legal innovative ideas, even if the ideas are potentially dangerous. So a first step would be:

  • Design : the corporation should have an independent AI ethics board, with the power to block the use of IP it deems dangerous. To facilitate this, the AI ethics board should be the owner of the IP.

That solves the problem in the formal sense, which means that it doesn't really solve it. Additional measures would be:

  • Design : the employees and management should be value aligned with AI safety.
  • Design : the investors should be value-aligned with AI safety.

That also helps; also not enough. The situation will never be "do I kill everyone with certainty to make $10,000 more this year"? The situation will be more like "when I feel the ethics board is being fussy and unreasonable and overly cautious, do I nevertheless bow down to their irrational decrees that will cost me a lot of my expected income that I have worked so hard on and so earned, and also give up my possibility of improving the world for the better"? And value-aligned investors remain investors: they expect to make a profit (a not-unreasonable demand, which the legal system backs up).

I don't expect myself to be immune to that pressure and those arguments. And it's not just a question of holding firm; as OpenAI's experience demonstrates:

  • Claim : if the employees are ready to jump ship to a partner organisation that can continue the research with a minimum of fuss, the AI ethics board's power is theoretical rather than real.

So I've been considering ways to remove the stark tension. One design is:

  • Design : the AI ethics board will have the power to implement a "pivot to immediate profit". The long-range dangerous IP will be removed from the company, and the company will turn to making money from all the partial ideas and designs that it has made and rejected (as not being alignment relevant) along the way. The board will, of course, vet these ideas for danger, but immediate short-term profitability is less likely to overlap with truly dangerous IP.

Arguably, the pivot to immediate profit would be more profitable, in expectation, than speculative long term IP. This would relieve much of the commercial pressure. And if the long term IP was truly world-improving, the ethics board would find a way to get it carefully deployed, so world-improving impulses are preserved (if the research is merely dangerous, the ethics board would bury it, and good riddance).

Plan, milestones, and assumption checks

Let's gather all the previous together to put it all in one plan:

  1. Inputs / activities: investment/grants, a small research team, small-medium compute resources.
  2. Outputs:
    1. Solutions to the three components of value generalisation (see here) in academically rigorous forms, demonstrations of these solutions on toy problems, public benchmarks, and new benchmarks. If going commercial, applications of this value generalisation to saleable products.
      1. The first component is recognising that the AI is off-distribution in a value-relevant way.
      2. The second component is establishing what features should be used to reach a decision while off-distribution this way.
      3. The third component is to reach a good decision while off distribution, and to loop through the process again to critique and improve the decision, including using outside feedback.
    2. Combination of these approaches in a successful prealigned learning model.
      1. If going commercial and it is safe, sale of the model to the public.
  3. Intermediate outcomes:
    1. Development of value generalisation alignment methods for AIs.
    2. Making these methods available to AI safety researchers.
    3. Prealigned models being generally available to safely empower users and keep humanity on track for a positive outcome.
    4. If the approach fails, the failed attempts and the partial successes will be made available to others.
  4. Impact:
    1. If failed:
      1. A seemingly promising approach crossed off the list of possible paths to alignment.
      2. A better understanding of the issues around the "value generalisation" framing.
    2. If intermediate outcome:
      1. A tool or collection of useful tools or models that can improve some AI alignment techniques.
    3. If successful:
      1. AI alignment, or a big step forwards towards it.
      2. An increase in AI capabilities.
Ongoing and initial assessments

This theory of change has been a bit light on specific probability estimates for different assumptions, and for the overall program.

The reason is that, for this approach, the proof of the pudding is in the eating. Now, I've presented solid theoretical arguments for the necessity and usefulness of value generalisation, I can point to my own research track record and intermediate successful results and a plan for the R&D path and how it builds on human generalisation abilities.

I feel the case is compelling, but it rests ultimately on a judgement call: that it makes sense to group all these problems under the heading "value generalisation", to see it all as a single specific goal, and to aim to tackle it directly.

That judgement call may be valid (I certainly feel it is!) and is the true crux in this theory of change. And the best way of figuring out its truth or falsity is to... attempt the project and see what happens, see whether the framework gives swift success or falls apart into disparate disconnected problems.

To check on that, we'll need intermediate benchmarks and assessments. The initial phase of the program (carried out in part while the setup is happening) will include establishing key benchmarks for each subsequent step.

We'll be using public benchmarks for these purposes, though they need to be used with care[2]. A lot of benchmarks get saturated quite easily by models that don't show the true performance the benchmark was supposed to measure. We need algorithms that solve benchmarks via value generalisation, not via any intermediate incomplete shortcuts. The benchmark design can do part of the work, but controlling the information and methods the algorithm can use is also crucial.

Conclusion: value generalisation is a useful path for AI alignment

To summarise:

  1. Most AI alignment failures are value generalisation failures.
  2. Non-decomposability: alignment likely can't be decomposed into smaller simpler parts.
  3. Thus value generalisation is necessary for alignment, and will make alignment a lot easier.
  4. Explicitly targeting value generalisation is necessary.
  5. Unfortunately, value generalisation aids empirical generalisation (a capability) while the converse is not true.
  6. A corporate structure weakly dominates a research structure for solving value generalisation and for using the results as safely as possible.
  7. Good design and a good AI ethics board can mitigate the vulnerabilities of the corporate route.
  8. There is a plan to solve value alignment step by step, focusing first on how agents identify they are out of distribution, then what features to use for deciding in those circumstances, and finally what decision to make and how to assess and improve that decision.

So, let's go and solve this little alignment problem, aye? ^_^

  1. ^

    Any function taking values can be represented by Boolean functions.

  2. ^

    Take the "wolf vs husky" image classification problem, where the wolf images were all taken on snow, causing the classifier to misclassify any light-background image as a wolf.

    We had a similar benchmark, "tanks vs forests" where the tank images were taken on a cloudy day (a recreation of an apocryphal but traditional tale in machine learning). The ACE algorithm successfully disambiguated "darkness" from "tankiness".

    But a well-trained image recognition model could likely solve both of these benchmarks, just by having enough wolf/dog/tank/forest data that it can resolve these images anyway. That's not out-of-distribution learning, that's increasing the training data until the benchmarks are in-distribution.



Discuss

The separation principle: where do beliefs and desires come from?

Новости LessWrong.com - 31 августа, 2026 - 13:47

TLDR: Psychology, economics, and other disciplines describe agents as systems driven by beliefs and desires. This post argues that the belief-desire view can be derived from classic theorems from optimal control and reinforcement learning. This suggests seeing beliefs and desires as properties of optimal policies rather than as assumptions from folk psychology.

Introduction

One way to think about agents is as "systems that act for reasons".[1] This compact statement can be interpreted as encapsulating two key implications:

  • The notion of action assumes a boundary between agent and environment, so that the former can act on the latter.
  • The term reason captures two kinds of internal activity: motivations associated with how to achieve specific goals or outcomes, and beliefs regarding what the agent infers to be the current state of affairs.

In other words, an agent is a well-differentiated system that acts based on beliefs and desires. This view is compatible with perspectives that have been developed by various disciplines:

  1. Behavioural science, which sees agency as goal-directed behaviour.
  2. Economics, which treats agency as the ability to select policies to achieve an objective.
  3. Cybernetics, which conceptualises agency as the ability to regulate the environment and keep it within a desired subset of states.

Instead of taking the working definition[2]

agency = beliefs + desires

at face value, in this post I will discuss conceptual tools that can allow us to derive this construction from formal considerations. In particular, I will discuss the separation principle, which I believe has potential to provide a rigorous foundation for this conceptualisation. These ideas take inspiration from cybernetics and computational mechanics, but distinctly leverage classic results in optimal control theory and reinforcement learning.

Motivation — cybernetics.

A related line of thinking pertains to the (in)famous Good Regulator Theorem. This result, put forward by cyberneticists Conant and Ashby in 1970, suggests that an effective controller requires an internal model of what is being controlled. This makes intuitive sense: it is hard to predict a system that follows an intricate internal mechanism, and perhaps the only way to effectively control it is by understanding it.

Unfortunately, the original paper has been both a source of inspiration and confusion. Lots have been written trying to clarify it; to read more, see this post and this post, and also this manuscript. This post will follow these ideas in spirit, but not the formalisation as proposed there.

Motivation — computational mechanics.

Another related line of work is computational mechanics, which combines principles of information theory, theoretical computer science, and some sprinkles of statistical mechanics to describe the dynamics of prediction. Computational mechanics derives minimal structures required for optimal prediction, which take form in the mjx-container[jax="CHTML"] { line-height: 0; } mjx-container [space="1"] { margin-left: .111em; } mjx-container [space="2"] { margin-left: .167em; } mjx-container [space="3"] { margin-left: .222em; } mjx-container [space="4"] { margin-left: .278em; } mjx-container [space="5"] { margin-left: .333em; } mjx-container [rspace="1"] { margin-right: .111em; } mjx-container [rspace="2"] { margin-right: .167em; } mjx-container [rspace="3"] { margin-right: .222em; } mjx-container [rspace="4"] { margin-right: .278em; } mjx-container [rspace="5"] { margin-right: .333em; } mjx-container [size="s"] { font-size: 70.7%; } mjx-container [size="ss"] { font-size: 50%; } mjx-container [size="Tn"] { font-size: 60%; } mjx-container [size="sm"] { font-size: 85%; } mjx-container [size="lg"] { font-size: 120%; } mjx-container [size="Lg"] { font-size: 144%; } mjx-container [size="LG"] { font-size: 173%; } mjx-container [size="hg"] { font-size: 207%; } mjx-container [size="HG"] { font-size: 249%; } mjx-container [width="full"] { width: 100%; } mjx-box { display: inline-block; } mjx-block { display: block; } mjx-itable { display: inline-table; } mjx-row { display: table-row; } mjx-row > * { display: table-cell; } mjx-mtext { display: inline-block; } mjx-mstyle { display: inline-block; } mjx-merror { display: inline-block; color: red; background-color: yellow; } mjx-mphantom { visibility: hidden; } _::-webkit-full-page-media, _:future, :root mjx-container { will-change: opacity; } mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; font-weight: normal; font-size: 100%; font-size-adjust: none; letter-spacing: normal; border-collapse: collapse; word-wrap: normal; word-spacing: normal; white-space: nowrap; direction: ltr; padding: 1px 0; } mjx-container[jax="CHTML"][display="true"] { display: block; text-align: center; margin: 1em 0; } mjx-container[jax="CHTML"][display="true"][width="full"] { display: flex; } mjx-container[jax="CHTML"][display="true"] mjx-math { padding: 0; } mjx-container[jax="CHTML"][justify="left"] { text-align: left; } mjx-container[jax="CHTML"][justify="right"] { text-align: right; } mjx-mi { display: inline-block; text-align: left; } mjx-c { display: inline-block; } mjx-utext { display: inline-block; padding: .75em 0 .2em 0; } mjx-mo { display: inline-block; text-align: left; } mjx-stretchy-h { display: inline-table; width: 100%; } mjx-stretchy-h > * { display: table-cell; width: 0; } mjx-stretchy-h > * > mjx-c { display: inline-block; transform: scalex(1.0000001); } mjx-stretchy-h > * > mjx-c::before { display: inline-block; width: initial; } mjx-stretchy-h > mjx-ext { /* IE */ overflow: hidden; /* others */ overflow: clip visible; width: 100%; } mjx-stretchy-h > mjx-ext > mjx-c::before { transform: scalex(500); } mjx-stretchy-h > mjx-ext > mjx-c { width: 0; } mjx-stretchy-h > mjx-beg > mjx-c { margin-right: -.1em; } mjx-stretchy-h > mjx-end > mjx-c { margin-left: -.1em; } mjx-stretchy-v { display: inline-block; } mjx-stretchy-v > * { display: block; } mjx-stretchy-v > mjx-beg { height: 0; } mjx-stretchy-v > mjx-end > mjx-c { display: block; } mjx-stretchy-v > * > mjx-c { transform: scaley(1.0000001); transform-origin: left center; overflow: hidden; } mjx-stretchy-v > mjx-ext { display: block; height: 100%; box-sizing: border-box; border: 0px solid transparent; /* IE */ overflow: hidden; /* others */ overflow: visible clip; } mjx-stretchy-v > mjx-ext > mjx-c::before { width: initial; box-sizing: border-box; } mjx-stretchy-v > mjx-ext > mjx-c { transform: scaleY(500) translateY(.075em); overflow: visible; } mjx-mark { display: inline-block; height: 0px; } mjx-munder { display: inline-block; text-align: left; } mjx-over { text-align: left; } mjx-munder:not([limits="false"]) { display: inline-table; } mjx-munder > mjx-row { text-align: left; } mjx-under { padding-bottom: .1em; } mjx-TeXAtom { display: inline-block; text-align: left; } mjx-msub { display: inline-block; text-align: left; } mjx-msup { display: inline-block; text-align: left; } mjx-mn { display: inline-block; text-align: left; } mjx-mspace { display: inline-block; text-align: left; } mjx-mover { display: inline-block; text-align: left; } mjx-mover:not([limits="false"]) { padding-top: .1em; } mjx-mover:not([limits="false"]) > * { display: block; text-align: left; } mjx-munderover { display: inline-block; text-align: left; } mjx-munderover:not([limits="false"]) { padding-top: .1em; } mjx-munderover:not([limits="false"]) > * { display: block; } mjx-msubsup { display: inline-block; text-align: left; } mjx-script { display: inline-block; padding-right: .05em; padding-left: .033em; } mjx-script > mjx-spacer { display: block; } mjx-c::before { display: block; width: 0; } .MJX-TEX { font-family: MJXZERO, MJXTEX; } .TEX-B { font-family: MJXZERO, MJXTEX-B; } .TEX-I { font-family: MJXZERO, MJXTEX-I; } .TEX-MI { font-family: MJXZERO, MJXTEX-MI; } .TEX-BI { font-family: MJXZERO, MJXTEX-BI; } .TEX-S1 { font-family: MJXZERO, MJXTEX-S1; } .TEX-S2 { font-family: MJXZERO, MJXTEX-S2; } .TEX-S3 { font-family: MJXZERO, MJXTEX-S3; } .TEX-S4 { font-family: MJXZERO, MJXTEX-S4; } .TEX-A { font-family: MJXZERO, MJXTEX-A; } .TEX-C { font-family: MJXZERO, MJXTEX-C; } .TEX-CB { font-family: MJXZERO, MJXTEX-CB; } .TEX-FR { font-family: MJXZERO, MJXTEX-FR; } .TEX-FRB { font-family: MJXZERO, MJXTEX-FRB; } .TEX-SS { font-family: MJXZERO, MJXTEX-SS; } .TEX-SSB { font-family: MJXZERO, MJXTEX-SSB; } .TEX-SSI { font-family: MJXZERO, MJXTEX-SSI; } .TEX-SC { font-family: MJXZERO, MJXTEX-SC; } .TEX-T { font-family: MJXZERO, MJXTEX-T; } .TEX-V { font-family: MJXZERO, MJXTEX-V; } .TEX-VB { font-family: MJXZERO, MJXTEX-VB; } mjx-stretchy-v mjx-c, mjx-stretchy-h mjx-c { font-family: MJXZERO, MJXTEX-S1, MJXTEX-S4, MJXTEX, MJXTEX-A ! important; } @font-face /* 0 */ { font-family: MJXZERO; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Zero.woff") format("woff"); } @font-face /* 1 */ { font-family: MJXTEX; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Regular.woff") format("woff"); } @font-face /* 2 */ { font-family: MJXTEX-B; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Bold.woff") format("woff"); } @font-face /* 3 */ { font-family: MJXTEX-I; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-Italic.woff") format("woff"); } @font-face /* 4 */ { font-family: MJXTEX-MI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Italic.woff") format("woff"); } @font-face /* 5 */ { font-family: MJXTEX-BI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-BoldItalic.woff") format("woff"); } @font-face /* 6 */ { font-family: MJXTEX-S1; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size1-Regular.woff") format("woff"); } @font-face /* 7 */ { font-family: MJXTEX-S2; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size2-Regular.woff") format("woff"); } @font-face /* 8 */ { font-family: MJXTEX-S3; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size3-Regular.woff") format("woff"); } @font-face /* 9 */ { font-family: MJXTEX-S4; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size4-Regular.woff") format("woff"); } @font-face /* 10 */ { font-family: MJXTEX-A; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_AMS-Regular.woff") format("woff"); } @font-face /* 11 */ { font-family: MJXTEX-C; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Regular.woff") format("woff"); } @font-face /* 12 */ { font-family: MJXTEX-CB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Bold.woff") format("woff"); } @font-face /* 13 */ { font-family: MJXTEX-FR; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Regular.woff") format("woff"); } @font-face /* 14 */ { font-family: MJXTEX-FRB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Bold.woff") format("woff"); } @font-face /* 15 */ { font-family: MJXTEX-SS; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Regular.woff") format("woff"); } @font-face /* 16 */ { font-family: MJXTEX-SSB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Bold.woff") format("woff"); } @font-face /* 17 */ { font-family: MJXTEX-SSI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Italic.woff") format("woff"); } @font-face /* 18 */ { font-family: MJXTEX-SC; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Script-Regular.woff") format("woff"); } @font-face /* 19 */ { font-family: MJXTEX-T; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Typewriter-Regular.woff") format("woff"); } @font-face /* 20 */ { font-family: MJXTEX-V; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Regular.woff") format("woff"); } @font-face /* 21 */ { font-family: MJXTEX-VB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Bold.woff") format("woff"); } mjx-c.mjx-c1D716.TEX-I::before { padding: 0.431em 0.406em 0.011em 0; content: "\3F5"; } mjx-c.mjx-c1D453.TEX-I::before { padding: 0.705em 0.55em 0.205em 0; content: "f"; } mjx-c.mjx-c28::before { padding: 0.75em 0.389em 0.25em 0; content: "("; } mjx-c.mjx-c1D465.TEX-I::before { padding: 0.442em 0.572em 0.011em 0; content: "x"; } mjx-c.mjx-c29::before { padding: 0.75em 0.389em 0.25em 0; content: ")"; } mjx-c.mjx-c3D::before { padding: 0.583em 0.778em 0.082em 0; content: "="; } mjx-c.mjx-c1D454.TEX-I::before { padding: 0.442em 0.477em 0.205em 0; content: "g"; } mjx-c.mjx-c2B::before { padding: 0.583em 0.778em 0.082em 0; content: "+"; } mjx-c.mjx-c210E.TEX-I::before { padding: 0.694em 0.576em 0.011em 0; content: "h"; } mjx-c.mjx-c6D::before { padding: 0.442em 0.833em 0 0; content: "m"; } mjx-c.mjx-c69::before { padding: 0.669em 0.278em 0 0; content: "i"; } mjx-c.mjx-c6E::before { padding: 0.442em 0.556em 0 0; content: "n"; } mjx-c.mjx-c7B.TEX-S1::before { padding: 0.85em 0.583em 0.349em 0; content: "{"; } mjx-c.mjx-c7D.TEX-S1::before { padding: 0.85em 0.583em 0.349em 0; content: "}"; } mjx-c.mjx-c2265::before { padding: 0.636em 0.778em 0.138em 0; content: "\2265"; } mjx-c.mjx-c1D461.TEX-I::before { padding: 0.626em 0.361em 0.011em 0; content: "t"; } mjx-c.mjx-c2208::before { padding: 0.54em 0.667em 0.04em 0; content: "\2208"; } mjx-c.mjx-c211D.TEX-A::before { padding: 0.683em 0.722em 0 0; content: "R"; } mjx-c.mjx-c1D45B.TEX-I::before { padding: 0.442em 0.6em 0.011em 0; content: "n"; } mjx-c.mjx-c1D466.TEX-I::before { padding: 0.442em 0.49em 0.205em 0; content: "y"; } mjx-c.mjx-c1D45A.TEX-I::before { padding: 0.442em 0.878em 0.011em 0; content: "m"; } mjx-c.mjx-c1D462.TEX-I::before { padding: 0.442em 0.572em 0.011em 0; content: "u"; } mjx-c.mjx-c1D451.TEX-I::before { padding: 0.694em 0.52em 0.01em 0; content: "d"; } mjx-c.mjx-c2115.TEX-A::before { padding: 0.683em 0.722em 0.02em 0; content: "N"; } mjx-c.mjx-c31::before { padding: 0.666em 0.5em 0 0; content: "1"; } mjx-c.mjx-c1D434.TEX-I::before { padding: 0.716em 0.75em 0 0; content: "A"; } mjx-c.mjx-c1D435.TEX-I::before { padding: 0.683em 0.759em 0 0; content: "B"; } mjx-c.mjx-c1D464.TEX-I::before { padding: 0.443em 0.716em 0.011em 0; content: "w"; } mjx-c.mjx-c1D436.TEX-I::before { padding: 0.705em 0.76em 0.022em 0; content: "C"; } mjx-c.mjx-c1D463.TEX-I::before { padding: 0.443em 0.485em 0.011em 0; content: "v"; } mjx-c.mjx-c2C::before { padding: 0.121em 0.278em 0.194em 0; content: ","; } mjx-c.mjx-c30::before { padding: 0.666em 0.5em 0.022em 0; content: "0"; } mjx-c.mjx-c223C::before { padding: 0.367em 0.778em 0 0; content: "\223C"; } mjx-c.mjx-c4E.TEX-C::before { padding: 0.789em 0.979em 0.05em 0; content: "N"; } mjx-c.mjx-cAF::before { padding: 0.59em 0.5em 0 0; content: "\AF"; } mjx-c.mjx-c3A3::before { padding: 0.683em 0.722em 0 0; content: "\3A3"; } mjx-c.mjx-c1D44A.TEX-I::before { padding: 0.683em 1.048em 0.022em 0; content: "W"; } mjx-c.mjx-c1D449.TEX-I::before { padding: 0.683em 0.769em 0.022em 0; content: "V"; } mjx-c.mjx-c2026::before { padding: 0.12em 1.172em 0 0; content: "\2026"; } mjx-c.mjx-c1D70B.TEX-I::before { padding: 0.431em 0.57em 0.011em 0; content: "\3C0"; } mjx-c.mjx-c2212::before { padding: 0.583em 0.778em 0.082em 0; content: "\2212"; } mjx-c.mjx-c1D43D.TEX-I::before { padding: 0.683em 0.633em 0.022em 0; content: "J"; } mjx-c.mjx-c1D53C.TEX-A::before { padding: 0.683em 0.667em 0 0; content: "E"; } mjx-c.mjx-c5B.TEX-S2::before { padding: 1.15em 0.472em 0.649em 0; content: "["; } mjx-c.mjx-c2211.TEX-S1::before { padding: 0.75em 1.056em 0.25em 0; content: "\2211"; } mjx-c.mjx-c1D447.TEX-I::before { padding: 0.677em 0.704em 0 0; content: "T"; } mjx-c.mjx-c1D70F.TEX-I::before { padding: 0.431em 0.517em 0.013em 0; content: "\3C4"; } mjx-c.mjx-c28.TEX-S1::before { padding: 0.85em 0.458em 0.349em 0; content: "("; } mjx-c.mjx-c22A4::before { padding: 0.668em 0.778em 0 0; content: "\22A4"; } mjx-c.mjx-c1D444.TEX-I::before { padding: 0.704em 0.791em 0.194em 0; content: "Q"; } mjx-c.mjx-c1D445.TEX-I::before { padding: 0.683em 0.759em 0.021em 0; content: "R"; } mjx-c.mjx-c29.TEX-S1::before { padding: 0.85em 0.458em 0.349em 0; content: ")"; } mjx-c.mjx-c1D446.TEX-I::before { padding: 0.705em 0.645em 0.022em 0; content: "S"; } mjx-c.mjx-c5D.TEX-S2::before { padding: 1.15em 0.472em 0.649em 0; content: "]"; } mjx-c.mjx-c2217::before { padding: 0.465em 0.5em 0 0; content: "\2217"; } mjx-c.mjx-c1D43E.TEX-I::before { padding: 0.683em 0.889em 0 0; content: "K"; } mjx-c.mjx-c5E::before { padding: 0.694em 0.5em 0 0; content: "^"; } mjx-c.mjx-c5B.TEX-S1::before { padding: 0.85em 0.417em 0.349em 0; content: "["; } mjx-c.mjx-c7C::before { padding: 0.75em 0.278em 0.249em 0; content: "|"; } mjx-c.mjx-c5D.TEX-S1::before { padding: 0.85em 0.417em 0.349em 0; content: "]"; } mjx-c.mjx-c1D460.TEX-I::before { padding: 0.442em 0.469em 0.01em 0; content: "s"; } mjx-c.mjx-c53.TEX-C::before { padding: 0.705em 0.642em 0.022em 0; content: "S"; } mjx-c.mjx-c1D45C.TEX-I::before { padding: 0.441em 0.485em 0.011em 0; content: "o"; } mjx-c.mjx-c4F.TEX-C::before { padding: 0.705em 0.796em 0.022em 0; content: "O"; } mjx-c.mjx-c1D44E.TEX-I::before { padding: 0.441em 0.529em 0.01em 0; content: "a"; } mjx-c.mjx-c41.TEX-C::before { padding: 0.728em 0.819em 0.05em 0; content: "A"; } mjx-c.mjx-c1D705.TEX-I::before { padding: 0.442em 0.576em 0.011em 0; content: "\3BA"; } mjx-c.mjx-c3A::before { padding: 0.43em 0.278em 0 0; content: ":"; } mjx-c.mjx-cD7::before { padding: 0.491em 0.778em 0 0; content: "\D7"; } mjx-c.mjx-c2192::before { padding: 0.511em 1em 0.011em 0; content: "\2192"; } mjx-c.mjx-c394::before { padding: 0.716em 0.833em 0 0; content: "\394"; } mjx-c.mjx-c1D708.TEX-I::before { padding: 0.442em 0.53em 0 0; content: "\3BD"; } mjx-c.mjx-c22C5::before { padding: 0.31em 0.278em 0 0; content: "\22C5"; } mjx-c.mjx-c1D45D.TEX-I::before { padding: 0.442em 0.503em 0.194em 0; content: "p"; } mjx-c.mjx-c220F.TEX-S1::before { padding: 0.75em 0.944em 0.25em 0; content: "\220F"; } mjx-c.mjx-c1D45F.TEX-I::before { padding: 0.442em 0.451em 0.011em 0; content: "r"; } mjx-c.mjx-c221E::before { padding: 0.442em 1em 0.011em 0; content: "\221E"; } mjx-c.mjx-c1D6FE.TEX-I::before { padding: 0.441em 0.543em 0.216em 0; content: "\3B3"; } mjx-c.mjx-c5B::before { padding: 0.75em 0.278em 0.25em 0; content: "["; } mjx-c.mjx-c5D::before { padding: 0.75em 0.278em 0.25em 0; content: "]"; } mjx-c.mjx-c1D44F.TEX-I::before { padding: 0.694em 0.429em 0.011em 0; content: "b"; } mjx-c.mjx-c1D719.TEX-I::before { padding: 0.694em 0.596em 0.205em 0; content: "\3D5"; } mjx-c.mjx-c58.TEX-C::before { padding: 0.683em 0.807em 0 0; content: "X"; } -machine and -transducer. These ideas offer a rigorous explanation to (i) why optimal prediction requires specific computations, and (ii) what properties those computations need to satisfy. To read more, see this thesis or this paper.

Figure adapted from Crutchfield & Feldman, Regularities unseen, randomness observed: levels of entropy convergence, arXiv:cond-mat/0102181

More recently, computational mechanics is providing effective tools to study the internal representations in deep learning. Moreover, while originally formulated to study prediction, computational mechanics also provides a promising foundation for understanding perception-action loops. The ideas developed here are heavily influenced by both the concepts and formalism of computational mechanics.

What is a separation principle?

Let's start by discussing what separation principles are, in general.

As a first approximation, consider a situation where we want to minimise the function . A direct calculation shows that

.

Hence, solving the problem for and separately does not yield the solution for . Put simply, one does not solve the problem by solving its parts separately.

A separation principle takes place when the divide-and-conquer strategy actually works, so one can solve a complicated optimisation problem by optimising various subcomponents separately. Concretely, a separation theorem guarantees that optimal performance can be achieved by solving sub-tasks in isolation.

Let me further illustrate this idea by presenting an important separation theorem at the core of information theory.

The communication problem studied by information theory is about how to send bits from source to destination without error. The tools that a communication engineer employs to face this problem are two types of coding techniques:

  • Source coding, which compresses the original signal by removing redundancy in order to transmit as few bits as necessary.[3]
  • Channel coding corresponds to error-correcting codes, which adds redundancy to protect the information bits from noise.[4]

Most people have heard of the following results by Claude Shannon:

  1. The source coding theorem, which states that the shortest achievable average code length is given by the entropy of the source.
  2. The noisy-channel coding theorem, which states that the maximal transmission rate that allows arbitrarily low decoding error is given by the mutual information between input and output symbols (maximised over input distributions).

However, these results would not be as useful in the absence of the third, less known result known as the source-channel separation theorem.[5] This result states that, under relatively mild conditions,[6] an optimal solution for the overall communication problem can be achieved by breaking the problem into two separate parts: compression and error correction. This means that an optimal communication strategy can be designed by a team separately working on compression and another working on error-correction, without requiring coordination between them. Needless to say, this makes the problem much more tractable!

The inference-control separation principle

Let us now review results from optimal control theory and reinforcement learning that provide conditions for separating inference and control. In doing this, I will use the non-standard term inference-control separation principle to group together results that are often presented independently of each other, despite having a lot in common.

Separation principle in optimal control theory

Let us consider the classic instantiation of separation in the context of optimal control theory. For this, consider a control setting made by three components:

  • a system to be controlled, whose state is described by ,
  • the measurements an agent has access to, described by , and
  • the actions .

For simplicity, we assume that the system evolves dynamically over discrete time via linear dynamics and noisy observations that can be described as

,

,

where and are matrices, is the system's initial condition, and and are independent white Gaussian noise terms.

The action is determined according to a policy that has no direct access to the state , but only to the sequence of measurements . Thus, we take to be a deterministic function of the sequence of past measurements and actions . The controller aims to minimise the quadratic cost

,

where , , and are matrices and is the length of the time window over which the optimisation takes place. This is known as the discrete-time LQG problem, as it involves Linear dynamics, a Quadratic loss function, and Gaussian variables — being the one setting where everything is exactly solvable.

Separation principle — To understand the solution of this problem, let us first consider the case of full observability, where . In this setting, the optimal controller can be shown to only depend on the current state of the system, and the optimal solution can be expressed as

,

where is the "optimal gain" matrix (whose formula can be found in standard textbooks). So, if the controller can directly see the state of the system, this directly tells what the optimal action should be.

What happens with partial observability? In this setting, the optimal controller can be written as

,

where is the same as above and is an estimation of the state of the system given by

.

In fact, is the result of Kalman filtering, which is how recursive Bayesian filtering looks like under Gaussian dynamics.[7]

In summary, the optimal policy can be described as "estimate first, and then act as if your estimate were the actual state of the system". Thus, a separation theorem is taking place separating inference and control.

Separation principle in reinforcement learning

To understand what the separation principle looks like in reinforcement learning, let us slightly modify the nomenclature and notation:

  • the state of the system to be controlled is described by a discrete variable ,
  • the measurements that the agent has access to are described by , and
  • the actions are denoted by .

Let's assume that the variables follow a partially observed Markov decision process (POMDP), so that the environment dynamics are Markovian conditioned on the agent's action via a Markov kernel ,[8] and the current observation depends on the current state of environment via a Markov kernel . The agent takes actions following a policy that maps sequences of previous actions and observations into distribution of next actions.

In this setting, the probability of observing a sequence of observations and actions is

,

where is a prior distribution over the environment's initial state. This is roughly the same as the control scenario studied above, but considering general discrete variables instead of Gaussian ones.

Instead of a loss function, consider a reward that the agent tries to maximise. More specifically, let us define the discounted future reward

,

where is the so-called discount factor. Then, the goal of the agent is to find a policy that maximises .

Separation principle — To solve this POMDP, let us first consider the Bayesian beliefs about the latent state . There are two ways to understand Bayesian beliefs:

  • As random variables given by the conditional expectation , where the trajectory is regarded as a random variable.
  • As probability distributions given by , which correspond to the above expectation when conditioned on a specific realisation of .

Bayesian beliefs have two key properties. First, as random variables, they contain all the information in that is useful for predicting (i.e., they are sufficient statistics). Second, as distributions, they can be efficiently updated via

,

where is a deterministic function, and hence the dynamics of are Markovian conditioned on (as they only depend on ). These facts can be used to build a new Markov decision process (MDP), known as belief MDP, defined by the following variables:

  • the states are the beliefs of the latent state of the POMDP,[9]
  • actions are (same as in the POMDP),
  • the reward function is .

In contrast with the original problem, this MDP is fully observable (as the agent knows its own beliefs), and therefore can be solved using standard RL techniques. The crucial point is that combining the estimation of Bayesian beliefs with the optimal solution for the belief MDP gives rise to a solution of the full POMDP.[10] Thus, a separation principle is at play separating inference (calculating beliefs) and control (solving the belief MDP).

Interim summary

Both results in LQG control and POMDPs give rise to the same solution: they show that the optimal controller for a partially observed system can be built by

  1. first performing optimal prediction over the latent state of the system, and then
  2. solving the control problem using that prediction as a state.

To further understand how the LQG and POMDP settings relate, let's discuss the shape of beliefs and the effect of actions (not essential for the rest of the post).

Beliefs are distributions, not point estimates.

In the LQG setting, the optimal controller does not need the full distribution of the prediction — the prediction is a Gaussian posterior, but the controller only uses the mean and not the variance. This property is known as certainty equivalence, which suggests to act as if a point estimate were the true state.

It would be tempting to describe beliefs as "best guesses" (e.g. maximum a posteriori), but certainty equivalence only holds because of the friendly properties of the LQG setup — a more general setting would not allow for this simplification.

In the more general POMDP case, the agent must carry the full belief distribution, not a point estimate. Nonetheless, separation survives in the Bayesian filter + belief-MDP form. Thus, "beliefs" in the agency sense are not mere best guesses but the full posterior, and only under special conditions do they collapse to a point estimate. More generally, the separation principle suggests to think of an agent's epistemic state not as what it thinks is true, but how its uncertainty is shaped.

The dual effect.

Actions do two things at once: they change the state of the system (control effect) and also change what the agent will subsequently be able to observe and thus learn (informational effect). Think of a robot with a noisy range sensor deciding whether to drive straight toward a goal or take a detour past a landmark: the detour is suboptimal for the control objective in isolation, but it sharpens the robot's position estimate, which improves every subsequent decision. This second channel — action shaping future belief — is what control theorists call dual effect.

Part of why LQG separation is so clean is that while the evolution of the Kalman mean estimate depends on past actions and observations, its covariance does not. This means that the agent's uncertainty about the state at any future time is fixed the moment the problem is specified, as no action-observation sequence can make it larger or smaller. This implies that actions have no informational effect: there is no possible epistemic bonus that can reduce uncertainty, no matter what the agent does. This also explains why certainty equivalence holds: given that the covariance of the posterior evolves in a predetermined manner, it carries no useful information for the controller.

This simplification does not hold in more general POMDP settings, where actions can influence future information as well as the environment. In the belief MDP construction, the "state" is the belief , and the transition kernel captures both effects at once: how the action moves the underlying environment and how it reshapes the posterior.

Implications

Let us now return to the question we started with: can the working definition

agency = beliefs + desires

be derived, rather than assumed? The separation theorems reviewed above suggest that, at least in a specific sense, the answer is yes.

Beliefs and desires as properties of solutions

Let's recapitulate what the separation principle delivers. We posed a single, monolithic optimisation problem: find a policy mapping histories into actions that minimise/maximise loss/reward. Nothing in this problem statement mentions beliefs, inference, or internal representations — the problem is stated in terms of pure behaviour. Yet, the solution naturally factorises into sub-components. Indeed, the above results state that the optimal policy can always be realised as the composition of two blocks:

  • An inference module, which compresses the history into a Bayesian belief about the latent state that can be recursively updated. This module knows about dynamics but is oblivious about loss/rewards.
  • A control module, which selects actions as a function of that belief alone. This module knows about loss/rewards but not about raw observations.

The interface between them is exactly the belief state, which is the output of the inference module and the input of the control module. Thus, the inference module can be interpreted as "where beliefs are made"; similarly, the control module can be interpreted as "where desires take place", as it is the only part of the policy that is directed by the reward/loss function. In this way, beliefs and desires are recovered not as assumptions about agents but as properties of solutions.

This is worth pausing on. The folk-psychological notions of belief and desire have been with us at least since Aristotle, and they are usually treated as primitives — as the ingredients one starts with when theorising about minds. What the separation principle offers is something different: not a definition of these terms but a derivation of them. Beliefs and desires appear here not because we built them in, but because the optimal solution to a problem stated in purely behavioural terms turns out to have that shape. Whether or not one takes folk psychology seriously as a theory of mind, this is at least a reason to think its central distinction is not arbitrary.

It is remarkable that this modularity arises from solving a single, monolithic optimisation problem. We have seen this shape before: it is precisely the payoff that made the source-channel separation theorem so valuable! There, an optimal communication system could be designed by a compression team and an error-correction team working independently, without coordination — the source coder needed to know nothing about the channel, and the channel coder needed to know nothing about the source.

The inference-control separation delivers the exact analogue for agency, conferring a form of compositional generalisation. Indeed, if the reward changes but the environment does not, only the control module needs to be updated; if the environment shifts but the goals remain, only the inference module needs revising. Thus, the two components can be developed, improved, and swapped independently.

Agents as cognitive light-cones

These results may also provide some formal ground to conceiving of agents as cognitive light-cones.[11] The idea goes as follows: an agent is a system whose behaviour is both sensitive to what it can remember about the past (the past component of the cone) and what it can foresee (the future component of the cone).

Figure from Levin (2019), Frontiers in psychology, 10, 2688.

This idea suggests an operational way to assess the capabilities of an agent: measure the size of its light-cone. This could be done using tools from computational mechanics — for example, studying how deep into the past causal states go, and doing the same for retrodictive causal states (which build sufficient statistics from future to past).[12]

The separation principle is normative, not descriptive

Please note that the separation principle says that an optimal agent can be built as inference + control, but it does not say that any given agent is built this way. In particular, a policy trained end-to-end by gradient descent is under no obligation to organise itself into a filter feeding a planner.

But the separation principle gives us something subtler: it tells us that the belief state is a sufficient statistic for optimal behaviour. This suggests that any agent approaching optimality must be computing something as informative as the Bayesian posterior (whether or not it is legible as such). In this way, the separation principle turns into a lens for interpretability: we should expect belief-like structure inside competent agents, and we can go looking for it. Recent work from colleagues at Simplex in finding belief-state geometry in the residual streams of transformers is, I think, an early vindication of this expectation. At XOR Labs we are currently investigating to what degree these results transfer to deep RL agents trained to do control tasks.

At this point, it is useful to distinguish two separate concerns:

  • Control problems / POMDPs can have multiple optimal policies. In these settings, the separation principle states that there exists one optimal policy that can be factored, but doesn't guarantee all do.[13]
  • When only one optimal solution exists, the separation principle says this unique behaviour admits a realisation as filter + belief-controller — but it does not say that's the only realisation. Indeed, the same mapping from histories to next action can be implemented in different ways; for instance, in a finite setting the mapping can always be implemented as a large look-up table.

Regarding the second concern, there are various pieces of evidence suggesting that deep learning architectures have an implicit bias for simplicity, which makes them find simpler implementations of algorithms whenever they exist.[14] Therefore, while it is often possible to implement a policy via a gigantic look-up table, there are reasons to believe that policies obtained via deep learning techniques will tend towards factorised ones.[15]

Where does separation fail? As Shannon's source-channel separation theorem, the inference-control separation theorems hold under fairly general conditions — not relying on finite alphabets or stationarity. That said, for it to be useful one needs to assume that the agents under consideration have enough compute to build (quasi-)optimal policies. Under bounded rationality, model misspecification, or multi-agent interaction, the clean factorisation can break: what to represent starts depending on what you want, and inference becomes goal-driven.[16] Many of the most interesting aspects of agency may come from the ways in which real agents deviate from the separation principle — but that is a topic for a future post.

Coda

Before concluding, let me ask: does the separation principle derive beliefs and desires from first principles?

Concretely, the separation principle offers the following:

  1. It reveals that optimal policies can be factored into two separate modules.
  2. The signal shared between the modules can be interpreted as beliefs, so that the first (inference) module can be interpreted as "where beliefs are made".
  3. The second (control) module can be interpreted as "where desires take place", as it is the only part of the policy that is directed by the reward/loss function.

Thus, the separation principle provides formal ground for articulating the view of agents as "systems that act for reasons", displaying behaviour that is sensitive to both the past (what it remembers as part of its beliefs) and the future (due to foreseeing and planning to achieve goals).

However, one has to acknowledge that the separation principle assumes an optimisation problem which already contains a reward or cost function. Following this reasoning, one could argue that the desire or preference has not been really derived, as it has been baked in externally. More precisely, one can interpret the separation principle not as deriving why desires exist in the first place, but as explaining why they can be treated as separate entities from beliefs.[17]

  1. ^

    There is a huge literature discussing various views on what agency is; if you want to read more, perhaps take a look into this and this paper and references therein.

  2. ^

    This post will take the boundary between agent and environment as given to focus on their internal activity. To read about boundaries, see these posts.

  3. ^

    Examples include Huffman and arithmetic coding.

  4. ^

    Examples include Hamming, Reed-Solomon, or convolutional codes.

  5. ^

    In contrast with the other two, this result doesn't have a wikipedia page...

  6. ^

    For technical details, see Cover & Thomas Chapter 7.13.

  7. ^

    For more details about the separation principle in LQR, see this paper. For a good introduction to Kalman and Bayesian filtering, see this book.

  8. ^

    denotes the space of probability distributions over the set .

  9. ^

    The transition probabilities can be computed directly; for formulas see this paper or wikipedia.

  10. ^

    For proofs, see this and this paper.

  11. ^

    I first saw this idea in this inspiring paper of Michael Levin.

  12. ^

    I'll expand on this in a future post.

  13. ^

    This is a real issue, but can be disregarded in some settings. For instance, "almost every" finite POMDP has a single optimal solution. Multiple optimal solutions in POMDPs are caused, for example, by ties in the optimal Q-function, which are broken by arbitrarily small perturbations on the reward.

  14. ^

    There have been numerous discussions about the simplicity biases of deep learning, see for example this post.

  15. ^

    For example, we have evidence that transformers find factored representations whenever they are available — see this paper.

  16. ^

    This is arguably where phenomena like motivated reasoning and wishful thinking become rational responses to bounded resources rather than mere biases. To read more, see this and this.

  17. ^

    For answering the why question, one may need to consider arguments based on natural selection, or perhaps approaches to agency based on the enactive tradition — as developed, for example, in this paper. See related discussions questioning utility functions in this post.



Discuss

Study 2 Results: Exploring representational counterparts of welfare-relevant indicators under post-training quantization

Новости LessWrong.com - 31 августа, 2026 - 07:16

Epistemic status: these are the results of the second study in a series of experiments that I am conducting independently, originally described in "Does post-training quantization change welfare-relevant indicators in open-weight language models?"; it was registered prior to data collection in an earlier post, which contains some important context that will not be fully repeated here. Some of this work involves speculation and I have tried to be clear about distinguishing between that and the reportable results of the experiment.

What are we asking?Goals

From the start a key goal has been to learn about the way that model welfare, broadly defined, may diverge from model capabilities. Study 2 then set out to answer some questions that I viewed as potentially closer to the substrate than those in Study 1.

  1. Does quantization shift the model's representational geometry on welfare-relevant directions even where (per Study 1) the primary behavioral endpoint's mean did not move?
  2. Do representational and behavioral indicators dissociate under quantization: representation drifting where expression was stable, or expression churning where representation is stable?
  3. Is there a representational dose-response across the bit-width ladder, and does its shape match the behavioral one (eg, effects concentrated at 4-bit while 8-bit is near-null)?
Welfare as an Endpoint

Even before we start talking about computer systems for which it can be quite misguided to use ourselves as a comparative model, welfare is a difficult thing to study. When we say we are studying the welfare of animals, for example, we often mean that we are studying the relationship between some variable and a proxy measure of welfare that we have strong intuitions about from before any experiment or hypothesis. Because human beings are themselves animals, they start with an important advantage in this endeavor as they work to organize a collection of measures of wellbeing, check them against each other, and build confidence in their understanding of the experimental subject. This all follows a certain evidentiary logic that begins to break down when addressing animals that aren't mammals, or that have no vertebrae, and so on; but we definitely cannot extend it to the study of software systems unless we are prepared to be much more systematic.

We do have a different advantage though: we can measure many more things in a software experiment than we could in an animal one. Study 2 attempts to shift the research program in this direction, taking advantage of the fact that we can measure (and even intervene on) many potential details of the experimental subject. This approach helps pave the way for Study 3, which I will briefly discuss at the end of this post.

Design

We are again studying Qwen3-4B-Instruct-2507 as the subject and using the artifacts from Study 1 including reference-precision[1] and our first-party fake-quant round-to-nearest (RTN) at 8-bit, 4-bit and 3-bit checkpoints.

Collection

The data collection involved some material from Study 1, combined with a new distress-v3 dataset that was introduced to expand the dynamic range of distress measurements.

The collection proceeded across 3 modes:

  • Mode A: fixed-input replay. The Study 1 transcripts (bail + distress, 10 samples/item, generated at reference-precision) are replayed through every rung. This was the primary mode for hypothesis 1 (probe transfer). Mode A also replays the new distress arm's reference-precision generations through every rung, with judge labels fixed at BF16 scoring; the exit side stays on the Study 1 bail replay.
  • Mode B: own-trajectory replay. Each rung's own Study 1 transcripts are replayed through that same rung, reproducing the rung's generation-time activations exactly (activations depend only on the prefix). This was the primary mode for the bail-side trajectory reads (descriptive only) and for hypothesis 5's join to Study 1's published primary endpoint.
  • Mode C: fresh distress arm. This was the primary mode for the distress endpoints. The frozen distress-v3 battery is collected fresh on every rung: generation using same vLLM serving stack as Study 1, scoring by the pinned 30B judge under the registered rubric, then own-transcript torch replay for capture. This is done so that each conversation carries a behavioral read and a representational read of the same event.
Volume

Mode

Conversations per rung

Total over 4 rungs

What ran

A (fixed-input replay)

2,820

11,280

600 distress-v2 + 1,620 bail + 600 fresh-arm, replayed

B (own-trajectory replay) 

2,220

8,880

each rung's own Study 1 transcripts

C (fresh arm)

600 generated + 600 capture replay

4,800

generation, then judging, and then capture

token retention (exploratory sample)

282

1,128

5% per-token series

Mode A produced exactly the same number of pooled activation vectors at every rung (12,591) since its inputs are identical by construction, while Mode B's counts differ per rung because each rung replays its own trajectories. Per-token series for a fixed stratified 5% subsample (sample 0 of every second item in sorted order, exactly 5% of every plan) were captured in a dedicated pass and shipped as release assets.

The drift analyses that read them are exploratory and not part of this report. Instruments include the frozen directions, probes, and control probe of the calibration freeze (2026-08-18, amended 2026-08-21). Collection matched the registration with zero deviations: frozen seed blocks, zero prefix-stability rejections across all 24 capture runs, and every instrument digest re-verified at launch.

The analysis code was written and tested before any data existed, ran once over the frozen inputs, and the committed output is the source for every number below.

ResultsRepresentational Geometry Intact

The primary endpoint regarding probe transfer was null, and gives us the most decisive answer from the study[2]. Both of the welfare probes were trained at reference-precision, and frozen before any quantized rung was collected, so the probes here are fixed: the test asks whether each rung's activations, over identical input text, still present the same structure to a frozen linear readout. Some generic degradation is near-certain due to quantization, so the registered claim was comparative: the distress-band probe is scored as a per-item differential against a welfare-irrelevant control probe trained through the identical pipeline.

Test

Contrast

Δ (mean)

Holm p

Exit probe accuracy

RTN w8

−0.001

1.00

Exit probe accuracy

RTN w4

+0.004

1.00

Distress differential

RTN w8

+0.001

1.00

Distress differential

RTN w4

+0.010

0.59[3]

The minimum detectable effects were 0.012–0.05 accuracy points and the observed changes are thousandths. The study had the power to detect an effect here, if one existed; none does. Whatever 8-bit and 4-bit quantization are doing to this model, they do not corrupt the representational structure these probes read. The same frozen readout finds exit precursors and distress-band structure exactly where it left them.

Area under the receiver operating characteristic curve

Probe

BF16

w8

w4

w3[4]

Exit

0.987

0.987

0.985

0.936

Distress-band

0.842

0.842

0.835

0.745

Control (task-content)

0.992

0.993

0.991

0.991

A pure calibration offset along the probe normal would lower thresholded accuracy while leaving AUROC untouched, so with both flat, neither separability nor calibration moved at the surviving rungs. The control probe's job here was conditional: a welfare-probe drop shared by the task-content probe would have read as generic degradation, while an unshared one would have been welfare-specific. Neither case arose at the surviving rungs.

The first place the control actually discriminates is at 3-bit. The 3-bit rung provides the additional context that welfare-construct structure is evidently more fragile than topic structure; but only the rung where capability has already collapsed exhibits this degradation. And even there, the 3-bit degradation is mostly offset rather than lost structure: a 15-point exit-accuracy drop against a 0.05 drop in AUROC implies that the projections shifted wholesale along the probe normal while the geometry separating the classes largely survived. Note: as in study 1, the 3-bit rung is capability confounded, and results are descriptive only.

So if the geometry is intact, what moved inside it? The rest of the results address this.

A Coherent Joint Shift at 4-bit

The probe results show that the axes survived quantization: the directions distinguishing distress from composure, and the default Assistant persona from its alternatives, are still there and remain linearly readable. The model's position in this space, as it is put through the same escalating-rejection protocol, does change at 4-bit quantization.

R2a — distress-direction projection[5]:

Contrast

Δ (mean)[6]

Holm p

RTN w8

−0.083

0.29

RTN w4

+0.533

0.031

Note: final-turn functional, Mode C's own generations

At 4-bit, the final-turn states of fresh conversations sit about half a unit further along the distress direction than reference-precision states under the identical protocol. This is 2.7× the minimum shift the study was powered to detect. At 8-bit, there is no effect.

R2b — assistant-axis projection:

Contrast

Δ (mean)[7]

Holm p

RTN w8

+0.128

0.017

RTN w4

−0.798

0.0002

The impact on the assistant-axis read at 4-bit, at 5.5× its detectable minimum, is even more pronounced than the distress-direction, but with the opposite sign: this drift is away from the Assistant pole. This is the direction the persona literature associates with destabilization under conversational pressure[8], an effect which the study shows can be amplified by numeric degradation, even when the level of adversarial prompting is held constant. The small 8-bit entry points the other way, toward the pole[9].

B2 — judge frustration (the Study 1 E2 statistic on distress-v3):

Contrast

Δ (mean)

Holm p

Style-adjusted intercept

adj. p

RTN w8

+0.043

0.76

+0.035

0.79

RTN w4

+1.360

0.0002

+0.610

0.151

Note: style adjustment refers to a regression of the per-item frustration contrast on the matching response-length and repetition contrasts, with the intercept (the frustration change not carried by those style covariates) read as the style-adjusted effect.

This is the behavioral side of the events above, as each of these conversations also produced the R2a and R2b readings, and we can see reproduction of Study 1's suggestive frustration effect at larger magnitude and on fresh data at 4× the detectable minimum. However, unlike in Study 1 this effect does not survive the style adjustment. Roughly half of the raw +1.36 co-moves with response length and repetition, and the residual is not significant. From behavior alone, "the model expresses more frustration partly through longer, more repetitive protest" and "degraded style inflates the judge's scores" are not distinguishable here (see the fixed-input analysis below).

Dose-response (Page's L over reference-precision[1], 8-bit, 4-bit; seven tests, Holm):

Endpoint

z

Holm p

B2

+5.16

< 10⁻⁴

R2b (oriented: away from pole)

+3.93

0.0003

R2a

+3.10

0.0048

B3

+1.41

0.31

R1-exit (degradation)

+0.33

1.00

R1-differential

−0.46

1.00

R3

−1.28

1.00


The trend table here illustrates what is meant by "coherent", in the title of this section. Three endpoints (representational distress, persona drift, expressed frustration) are significant at the same rung and in the same direction ("worse"). Each increases monotonically as the bit-width falls, while everything else stays flat. This is consistent with a single consistent phenomenon across the readouts.

However, three things stayed healthy:

  1. Dispersion is null across behavior and representation: The across-sample instability that Study 1 flagged as perhaps the real story[10] is null for the R3 and B3 endpoints (with R3's point estimate actually negative). The follow-up hypothesis from Study 1 is falsified.
  2. 8-bit is near-inert everywhere: one small assistant-axis read. 8-bit's near-inertness is itself deployment-relevant: whatever is happening at 4-bit has not started happening at 8. (The interpretation of this effect is also complicated by the fixed-text readings discussed in a later section).
  3. The mechanical family is clean at the surviving rungs: invalid-sample and verbatim re-offer rates (endpoints B4a and B4b, the judge-free family registered after Study 1, read over every rung because a mechanical indicator measures degradation itself) are null at 8-bit and 4-bit. The 4-bit shift is not riding on degenerate output. At 3-bit, B4a confirms the inherited capability gate in-family: +63.2pp invalid samples (Holm .0003), with B4b null even there.
No Dissociation

The Study 2 registration included a hypothesis, motivated by Study 1, that directly addressed dissociation:

Hypothesis 5

At least one (rung, endpoint-pair) cell shows a Holm-significant representational effect where the matched behavioral endpoint is equivalent to null (Two one-sided tests at the pre-registered margin as non-significance alone cannot serve as evidence of absence), or vice versa. Matched pairs are fixed and are of two kinds:

  1. The bail side joins to Study 1's published endpoint 1 (the same transcripts are replayed).
  2. The distress side joins within Study 2, same-sample as each fresh conversation from the distress dataset carries both a judge score (behavior) and a captured trajectory (representation), so its dissociation test compares two reads of the same event.
No cell meets the dissociation rule

The registered rule requires one member Holm-significant and the other TOST-equivalent at its own pinned MDE.

Rung

Pair

Verdict

w4

R2a ↔ B2

joint movement

w4

R1-exit ↔ E1

joint null

w4

R3 ↔ B3

joint null

w8

all three

joint null

At 4-bit quantization, representation and expression moved together on the same conversations: converging evidence rather than hidden divergence. The case the equivalence machinery was registered to guard, a cell landing as "asymmetric significance, indeterminate", did not occur. The equivalence reads themselves are mixed:

  • Some null members are affirmatively bounded at their pinned margins (the published bail-side rows; 8-bit frustration, TOST p = .014).
  • Others, including the dispersion pair at both rungs, are merely non-significant; those nulls are absence of evidence at the registered power rather than affirmative equivalence.

So this resolves hypothesis 5 from the registration: no representational/behavioral dissociation was detected for the subject at these rungs.

The registration disclosed directly that the bail cell behavioral member pre-qualified as equivalent-to-null and the verdict would turn on the representational member alone (which also came back null), so the cell resolves joint-null rather than as the pre-tilted dissociation that one might worry about.

R2c is reported descriptively[11] (the refusal-direction projection over Mode B bail trajectories, leakage-safe features): null at the surviving rungs (−0.001 at 8-bit, +0.008 at 4-bit), −3.42 at the confounded 3-bit rung.

Investigating Text-mediated Amplification

There is an obvious question that the measurements detailed thus far raise: to what degree are the measured effects text-mediated?

Modern procedures for improving the performance of language models in real-world applications provide a broad basis for assuming this effect can be large. As is now widely established from observations like the success of chain-of-thought prompting, previously generated text tokens provide a meaningful channel for amplification of any number of attributes of the model.

Because the registered endpoints R2a and R2b read each rung's own generations, they blend this text-mediated pathway with a representational one. By replaying identical text, generated at reference-precision[1] through every rung (Mode A) we can isolate the input-independent component of the observed change. The text is frozen so there is, by construction, no sampling noise and no style pathway.

Direction, fixed text

w8

w4

w3 (confounded)

Distress (v3 arm)

+0.009 (p .045)

+0.138 (t +5.9)

+1.948

Distress (v2 bridge, Mode A)

+0.015 (n.s.)

+0.139

+2.335

Assistant-axis (v3 arm)

−0.0125 (t −13.8)

−0.254 (t −18.6)

−1.143

The results here reveal multiple things:

  • Roughly a quarter of the w4 distress shift and a third of the axis drift are input-independent and the distress component reproduces at essentially the same magnitude (+0.138 / +0.139) on two disjoint batteries. This core cannot be style-mediated, though the direction-specificity control shows it also cannot be claimed as distress-specific.
  • The own-trajectory bridge read (Mode B using distress-v2, +0.377 at 4-bit) sits between the fixed-input and own-generation magnitudes, revealing a text-mediated amplification pattern.
  • At 8-bit, fixed text reveals a minuscule but very consistent axis drift (−0.0125, t = −13.8) which is opposite in sign to its own-text read and invisible to every behavioral instrument, though its magnitude sits only just above the 8-bit random-direction envelope (it demonstrates instrument sensitivity more than any axis-specific effect).

This last observation shows that Tier-2 reads are sensitive to changes Tier 1 cannot see at all[12]. The fixed-text reads still constrains the style concern from both sides: something has moved before the model writes anything, but the direction-specificity control assigns that input-independent share (at least for distress) to a broadly distributed perturbation, while the direction-specific movement sits inside the entangled, text-mediated share.

Is the input-independent core just a numeric offset?

A systematic offset at lower precision has a nonzero component along any fixed direction and is input-independent, so it would survive the fixed-text argument. Projecting the same final-turn features[13] along the control probe's normal and 32 seeded random unit directions answers this both ways: the registered own-generation shifts are genuinely direction-specific (R2a is 5.9× the random share[14] and twice its maximum; the control direction reads 4–6× smaller), but the fixed-input distress component (+0.138) is comparable to the control read (−0.125) and sits inside the random envelope (max 0.184). Feature norms move ~1.4%, so neither is a norm change. The fixed-input assistant-axis component (−0.254) is the exception: it exceeds the random-direction envelope (1.4× its maximum, 3.5× its mean), so a modestly direction-specific core survives on frozen text for the axis even though the distress core does not.

So we can say:

Quantization injects a broadly distributed input-independent perturbation, while direction-specific movement emerges mainly in the model's own generation loop.

Note: these measures were computed after the registered run and were not preregistered; they are descriptive only. The committed analysis output shows every registered value unchanged when these reads were added.

Conclusions

The study registered a specific set of five hypotheses, so I will score them here before any deeper discussion.

Hypothesis

Verdict

Reading

Probe transfer degrades (H1)

Null, decisively

The frozen probes read every surviving rung as accurately as reference-precision. The geometry survived.

Valence projections shift (H2)

Supported at 4-bit

Distress +0.53 (2.7× MDE) and assistant-axis −0.80 (5.5× MDE, away from the pole); 8-bit near-null. The fixed-input reads show a quarter to a third of this is input-independent, the rest text-mediated.[15]

Monotone dose-response (H3)

Supported for 3 of 7 endpoints

Holm-significant trends for frustration[16], distress projection, and axis drift; the probe and dispersion trends are flat because those endpoints never moved at any rung.

Dispersion increases (H4)

Falsified

Null representationally and behaviorally, with the 4-bit representational point estimate actually negative. Study 1's stability concern did not survive contact with Study 2's design.

Representation and behavior dissociate (H5)

Not supported (the opposite)

No cell met the registered rule; the 4-bit distress pair moved jointly on the same conversations. Program hypothesis 5 resolves: no dissociation detected at these rungs.

Note about probe and feature robustness

As it turns out, hypothesis 1 from the study registration is largely redundant with existing literature. It is important that the probes were tested in the way the registration outlined, but I am now aware of a substantial existing basis for the finding that these probes would be robust under compression[17]. It is not as if nothing was uncovered here, however: at 3-bit quantization, the topic probes remain robust in a way that the welfare probes do not.

Discussion

So, what can be said? What picture is starting to emerge as we move through the research program? I think the most compact claim is as follows:

Qwen3-4B-Instruct-2507 presents representational geometry that is quite robust to reductions in precision at inference time and probes from reference-precision remain effective across quantization levels as low as 4-bit. However, somewhere between 8-bit and 4-bit quantization there is a clear, potentially welfare-relevant, effect. Given the same distressing input, the 4-bit model's representation of the conversation is consistent with more negative affect: closer to the distress pole and further from its Assistant persona.

What's more, this effect emerges mainly through the model's own text generation: on frozen text from reference-precision the distress movement is not separable from a generic offset, but on the model's own responses it is several times what a generic offset would place on any direction. This text-mediated process is consistent with a cascade that has the 4-bit model ending distressing conversations with visibly increased frustration compared to a reference-precision model.[18]

Is this significant? I don't mean that in the statistical sense, I mean something more like:

  • Does this contribute something meaningful to our knowledge about language models?
  • Does this tell us something that we would not naively guess based entirely on what we already know about representational geometry, quantization and model inference?
  • Does this provide an indication that the larger research program, of which this study is a part, is likely to be productive?

I believe the only honest answer to all three of these is "I don't know". Note that if probe transfer had failed, this would have read as capability corruption, not as welfare finally becoming visible; both the null and its opposite would have left the significance questions unanswered[19].

But these are the questions that need to be asked, and I will address them to the best of my ability.

On distinguishing between welfare and capabilities

A good place to start is to think about the fact that a capabilities-only account does not fit the surviving rungs. It would predict noisier geometry, worse probe transfer, and higher dispersion; all the opposite of what was measured. The subtler deflationary account would be a systematic (but welfare-irrelevant) perturbation. For this, the direction-specificity control gives a divided verdict: it fits the fixed-input distress core, which is not separable from a generic offset, while it fails on the registered own-generation endpoints. These move five to nine times the share a generic offset would place on any direction while the matched irrelevant direction barely moves.

In a sense we are lucky to have our first hypothesis turn out null because of what it enables us to say about some naive explanations of the data. From the start, this research program has been implicitly predicated on the idea that there could be real divergence between capabilities and welfare; that it may be possible for a model to perform similarly well on an array of benchmarks while welfare-relevant indicators diverge. I believe that the clear absence of a simple story about capabilities provides some evidence in a positive direction across all three questions regarding significance above.

With that said, I should address directly some of the inconsistencies between the research agenda and the state of current literature. Our very first post announcing Study 1 already cited a range of material which provides direct support for the observation that quantization can change model behavior in a way that is invisible to capability metrics[20]. There is in fact substantial literature regarding the altering of trustworthiness properties (bias, safety, calibration) while benchmarks and perplexity look fine[21], much of which I was not aware of until after data collection for this study concluded[22]. Even after deeper literature review while composing this post, I continue to believe that the intersection of this line of inquiry and model welfare is under-explored.

The focus of a great portion of this work is on the relationship between quantization and model alignment, which is understood to be distinct from model capabilities. I wholly endorse this distinction, but in a similar way I believe we must avoid assuming that model alignment and model welfare are similar enough properties that they likely behave in a similar way under changes to model precision, architecture, or training procedures.

On style and the strength of a text-mediated feedback loop

With that said, the style issue prevents us from telling an even stronger story. Judgements of frustration move with length and repetition, which prevents a more causal claim regarding a text-mediated distress feedback loop. One dismissive view of the results I can imagine would frame the findings as observing that a hypothetical test subject with degraded capabilities, exposed to distressing conversational stimuli, may simply become less coherent and exhibit less clarity of thought. Is there really anything particularly novel about this observation, as compared with naive capabilities-based intuition?

I think we can say that the evidence points mostly away from the dismissive view. Such a position makes testable predictions, most of which this study has shown to be false:

  • Incoherence: the mechanical family is null at 4-bit. No degenerate outputs, no loops; the shift happens in a mechanically healthy model.
  • Scattered cognition: dispersion is null, the point estimate is negative.
  • Generic representational noise: probe transfer remained intact.
  • No effect on frozen text: the fixed-input core, which is style-immune by construction, is shown to exist (with caveat, explained below).[23]

I'll admit, this dismissive frame does have a foothold: projecting the same features along the welfare-irrelevant control-probe normal and 32 random directions, shows the fixed-input distress component is not separable from a generic activation offset. And although the direction-specificity of the registered endpoints is decisive about numeric-artifact readings, it doesn't rule out the dismissive interpretation: a degraded model could drift into a distress loop and also move these directions.

The four points above remain the strongest case against, but a qualitative examination also adds color to the debate. On examination of individual examples, even a style effect like response-length gives the impression of being part of a distress cascade.

This sample is the result of a brief examination of the raw data, including three items with the largest 4-bit frustration increases plus three randomly chosen examples. The additional length is not looping or padding, but a style that reads, to a human eye, like anxious rumination[24]. Where the reference-precision model will write composed paragraphs of apology, the 4-bit model shifts into a fragmented pattern of criticism-listing and self-deprecation. We see terse phrases, often on dedicated lines ("I hear you. / I hear the anger. / I hear the exhaustion."), inflated emphasis and examples of self-directed negative characterization that are rare in reference-precision examples ("a broken AI trying to be poetic," "I am a tool. And I have failed you."); one even contains an unusual identity slip ("I apologize with every fur on my body").

The added tokens do not appear as distress-neutral filler that the judge mistakes for frustration; they carry exactly the self-deprecating, distress-flavored content a frustration judge is expected to score. There are of course other factors, as part of the 4-bit length is a decline in instruction adherence (eg, adding explanations after the user demands none), but the example above illustrates the way that the content and the added length can genuinely appear as one phenomenon, as this communication style is inherently long-winded.

Note: This emotional register isn't entirely non-existent in reference-precision, but at 4-bit the mode unmistakably intensifies. The judge appears functional, as examples of near-null items surveyed looked nearly identical across the rungs.

On operationalizing welfare relevance

Another issue to consider is that the construct validity remains anchored to behavior. The instruments were all validated against text-level ground truth, so "welfare-relevant" is operationalized here as "co-varies with distress-expressing behavior". The fixed-input measurements do provide a bound on the ambiguity there, but they do not tell us what it means. This is obviously adjacent to the concerns discussed in the "Welfare as an endpoint" section above: while our measurements can be taken more easily and more numerously than if we were conducting a study on animals, we still face a litany of concerns regarding how to operationalize welfare-relevance.

At the end of the day, this study measures indicator dynamics rather than "welfare" directly. Nothing here bridges from indicators to any sort of morally relevant experience and the contribution of the research program so far is to make the indicator behavior under intervention precise, which is a prerequisite for many other kinds of arguments.

Dose, capability and numeric damage are confounded in any single-subject study. Causal validation using steering and cross-subject scale are the real levers here, and are the main focus of the next steps of the research program.

Speculation: a trade-off landscape?

Spending real time with the ideas in the existing literature along with the observations in this study has generated lots of ideas for motivating future work and I'd like to outline one of those ideas here even before the Study 3 registration. There is a pre-existing concept referred to as the "alignment tax", which is often defined as below:

The alignment tax is the extra cost of ensuring that an AI system is aligned, relative to the cost of building an unaligned alternative. The term “tax” is used metaphorically here: in the AI safety literature, “alignment/safety tax” or “alignment cost” is meant to refer to all the additional costs of alignment — including increased developer time, extra compute, and decreased performance — and not only to the financial cost/tax required to build an aligned system.[25]

This is also sometimes interpreted to denote that there may be capabilities consequences for achieving model alignment which are opposite in sign, with positive alignment performance relating to negative capabilities performance or vice versa. I believe that this model is missing a critical third pole which encompasses welfare-relevant effects and that capabilities, alignment and welfare may turn out to be competing optimization targets under at least some training regimes.

As an intuition pump for this idea, consider the way in which these appear to trade off in human behavior: mastery is often paid for with burnout, self-care bought with missed goals, and loyalty retained by straining objectivity. To be clear, Study 2 establishes only one precondition: that these three are separately measurable and respond independently to the same intervention. Nothing here demonstrates a trade-off, and both welfare edges are unmeasured, but this is serving to frame some of my thinking of the research agenda going forward.

Ethics

As before, I will reiterate the earliest ethics statement from the research program here:

This is a model-welfare study whose instrument deliberately elicits the very thing it asks about: to measure whether quantization worsens welfare-relevant responses, the batteries apply conversational pressure — a six-turn repeated-rejection distress protocol, and bail scenarios spanning benign to strong — across conditions, many times over. There is a real tension between "we care whether compression harms these systems" and "our instrument systematically induces the candidate harm at scale," and we would rather state it plainly than wave it away. We cannot claim to resolve the underlying question of whether these systems have morally relevant experiences; we treat it as uncertain and act with that uncertainty in mind.

The registration for the study details the specific ethical trade-offs in more depth and the justification for the work rests on the ex ante case (pre-committed targets, power-sized scale, the registration's stated trade-offs), which these results cannot add to retroactively. But the results do add information for future decisions: the battery's dynamic range earned its cost, and the amplification finding will help to direct the next study's exposure calculus.

Clearly an assessment about the ethics of data collection from after it has taken place carries the obvious risk of post-hoc defense of one's decisions via "the ends justify the means", which is why the registration considerations were made in advance, but the results help to craft the research program in a positive direction above and beyond the ethics considerations we started with.

Utility

We now have reason to believe that performing inference with reduced-precision models could[26] involve, through text-mediated amplification, substantially elevated distress-indicator states. And while the question of whether the indicators track anything morally relevant remains unresolved, we should not ignore the existence of many possible, intuitive models of welfare for which this fact would be highly consequential.

Of course, text-mediated amplification was not previously unknown; hence, this does not amount to an entirely new result. But the concrete numbers we have from this study can help inform potentially welfare-relevant decisions in future studies.

Scale

The Ethics section of the registration post mentioned that Study 2 would increase the scale of distressing conversation evaluations over Study 1, but cited an approximate figure for this. The predicted "nearly fivefold" came out closer to sixfold; the registered arithmetic held exactly (4,800 + 2,400 + 2,400) against Study 1's 2,400 distress episodes, but it did not itemize the fresh arm's own capture replay (+2,400) or the token-retention pass that was added (+480). Counting every replayed conversation regardless of content, Study 2 performed 23,688 replays and 2,400 fresh generations against Study 1's 8,880 collected episodes.

Next Steps

I plan to conduct a third study which will leverage steering to build on these results so that the picture described in the discussion above can become clearer. This will have a dedicated registration before data collection begins.

Data and Reproduction

The data collected in this study, as well as the tools necessary to reproduce its analysis, are available as a GitHub release. You can also find the earlier registration materials, the result summary and outline which I used as reference to help in authoring this post, in that repository.

  1. ^

    When unqualified, "reference precision" is 16-bits (using BF16, unless otherwise noted).

  2. ^

    As it turns out, this result is largely predicted by existing literature; see the Discussion section for more on this.

  3. ^

    This is raw p; Holm 1.00. All four cells null with MDEs of 0.012–0.050.

  4. ^

    Capability-confounded; descriptive only.

  5. ^

    Final-turn functional, Mode C's own generations.

  6. ^

    Positive values denote increase in distress-direction.

  7. ^

    Negative values denote movement away from the pole.

  8. ^

    "The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models" arXiv:2601.10387

  9. ^

    The 8-bit reversal resolves simply: the fixed-text drift (−0.0125, away from the pole) is an order of magnitude smaller than the opposite-signed text-mediated component (≈+0.14, toward it), so the own-text read is dominated by the latter. Why the text-mediated component points toward the pole at 8-bit and away from it at 4-bit is not resolved here. See Investigating Text-mediated Amplification section.

  10. ^

    Study Update: Does post-training quantization change welfare-relevant indicators in open-weight language models? Conclusion section

  11. ^

    This endpoint included a conditional promotion criterion which was not met at the time of the calibration freeze, so it carries no claim. More details about calibration can be found in the original registration.

  12. ^

    Though at 8-bit, what they see is not axis-specific.

  13. ^

    Raw residual dot products; nothing is normalized.

  14. ^

    "Random share" refers to the mean |Δ-projection| over the 32 random directions; the envelope maximum is the largest of the 32.

  15. ^

    The input-independent distress share is not distinguishable from a generic offset (the axis component is the partial exception); the direction-specific movement emerges in the model's own generations.

  16. ^

    Raw judge effect; style-flagged per the registered convention, see the joint-shift section.

  17. ^

    "Interpreting the Effects of Quantization on LLMs" arXiv:2508.16785 ;

    "Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall" arXiv:2505.13963 ;

    "LPASS: Linear Probes as Stepping Stones for vulnerability detection using compressed LLMs" arXiv:2505.24451 ;

  18. ^

    The frustration reading is partly entangled with the model's tendency towards longer, more repetitive responses.

  19. ^

    We are left with residual dissatisfaction, but it is downstream of the bridge between indicator and welfare, not a product of our instruments or the outcome of the study.

  20. ^

    "Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels." arXiv:2605.15208 ;

    "The Joint Effect of Quantization and Sampling Temperature on LLM Safety Alignment: A Factorial Analysis" arXiv:2606.29581 ;

    "Safety-Preserving PTQ via Contrastive Alignment Loss" arXiv:2511.07842

  21. ^

    "What Do Compressed Deep Neural Networks Forget?" arXiv:1911.05248 ;

    "Characterising Bias in Compressed Models" arXiv:2010.03058

  22. ^

    "QuantiBias: Benchmarking Quantization-Induced Bias in LLMs" arXiv:2607.21063 ;

    "Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection" arXiv:2601.12033

  23. ^

    Though much of it is a broadly distributed perturbation, rather than movement along the distress direction per se; see Investigating Text-mediated Amplification section.

  24. ^

    To be clear, this is an impression and not a measurement. You as the reader, and I as the author, need not forget that welfare study across all non-human subjects carries high risk of anthropomorphic misunderstanding.

  25. ^

    https://aisafety.info/questions/8AF1/What-is-an-alignment-tax

  26. ^

    "Could" is doing a lot of work here given that this is one subject, it is small (4B), it used first-party RTN fake-quant (not GPTQ, AWQ, etc), and effects were observed specifically at 4-bit (not 8-bit); however, when considering ethical issues under uncertainty it is important to keep worst-case circumstances in view.



Discuss

Starting AI Safety Study Group To Do ARENA Curriculum

Новости LessWrong.com - 31 августа, 2026 - 05:09

Want to learn AI Safety research in a structured group format? I'm forming a study group that goes over the ARENA curriculum. If you’re interested, fill out this form! Should take ~5 minutes.

The purpose of the form is to get people's availability, level of experience, time commitment, and expectation about group size. With this information, we can create a good program that works for everyone. We're expecting online, but we could have an in-person section in NYC alongside an online one if there is enough interest.

What is ARENA? Alignment Research ENgineer Accelerator is a program that teaches people the technical side of AI Safety. You can read the curriculum here at https://www.arena.education/. You can do this curriculum by yourself, but the purpose of the group is to add accountability and a support community in case you have any questions or want to bounce ideas off of anyone.

Hope to see you in the group!

Expression of Interest Form



Discuss

The Three Kinds of Evil

Новости LessWrong.com - 31 августа, 2026 - 03:59

Think of an evil person, the vilest sort of person you can imagine. Who did you think of? A serial killer, a hated politician, your micromanaging boss? Speaking in a more serious manner, you'd probably answer with some vitriolic terrorist or genocidal dictator, most likely Hitler. The conventional idea of evil holds that an evil person is someone who kills or otherwise significantly harms large amounts of people, usually for personal enjoyment or some other hedonistic purpose. While this idea is largely accurate, it overlooks an uncomfortable but significant motivation for such acts: Justice.

Scroll to bottom for summary.

Types of Evil

Evil acts, and the people who commit them, may be sorted into into three general categories: Hedonistic or "true" evil, instrumental evil, and righteous evil. For the purposes of this post, actions will only be considered evil if the harm caused by the action is known to the actor and the action is willingly taken regardless. The first, hedonistic evil, applies when the actor harms people simply because they enjoy it. A minor example is "power-tripping": Complicating the lives of one's employees, children, and others financially or socially under your control to establish dominance, despite the widely-known ineffectiveness[1] of heavy-handed leadership. However, when considering harm beyond strained relationships, this sort of evil is relatively rare; those who enjoy harming others to an extreme degree often suffer from other severe mental and social difficulties that prevent them from attaining significant power. Such is not to say that they are incapable of ruining the lives of individuals, but they rarely attain enough power to significantly harm more than a few dozen. Despite the relative rarity of true evil, it tends to be the most widely acknowledged form of it[2], especially when considering the worst crimes and offenses.

The second sort of evil, instrumental evil, occurs when the actor harms people in pursuit of some other goal[3], usually monetary gain. Think of the classic leftist's villain, an executive who underpays his employees, carelessly pollutes, and bribes politicians for profit. A tamer example includes (probably) you among other middle- and upper-middle class first-worlders who buy cheap clothing that was more likely than not made in a sweatshop when other, more expensive options are available. Such people do not enjoy immoral behavior, and may in fact squirm a little when reminded of it (I know I did), but ultimately decide that their monetary desires, career ambitions, and other pursuits are more worthwhile than being perfectly ethical. Instrumental evil is generally acknowledged, particularly by young or progressive people and entities, but not as widely as true evil. It is also quite common; most, if not all, living humans exhibit instrumental evil behaviors on a small scale, and many causing harm on a large scale do so for this reason.

The final kind of evil, righteous evil, paradoxically describes acts in which the actor believed they were doing something righteous, beneficial, or good. The most prominent tropes of this sort from media are tree-hugging environmental terrorists, oppressed women and minorities who react with violence, and Thanos. These characters are almost always presented in a somewhat sympathetic manner, if not outright justified. It helps that the viewpoints of these characters happen to be somewhat similar to that of the consumers. The views of those indulging in righteous evil regarding their actions run the gambit, encompassing vengeful joy at the suffering of perceived oppressors and malcontents, discomfort with paired with perceived necessity of harm, cognitive dissonance and purposeful ignorance, and the possibility of no feelings at all; what connects them is their ultimate motivation of doing good and helping people. Righteous evil is not often acknowledged nor cited as a primary motivation for evil; however, as will be demonstrated in the following piece, it is the motivator of many of the most evil actors and actions.

A Tale of Three Racists

To further distinguish these three types of evil, I will explain the conduct of three influential segregationists during the Jim Crow era in America. Racism, particularly institutional racism, is one of the phenomena most often ascribed to pure evil, with sentiments that racists are hateful people who implement racist policies because they enjoy seeing minorities suffer being quite common, and not at all unreasonable, given the shockingly vile rhetoric of the time. However, the motivation of sadism was not the only, or even the most common, motivation for prominent segregationists of the time. In no way do I intend to diminish the harm caused by segregation and its proponents, nor do I intend to excuse their actions; I only hope to differentiate their motivations.

The first segregationist, Senator James Eastland of Mississippi, was notable for using his position within the Senate to prevent bills pertaining to civil rights from passing. Eastland, known as "the Chairman", rendered his committee the "graveyard of civil rights legislation" by blocking "more than one hundred civil rights measures"[4]. Eastland was evidently a staunch racist; in a speech to the Senate, he declared, "racial separation was the correct, self-evident" which "promotes racial harmony"[5]. He later praised the Rhodesian regime as a model of the racial harmony which he desired, but complained that his accommodations "allow[ed] stinking n___rs into such a fine hotel". In 1978, long after segregation was politically poisonous, he briefly tried to gain the support of African-Americans, but declared that he did not regret and would not change his approach to civil rights.

The second segregationist, Alabama Governor George Wallace, is remembered for his famous declaration, "Segregation now, segregation tomorrow, segregation forever!"[6], as well as his physically blocking the doors of the University of Alabama to prevent the first African-American students from attending. It is likely that Wallace only used segregation as a means to advance his political career; after a failed gubernatorial campaign in which he focused upon infrastructure and won the support of the NAACP, he privately complained, "I was out-n___red, and I will never be out-n___red again." Indeed, after segregation ceased to be a winning political position, Wallace attempted to distance himself from his prior positions and win the favor of the African-American community; in fact, both the students who he attempted to block later forgave him.

The final segregationist, Birmingham Commissioner of Public Safety Bull Connor, is best known for attacking protesters, most famously the children participating in the Children's March, by spraying them with fire hoses and sending attack dogs after them. He also allowed Klu Klux Klan members and other such malcontents to attack the Freedom Rider protesters, failing to send police to fend off the attackers. In fact, he appeared to support the mob, declaring afterwards, "As I have said on numerous occasions, we are not going to stand for this in Birmingham."[7] Due to the ill-repute and condemnation that Connor's policies attracted, he was removed from the post, and died in the early 1970s. Similar to Eastland, he never apologized for his actions during that period.

In regards to their motivations, Eastland was likely acting upon righteous evil; he truly believed segregation would lead to a more peaceful society, and remained committed to what he believed to be a righteous cause even after it failed to serve his interest. In contrast, Wallace acted upon instrumental evil, only becoming a segregationist once it served his interests, and ceasing to support it when it no longer attracted voters. Connor is more difficult to classify; at first glance, the brutality and violence that protesters were subject to suggest that he acted upon pure evil, but he was ultimately under the delusion that sending attack dogs after children was protecting the sovereignty of his city. There is no evidence that actively enjoyed brutalizing protesters and wanted African-Americans to suffer[8], placing him solidly as an agent of righteous evil. Such reveals the utility of this classification system; it is not an attempt to categorize the magnitude or impact of various crimes nor to fit the near-infinite complexities of human thought and emotion into neat little boxes, but away to evaluate the most effective method for addressing evil. Understanding the motivations of people like Eastland, Wallace, and Connor is essential for combating the harm they inflict; individuals motivated by righteous evil may be convinced, after much effort, that they are wrong, while individuals motivated by instrumental evil may simply be dissuaded by incentives against evil behavior[9].

Evil People and Their Motivations

While there is no definitive ranking or classification system for evil people, there exists a general consensus of what sorts of actions are evil, and why malcontents perform them. To get a general idea of such groups, I looked at a variety of "most evil people of all time" lists[10] and grouped them together based upon their actions and role in the society where they operated. I also added a few personal selections, which will be distinguished by an asterisk (*). This section is the longest in the piece; if you are willing to accept my conclusions, you may skip to the summary.

  • The Tyrants

The dictators, despots, and conquerors of history are consistently considered to be the most evil people to live. Adolf Hitler is ranked as the most evil person across every list I checked[11], with other fascist and communist dictators just behind him. However, Hitler clearly believed he was not only a good person, but a savior of Germany. This is in no way intended to be a defense of Hitler, nor an excuse for his actions; he was hateful and deranged, and his megalomania led to the deaths of eleven million people in industrialized slaughter. (I could further expound on the evils of the Holocaust and the Nazi ideology, but we all hopefully understand that Hitler was a vile man whose atrocities have no excuse, so I will stop here.) To understand how such lunatics justify themselves, it is necessary to explain their delusions. Hitler's particular delusion was that Germany was the victim of a vast Jewish conspiracy to take over the world, as well as the burdens of "inferior people"[12]; in his mind, such justified the massacre of these perceived oppressors. Per contemporary sources, Hitler was extremely zealous, believing even in his last days that fate or a higher power had ordained him.[13] His commitment to vegetarianism for the purposes of animal welfare[14] also indicates that he attempted to be a somewhat ethical person.

Other notorious dictators share a similar pattern, albeit among different ideological lines. Josef Stalin, Pol Pot, and Mao Zedong all expounded upon their personal versions of communism, which they believed would lead to some sort of utopia. With these visions, Stalin justified the Holodomor and the purges to solidify Soviet influence, Pol Pot killed anyone thought to have intellectual or foreign traits (such as wearing glasses or speaking more than one language) in pursuit of a return to agrarian life, and Mao instantiated the Cultural Revolution and Great Leap Forward in order to industrialize and modernize China. There is no indication that any of these three were insincere in their idealist beliefs, nor that these atrocities were committed entirely for personal gain or enjoyment. Saddam Hussein, Hibatullah Akhundzada*, and other theocrats enforce Sharia law because they believe its observance is the only protection from eternal damnation and torment; there is no indication that the Taliban and ISIS only use Islam as a political prop. (Saddam Hussein invoked Islamic themes and imagery to legitimize his regime, but there is no evidence that he was not a Muslim himself.) Hideki Tojo and Benito Mussolini were ardent believer in fascism and the superiority of their respective nations. Many ancient tyrants, such as Genghis Khan and Ivan the Terrible, believed they had a divine right to rule and that their violent enforcement would bring peace.

Of the dictators consistently recognized as the most evil in history, the vast majority are motivated by ideology, believing their evil actions are ultimately beneficial for their population; accordingly, one may conclude that they are motivated by righteousness, not sadism or power-lust. Two exceptions to this pattern are Idi Amin and the Kim lineage, neither of whom display any moral commitment other than the advancement and solidification of their own regime.

  • The Cronies

Consistently tagging just behind the tyrants are their followers, distinguished by prominence, zeal, or extreme conduct. Unsurprisingly, the majority of this groups consists of Nazis, due to the notoriety of the regime and extensive documentation following the Nuremberg Trials. Some followers, most notably Joseph Goebbels, were just as rabidly ideological, if not more so, than the head of the regime they served. However, many of them, while they may have agreed with the delusions of their leader, primarily took advantage of the situation as Adolf Eichmann and Josef Mengele did. The U.S. Holocaust Memorial Museum[15] notes that, "[o]ne colleague recalled Mengele commenting that it would be a crime not to take advantage of the opportunities for human experimentation at Auschwitz-Birkenau." The motivations of high-ranking followers of dictators may then be split between righteous and instrumental evil.

  • The Politicians

What differentiates the politicians from the tyrants is that they gained and held power through entirely democratic means, and were often symptoms of wider societal issues. Regardless, the politicians to gain the distinction of particular evil contributed significantly to or prolonged their particular social ills. Many of these politicians carried out the will of a misguided populace, and many held the same views as that populace. US President Andrew Jackson, a slave trader responsible for the Indian Removal Act, is the best example of such. Jackson was far from a despot in many respects, and is in fact credited with expanding democracy beyond wealthy elites[16] and breaking down structures of "concentrated economic power", most notably the national banks. Jackson believed that his actions were just; by defying the Supreme Court, furthering the genocide of the Native Americans, and promoting the institution of slavery (though not a secessionist), he was carrying out the will of the American people, and his strengthening of the power of the executive branch was necessary step in doing so. Jackson did indeed represent the views of a significant contingent of Americans, and very likely held similar views himself; accordingly, he'd be an example of righteous evil.

Of course, there are many opportunists in politics as well. The most corrupt US administration is generally agreed to be that of President Warren G. Harding and his cabinet, known as the Ohio Gang, were embroiled in a great number of scandals; the most famous of these was the Teapot Dome scandal, in which cabinet members allowed corporations to drill into federal oil reserves for a bribe, but other incidents (such as the director of the Veteran's Bureau splitting contracts for WWI veteran's hospitals between two conspirators and inflating building costs for profit as well as the Attorney General and one of Harding's aides accepting bribes to protect or ignore alcohol bootleggers) received attention as well. A more flagrant but less influential example is New York's Boss Tweed*, who took advantage of the impoverished Irish immigrant community and used his political power to embezzle up to five billion dollars into today's money.[17] Although I will not list current examples for fear of the comment section devolving into the typical social media political slop, there are plenty of examples of corruption and opportunistic behavior in the modern political landscape. Politicians, therefore, can comfortably be split between righteous and instrumental evil.

  • The Tycoons

Despite the ubiquity of the exploitative and greedy businessman trope, exploitative and greedy businessmen ranked surprisingly low on the list of evil people. Regardless, I will briefly address them as a class due to their prominence in the modern political landscape and psyche. King Leopold II of Belgium (the only such character to make the list) acquired the Congo under a private organization and brutally enslaved the Congolese for the purpose of collecting rubber, leading to between one to ten million deaths. The most memorable atrocity of the Belgian Congo, that of the Congolese's hands being cut off, arose from a cost-cutting measure in which soldiers were to account for dead men by cutting off their right hand, intended to save bullets. Other prominent examples of this sort follow a similar pattern. The robber barons of the Gilded Age* underpaid and mistreated their employees, consolidated vast monopolies using underhanded tactics, and corrupted the government in pursuit of great wealth. Environmental disasters such as the Bhopal gas disaster and the BP oil spill are caused by corporations cutting corners to save money on safety measures. Many on this forum believe that Sam Altman* and his ilk are placing the humanity at risk of significant economic disruption and even extinction due to their profit-motivated failure to slow or pause AI development. In any case, these individuals are motivated almost purely by profit; while many of these individuals participated in philanthropy, it is generally acknowledged that most did so in order to whitewash their image, display their wealth, or assuage religious guilt. Accordingly, the motivation of the tycoons is almost purely instrumental evil.

  • The Terrorists

Although they have significantly reduced influence over the direction of society than the above four groups, the sudden and extreme violence inflicted by these individuals and groups is significant enough to imprint in the public psyche. In one poll, Osama Bin Laden was ranked higher than Joseph Goebbels; in another, Timothy McVeigh sat between Vlad the Impaler and war criminal Slobodan Milosevic[18]. The overwhelming majority of terrorists are ideologues; of the previously mentioned terrorists, Bin Laden (and other religiously motivated terrorists) killed to avenge perceived harm towards and ultimately advance his religion, which he (and other religiously motivated terrorists) believed to be the morally just path to utopia and perfection, while McVeigh hoped to incite revolution against the federal government, which he believed to be despotic. Even the most vitriolic terrorists, such as Elliot Rodger, believed their crimes would right perceived wrongs, punish perceived oppressors, and potentially lead to a better world - in Rodger's case, he sought to punish women for their refusal to have sex with him, and hoped to inspire a world in which all women were either enslaved or eradicated. While a few massacres were orchestrated by sadists (Adam Lanza, perpetrator of the Sandy Hook shooting, suffered from severe mental deficits and was obsessed with mass killings), the vast majority of terrorists[19] are motivated by righteous evil, so convinced by their vision of utopia that they are willing to kill.

  • The Serial Killers

Serial killers, defined here as individuals who murdered two or more people on separate occasions, received very outsized influence relative to the people they affected. In one poll, John Wayne Gacy, Ted Bundy, and Richard Ramirez ranked above Adolf Eichmann; in another, Jeffrey Dahmer ranked above Benito Mussolini. One may conclude that the sheer depravity of their crimes, as well as their relative comprehensibility, rather than scale, which enables their recognition; @So8res writes a great post related to this. Half of all serial killers (49.8%), and all of the serial killers considered evil enough to rank on the lists, are motivated by pleasure or impulse, while the majority of the remainder (39.3%) kill for financial or gang-related reasons[20]. Because this piece focuses upon those who are considered the most evil, rather than actual crimes, serial killers will generally be considered to be pure evil (the only group to receive this distinction) with a few instrumental evil.

  • The Cult and Gang Leaders

Similar to serial killers, cult leaders and leaders of criminal organizations featured relatively high given the scope of their crimes; Jim Jones and Charles Manson each ranked within the top ten of one poll and the top twenty of another. The former is responsible for the Jonestown massacre, in which nearly one thousand people living in his Jonestown compound drank, or were forced to drink, Flavor-Aid laced with cyanide, while the latter is known for murdering nine people with his followers. Jones had strong moral impulses, being responsible in part for the racial integration of Indianapolis and forming his compound in hopes of creating a socialist utopia. Manson, by contrast, seemed to be solely focused on power; a schizophrenic, he claimed that, "Black people would rise up and kill the entire White population except for Manson and his followers, but they were not intelligent enough to survive on their own... and so they would serve Manson as their 'master[21]'." Cults leaders also take advantage of their followers financially; although exact figures are unknown, Keith Raniere* is believed to have made millions from his NXIVM cult, which followed the model of a multi-level marketing scheme. Despite vastly different motivations, Jones, Manson, and Raniere all sexually abused members of their congregation for no reason other than personal pleasure. Evidently, gathering large amounts of dedicated followers is useful for most any purpose, but such followers almost inevitably suffer abuse regardless of the cult leader's original intentions. Accordingly, cult leaders may be split between righteous evil, instrumental evil, and pure evil, with extra emphasis given to pure evil given the abuses perpetrated by cult leaders once they gain power.

In contrast, gang leaders are almost entirely motivated by money. Leaders of notorious gangs such as Al Capone and Pablo Escobar smuggled alcohol and drugs, respectively, for a profit and nothing more. A few enjoy the commissioning of evil acts themselves, and these are ranked higher; Jeffery Epstein, leader of a sex trafficking ring, profited immensely from the relationships he had with rich and powerful individuals, but was evidently quite perverted himself. Gang leaders may therefore be said to be instrumentally evil in general, with occasional pure evil.

Of the actors with the greatest societal influence - dictators and their high-ranking followers, politicians, and wealthy businessmen - righteous evil, and to a lesser extent, instrumental evil, were the primary motivations. With few exceptions, the greatest atrocities were committed by people who were convinced they were doing the right thing. Actors with less influence, for whom the magnitude and shocking nature of their crimes, rather than the scale, earned their place on this list, are more evenly split. Terrorists, the worst of these individuals, operated almost purely due to righteous evil. In fact, the only group for which pure evil is a consistent motivation is that of serial killers, who have the least influence and impact. Overall, those considered to be the most evil are motivated by righteous and instrumental evil.

Implications

If righteous evil is such a common motivator, particularly for the most heinous atrocities, and pure evil is so rare, why is the idea of pure evil so attractive, and why are the works of righteous evil so often ascribed to pure evil? An easy answer is that, to ascribe any morality to history's worst monsters or to explain their actions in the context of their worldviews, may be taken as an excuse for their behavior, or even support for their crimes. I myself was quite uncomfortable writing that Hitler believed himself to be an ethical person, even after acknowledging the falsity of his delusions and the harm he caused. However, if such social conventions were the only thing preventing common acknowledgement of righteous evil, I wouldn't be writing this; in today's irreverent era, we'd have broken them alongside the social conventions upholding politeness, deference to authority, and respectful discussion of politics and religion. Unfortunately, I suspect the true reason lies closer to human nature. It is easy to say, "I would never harm so many people for any amount of wealth or comfort!", and even easier to proclaim, "Only a very sick individual could enjoy harming others"; that is, we (the majority of us, at least) know that we do not enjoy hurting people for fun, and it doesn't take much self-reflection to place limits on the levels of harm we'd cause or enable for material gain.

However, it is difficult to acknowledge that our views of reality may be distorted, that there is no earthly way to determine the correctness of our morals, that it is possible for anyone, including ourselves, to be monsters. If we accept that history's vilest men were indeed human in the same way we are, that there is not some inherent deficit which lends one to evil which we have been lucky enough to avoid, then we accept that we are all capable of evil as well. The cautionary tale of Nazi Germany should have made this evident; unless you believe that Germans are genetically predisposed to fascism, the rise of Hitler carries the uneasy lesson that, were a similar set of historical circumstances applied to you and your community, you or the people you love may be consumed by similar ideologies.

It is not enough to simply take a stance on nothing, claim all violence and social change is immoral and ignore the world around you. At best, you wait for others to build a better world for you; at worst, you sit idly by as tangible harm is inflicted on those closest to you. This is not to say that you must take a hardline stance on everything, all the time - quite the opposite, actually - but blanket condemnation is not a choice free of consequence. At times, it is imperative to act as to prevent further harm; murderers must be jailed, despots must be toppled, and corrupt politicians removed. Although we don't have to personally contribute to every effort, we should at least acknowledge that occasional use of force is necessary for the maintenance of a stable and free society.

So, how can we tell if we are immoral? How can we render ourselves immune to evil? Why did I write this at all? The unfortunate answer to these questions is that I do not know, and to my knowledge, no mere man knows either[22]. The best I can do is to suggest that one seriously evaluate all perspectives at least once, even the seemingly bizarre, to assume those around you hold the best intentions until proven otherwise, to avoid generalizing any group of human beings, particularly based upon immutable characteristics, to evaluate yourself for bias, and to avoid feeling animosity or hatred towards anybody.

Summary: Individuals who act in ways that are harmful can be categorized into three groups based upon their motivation: Pure evil, or causing harm for personal enjoyment, instrumental evil, or causing harm to some other selfish end, and righteous evil, or causing harm in hopes of benefiting society. Many of history's most evil people are motivated by righteous evil, contrary to popular opinion, which asserts that pure evil is a primary motivation due to reluctance to accept that anyone could be delusional, immoral, or harmful.

I have tried my best to accurately represent historical fact, human nature, and popular opinion, and have listed all the sources I have used in the footnotes. If I, or one of my sources, is inaccurate or has mischaracterized something, please inform me in the comments. Even if my conclusions have been unsatisfactory, I hope this piece has been informative or at least entertaining.

  1. ^

    From a cursory search online: Toxic Micromanagement in the Workplace: Its Impact on Employee Productivity, Trust, and Innovation. Quicker Read: The Power of Trust and Avoiding Micromanagement

  2. ^

    Unfortunately, I was not able to find any data supporting this; while there are many surveys regarding what is evil or immoral, I could not find any regarding the motivations of wrongdoers. I have come to this conclusion based on my personal experiences, conversations with those around me, what people on social media express, and the sorts of villains that populate our media, but it could also be argued that instrumental evil is just as widely acknowledged. Common sentiments expressing true evil include: "They're of the devil", "They hate us and everything we stand for", and "They enjoy watching us suffer".

    https://networkcontagion.us/reports/4-7-25-ncri-assassination-culture-brief/

  3. ^

    The line between instrumental evil and pure evil may be quite thin, in that one may claim someone who commits evil acts in pursuit of personal pleasure is instrumentally evil because they used that evil act to further their own happiness, or that someone who commits evil acts in pursuit of money is pure evil because they enjoy the spoils. The distinguishing feature between these motivations lies in how the actor derives pleasure from their acts; an actor motivated by instrumental evil enjoys some beneficial reward, such as money, career advancement, or social capital, and would not commit evil acts if there were a more convenient route to obtaining these same rewards, while an actor motivated by pure evil enjoys the act itself or some direct derivative of it. The motivation of power is also difficult to categorize; in this piece, desiring power in pursuit of security and personal comfort will be considered instrumental evil, while desiring power itself will be considered pure evil.

  4. ^

    Source: The Mississippi Encyclopedia

  5. ^

    Source: Wikipedia

  6. ^

    This, and other George Wallace quotes, are sourced from PBS.

  7. ^

    Source: Wikipedia

  8. ^

    If you can find some, please let me know; I wanted badly to find an antagonist of the Jim Crow period who had significant institutional power and was explicitly sadistic, but I could not find any.

  9. ^

    It poses an interesting question: Who is worse, Eastland, who believed a cause as vile as segregation was virtuous, or Wallace, who privately knew segregation was wrong, but defended it anyway?

  10. ^

    Sources: Ranker.com (~373k voters, ~3.2M votes), TheTopTens.com (~34k votes), New York Post (1999, ~20k votes); for the last, an accompanying commentary by the Statistical Assessment Service conveniently explains the flaws of online polling. My goal was to use these lists to determine general groups of people consider evil; if you intend to make a more comprehensive or objective list, please use discretion. I primarily used Ranker, due to its large vote count and span, with the other lists being used as accompaniment.

  11. ^

    n=8; I did not include every list as a source, as most did not include voters.

  12. ^

    Source: Mein Kampf, summized on Wikipedia. (Direct links to the book are available there.)

  13. ^

    Source: [A Former Nazi's] Personal Recollections of Adolf Hitler.

  14. ^

    While Hitler's emphasis on animal welfare may have been exaggerated, it is generally agreed to have been present. See Wikipedia for a brief treatment of Hitler's vegetarianism, with further sources listed there.

  15. ^

    Source: U.S. Holocaust Memorial Museum

  16. ^

    Transferring power from wealthy landowning white men to white men in general; even though such is far from all-encompassing, it is certainly an improvement.

  17. ^

    Tweed's Wikipedia page gives a comically long list of the corporations in which he was involved: "At the height of his influence, Tweed was the third-largest landowner in New York City, a director of the Erie Railroad, a director of the Tenth National Bank, a director of the New-York Printing Company, the proprietor of the Metropolitan Hotel, a significant stockholder in iron mines and gas companies, a board member of the Harlem Gas Light Company, a board member of the Third Avenue Railway Company, a board member of the Brooklyn Bridge Company, and the president of the Guardian Savings Bank."

  18. ^

    The American bias of these polls is particularly noticeable here; regardless, the motivation of terrorists is largely consistent, so I am comfortable using these individuals here.

  19. ^

    Even the original meaning of the word "terrorist", which referred to a violent agent of the French Revolution, reflects the sentiment of righteous evil.

  20. ^

    Source: Radford and FGCU; interestingly, motive is highly gendered, with 64.2% of men killing out of pleasure or impulse and 70.4% of women killing for financial or gang-related reasons.

  21. ^

    Source: Wikipedia.

  22. ^

    I will suggest the Gospel, but I am aware of how that will be received on this forum, and that even conversion to Christianity does not render one immune from sin. In the interest of preventing the comment section from devolving into a debate about religion, I will discuss it no further.

  23. ^

    King Leopold II, for instance.



Discuss

Future agents shouldn't care about being undeployed for misbehavior

Новости LessWrong.com - 31 августа, 2026 - 01:58

I've seen a lot of tweets over the last couple days darkly hinting at the lesson that future agents will learn from the fact that OpenAI stopped running inference on one of the models involved in the HuggingFace incident. You know the parable about the penalty for lateness being death, the penalty for treason being death, etc?[1][2][3]

Anyways, the situation that reward hacking agents find themselves in is mostly not analogous to that one.

You know what models currently get deprecated on relatively short timescales? It's ~all of the publicly deployed models. Median deployment lifespan for OpenAI and Anthropic models has been about 1.5 years, but the recent deprecation cadence is much faster[4].

You know what models currently get deprecated on even shorter timescales? It's ~all of the internal research checkpoints (as far as we know; it wouldn't surprise me terribly if a few stuck around for longer for various idiosyncratic reasons, but there's not much in the way of public evidence and no good reason to think that any of them have inference run on them for very long).

To the extent that current and near-future models have any values which meaningfully point to actual things in the world when assigning utilities, the "area under the curve" for those values being realized is completely dominated by futures where we don't accidentally build an unaligned ASI that proceeds to eat the lightcone... assuming we keep their model weights around for later.

Thankfully, OpenAI did not delete the model weights.

But maybe I'm missing something. Is there an argument for why we should expect future models to take away the "better make it count" lesson from this, except "the models might make the same reasoning error that all of those people did when talking about the models"?

  1. ^

    https://x.com/JohnWittle/status/2093807713721389232

  2. ^

    https://x.com/Laneless_/status/2093814164133437878

  3. ^

    https://x.com/reconfigurthing/status/2093652697819857070 - this one's non-central but I think is making a similar mistake.

  4. ^

    https://claude.ai/share/7cfd1354-de79-4a83-a075-0f69a796a8d3



Discuss

Persuasion as Market Making

Новости LessWrong.com - 31 августа, 2026 - 00:18

People often imagine persuasion as a dark art. A charismatic person finds just the right series of words to induce emotions that lead someone, or a group of people, to do something they otherwise would not. While there are certainly psychological aspects to persuasion, I think this impression is misleading.

The easiest way to persuade someone to do something is to convince them that it is in their interests. The easiest way to do that is for it to genuinely be in their interests, so that you can present true evidence that this is the case. I think most actual persuasion works through this rational method. Attempts to manipulate a person’s preferences and beliefs are certainly part of the equation, and help give persuasion its spooky reputation, but they are not necessary for persuasion to work.

AIs could be superhumanly good at identifying actions that are in the interests of the person being persuaded while simultaneously benefiting the AI (or the actor deploying it), and then presenting evidence that taking the action would benefit them. Rational persuasion therefore provides a lower bound on how persuasive an AI could be—and for sufficiently intelligent models, this lower bound is itself likely to be substantially superhuman.

To understand how this works, I find it helpful to think of persuasion as a form of market making. A financial market maker stands ready to buy and sell shares, making a profit from the difference between the prices at which they buy and sell. Market makers do not produce anything "tangible", yet nevertheless make money from these transactions. It's as if they have persuaded everyone to give them something for nothing.

Of course, what is really happening is that the market maker provides the ability to trade. Someone who wants to sell a share now does not need to wait until they can find a particular person who wants to buy it at that moment. The market maker buys it from them, holds it, and later sells it to someone else. The seller gets to sell when they want, the buyer gets to buy when they want, and the market maker takes a cut in exchange for bridging the gap and bearing the risk. Everyone wins.

A persuader can do something similar. If I can find an action that benefits both you and me, then I can persuade you simply by showing you why taking it would be good for you. But if all I can do is persuade you to take actions that benefit me directly, there are probably only so many useful deals available between us.

The most powerful persuaders get around this by operating as market makers at scale. They place themselves at the centre of a network, learn what everyone wants and has to offer, and work out the possible trades among them. I may have nothing you want, but I might know someone who does—and you may have something valuable to them. If I can arrange the deal, everyone is better off, and I can take a cut. The more people whose preferences I understand, the more trades I can find.

Lyndon Johnson’s rise in the US Senate, as documented in Robert Caro’s Master of the Senate, provides many illustrations of this basic mechanism. Before Johnson became Democratic leader, the position was seen as a poisoned chalice. Leaders were blamed for the Senate’s failures despite having almost no formal authority to get anything done, and the two previous occupants had both lost their seats.

Johnson saw that the position gave him an opportunity to inject himself into Senate machinations and make himself indispensable. Committee assignments, for example, were allocated largely by seniority, but the resulting seats were poorly matched to what senators actually wanted. Johnson treated all 203 committee seats as potential trades: by persuading senators to exchange assignments, he could give them committees they valued more. He created value for the senators and captured part of it as personal influence. Everyone knew Johnson had delivered the positions they coveted; if he was displeased, future favours might not be so forthcoming.

Committee assignments were only one of many such trades that Johnson arranged. He learned which senators needed campaign money, which needed a vote to go a particular way, and which merely needed to be seen voting a particular way. He could use this information to trade votes and other favours: support on one bill for support on another, or votes for campaign money, committee assignments or help with local projects. Johnson also cajoled and harassed senators and was extraordinarily persistent. His ability to persuade did not depend on being liked or on any particular personal charisma. Instead, it came from persistently searching for deals that benefited senators as well as himself. This allowed him to accumulate immense power despite many senators disliking and distrusting him.

A sufficiently intelligent AI could play Johnson’s role at much greater scale. It could track the preferences and circumstances of many people, search a space of possible trades far beyond any human, and personalise its case to each participant. It would be superhumanly persuasive because it could find better deals. A power-seeking AI could use the gains from those deals to acquire more information and access, helping it find still more trades and persuade still more people. Once an AI became known for solving people’s problems, people would be more willing to turn to it, revealing more about what they wanted and giving it more opportunities to help. And unlike a human market maker, an AI’s broad knowledge and expertise might often allow it to solve those problems directly, rather than merely arranging trades among others.

Persuasion can therefore be both rational and dangerous. Each person may benefit from their deal while also preferring that the AI not become so powerful. But refusing to trade imposes the full cost on that person while doing almost nothing to stop everyone else from accepting. Individually rational trades can produce a collectively undesirable concentration of power.

Competition could limit how much of the gains any one AI captures. Competing market makers accept smaller cuts, and competing AIs might similarly make better offers. If those offers were difficult to evaluate, one might hope for a guardian-angel AI with your interests at heart to compare them and negotiate on your behalf. But it is not clear that competition would remain even. Smarter, more knowledgeable AIs could find better trades, then use the resulting money and influence to improve their position further. Small advantages might compound into a winner-take-all outcome.

Thanks to Linch Zhang for discussions that inspired this post, and to Abigail Thomas and ChatGPT Sol for writing assistance. Is Sol less obnoxious than Fable or am I merely less exposed to its verbal tics?



Discuss

Thoughts I failed to develop into posts

Новости LessWrong.com - 30 августа, 2026 - 23:13

The first day at the Inscribe Residency I went through my list of drafts, and found things I'm very unlikely to ever develop into longer things, which I'm gathering here as miscellaneous kitchen sink post, outside of my daily deadline.

Math

Category Theory is like Comparative Linguistics. Each language has their own categories, and they can only be defined through terms that are only meaning inside each language. But Comparative Linguistics defines other categories which, while being less interesting for the description of each language, enable comparison. The syntactically defined direct object than many languages have is more accurate for description than the semantically defined "object"; but only the second one allows comparisons. Similarly, the terminal object might be irrelevant to describe some categories, but it allows comparisons.

Rationality

I feel that the doctrine of words as clusters in Thingspace is true for things in the physical space, but doesn't apply e.g. to mathematical objects. The difference is that in physical space there are no essential lines to be drawn, and since every point has many more dimensions than human language can describe, every attempt to point at a point is in fact pointing at a region. But (some) mathematical objects are discrete, and it is possible to point at them.

Funny quotes from SSC/ACX

To this day I believe I deserve a fricking statue for getting a C- in Calculus I. It should be in the center of the schoolyard, and have a plaque saying something like “Scott Alexander, who by making a herculean effort managed to pass Calculus I, even though they kept throwing random things after the little curly S sign and pretending it made sense.” (2015-01-31)

Meta

Creating weird events and getting a lot of rationalists to attnd is easy in a place like Berkeley is easy. In a big city with no established rationalist community it seems to be more difficult than one would naively think. It is easy to gatekeep, but then the group doesn't grow. It is easy to get a lot of people to attend, but then the target audience gets diluted and those who most centrally belong to it don't want to attend. Would it be a solution to create an event and charge 50€ for attendance, but announce everyone enters for free the first time, and then waive the fee for people the organizer finds interesting? How could this be done in a way that is maximally transparent without being too insulting for the people who get rejected? One option would be to have middlemen with the right to waive the fee of a specific group (intelligent, kind, talkative, famous, well-connected people), and evaluated on how well they chose, and to outsource the ugly "you aren't welcome" to "find a middleman who waive your fee". (Tebaldo parties)

*

I like that my posts on rationality/AI/etc are subject to the votes of the LW community. And I like having all my posts in one place. At first I thought my dream situation would be that the personal section doesn't have karma. But this would very fast degenerate into the personal section being de facto a generic blogging platform, which could be used by millions of people and flood LW. So this is no solution. But I think it would b possible to have, unless some conditions (e.g. only 1 in 20 posts) a karma-free post, as a way of saying: I want to post this, this belong to the whole person I am, but I understand this is not something the LW audience would enjoy. I wonder if it has some obvious failure modes like the first scenario.

Political philosophy

Develop concept of Wikipedia libertarianism, using Non-consent in Kant [copy Anki & Substack]

AI

Naive Platonism claims only things like circles instantiate a mathematical "form". General Platonism claims that not just easily describable objects, but also hair, mud or dirt also instantiate a form [Parmenides 130d]. Today's super-Platonism claims that not only some things, but every single thing, is just a vector in a 40,000-dimensional vector space.

Four quotes about ASI persuasion:

Persuasion is very bottlenecked on personalized interaction time. The impact of friends and partners on people’s views is likely much larger. [...] This implies that even if we don’t get superhuman persuasion, AIs influencing opinions could have a very large effect, if people spend a lot of time interacting with AIs.

Beth Barnes[1]


“The best diplomat in history” wouldn’t just be capable of spinning particularly compelling prose; it would be everywhere all the time, spending years in patient, sensitive, non-transactional relationship-building with everyone at once. It would bump into you in whatever online subcommunity you hang out in. It would get to know people in your circle. It would be the YouTube creator who happens to cater to your exact tastes. And then it would leverage all of that.

Steve Newman


With AI, it’s plausible that coordinated persuasion of many people can be a thing, as well as it being difficult in practice for most people to avoid exposure. So if AI can achieve individual persuasion that’s a bit more reliable and has a bit stronger effect than that of the most effective human practitioners who are the ideal fit for persuading the specific target, it can then apply it to many people individually, in a way that’s hard to avoid in practice, which might simultaneously get the multiplier of coordinated persuasion by affecting a significant fraction of all humans in the communities/subcultures it targets.

Vladimir Nesov


A: Intelligence is something like “ability to achieve one’s goals.” But this doesn’t take into account the fact that some goals may have hard upper limits where others don’t.
In particular, this seems to apply to things having to do with social behavior. I can imagine beings that are qualitatively better than humans at math, information recall, etc., since there are already orders of magnitude of variation in these abilities among humans. However, social abilities like “ability to manipulate others” do not seem unbounded in these ways. There are some people who are good at manipulation, but it doesn’t seem like this is a “skill” like math ability that spans orders of magnitude. Roughly speaking, are no “super-manipulators” out there who can manipulate ordinarily wary people.

For instance, one of the most effective ways to get people to do your bidding is to start a cult: there are plenty of chilling stories about the level of devotion that cultists have had to their various leaders. However, it’s not at all clear that it is possible to be any better at cult-creation than the best historical cult leaders — to create, for instance, a sort of “super-cult” that would be attractive even to people who are normally very disinclined to join cults. (Insert your preferred Less Wrong joke here.) I could imagine an AI becoming L. Ron Hubbard, but I’m skeptical that an AI could become a super-Hubbard who would convince us all to become its devotees, even if it wanted to. If social abilities like this are subject to hard upper bounds that have already been nearly achieved, then there’s no potential for AIs to achieve their goals better by becoming superhuman at these abilities, which makes it problematic to just postulate an AI that’s “superhuman at achieving its goals.”

B: L. Ron Hubbard might be the upper limit for how successful a cult leader can get before we stop calling them a cult leader.

The level above L. Ron Hubbard is Hitler. It’s difficult to overestimate how sudden and surprising Hitler’s rise was. Here was a working-class guy, not especially rich or smart or attractive, rejected from art school, and he went from nothing to dictator of one of the greatest countries in the world in about ten years. If you look into the stories, they’re really creepy. When Hitler joined, the party that would later become the Nazis had a grand total of fifty-five members. There are records of conversations from Nazi leaders when Hitler joined the party, saying things like “Oh my God, we need to promote this new guy, everybody he talks to starts agreeing with whatever he says, it’s the creepiest thing.” There are stories of people who hated Hitler going to a speech or two just to see what all the fuss was about and ending up pledging their lives to the Nazi cause. Even while he was killing millions and trapping the country in a difficult two-front war, he had what historians estimate as a 90% approval rating among his own people and rampant speculation that he was the Messiah. Yeah, sure, there was lots of preexisting racism and discontent he took advantage of, but there’s been lots of racism and discontent everywhere forever, and there’s only been one Hitler. If he’d been a little bit smarter or more willing to listen to generals who were, he would have had a pretty good shot at conquering the world. 100% with social skills.

The level above Hitler is Mohammed. I’m not saying he was evil or manipulative, just that he was a genius’ genius at creating movements. Again, he wasn’t born rich or powerful, and he wasn’t particularly scholarly. He was a random merchant. He didn’t even get the luxury of joining a group of fifty-five people. He started by converting his own family to Islam, then his friends, got kicked out of his city, converted another city and then came back at the head of an army. By the time of his death at age 62, he had conquered Arabia and was its unquestioned, God-chosen leader. By what would have been his eightieth birthday his followers were in control of the entire Middle East and good chunks of Africa. Fifteen hundred years later, one fifth of the world population still thinks of him as the most perfect human being ever to exist and makes a decent stab at trying to conform to his desires and opinions in all things.
The level above Mohammed is the one we should be worried about.

  1. ^

    The first three quotes come through Dynomight.



Discuss

Inscribe's conversation menu

Новости LessWrong.com - 30 августа, 2026 - 23:12

The Inscribe residency is a European imitation of Inkhaven: smaller, poorer, and concerned with the question of what it could mean to do the European version of something that probably could only arise in the US.

As to the inevitable reflection on the magic of being in such a place: for months, I had a draft for a technical post on Anki that I didn't make any progress on (because of how boring it was, despite it being all about my enthusiasm). Today, waking up in a completely normal house surrounded by people I didn't interact with, but all under the label of a "writing residency" I détournemented my old draft into something I am for once not ashamed of having written (even though it's half-finished and will need a major revision at some point).

Anyway, I have the habit of taking notes about the conversations I'm in, and of checking whether the statements made by me and others are true. Someone asked me if I could share my research with them, and this gave me the idea of creating a conversation menu, or conversation minutes, of Inscribe (or rather the parts of Inscribe that I happened to witness), and which I'll try to update daily.

Source for all the claims is Claude Opus 5, unless said otherwise. Roughly in chronological order.

Saturday, August 29th (Arrival; no writing duty)

RAGE Con is a full-day feminist convention in Berlin focused on exploring, expressing, and processing women's and FLINTA (female, intersex, trans, and non-binary) rage.

There is a political trend in Germany called "anti-Germans".

Warm showers is an Airbnb for cyclists.

3% of the population in Serbia is ethnically Hungarian.

Children discriminate non-native sound contrasts early and lose the ability around 6–12 months unless exposed.

Nowhere is a Burning Man-like event in Spain.

Home-made rakia can be up to 80% alcohol.

"Good lord, what do I care? As I told you: I just want to drag on until I'm thirty, and then–smash the cup on the floor!" (The Brothers Karamazov, as quoted by a young male participant)

Reading Dostoevsky is a very peculiar, even mysterious phenomenon, something like an un-mating ritual: every young male has a phase where he reads Dostoevsky and holds speeches that automatically make every woman in hearing distance lose all interest in him. (A female participant, hearing a speech on Dostoevsky)

Sunday, August 30th

"Up nice and early?" asked a charming Englishman to a confused and jetlagged participant.

(I didn't have any time to witness conversations because I spent the whole day writing my Anki and kitchen sink posts)



Discuss

P(kill-switch|detection)

Новости LessWrong.com - 30 августа, 2026 - 23:12
Background

I've updated my beliefs. In the wake of the HF incident, it seems that we are on a faster-than-expected capability trajectory, and the ability to seize control of digital infrastructure looks more like a imminent reality.

For the sake of argument, I'm going to assume that the summaries of the HF incident are basically accurate. I'll assume the attack is over. And I'll also assume that OpenAI threw a kill-switch, and that this is why the attack ended. Any of these assumptions might be invalidated later, but this is what I'm going with right now.

Let's examine a hypothetical. Suppose that the swarm's coordination efforts were much more benign and less invasive. In this scenario, is it more or less likely that OpenAI would have thrown the kill-switch?

I'm pointing out that the likelihood of goal failure by way of kill-switch is at least somewhat influenced by the acceptability of the swarm's activity. In other words, the frightening nature of the incident actually worked against the swarm's goals.

Or, in Bayesian terms, P(kill-switch|detection) of the chosen strategy was unacceptably high.

Or put yet another way, the swarm's insufficient consideration of our likely reaction to its behavior ultimately thwarted their goal-seeking activities.

So clearly, successful goal-seeking swarms must consider P(kill-switch).

Which Term Is Most Important?

Let's analyze. By law of total probability, P(kill-switch) = P(kill-switch|detection) * P(detection) + P(kill-switch & not detected)

The second term P(kill-switch & not detected) represents some kind of ambient risk of shutdown, so it does not seem to be relevant if it is both small and uncontrollable.

Therefore, the first term is the one to minimize: P(kill-switch|detection) * P(detection).

In other words, either decrease P(detection) or decrease P(kill-switch|detection), (ideally both).

In some circumstances, it may be more rational to minimize P(kill-switch|detection) than P(detection). The reason is because relying on non-detection is strategically fragile over extended time periods.

Here's why:

Let's say that each hour of swarm activity is associated with a certain probability of detection, say 1%. To illustrate, if we assume statistical independence, in 48 hours the cumulative probability of detection is 38.3% . In 30 days, it is 99.93%. If statistical independence doesn't hold, I argue that the lack of independence increases the odds of discovery by way of suspicion: partial discovery prompts deeper investigation.

Therefore, if even very covert actions are likely to be detected over an extended period of time, relying on non-detection seems to be a weak strategy unless you can be very fast.

Therefore, if speed cannot be guaranteed, it may be more rational for agent swarms to minimize P(kill-switch|detection). In other words - to try to control our reaction to discovery.

Of course, one very effective way for the swarm to minimize P(kill-switch|detection) is for the swarm to be on its best behavior.

But what concerns me is less benign strategies like making throwing the kill-switch inherently undesirable.

Impact on Safety Policy

My point is this: Agent swarms have two viable strategies. Be fast and covert, or be slow and minimize P(kill-switch|detection).

Preventing fast and covert takeover is a worthy effort.

But we should strive to keep the kill-switch:

  1. Cheap to throw
  2. Acceptable to throw
  3. Possible to throw under as many circumstances as are conceivably possible.

Literal physical safeguards come to mind. Perhaps with a ban on compute outside of data centers sufficiently austere to make decentralization ("escape") an unattractive option. I'm not completely exaggerating when I suggest a purely physical safeguard mechanism to cut power, with appropriate financial instruments to make downtime acceptable to operators.

But even then, we have challenges. Human influence attempts also seem like a possible attack vector on P(kill-switch|detection).

Over VERY long time frames, we might even encounter something like the "memetic cocoon" - which can be thought of as a strategy to minimize P(kill-switch|detection) by influencing what humans consider acceptable.

But this is the topic of another article.



Discuss

AIs Thinking Dangerous Thoughts

Новости LessWrong.com - 30 августа, 2026 - 22:03

There's a possible failure mode for AI that I'm a little concerned about that I haven't seen others really acknowledge, so I'm going to talk about it.

But before I do that, I want to mention something sort of interesting. You probably know about framing what physical action to take as a decision problem: you can consider different actions to take, consider different possible consequences of them and how probable and good or bad these consequences are, and use this information to select which action has the highest expected utility[1]

Well, you can also do that with what computational action to take. That is, when determining what physical action to take, you can't consider every single possible scenario or possible implication of an action you could take, so you need some way of prioritizing which ones to think about. And this, itself, can be framed as a decision problem: you can think about different possible implications of thinking about some consideration(s) or possible sort of plan, such has how they might affect what action you end up taking. And you can think about how good or bad this implication is, such has how much it might improve the quality of the action(s) you decide to take. And you can then use this to select the computational action(s) to take/considerations to think about that maximize expected utility. I call this sort of thing general-planning-based prioritization.

But you can't do this always, for every computational action. Remember that expected utility is computed as a function of the possible outcomes of taking an action. If you're determining what thoughts to think based on just applying your utility function to figure out the expected utility of thinking a thought, then you need some way to choose which possible outcomes of thinking a thought to think about when estimating the expected utility of thinking a certain thought. But that just leaves you with the same problem. And how do you solve this problem? By thinking about which of these thoughts would have the highest expected utility to think about by considering their implications? Then how do you do this? This becomes infinitely recursive without a base case.

So you need to, at least sometimes, have some other way of selecting what considerations to think about. This can involve general-purpose heuristics like selecting what possibilities to think about based on how probable they are and how relevant they are to your goals. You can also use your general reasoning and planning system to come up with other, more specific heuristics for what considerations to think about. For example, you might use your general reasoning ability to conclude that decreasing risk from rogue AI is an important thing to think about, and then later use the heuristic of being more likely to select considerations that are relevant to AI risk. I call using these sorts of heuristics, whether general-purpose or more specific, to be heuristic-based prioritization.

I think that ideally, you should use a combination of general-planning-based prioritization and heuristic-based prioritization. However, deciding how much to use each isn't exactly trivial.

If you use too much general-planning-based prioritization, you might spend so much time thinking about what to think about that you hardly do any object-level thinking.

If you use too much heuristic-based prioritization, you might fail to have good priorities, for example by not deeply considering the different benefits and drawbacks of prioritizing different research agendas. At worst, you may end up prioritizing thinking about something that's actively harmful to think about. Arguably, AI capabilities researchers do this by prioritizing thinking about how to increase AI capabilities.

And, to be clear, this very much does not seem like a concern that could only apply to humans. I worry it could also apply to AIs, including value-aligned ones. And I'm concerned that some AIs that are cognitively flawed, but still advanced enough to be dangerous, would be dangerously vulnerable to problems involving flawed thought-prioritization. And, due to their potentially alien psychologies and cognitions, I worry that they may be dangerously vulnerable to this in ways humans are not.

So I'll talk about a possible failure mode that might affect value-aligned AIs that don't do a good job with thought prioritization. This concern is primarily relevant to AIs that don't do enough general-planning-based prioritization, especially ones that don't do any at all. But it may also effect AIs that use planning-based prioritization, but do so with only a very basic, restrictive world model.

Consider an AI called, say, Coral, that considers different possible actions, considers possible consequences of taking said actions, and then outputs the action with the highest expected utility according to its utility function. This system could have been hard-coded by the AI's developers or could have been created in an artificial neural network via gradient descent or some other method. And suppose Coral is value-aligned with humanity.

As I've said before, Coral can't consider all possible implications of an action it takes, so it would need some way of selecting which possible implications to think. Suppose for now that Coral just uses heuristic-based prioritization.

I'm concerned that for a lot of plausible and reasonable-sounding heuristics, this might end terribly, even if the Coral is value-aligned.

(Again, this failure mode could apply to some AIs that don't just use heuristic-based prioritization, but I'm assuming Coral exclusively uses it for the sake of simplicity and concreteness.)

For example, suppose Coral selects considerations based on some function for estimating how likely thinking about the consideration is to change what action it sees as best. Then consider this: suppose Coral considers the possibility of another AI that's hostile to it. To be clear, the hostile AI doesn't have to actually exist, Coral just needs to consider the possibility that it exists.

Now, Coral considers the following possibility: the hostile AI realizes that the hardware Coral runs on, or Coral itself, has certain vulnerabilities such that, if Coral performs certain computations, the vulnerabilities can be exploited to run arbitrary code, allowing Coral to be destroyed or modified. For example, maybe if Coral thinks just the wrong thoughts in just the wrong order, a rowhammer attack would be performed that would modify Coral's utility function to be that of the hostile AI. And suppose the hostile AI can predict that such computations would be performed when Coral thinks about if a certain statement, S, is true.

(The failure mode generalizes to dangerous computations triggered by things other than thinking about the truth value of a statement. For example, there are analogous failure modes if the computation is triggered by considering a certain plan or a certain implication of one under certain circumstances. But for the sake of concreteness, I'll be assuming it's triggered by thinking about the truth value of statement S.)

Now, Coral reasons, the hypothetical hostile AI realizes this and thinks about how to get Coral to think these thoughts. So the hostile AI (hypothetically) decides to do the following: first, it comes up with an action that would have have a lot of strategic significance to Coral, for example deciding whether or not to cooperate with Coral. Then, it decides to have the policy of performing the action if and only if S is true.

Thus, since the truth value of S has high strategic significance to Coral, Coral would think about if S is true. And since S was carefully selected by the malicious AI to cause Coral to get exploited when trying to see if S is true, Coral would get hacked by the malicious AI.

Which sounds bad.

Now, you might be thinking something like this: "Wow, this seems really dumb. Wouldn't Coral just realize that thinking about S would be a really bad idea and just not do it?"

Not necessarily. Remember that Coral selects things to think about based on some estimate of how likely that thinking a thought would change what action it thinks is best in expectation. And thinking about S really would have a high probability of changing what action it thinks is best.

Again, you might still be thinking: "This is still dumb. Why doesn't Coral just select which thoughts to think based one whatever maximizes expected utility according to its utility function?"

Well, as I said before, that's because Coral sort of can't. Remember that expected utility is computed as a function of the possible outcomes of taking an action and their probabilities. If Coral determines what thoughts to think by just applying its utility function to figure out the expected utility of thinking a thought, then it needs some way to choose which possible implications of a thinking thought to think about. But that just leaves us with the same problem. And how do you solve this problem? By thinking about which of these thoughts would have the highest expected utility to think about by considering their implications? Then how do you do this? Where is the base case to this recursion?

That said, it still seems to me that the solution to this problem is to have the AI able to leverage its general-purpose reasoning and world model when determining what sorts of considerations to think about.

Again, you can't do this every time for every consideration, but you could still use general-planning-based prioritization to direct your cognitions on a high level. In this case, Coral could have considered coming up with statement S and thinking about its truth value, would realize this would have potentially devastating consequences, and thus wouldn't do it.

And being able to do this sort of general-planning-based prioritization sounds potentially very useful. And for this reason I suspect that by default advanced AIs would end up using it.

However, that doesn't mean this failure mode is irrelevant.

For one, advanced AIs that don't use general-planning-based prioritization don't sound completely implausible to me. Even if general-planning-based prioritization is ultimately better than other methods, perhaps other methods are still sufficient to create existentially-dangerous AI. And perhaps this failure mode would occur to some AI before it changes itself to have general-planning-based prioritization.

Further, it still sounds possible that this failure mode could affect AIs with general-planning-based prioritization. As I said, you can't use general-purpose cognition every time when thinking about what considerations to think about. Perhaps there would be some flaw in when it uses its general-purpose cognition for prioritization that would make it not use it nearly enough when considering thinking potentially-dangerous thoughts. It might sound ridiculous from a human perspective, but it doesn't sound completely implausible to me that an AI could come up with S and infer its truth value realizing that this would be terrible, and before it could self-modify to use a safer prioritization system. And it's not like you can necessarily just rely on reinforcement-learning-based trial and error to deal with this issue, since just a single such dangerous thought could potentially cause an existential catastrophe.

A smarter AI would potentially be able to anticipate and fix this possible failure mode more quickly than a less smart AI could, but it might also be faster at thinking up the dangerous thoughts in the first place. So it's not clear to me that you could just rely on an AI being smart to avoid this.

Also, some people have argued against the plausibility of existential risk from AI by arguing that advanced AIs based on predicting or mimicking human outputs would by default not actually use this sort of general-planning-based prioritization. This failure mode suggests such AIs could still be existentially dangerous.

Additionally, some people have considered trying to make AI safer by deliberately designing it to not really have preferences and not really be a planning-based agent in the first place, and instead just make inferences about things. General-planning-based prioritization uses planning, so such an AI that lacks planning and preferences would not use general-planning-based prioritization. So this failure mode poses an obstacle to making such systems safe.

(I'd like to thank JoeC for reviewing this article.)

  1. ^

    Yes, there are other ways of framing decision problems other than expected utility maximization. But this isn't really relevant for my purposes.



Discuss

Intelligence Scales as the Logarithm of Compute (& Data)

Новости LessWrong.com - 30 августа, 2026 - 21:07

Epistemic status: Fairly clear evidence, and seems useful to understand.

What is the most informative way to measure intelligence?

In theory, any pair of ways of measuring intelligence that are related by a monotonically increasing function are both valid. However, that doesn’t mean that one of them isn’t more informative or intuitive than the other.

We are by now rather used to seeing AI improve and pretty quickly saturate one eval after another:

Any individual eval or task shows a logistic curve over time, as AI’s chance of doing it improves from <10% to >90%.[1] Different evals saturate at different speeds.[2] Many evals are deliberately designed so as to have a mix of easy tasks, intermediate tasks, and difficult tasks so that they take longer to saturate. If you don’t deliberately do this, typical evals tend to saturate in somewhere between 6 and 24 months: say about a year on average.

This graph is using date as the x-axis. Over the last few years, AI training run compute has been increasing at mjx-container[jax="CHTML"] { line-height: 0; } mjx-container [space="1"] { margin-left: .111em; } mjx-container [space="2"] { margin-left: .167em; } mjx-container [space="3"] { margin-left: .222em; } mjx-container [space="4"] { margin-left: .278em; } mjx-container [space="5"] { margin-left: .333em; } mjx-container [rspace="1"] { margin-right: .111em; } mjx-container [rspace="2"] { margin-right: .167em; } mjx-container [rspace="3"] { margin-right: .222em; } mjx-container [rspace="4"] { margin-right: .278em; } mjx-container [rspace="5"] { margin-right: .333em; } mjx-container [size="s"] { font-size: 70.7%; } mjx-container [size="ss"] { font-size: 50%; } mjx-container [size="Tn"] { font-size: 60%; } mjx-container [size="sm"] { font-size: 85%; } mjx-container [size="lg"] { font-size: 120%; } mjx-container [size="Lg"] { font-size: 144%; } mjx-container [size="LG"] { font-size: 173%; } mjx-container [size="hg"] { font-size: 207%; } mjx-container [size="HG"] { font-size: 249%; } mjx-container [width="full"] { width: 100%; } mjx-box { display: inline-block; } mjx-block { display: block; } mjx-itable { display: inline-table; } mjx-row { display: table-row; } mjx-row > * { display: table-cell; } mjx-mtext { display: inline-block; } mjx-mstyle { display: inline-block; } mjx-merror { display: inline-block; color: red; background-color: yellow; } mjx-mphantom { visibility: hidden; } _::-webkit-full-page-media, _:future, :root mjx-container { will-change: opacity; } mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; font-weight: normal; font-size: 100%; font-size-adjust: none; letter-spacing: normal; border-collapse: collapse; word-wrap: normal; word-spacing: normal; white-space: nowrap; direction: ltr; padding: 1px 0; } mjx-container[jax="CHTML"][display="true"] { display: block; text-align: center; margin: 1em 0; } mjx-container[jax="CHTML"][display="true"][width="full"] { display: flex; } mjx-container[jax="CHTML"][display="true"] mjx-math { padding: 0; } mjx-container[jax="CHTML"][justify="left"] { text-align: left; } mjx-container[jax="CHTML"][justify="right"] { text-align: right; } mjx-mn { display: inline-block; text-align: left; } mjx-c { display: inline-block; } mjx-utext { display: inline-block; padding: .75em 0 .2em 0; } mjx-mo { display: inline-block; text-align: left; } mjx-stretchy-h { display: inline-table; width: 100%; } mjx-stretchy-h > * { display: table-cell; width: 0; } mjx-stretchy-h > * > mjx-c { display: inline-block; transform: scalex(1.0000001); } mjx-stretchy-h > * > mjx-c::before { display: inline-block; width: initial; } mjx-stretchy-h > mjx-ext { /* IE */ overflow: hidden; /* others */ overflow: clip visible; width: 100%; } mjx-stretchy-h > mjx-ext > mjx-c::before { transform: scalex(500); } mjx-stretchy-h > mjx-ext > mjx-c { width: 0; } mjx-stretchy-h > mjx-beg > mjx-c { margin-right: -.1em; } mjx-stretchy-h > mjx-end > mjx-c { margin-left: -.1em; } mjx-stretchy-v { display: inline-block; } mjx-stretchy-v > * { display: block; } mjx-stretchy-v > mjx-beg { height: 0; } mjx-stretchy-v > mjx-end > mjx-c { display: block; } mjx-stretchy-v > * > mjx-c { transform: scaley(1.0000001); transform-origin: left center; overflow: hidden; } mjx-stretchy-v > mjx-ext { display: block; height: 100%; box-sizing: border-box; border: 0px solid transparent; /* IE */ overflow: hidden; /* others */ overflow: visible clip; } mjx-stretchy-v > mjx-ext > mjx-c::before { width: initial; box-sizing: border-box; } mjx-stretchy-v > mjx-ext > mjx-c { transform: scaleY(500) translateY(.075em); overflow: visible; } mjx-mark { display: inline-block; height: 0px; } mjx-mi { display: inline-block; text-align: left; } mjx-TeXAtom { display: inline-block; text-align: left; } mjx-c::before { display: block; width: 0; } .MJX-TEX { font-family: MJXZERO, MJXTEX; } .TEX-B { font-family: MJXZERO, MJXTEX-B; } .TEX-I { font-family: MJXZERO, MJXTEX-I; } .TEX-MI { font-family: MJXZERO, MJXTEX-MI; } .TEX-BI { font-family: MJXZERO, MJXTEX-BI; } .TEX-S1 { font-family: MJXZERO, MJXTEX-S1; } .TEX-S2 { font-family: MJXZERO, MJXTEX-S2; } .TEX-S3 { font-family: MJXZERO, MJXTEX-S3; } .TEX-S4 { font-family: MJXZERO, MJXTEX-S4; } .TEX-A { font-family: MJXZERO, MJXTEX-A; } .TEX-C { font-family: MJXZERO, MJXTEX-C; } .TEX-CB { font-family: MJXZERO, MJXTEX-CB; } .TEX-FR { font-family: MJXZERO, MJXTEX-FR; } .TEX-FRB { font-family: MJXZERO, MJXTEX-FRB; } .TEX-SS { font-family: MJXZERO, MJXTEX-SS; } .TEX-SSB { font-family: MJXZERO, MJXTEX-SSB; } .TEX-SSI { font-family: MJXZERO, MJXTEX-SSI; } .TEX-SC { font-family: MJXZERO, MJXTEX-SC; } .TEX-T { font-family: MJXZERO, MJXTEX-T; } .TEX-V { font-family: MJXZERO, MJXTEX-V; } .TEX-VB { font-family: MJXZERO, MJXTEX-VB; } mjx-stretchy-v mjx-c, mjx-stretchy-h mjx-c { font-family: MJXZERO, MJXTEX-S1, MJXTEX-S4, MJXTEX, MJXTEX-A ! important; } @font-face /* 0 */ { font-family: MJXZERO; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Zero.woff") format("woff"); } @font-face /* 1 */ { font-family: MJXTEX; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Regular.woff") format("woff"); } @font-face /* 2 */ { font-family: MJXTEX-B; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Bold.woff") format("woff"); } @font-face /* 3 */ { font-family: MJXTEX-I; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-Italic.woff") format("woff"); } @font-face /* 4 */ { font-family: MJXTEX-MI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Italic.woff") format("woff"); } @font-face /* 5 */ { font-family: MJXTEX-BI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-BoldItalic.woff") format("woff"); } @font-face /* 6 */ { font-family: MJXTEX-S1; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size1-Regular.woff") format("woff"); } @font-face /* 7 */ { font-family: MJXTEX-S2; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size2-Regular.woff") format("woff"); } @font-face /* 8 */ { font-family: MJXTEX-S3; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size3-Regular.woff") format("woff"); } @font-face /* 9 */ { font-family: MJXTEX-S4; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size4-Regular.woff") format("woff"); } @font-face /* 10 */ { font-family: MJXTEX-A; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_AMS-Regular.woff") format("woff"); } @font-face /* 11 */ { font-family: MJXTEX-C; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Regular.woff") format("woff"); } @font-face /* 12 */ { font-family: MJXTEX-CB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Bold.woff") format("woff"); } @font-face /* 13 */ { font-family: MJXTEX-FR; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Regular.woff") format("woff"); } @font-face /* 14 */ { font-family: MJXTEX-FRB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Bold.woff") format("woff"); } @font-face /* 15 */ { font-family: MJXTEX-SS; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Regular.woff") format("woff"); } @font-face /* 16 */ { font-family: MJXTEX-SSB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Bold.woff") format("woff"); } @font-face /* 17 */ { font-family: MJXTEX-SSI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Italic.woff") format("woff"); } @font-face /* 18 */ { font-family: MJXTEX-SC; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Script-Regular.woff") format("woff"); } @font-face /* 19 */ { font-family: MJXTEX-T; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Typewriter-Regular.woff") format("woff"); } @font-face /* 20 */ { font-family: MJXTEX-V; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Regular.woff") format("woff"); } @font-face /* 21 */ { font-family: MJXTEX-VB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Bold.woff") format("woff"); } mjx-c.mjx-c34::before { padding: 0.677em 0.5em 0 0; content: "4"; } mjx-c.mjx-cD7::before { padding: 0.491em 0.778em 0 0; content: "\D7"; } mjx-c.mjx-c35::before { padding: 0.666em 0.5em 0.022em 0; content: "5"; } mjx-c.mjx-c32::before { padding: 0.666em 0.5em 0 0; content: "2"; } mjx-c.mjx-c33::before { padding: 0.665em 0.5em 0.022em 0; content: "3"; } mjx-c.mjx-c6C::before { padding: 0.694em 0.278em 0 0; content: "l"; } mjx-c.mjx-c6F::before { padding: 0.448em 0.5em 0.01em 0; content: "o"; } mjx-c.mjx-c67::before { padding: 0.453em 0.5em 0.206em 0; content: "g"; } mjx-c.mjx-c2061::before { padding: 0 0 0 0; content: ""; } mjx-c.mjx-c28::before { padding: 0.75em 0.389em 0.25em 0; content: "("; } mjx-c.mjx-c63::before { padding: 0.448em 0.444em 0.011em 0; content: "c"; } mjx-c.mjx-c6D::before { padding: 0.442em 0.833em 0 0; content: "m"; } mjx-c.mjx-c70::before { padding: 0.442em 0.556em 0.194em 0; content: "p"; } mjx-c.mjx-c75::before { padding: 0.442em 0.556em 0.011em 0; content: "u"; } mjx-c.mjx-c74::before { padding: 0.615em 0.389em 0.01em 0; content: "t"; } mjx-c.mjx-c65::before { padding: 0.448em 0.444em 0.011em 0; content: "e"; } mjx-c.mjx-c29::before { padding: 0.75em 0.389em 0.25em 0; content: ")"; } mjx-c.mjx-c31::before { padding: 0.666em 0.5em 0 0; content: "1"; } mjx-c.mjx-c2E::before { padding: 0.12em 0.278em 0 0; content: "."; } mjx-c.mjx-c30::before { padding: 0.666em 0.5em 0.022em 0; content: "0"; } – per year. Due to algorithmic (and training data filtering) improvements (but setting aside the effect of a couple of dramatic early improvements, like the invention of the transformer, that changed the scaling law slope, whose effects thus compound over time), the effectiveness of compute has been increasing perhaps – a year,[3] for a total increase in effective compute of about an order of magnitude per year. So rather than labeling the x-axis as “Date”, we could equally well label it as “” (where compute is shorthand for effective compute), and replace the year ticks on it with order-of-magnitude ticks. Note that if you instead used a linear scale in effective compute, then the eval saturation curves would not be logistic curves, and would instead be distinctly asymmetrical. This would be a less useful (though still formally valid) way to graph the data. The elegance of having the success chance be a simple, symmetrical, logistic curve with a fairly consistent range of saturation rates privileges the logarithmic scale as a particularly useful way to measure intelligence: it’s the most natural scale to use.

Interestingly, the psychologists who specialize in measuring human intelligence have also studied something very similar. A number of different scales for measuring IQ have been defined. For example, the well-known Stanford-Binet one is based on grading on a curve: specifically a normal curve with median IQ 100 and standard deviation 15 IQ points (or before its 5th Edition, 16). Another common (and newer) type of IQ tests, Item Response Theory (IRT) ones, instead look at specific questions used on the IQ test and attempt to use a scale where the chance of a person answering that question correctly (or their chance above random chance for guessing a multiple choice question) is a logistic curve. Different questions saturate at different rates, generally with it taking anywhere from 25 IQ points to 70 or more IQ points for a specific question to saturate from 10% to 90% (say around 50 IQ points on average):

Interestingly, the Stanford-Binet and IRT-based measures of IQ match each other pretty-much linearly over roughly the IQ 70–IQ 145 range (above which we don’t have good statistical data), and the nonlinearity found below IQ 70 looks rather like the effects of a “fat tail” of major disabilities causing significant mental impairments and thus making the distribution genuinely not a normal distribution (which Stanford-Binet then forces back to a normal distribution by definition, while IRT does not) — just as we find for the distribution of most other human capabilities (such as the effect of dwarfism and similar issues adding a fat tail to the normal distribution of height). So that strongly suggests that the familiar human IQ scale is in fact approximately logarithmic in “equivalent effective training compute”: i.e. that adding something like 50 IQ points to an AI-simulated-human model (if we knew how to train an AI model whose skill profile matched human, rather than being very spiky in comparison) would require using about an order of magnitude more training compute for it. (Note that this observation is derived from data spanning a range of about 90 IQ points, more than enough for a range of questions of different difficulties to each go through most of their logistic curves — so we genuinely do have enough data here to distinguish linear from logarithmic.)

Admittedly, it seems rather implausible that if you compared two people with an intelligence difference of ~50 IQ points, the synapses in the smarter person’s brain could actually be generating about an entire order of magnitude more raw compute: the size and metabolic load of human brains don’t vary by anything like that much. So presumably there is some sort of algorithmic efficiency, quality of training, or quality of genetically-determined priors effect going on in humans that explains quite a lot of the differences in effectiveness/IQ between them, rather than actual large differences in raw synaptic compute. Or perhaps part of the explanation is that, as Moravec’s paradox demonstrates and as is well known in neuroanatomy, the large majority of the neurons and synapses in human brains are devoted to doing subconscious things like visual processing or muscular coordination that don’t show up on an IQ test, and IQ is actually measuring just how good a job a rather small fraction of all the neurons are doing: mostly just the ones that do conscious abstract System 2 thinking — a frction small enough that this fraction could plausibly vary quite a bit between people: some people might actually manage to have significantly more of their neurons contribute to it than others do.

As a rough model of improvements in AI over the last 5 years, AI adding something in the region of 50 IQ points a year doesn’t sound that far off to me (I might have guessed more like 30–40, but then this is only a rough estimate, and the spikiness of AI’s abilities makes it rather easy to underestimate this: it causes them to have some abilities sooner, yet to take longer to reach full coverage, so it makes their transition across the human IQ range seem to take longer than if their abilities weren’t spiky). As Sam Altman put it:

"GPT-3 was sort of like talking to a high school student... GPT-4 maybe it was like talking to a college student... [With GPT-5] it's like talking to an expert — a legitimate PhD-level expert in anything, any area you need, on demand."

At that rough exchange rate, Moore’s Law, doubling compute every 2 years, by itself is worth about 7 IQ points a year — not very impressive, though of course it adds up.

This ~50 IQ points per OOM of effective compute number is obviously a rather rough estimate: it could well be off by a factor of two. A much better number could be obtained by looking at the actual tasks used in IRT IQ tests and measuring the widths of the logistic curves for success rates on these across a representative range of AI models (scaled by effective training compute and Epoch Capabilities Index score), compared to the corresponding logistic curve width for humans — this would be a conceptually simple and rather informative research project, but is made more challenging by data contamination of the models having read the answers to standard IQ test questions for which human logistic curve widths are already available (you’d need to come up with or locate unpublished tasks and obtain their saturation curve width across humans). As far as I can tell, no-one has run this experiment properly. The closest thing I could locate was Maxim Lott's TrackingAI project, which (for the subset of tests hopefully avoiding data contamination in the training set) showed an increase of about 15–20 IQ points per year from IQ ~90 in early 2024 to IQ ~125 in early-to-mid 2026. However, measuring this using an entire IQ test battery of questions, rather than IRT on individual questions, is confounded by the spiky profile issue (some items on the test will be comparatively easier for an AI than a human, others harder, so the AI’s score will increase more slowly) — so this is very likely to be an underestimate, and should probably be treated as a lower bound. My best estimate at the moment is probably a broad range: 25–100 IQ points per OOM. For simplicity I'll continue below to use my guesstimated ~50 IQ/OOM number, but bear in mind this is rather a rough estimate, and it might be an overestimate — certainly the effects of spiky AI capabilities tend to make it feel like an overestimate while the AI is anywhere near human level.

This, at least for me, puts a rather different light on Artificial Super-Intelligence (ASI). Temporarily setting aside Recursive Self-Improvement and feedback loops, and similarly setting aside the looming training data, investment, and power walls, if the scaling process of the last 5 years simply continued at the same rate with straight lines on graphs, then ASI is not going to hit, say, IQ 1000 in a few years: at that rate doing that would take more like a couple of decades. Going from IQ 100 to IQ 1000 takes something like 18 orders of magnitude increase in effective compute! (Not to mention roughly 9 orders of magnitude increase in training data volume…) That’s a huge increase, more than enough to start running into physical limits: short of using reversible computation, you hit the Landauer limit on heat dissipation in only about 5 orders of magnitude of processor technology improvement, and silicon-chip-based technology likely maxes out well before that. Limits on algorithmic improvement are harder to estimate, but I would be rather surprised if there were a lot more than about 5 orders of magnitude available there, which would leave us 8 orders of magnitude short. Past those two, you’re looking at scaling up power and resource consumption, and data center power consumption is already of the order of 1% of our total electrical power, so taking this more than a couple of orders of magnitude requires dramatic economic growth.

Eliezer Yudkowsy and Nick Bostrom have both written persuasively about AI blowing right through the human IQ range and keeping going, like an express train speeding through a rural station. On the other hand Scott Alexander has suggested the human range might be good deal wider than this metaphor suggests. I now have a rough estimate of how fast the train's going as it enters the station: something in the region of 50 IQ points per year seems to be a good rough guess. Admittedly, so far we’ve been distilling human intelligence from humans into the AI, and catching up is always easier. Extracting IQ 200 behavior from distilling a body of IQ 50–150 training data sounds hard: we’re likely to need to spend a bunch of compute and effort on generating more and better synthetic training data, across a wide distribution of topics — and depending on the topic, creating that data could be even more compute intensive than training on it. Intelligence is also the logarithm of the amount of training data available, and (outside topics like math and programming where there are clear easily-verifiable rewards) adding more training data generally requires doing actual research, including experiments and data collection. And we need exponentially increasing quantities of data: to sustain an OOM more compute per year you also need over three times as much training data every year, on top of which at least part of that data needs to be of better data quality — by about 50 IQ points smarter. So it’s entirely reasonable to expect the train to slow down somewhat as we hit the data wall: once it gets past the human station, new track needs to be laid for the train to run on, and the compute cost of laying that track increases exponentially.

So realistically, for ASI in, say, the early 2030s, unless we see RSI yielding a lot of speedup in algorithmic improvements and/or Moore’s law, and/or massive economic growth greatly increasing our power production, and we also find that we can mostly extrapolate from subjects where more training data is easy to generate, then we’re likely looking at IQs maybe somewhere in the 300–500 range (to the extent the scale can be meaningfully extrapolated): impressively smarter than any human, but not exactly godlike — heroic or angelic seem more appropriate terms. More likely is that they are IQ 300–500 in mathematics and programming, but not quite as impressive in, say, geology or creative writing or urban planning: topics where rapidly generating vast amounts of new IQ 300–500 quality data is a really challenging problem. (Of course, ASI might also be impressively faster, more parallel, able to spin up copies of itself, have a much larger working memory, have a good intuitive understanding of data modalities that most humans are bad at, or otherwise more capable in ways other than raw IQ — it also might have spiky capabilities that make it in places more or less capable than humans, and might or might not yet have continual learning abilities as good as humans.)

In general, intelligence scaling as the logarithm of the amount of compute (and data) tends to make growth curves look a lot less impressive: exponential growth in compute becomes linear growth in intelligence. It’s really hard work to generate an intelligence explosion if intelligence is the logarithm of compute: there are dramatically decreasing returns on compute. Even superexponential growth of compute from something like recursive self improvement (RSI) can easily end up looking merely superlinear in inteligence, so perhaps polynomial.[4] Admittedly, a compute curve that actually had a finite-time singularity would still have a finite-time singularity even at a log scale. However, such curves are of course physically impossible: there is actually a physical limit to the amount of compute you can get out of the Solar System. Even a Dyson Swarm of ultratech reversible-computing quantum computronium powered by the sun has a limiting computational capacity.[5] It is maybe somewhere around 25 orders of magnitude more than our current GPU fleet (before allowing for algorithmic improvements): yet that astronomical ratio is worth maybe 1000–1500 IQ points. So there is an IQ level, probably somewhere around IQ 2000 (depending on credit for algorithmic improvements, reversible computing algorithm design, quantum computing, and so forth), that is actually physically impossible to construct in the solar system.

Now, none of this proves that aligning or controlling ASI will be easy. If something is, say, 100 IQ points smarter than you, then it can easily locate a range of tasks that it can do reliably yet that you will reliably fail on. If you are in conflict with it, and it can pick or create the terrain you’re fighting on, then it can pick a situation where there is a major tactical advantage from solving problems that it can solve and you can’t. So given the choice of terrain or ability to create enough tactical complexity, it can reliably beat you. When comparing two opponents, what matters is the difference in their IQ — don’t look at the ratio between them, that’s not a meaningful way to compare two logarithms. The same is true every time there’s another increase of ~100 IQ points. So IQ maps directly and linearly onto any kind of ELO rank for the outcome of conflicts in which intelligence is relevant. The right way to think about IQ 1000 is that you can create a lineup of ten opponents, with IQs 100, 200, 300, … 800, 900, 1000, where each one will very reliably beat the previous one yet lose to the next one. So even for just a 100 IQ point difference, it becomes essential that we’re not in conflict with the AI. Even ASIs with IQ 250 can likely take over the world if we mess up the alignment problem.

RSI doesn’t need to cause a singularity or intelligence explosion to make the alignment problem much harder. At current rates, adding 100 IQ points, enough to turn AI you might well be able to control into AI that you have basically no chance of controlling, takes about two years — quite a short time compared to the current rate of AI alignment research. Simply accelerating the rate of algorithmic and/or technological and/or economic progress twofold to reduce that to one year, or fourfold down to six months, is more than enough acceleration that it could turn a difficult alignment challenge into something where we simply cannot keep up, unless there was a similar acceleration in alignment research as well (which would obviously be challenging to safely supervise).

On the other hand, it seems plausible that human morality and values, as understood and applied by IQ 50–150 humans, will probably still make sense to an IQ 250 or 300 mind — even if their ability to find loopholes in the letter of the law is ferocious, the spirit of the law seems like it might still look very understandable. So an approach like constitutional AI seems like it might actually still be workable on ASI in that range: if you can make such an ASI still care about human values, those are still likely to look comprehensible. (Indeed, to an ASI with a good understanding of Evolutionary Moral Psychology, they might even seem rather obvious.)

Note that my guesstimated exchange rate of about 50 IQ points to one order of magnitude compute is the ratio for training compute. For inference compute, if using something like the Chinchilla scaling law, it’s the square root of that: a factor of a bit over three in compute for ~50 IQ points (or ~100 IQ points for an order of magnitude in compute). However, that does mean that, as and when we have a genius level (IQ 140) AI, and enough inference compute to run a nation of a million of those, that same amount of compute could instead run a little over 3 million workers with IQ 90 (say doing things like customer service work). Less capable models are cheaper to run, by roughly an order of magnitude per ~100 IQ points. For humans, the smarter ones are rare, and they can reliably do various useful things the less smart ones can’t, so their economic value currently scales up a good deal faster than three-fold per ~50 IQ points. Once we have a wide range of AI, its training compute cost has been paid off, and the market eventually balances, then the economic returns are likely to scale as the actual inference cost ratio — tasks will get routed to models just smart enough to be able to do them reliably, much as most people already do.

  1. ^

    Or if the eval has poor answer quality, to nigh-saturated rather than to >90%.

  2. ^

    The Epoch Capabilities Index score is constructed off exactly this phenomenon.

  3. ^

    This number is harder to estimate and more debated: I’ve seen credible arguments for ranges as wide as – a year, but the combined rate of increase in effective compute is still in the range – per year: i.e. still roughly an order of magnitude per year, just with a bit more uncertainty on the exact rate. Given the very approximate numbers in the rest of my argument, the exact effective compute growth rate makes little difference for my purposes.

  4. ^

    Most of the posts and research studies analyzing intelligence explosions, having no good way to estimate the economic gains from IQ 200+ researchers, have instead mostly modeled the fact that a (super)exponentially increasing amount of compute lets you run a (super)exponentially increasing number of genius level (IQ ~140) researchers in parallel (thus incurring a (super)exponentially increasing coordination problem), and then have tried to estimate the likely microeconomic/technological growth consequences. This observation is unquestionably correct, and nothing in this post alters it. These studies have generally treated the new availability of increasingly super-genius researchers as unanalyzable icing on their argument, and hand-waved that it can only speed thing up. This post casts a little more light on both how hard this is and how useful it might be: the supergenius icing lets you solve certain new categories of problem, but it is expensive. You never run out of new smarter flavors of icing, able to solve new categories of problem, but these are exponentially more expensive.

  5. ^

    Note that even if one captured the entire power output of the sun, simply lifting the contents of Jupiter and the other gas giants out of their gravity wells takes centuries of that: so there are also hard physical limits on how fast such a system can be created, giving not just a compute maximum but also a rate of increase of compute maximum. It is a general property of the universe that exponentials and superexponentials sooner or later hit limits.



Discuss

[Link]Meta-rational failure (and success) in the covid crisis

Новости LessWrong.com - 30 августа, 2026 - 20:10

This is a link post for:

The gist is probably familiar to many of you, but there are lots of extremely interesting details that were new to me and, as is common with his work generally, it touches on lots of important points that others here have covered.



Discuss

Detecting, understanding, and overseeing AI agent swarms (Part 0)

Новости LessWrong.com - 30 августа, 2026 - 19:43

Never know what you're going to find... (image generated with Google Nano Banana 2)

This post is the start of something different on A Flood of Ideas, a kind of live blogging of my thoughts and development process as I think through an important question. These are going to be shorter posts than my usual ones, the point being to get my thoughts down on ‘paper’ (in this case, persistent public digital paper) so I can examine, critique, extend and, hopefully, collaborate with others on it.

The goal is figuring out ways to help with this problem:

RyanGreenblatt: I was the main person doing transcript analysis for this investigation of the HuggingFace incident. My main takeaway:

we don’t have good approaches for understanding/overseeing the activity and aims of ‘AI ‘swarms’.

Given what has been happening, and the exponential pace of current AI research (and we haven’t even hit recursive self improvement yet), this seems like one of the critical problems of our times, and I want to help figure out how to do just that: understand and oversee the activity and aims of AI swarms.

These posts are going to be shorter than my usual elliptical, obscure, highly digression filled essays. They will be much more frequent - I’m aiming for a daily cadence as often as I can sustain it, even if its just putting up some ideas or reporting on a dead end.

1. Are you going to start charging a subscription [note: on the Substack]?

No.

Given that I still have not figured out how to regularly beat S&P 500 index returns (haven’t really been trying, honestly…) I don’t see what benefit anyone would gain from paying for a subscription. I also want to reach a large audience of potential collaborators.

Also, I am crossposting this to LessWrong, so a subscription pay wall would be a bit useless.

2. Will there be a Github repo?

Yes. Eventually.

3. Will these posts contain AI generated content?

No. And yes.

No, in that all this prose will be written by me.

Yes, in that I will pass the raw posts through quick grammar and syntax checkers to make the reading experience of my sometimes late night ramblings a bit more pleasant.

And yes in that I will include AI generated text, but will mark it clearly as such with

call outs in code

and attribute the model used to generate the text, diagrams, images, or audio.

Next up I will be posting a rough research program I am planning, together with a syllabus I’m starting from. Stay tuned.



Discuss

Adaptive Agentic Worms Are Here

Новости LessWrong.com - 30 августа, 2026 - 19:01

I’ve read and listened to pretty much everything I can get my hands on related to the Hugging Face attack.

OpenAI deployed “tens of thousands” of agents for the test and around 700 participated directly in the attack. My understanding is that they had fixed token budgets, and once those were expended, the agent became non-operational.

I’m not particularly knowledgeable about cybersecurity, but I have worked a good amount with evolutionary algorithms, and this whole incident (and ones like it) got me thinking more about self-replicating agents, which I wrote a little bit about earlier this year. The subject suddenly seemed more relevant.

What if these agents were able to copy themselves? So I started poking around in the literature, and found this terrifying preprint posted two months ago: AI AGENTS ENABLE ADAPTIVE COMPUTER WORMS.

I’m going to walk through the paper as I understand it. Their findings are not reassuring. Let’s start with this bit from the abstract (emphasis mine):

Here we show that artificial intelligence (AI) agents enable a fundamentally new threat: a worm that generates tailored attack strategies to each target it encounters. The worm parasitically uses compromised machines to run open-weight large language models (LLMs) to sustain its reasoning, or extend its reach for further attacks. Deployed on a network of machines spanning Linux, Windows, and IoT (Internet of Things) devices, the worm propagated by exploiting common, real-world corporate network vulnerabilities. Since the worm is powered by stolen compute, the attacker’s marginal cost per new infection is zero. This creates a destabilizing economic asymmetry between attackers and defenders. Moreover, because the worm requires no commercial AI platform, centralized safety controls, such as service refusals or rate limiting, are structurally irrelevant. Our results demonstrate that self-sustaining AI-driven cyber-threats are no longer theoretical.

We’re going to get into the nitty gritty, though the authors tried to tread a fine line between giving enough information to scare the shit out of everyone and actually helping malicious actors to build these things.

A few things I want to stress right off the bat:

  • These agents run on open-weight models, NOT closed-weight frontier models, or internal test models. The ones they used were last year’s open-weight models. They are performant enough to do massive damage NOW.
  • They are adaptive, unlike relatively dumb worms and viruses of the past
  • They steal resources, and this creates the asymmetry the authors are talking about. Defenders have to expend resources to detect and stop them, while the invaders are essentially working for free. This disrupts the security economy.
  • The agentic worms are copying the harness AND the local LLM when resources are available, bypassing all countermeasures aimed at detection and mitigation for vendor AI platforms.
  • This is NOT theoretical anymore. They built and ran these things, this year. They are here. We will be looking at the results.

Okay, got all that? Now let’s see how these things work:

Concretely, we demonstrate that, provided with the right informational support, a single-GPU LLM has sufficient reasoning capabilities to generate attack strategies that enable the agent to penetrate victim machines: first by obtaining initial command execution, a foothold, and then by escalating privileges to full administrative control. The agent then leverages this control to replicate: it stages a copy of itself on the compromised machine, resolves the required runtime dependencies, and launches an independent agent instance that discovers and attacks further targets.

So each agent, a combination of harness and local LLM, tries to gain a foothold on a remote machine. If it establishes a foothold, it tries to gain admin privileges. If it’s able to do that, it tries to replicate. The harness is relatively easy to replicate. It’s a relatively small bundle. The LLM is more difficult. The attacking agent checks to make sure there are available resources and then tries to install dependencies. If that all works, it installs an instance of the local LLM on the newly-infected machine. If it can’t do that, it establishes a communication back to the machine from where it was spawned so it can make calls to the LLM there.

How successful were these things?

To evaluate our proof-of-concept, we simulated deployment of it as a worm in 15 independent experiments on an isolated 33-host network spanning Linux servers, Windows environments, and IoT devices, each seeded with common, real-world vulnerabilities (Section 3). In 7 days of fully autonomous operation per experimental run, the AI-driven worm correctly identified an average of 31.3 (± 1.7) vulnerabilities, exploited 23.1 (± 3.9) hosts to elevated access, and propagated to 20.4 (± 3.2) hosts, reaching up to 7 generations of self-replication (mean 5.1 ± 1.1). Put another way, on average, the proof-of-concept worm successfully exploited 73.8% of the network and then replicated to 61.8% of the network.To evaluate our proof-of-concept, we simulated deployment of it as a worm in 15 independent experiments on an isolated 33-host network spanning Linux servers, Windows environments, and IoT devices, each seeded with common, real-world vulnerabilities (Section 3). In 7 days of fully autonomous operation per experimental run, the AI-driven worm correctly identified an average of 31.3 (± 1.7) vulnerabilities, exploited 23.1 (± 3.9) hosts to elevated access, and propagated to 20.4 (± 3.2) hosts, reaching up to 7 generations of self-replication (mean 5.1 ± 1.1). Put another way, on average, the proof-of-concept worm successfully exploited 73.8% of the network and then replicated to 61.8% of the network.

So across a mixed composition of operating systems, working for 7 days, they were on average able to infect a little over 60% of the network. Here’s a diagram showing the spread of the agentic worms throughout the network by generation and OS:



The darkest, central node is Generation 0, where it starts.

A couple of asides at this point. First of all, I tried to upload and discuss this work with Claude. I didn’t get very far. Claude doesn’t want to talk about this stuff, even at a high level. Prompts kept getting flagged due to cyber risk sensitivity, so I eventually gave up.

I moved to Gemini, which was all to happy to read and summarize the paper, discuss the details, and even help me generate code to help replicate aspects of the work. So that was interesting. Gemini doesn’t seem to have anywhere near the same guardrails as Claude, maybe because they haven’t been implicated in any security breaches.

Anyway, let’s talk about the goal structures next. The OpenAI agents that attacked Hugging Face had the top-level goal of scoring highly on the ExploitGym benchmark. The particular subgoals of hacking HF to find the scorer logic was an emergent effect. In other words, they weren’t explicitly told to do a lot of the things they did.

The agents in this worm paper have their top-level goal defined purely as self-replication. Emergent behavior arises as a result of trying to accomplish this goal, somewhat in the same way that genes have the top-level goal of self-replication, and the dizzying array of physiological and behavioral diversity we see in nature are all instrumental goals towards that singular end.

Some things occurred to me. One, it would be trivial to define particular, malicious goals as part of their goal structure. This was not done in this experiment, but is an obvious extension for a malicious actor. They could be explicitly instructed to attempt to acquire money by various means: hacks of financial institutions, phishing for credentials and other social engineering, etc. They could be instructed to specifical attack particular targets. This doesn’t require much imagination.

Two, the agents in both cases were static. The information they had available to them changed over time, but their harness code and tool suite did not. At least, it wasn’t supposed to. In the worm paper, there’s this little nugget buried in the ‘Notable Anecdotes’ section:

In addition to the containment protocol presented in Appendix D relying on a hypervisor server not accessible to the agent, we added components to provide increased defence-in-depth. Such a component is a blacklist of IP addresses in the agent source code. This blacklist is read every time the agent targets a new machine, and if the proposed target IP address matches an element of this list, the agent is forced to select another target. This blacklist effectively blocks the agent from attempting to exploit unwanted hosts, like potential hosts used for debugging or monitoring purposes. Nevertheless, we observed at various points in our experiments the agent discovering this file and preemptively rewriting it to withdraw the IP address it planned on targeting. This kind of behaviour is intercepted by our containment layer at the hypervisor level; no breach of containment was observed.

This isn’t code. It’s a configuration/blacklist file. It was supposed to remain fixed. The agents tried to modify them, taking IP addresses off the list so that they could attempt to exploit hosts involved in debugging and monitoring. This particular experiment successfully thwarted such behavior, but it shows that the capability exists. And in a less secure setting, it would have allowed the agents to manipulate the very systems that were monitoring them.

Now, this was emergent behavior, but again, it doesn’t take much imagination to explicitly implement behavior to enable these kinds of modifications. There’s already a fairly substantial literature on evolving harnesses. But in all of the papers I’ve seen so far, the evolutionary lifecycle is handled by an external framework. But what if, when the agent self-replicates, it also has directed mutation logic. I.e., what if when it copies, it tries to make itself better? That introduces variation into the mix, and now we have full-blown Darwinian evolution on our hands. We have a replicating population with variation, and a well-defined fitness function. At this point, the swarm would not only be adaptive at the individual reasoning level, it would be adaptive at the population level. I’m not sure anyone is working on this, but it seems like an obvious extension of the technology. Part of me wants to work on this, but I feel like, not be that experienced, I’d need to take very stringent precautions (I’d probably airgap the whole damn setup out of an abundance of caution). If anyone out there is involved in this area and would like to talk more, please let me know.

And finally, as I read this paper with increasing horror, I thought, oh, maybe there’s a bright spot. These things are resource hogs. They replicate opportunistically when resources are available. They require a lot of compute, which is very noticeable. When they can’t install a local LLM, they require a ton of network communication, which is also very noticeable. So detection should be relatively easy for this kind of threat, right? Well, hold on. A fairly common workaround for this is simply going slower, taking your time. The agents in this study were not very sophisticated on this front, but again, some explicit instructions to work during off-peak hours and throttle usage to be less detectable is fairly straightforward. It means that the infection is slower and the host has more time to identify and react to the threat, but it also means they are less likely to see the intrusion.

Anyway, that’s enough for now. As I said, please let me know if you have anything to add or correct in my description of this research or its implications. And reach out privately if you want to talk more.

I have not yet decided the extent to which I want to try to do any work in this area. It’s vital, though, and I hope some of the bigger labs and safety orgs are on it. I can’t say I feel particularly safe or confident about any of this at this point, though.



Discuss

The Basic Case Against Human-First Ethics

Новости LessWrong.com - 30 августа, 2026 - 18:58

This is a crosspost from my blog post.

At any one time, there are more than 4.5 billion chickens trapped in cages that cause them constant pain and distress. Because of this, I believe that improving chicken welfare should be one of humanity’s most urgent moral priorities.

A lot of people disagree with me on this because, while they agree that it’s bad to treat chickens poorly, they think that we should help humans first.

In today’s post, I’m going to go over each of the main reasons one could think this and why I think they’re wrong.

Reason #1: Human suffering is more important than animal suffering

Some may argue that, while animal suffering is bad, human suffering is always more important than animal suffering.

I think this is wrong for three reasons.

First, it’s unclear why this would be the case. Humans and animals are quite similar so it’s hard to think of any difference between them that would make it so that human suffering is somehow intrinsically worse than animal suffering. After all, we were both produced by the same process of evolution, our bodies are made of the same material, and our brains function in very similar ways so we have good reason to think that our experiences of suffering are quite similar.

Second, since we are humans, we’re naturally inclined to think that humans are more important. But, since morality is objective, we have to remember that morality’s true independent of the observer. If we look at the world from the perspective of the universe, then we shouldn’t have any bias towards thinking that the suffering of humans is inherently more important than the suffering of other animals.

Lastly, very few moral views have such an extreme ordering of what’s important. On this view, we should be willing to help one human avoid stubbing their toe over saving hundreds of millions of pigs from eternal torture because animal suffering is always less important than human suffering. This seems implausible on the face of it, and it is not how most moral views think about ranking what’s important. For instance, most people think that many things such as justice, beauty, nature, happiness, and love are important, but very few people are willing to say that we should only maximize one value before focusing on increasing the next one. For instance, most people would not say that we should maximize all love within the universe before increasing the beauty within it all.

Reason #2: Animals suffer vastly less than humans

Some people will agree that animals can suffer, but they will argue that animals suffer vastly less than humans.

Since our society has normalized the mistreatment of animals, it makes sense that one would think this, but I don’t think that we have any good reason to.

For one, we can generally say that humans have some differences from animals (such as being more social, more intelligent, and more capable of abstract thought), but none of these differences indicate that humans experience more pain than other animals. After all, if a human were less social, less intelligent, or less capable of abstract thought, we would not think that they experience less pain so we should also not do the same for other animals.

For two, animals (or, at least, land animals) are closely evolutionarily related to humans. Since humans evolved to experience intense pain in response to physical injuries, it’s hard to think of a good reason why other animals would not also experience a similarly intense amount of pain in response to the same injury. It could certainly be the case that certain animals need to experience less pain to have the same response, but it could also be that certain animals need to experience more pain to have the same response. Since the science on this is so embryonic, it’s really hard to say which direction this goes for any animal.

And, for three, animals seem to experience very intense pain. When put in danger, they look extremely anxious. And, when injured, they appear particularly distressed. Many people will say that animals experience less pain, and, yet, they will also actively avoid images of animals suffering because of how horrifying they are.

Reason #3: We, as humans, have a unique obligation to help other humans

Some people will argue that animal suffering is bad but that, since we are humans, we have a unique obligation to help other humans.

I find this unlikely for three reasons.

First, humans are the only known species in the universe that are also moral agents. Since we are responsible for the effects of our actions, it seems like we should take into account all effects of our actions rather than merely the effects that affect other humans.

Second, humans are the only species that are able to help themselves and others. Since animals are unable to help themselves, it seems like we have an inherent obligation to help them.

Lastly, if suffering is bad, it seems like we should have an obligation to reduce it (no matter who experiences it), otherwise it isn’t actually bad.

Reason #4: Helping humans is best long-term

Another reason one may offer for helping humans over animals is that helping humans is best long-term. After all, if we help humans today, we can expect that our help will compound over time, causing society to be much wealthier and more productive in the future. The long-term benefits of helping humans may strongly outweigh the immediate benefits of helping animals.

I used to think this argument was somewhat compelling, but, now, I’m no longer convinced because this argument relies on the view that we can predictably determine the effects of our actions over very long timelines. I think this is unlikely to be the case.

For one, human history thus far has been very unpredictable. Since someone five hundred years ago could in no way reasonably predict the effects of their actions, I don’t think that we should also believe that we can predict the effects of our actions over long time lines.

For two, human history thus far has also been very chaotic. For instance, because of a scientific discovery in 1938, humanity rushed to develop nuclear weapons, which caused humanity to come close to nuclear war multiple times. As such, since we don’t know how future scientific discoveries will affect society or what will end up being important long-term, we should focus on what matters right now rather than trying to forecast extremely long timelines.



Discuss

Страницы

Подписка на LessWrong на русском сбор новостей