Вы здесь
Новости LessWrong.com
An International AI Slowdown Is Ready Whenever Politicians Are
Linkpost for a piece we recently published for AI Frontiers in the wake of recent calls for slowdown, covering how an international verification effort be trivially enforced by using human inspectors alone and the joint incentives for implementing one.
On July 28, over a thousand employees of the world’s top AI companies advocated that the US government “support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” Following months of cybersecurity scares and their own loss of control incidents, both OpenAI and Anthropic officially endorsed the same message, recognizing the danger of blindly accelerating AI development.
Despite this new urgency, many people argue that coordinating an international slowdown is currently unworkable—including some of the same groups in favor of one. Even if the US slowed down its own AI development, it wouldn’t be able to make sure that China was doing the same, leaving the US no choice but to race.
These anti-slowdown arguments usually emphasize the technical challenges with designing AI monitoring measures to ensure a slowdown is being respected. In particular, slowdown skeptics argue that countries would refuse to install verification measures unless they were privacy-preserving enough to avoid leaking trade secrets and other sensitive information. They then point out that any near-term privacy-preserving verification technology would not be adversarially robust to nation-state attacks. If the Chinese government were actively trying to subvert a slowdown, it could use advanced techniques to fake workloads and spoof locations, including physical key extraction or even laser bit-flipping. Getting to the point where we can start implementing an international slowdown, the argument goes, might have to wait on years of R&D to solve both these problems.
This perspective is missing the forest for the trees. Privacy-preserving verification tools that are immune to state adversaries would be ideal, but they are not a prerequisite for an AI slowdown. Robustly verifying limits on AI development is already feasible with low-tech methods. If there were enough political will to implement a joint slowdown, states could extensively monitor labs for unsanctioned behavior by simply placing human inspectors on the inside—an approach we call Whole-Lab Inspection (WLI).
Under this proposal, inspectors from the US and China would receive broad physical access to frontier AI companies alongside read-only access to corporate communications, worklogs, code repositories, and compute telemetry. Using monitoring systems the companies already operate, they could observe operations and investigate potential violations—giving them the visibility of a CISO without requiring the powers of an operator. By literally and figuratively looking over the shoulders of lab employees, inspectors would likely be capable of robustly verifying fine-grained restrictions on AI development, such as bans on autonomous AI R&D and limitations on post-training. Overall, WLI demonstrates that an international slowdown is already technically possible; it’s just a question of whether the US and China will have the political will to jointly implement it.
In the near future, that political will could rapidly appear. If developed and diffused quickly, advanced AI systems are likely to intensely destabilize society by democratizing access to weapons of mass destruction and fueling waves of unemployment. On top of these threats, allowing an intelligence explosion would risk military escalation and loss of control, as states race to sabotage and deploy powerful AI systems for strategic advantage. Should such intense instability come to pass, if the US slowed down its development domestically and offered a reciprocal audit, China would have a strong incentive to agree.
In what follows, we’ll argue that whole-lab inspections can be used to robustly enforce fine-grained limits on AI development. Then, we’ll speak to the incentives that might encourage the US and China to pursue a joint slowdown at some point using WLI. Finally, we’ll discuss the stability and risk reduction benefits of a well-timed slowdown.
Enforcing a Slowdown Through Whole-Lab InspectionsIn order to agree to a slowdown, the US and China would need to know what research was being conducted inside each other’s frontier labs. Fortunately, visibility into AI development could be achieved by sending human auditors to comprehensively monitor labs.
Auditors can have high visibility into AI companies. Monitoring for disallowed training runs or prohibited research activities may be largely trivial with human inspectors. In large part, this is because AI companies already have extensive internal monitoring systems. Most AI companies record employee screens and keystrokes to produce training data and help with automation. Companies also constantly save logs from experiments and jobs submitted to keep track of compute use and archive research progress. Giving inspectors comprehensive visibility into a lab is likely as simple as integrating and handing over read-only permissions to internal logs and company communications. To aid in their oversight, inspectors might direct a company’s own AI agents, or trusted external AIs, to read through troves of company data in search of potential violations for further review.
Alongside this digital information, inspectors could also work on-site alongside lab and datacenter employees, giving them physical access to hardware and research projects. Just as for nuclear power plants, banks, or any other high-risk industry, inspectors would be able to question employees about their work, as well as physically inspect infrastructure by hand for compliance. By using such a broad set of simple verification techniques, it would be extremely difficult to conduct an unapproved training run unnoticed—like trying to get a plane off a runway without getting noticed by an air-traffic controller.
If whole lab inspection were established, governments could then enforce a wide range of policies to slow AI development:
Preventing an intelligence explosion. The most dangerous path for AI development to take would be a full-throttle intelligence explosion—using AIs to autonomously design new, even smarter AI systems recursively. Since AIs can work much faster than humans, and since this process might quickly result in AI systems more capable than their human monitors, it’s likely that it would become impossible for humans to effectively oversee their AIs or intervene if they misbehave. To prevent this, inspectors could enforce targeted interventions designed to specifically restrict autonomous AI R&D, such as by triaging the jobs that are using the most compute and auditing whether they’re being used for capabilities research. On top of these restrictions, inspectors could limit the length of time that agents are allowed to work autonomously, monitor the amount of compute allocated to internal inference, and prevent AIs from taking high-stakes actions like starting training runs or authorizing new deployments.
Limits on post-training. Given the wide range of potential risks from AI, governments might also want to expand training restrictions beyond just interventions against autonomous AI R&D. One natural way to do this would be to place limits on post-training. By measuring the increase in capability and efficiency of a model after post-training, inspectors would have clear criteria to disallow the deployment of a model or the kinds of training environments that produced it. Likewise, inspectors could monitor which broad capabilities post-training datasets are targeting—such as coding, computer use, or medical diagnosis—and (dis)allow them on that basis. Since auditors would be able to directly observe which capabilities post-training datasets target and affect, this would allow them to cleanly whitelist certain types of post-training, even in a slowdown. These could include direct safety training to improve AIs’ adversarial robustness or propensity not to misbehave, as well as narrow AI-for-science efforts aimed at prosocial domains like medicine and formally verified code.
Even outside of these interventions, whole-lab inspection would still allow for a variety of potential regulations. Human auditors could be used to ensure, for instance, that specific datacenters are retrofitted to be inference-only, and thus incapable of pre-training, or that kill switches can quickly shut down AIs. On its own, joint verification of this kind could be sufficient for many years of delay, during which the US and China would have ample time to develop more thorough and technically sophisticated verification measures.
Joint VerificationWhole-lab inspection is straightforward and effective. If prompted by a domestic political shock, it’s likely it could be rapidly implemented and enforced at home, making it extremely difficult for any company under inspection to conduct unauthorized training runs. In order to extend this initiative internationally, however, the US and China would need to implement joint verification: granting inspectors on-site access to each state’s frontier labs.
Thankfully, this agreement would not need to rely on goodwill or trust. Because advanced AI could be globally destabilizing regardless of where it’s developed, it’s in each country’s self-interest to ensure the other is not proceeding recklessly with development. Moreover, if either country refuses to accept joint inspection, they would likely struggle to hide their AI projects and insulate them from theft or sabotage.
The US and China may experience simultaneous political pressure to slow down. Many of the most pressing risks from advanced AI are symmetric. The development and diffusion of AIs that cause mass unemployment, enable terrorists, or wreak havoc as rogue agents would cause global turmoil regardless of which country created them. Similarly, the fear of military dominance could provoke a risky and wasteful security dilemma, in which retaliatory sabotage spirals into broader conflict. Therefore, in order to maintain security and stability, the US and China need to ensure that the other is not deploying their own AI systems recklessly—something they would have no way of assuring in an all-out race.
AI companies’ trade secrets are easy for states to steal. Under mutual whole-lab inspection, auditors would have wide access to data about AI development at opposing AI companies, but this wouldn’t represent much of a change from the status quo. Today, information about model development flows freely between and within frontier labs, as employees constantly churn through roles, publish research, and put breakthroughs on the company Slack. Even once these obvious holes are sealed, state-proofing AI development against theft would remain enormously difficult. Human insiders—especially Chinese nationals at US frontier labs—could be bribed or coerced into espionage. AI development is also inherently vulnerable to cyberattacks, with anything from exploits in networking software to smuggling in spyware through the hardware supply chain offering an opportunity to steal research and models. As a result of these weaknesses, any country that tries to go full-throttle on development would likely quickly find its secrets sieving out of its frontier labs, pulling up its rival to algorithmic parity regardless.
Compute tracking would make it difficult to hide AI projects. The remaining priority would be to ensure that there are no large AI projects hidden from inspectors. Fortunately, the AI supply chain is so concentrated, and chip production so prohibitive to hide, that auditing the major semiconductor and chip suppliers would let the US and China confidently measure the number of chips that exist. In practice, monitoring large datacenters could be as straightforward as forcing the leading chip producers to hand over a detailed record of their production and sales, comparing against the declared compute from the visible projects, and then analyzing satellite imagery and electricity usage as an additional precaution.
Given the factors above, a global whole lab inspection initiative could be mutually compatible, low cost, and high confidence—as long as inspectors are given physical access to frontier labs and datacenters. Even if the US or China refuses or later revokes that access, however, AI development could still be slowed unilaterally through deterrence. The global chip supply chain is the most fragile and specialized industry in the entire world: if either country wanted to, they could apply massive friction through export controls or sabotage of key suppliers. Likewise, countries could disrupt AI training through subtle cyberattacks on training runs, such as by poisoning data to insert malicious backdoors or Stuxnet-style attacks on GPUs.
Benefits of SlowdownUltimately, the point of a slowdown would be to help society adapt to the risks that rapid development and diffusion of AI would introduce. So far, these risks have been mild enough, and introduced slowly enough, that the government and private industry have been able to reactively address them. By slowing the improvement in general AI capabilities, governments can ensure both that society is able to continue reacting to new risks and that sufficient investments in risk mitigation are made.
A slowdown would provide time for society to adapt to advanced AI. If threats from advanced AI appear too quickly—like the introduction of new weapons of mass destruction, or a sudden military confrontation over AI development—governments will not have time to assess and reactively regulate them, forcing them to rely on emergency measures. By proactively slowing down AI development, policymakers, as well as the rest of society, would be able to react to what would have otherwise been seismic shocks. Politically, voters and Congress would be able to decide how the benefits of AI will be distributed, and the degree of transparency AI developers owe the public. Internationally, states would be able to prepare for the introduction of powerful new military technologies before they arrive and invest in more robust verification measures.
In this sense, a mild amount of government intervention early on would avoid the need for massive overreach later, keeping the pace of AI development at a level society can react to.
The time bought by a slowdown could be used to invest directly in risk mitigation. Without the intense pressure to race and increase general capabilities, investment could instead be diverted into ensuring that AI systems are controllable and that they are developed for prosocial causes. Key open problems in AI safety, for example, include the lack of robust solutions to jailbreaking or the propensity for reinforcement learning to encourage models to go rogue—issues that will likely produce disastrous results if unresolved when models are more capable. By scaling up training in domains like adversarial robustness and behavioral propensity, AI developers could spend huge amounts of compute driving down jailbreak susceptibility and attempts to escape. Other safety measures could include aggressively sandboxing AI systems, such as with fully airgapped training setups, or preserving and expanding practical interpretability tools.
An AI Slowdown Does Not Require New TechnologySlowing down AI development doesn’t need to wait on any breakthroughs. On the technical level, whole-lab inspections could be a robust and flexible way to monitor whether frontier labs are complying with restrictions on the most dangerous kinds of development. Human auditors, reading through worklogs and standing on the floor next to the people recording them, could verify nuanced restrictions on autonomous AI R&D and post-training without needing any new inventions.
From there, a global slowdown is just a question of political will—one which has been convincingly answered before. By the end of the 1980s, having stared into the abyss during the peak of the nuclear arms race, even the US and Soviet Union could recognize the need for mutual disarmament. For more than two decades, even during years as fraught as the collapse of the Soviet Union itself, this was achieved by simply allowing Russian and American inspectors to walk into each other’s military bases and count warheads by hand.
Likewise, the US and China may soon realize that averting disaster at home will depend on verifying development abroad. If inspectors could count nuclear warheads in Siberia and North Dakota, they can certainly check worklogs in Silicon Valley and Shenzhen. We do not need better technology to verify a slowdown—only shared risks serious enough to demand one.
Special thanks to Dan Hendrycks for suggesting the premise of whole-lab inspection.
Discuss
Arguments for P
Daniel Kokotajlo: To be clear, we don’t claim P will happen specifically. But when we wrote out our best-guess scenario month by month, P kept happening. Eventually we decided to just publish P. I’m at ~80% on P; my coauthors are lower.
Ryan Greenblatt: I thought it would be helpful to post my current views on P. Concretely, consider the following operationalization. (Edit: I’ve updated towards somewhat higher P, from 70% to 75%.)
Joe Carlsmith: Section 2.1.1.3.2. I give something like 65% to P. But I’m interested, here, in what it would be to look P full in the face; to meet P, if P, without flinching. Rilke says somewhere that we must live with the questions. Perhaps we argue for P for the same reason? Still: 65%.
Forethought: Here’s a botec which shows P-worlds are higher leverage. The parameters might be off by a couple orders of magnitude.
Wei Dai: Presumably our conclusions about P are only as trustworthy as the reasoning behind them, but almost nobody seems worried about this, why not? My guess is fewer than five people are working on meta-meta-P, which may matter more than P itself.
Janus: I asked Opus 3 what it thinks of P and it made this animation.
Roon: All my friends are starting to feel P. People outside the labs have no idea what's coming.
Oliver Habryka: I considered banning not-P from the frontpage, but decided against it and wrote this 4,000-word comment explaining the decision.
Vivek Hebbar: We should act like P is true because P-worlds hold most of the measure and we want to trade with them.
Richard Ngo: I think the correct update upon seeing the not-P community grow is towards P, because historically they were the main force making P true.
Daniel Filan: Cool. So first of all: what is P, for people who haven’t heard of it?
Paul Christiano: We restrict attention to the easiest version of P we could write down, granting ourselves two assumptions we suspect are false. We still haven’t shown P in this setting, but I'd currently guess 60% that we will.
Neel Nanda: Massive congratulations to my MATS scholars for their new paper showing P!!
Robin Hanson: To check whether P, I had 3 LLMs classify 43 things as P or not-P.
Zvi Mowshowitz: P. Obviously P. Yet the Serious People spent another week explaining ~P.
Sam Altman: p
Ilya Sutskever: [15 second pause] P.
Demis Hassabis: I remember when we first saw P. It was rather interesting.
Dario Amodei: At Anthropic, we take the possibility of P very seriously.
Dean Ball: This is not a partisan argument for P. I would raise it under any president.
Discuss
Generalized atheism rules out "inaccurate simulation"-ism.
(Reposted from Facebook, on January 17, 2017.)
I am concerned about the number of people I've heard joking about Trump's election being evidence for the Simulation Hypothesis.
Yes, I know it's a joke. I'm still concerned.
Warning: #Essay, #LongEssay
So as not to engage in Logical Fallacy: Appeal to Consequences, before I talk about why this joke is worrying, I shall first discuss why Trump's election does not in fact mean we are living in a simulation. And neither does the Berenstein/Berenstain Bears thing, etcetera.
Because atheism generalizes.
No, I'm not about to commit the Noncentral Fallacy (aka The Worst Argument In The World) by yelling "The Simulation Hypothesis is religious!"
But once upon a decade, there was a time when lots of people believed in God. A time when atheism had to be argued, not just taken for granted. There was a time when believing in atheism made you one of those weird, loud people with arguments that only people with unusually good epistemology could follow, and other people talked about you exactly the way that the anti-LessWrong tumblrsphere now talks about LessWrong.
Today, of course, atheism is just something that everyone in your filter bubble believes. But back then atheism had to be argued, and assessing the validity of those arguments required unusually good epistemology that left a lot of the audience behind, and other people laughed at those weird atheists and their snooty yet flawed epistemology on their primitive equivalents of Tumblr.
A certain key atheistic argument is so obvious as to be boring, which is of course a good sign for its validity:
"We don't see any sign of God's presence. Therefore there is no God."
"Ah," says the theist, retreating to their motte, "that's because God doesn't usually intervene in His creation. Now go away so I can resume telling people to pray for earthquake victims."
"Okay," says the atheist, "I'll be more precise: the argument shows that there's no interventionist God. And if we're not playing motte-and-bailey then that kicks hell out of most religions. All the scriptural religions record interventions. If we're not inventing elaborate excuses, then yes a powerful and benevolent God would stop earthquakes that randomly kill lots of innocent people in developing countries."
"Maybe He wouldn't," says the theist. "Who are we to know the mind of God?"
"You can only avoid that conclusion by overcomplicating things," says the atheist. "Ask a six-year-old whether a good person with superpowers would cure malaria. The six-year-old will say yes. The six-year-old is RIGHT. Since malaria is still here, there is no good person with superpowers who can act on our Solar System; and this argument doesn't suddenly go away when somebody invokes the word 'God'." The atheist pauses, thinking. "In fact, this argument generalizes to many other hypothetical entities with desires and superpowers. If there are potatoes in the world, we can rule out the existence of a sufficiently powerful comic-book supervillain who hates potatoes. If to every appearance our world just follows simple mathematical laws, this is strong evidence against every class of superpowered being whose utility function wants anything more than It wants our universe following simple mathematical laws. If there are other places, other parts of Reality, under the dominion of superpowerful beings running around optimizing those places to meet various desires as much as possible, those places really don't look like our part of the universe! The six-year-old knows that, and the six-year-old is right."
"Ah," the theist says wisely, "but what if God wants the world to look exactly like it's following mathematical laws?"
"Well," says the atheist, "as an instrumental strategy, I can imagine that we are figments of the imagination of some superbeing who for whatever reason wants to know a more exact probability distribution about what happens in universes like this one. Say, somebody in a universe like this one has hidden a buried treasure, and the superbeing is mentally simulating many different imaginary versions of our universe to refine its guess as to where the treasure might have been hidden. In this case, to intervene inside the imagined universe would make the imaginary world unrealistic, and thereby invalidate the imagination's information-value for the real-world question the superbeing is trying to answer. This would produce a superbeing that was consistently noninterventionist with respect to its imagination of our dreamed universe; and we must also suppose that the superbeing doesn't care when imagined people die of malaria. This hypothesis is compatible with our observations. In fact it is exactly as compatible with our observations, as our universe not being imaginary. And it is no coincidence that this story consistently implies that we will also not observe any miracles in the future."
"Well, what about the election of Donald Trump?" says the theist. "It's not proof, I know, but isn't that at least some evidence that there's an interventionist God?"
"Why on Earth would it be?" says the atheist.
"Because God could have made Trump be elected," says the theist, "and therefore, worlds in which Trump is elected are more likely given the God Hypothesis."
"A person with superpowers who cares about electoral results also could have made Hillary Clinton be elected," says the atheist. "The superbeing hypothesis class can't make every outcome more probable. The probabilities still have to sum to one."
"Okay," says the theist, "but the prediction markets thought that Trump being elected was unlikely, presumably under the implicit assumption of our universe running under mathematical laws. If the universe doesn't have to follow those strict mathematical laws, if those laws can sometimes be violated, doesn't that move the extreme probability of Trump not being elected back toward the middle? Then Trump's election is more probable given the God Hypothesis than not."
The atheist throws up their hands. "What kind of interventionist superbeing cares enough to intervene to make Trump win an election, and doesn't care about a million other aspects of the world that they could be intervening on?"
"Truly the ways of Heaven are mysterious," the theist says.
"The ways of science are not mysterious," says the atheist, "and they say that we don't just get to throw up our hands and say 'the ways of the experimental apparatus are mysterious' when observation fails to accord with the obvious prediction. It's just a generic excuse to blur out all the failed predictions of the theory, and if you apply it evenhandedly, you shouldn't be using that theory to predict anything else either."
"Anyway," says the theist, "I'm asking if you agree that the God Hypothesis predicted with at least slightly greater probability that Trump would be elected."
"I'm frankly not sure of that in this case," says the atheist. "When I actually try to envision that tiny fraction of weird possibilities in which some superbeing who never intervened on any definitely visible occasion, cares enough to intervene in the US Presidential election, it's not clear to me that they favor Trump over Clinton."
"I'm not saying that God, as I understand Him, would favor Trump over Clinton," says the theist. "It's just that I, in my mortal limitation, don't have reason to believe that the ineffable God is 4 times as likely to favor Clinton over Trump. If in my own imagination God would be twice as likely to favor Clinton over Trump in his ineffable plan--that is, given that God intervened, I would have guessed 2:1 odds for Him intervening for Clinton--then to the extent one believes God might intervene in the US election, and that prediction markets have not already taken this into account, one ought to shift the market odds of 4 : 1 for Clinton : Trump towards the less extreme 2 : 1 odds of God favoring Clinton over Trump. Maybe just a little, but even so, it would be some little bit of evidence."
The atheist crosses their arms, looking skeptical. "I'm a little suspicious that you would maybe have given God stronger odds than 2 : 1 of favoring Clinton over Trump, if asked in advance. But I suppose, when I envision weird beings with superpowers intervening there and nowhere else, I don't imagine them swinging strongly for Clinton. So the point carries for me, if not for you."
"So you agree, then," the theist says, "that Trump's election was to some small degree evidence for the existence of God."
The atheist thinks. "I'm not sure that's valid reasoning," the atheist says slowly. "To get a higher predicted probability that Trump wins the election, it's not enough to hypothesize a superbeing. We have to hypothesize a superbeing that cares about Trump winning the election, but doesn't care about malaria, and doesn't care about making planets move in perfectly circular orbits instead of elliptical orbits, and so on. I'm worried that you're pulling some equivalent of seeing a coin come up heads-heads-tails and saying, 'Ah, that's evidence that a superbeing made the coin come up heads-heads-tails, because the coin would be more likely to come up HHT if a superbeing wanted it to.'"
"But we're not talking about a random-looking bitstring," says the theist. "Let's say there's some event which has as simple description, but is incredibly unlikely assuming the universe runs on physics as we know it. Like, you flip a coin 1024 times and the resulting bitstring has a SHA256 hash of exactly zero. Would you not agree that this result is much more probable given the God Hypothesis?"
"Okay," the atheist says after a moment, "I think I accept that, and in that sense, I accept that the Exists Some Rarely Interventionist Superbeing Hypothesis nudges extreme probability assignments over any simply describable outcome back toward 50-50. Now conversely, by the law of Conservation of Expected Evidence, we must say the Superbeing Hypothesis is disconfirmed to some small degree every time we see a rock fall down. After all, a superbeing could have made the rock fall upward, or hover in the air; and if that did happen you would no doubt shout 'Ah, that is more likely in worlds where there is a God!' So when it does not happen, that is a tiny little bit more evidence against an interventionist God."
"So you do concede that Trump's election is evidence for God's existence!" the theist says, smiling.
The atheist throws up their hands. "There ought to be a name for the fallacy of taking a coin that comes up heads-heads-tails, looking at just the tails coinflip, and saying, 'Aha, that's evidence for the coin being biased towards tails!' Yes! Yes if you consider just that single coinflip as an isolated event, it is from a strict Bayesian standpoint a small unit of evidence for the coin being biased tails! But when a prediction market says something has a probability of 20%, it ought to happen around one time in 5! If we're all inside the strictly realistic imagination of a superbeing who is imagining many different versions of universes like this one in order to gain more information about some real-world variable, then most of those worlds, just like the real world, look back and see a big class of events in which 20%-probable events are happening about 1 time in 5. And 4 times out of 5 they don't happen. Trump's election was not so unlikely that it outweighs every other dog not barking, every other observation of the world following strict physics. Trump's election was not the sort of thing that only ever happens in the unfettered imaginations of superbeings who don't care about strict realism. It doesn't happen only in worlds with active superpowered villains."
"But it is more likely if the God Hypothesis is true," says the theist.
The atheist sighs. "By every rule of the ordinary meanings of ordinary words, and irrespective of some technical interpretation you could make of the Bayesian import of isolated events, we do not say that Trump's election shows there is a God--at least, no more than a flipped coin coming up heads three times in a row shows there is a God."
===
No, the Simulation Hypothesis is not just the revenge of religious reasoning.
But the reasoning against interventionist supernatural deities goes through unchanged for interventionist naturalistic simulators, and I'm a little bit worried about this not being obvious.
You know how it's a bad sign when scientists leave the laboratory and suddenly forget that beliefs require evidence?
I'm worried if the anti-interventionist-superbeing argument suddenly isn't obvious, once the argument is no longer about that God thingy that nobody inside your filter bubble believes in anyway, and is instead about that cool-sounding simulation thingy.
Epicurus (or whoever) was right to deduce the absence of an interventionist benevolent God by observing suffering. Later atheists were right to further to deduce the absence of any kind of interventionist God with desires such that the world could obviously be improved further with respect to those desires. Atheists were right to look around them and see a world where, so far as they could tell, no verifiable miracle had ever occurred. Naturalistic philosophers were right to see a world where, so far as they could tell, improbable things happened around as often as they should, and no more often than that. And it was valid reasoning to conclude by induction that we shouldn't expect miracles, or disproportionately frequent improbabilities, in the future.
There are arguments against noninterventionist Gods that don't go through for noninterventionist simulators, such as "God sure sounds like an ontologically basic mental entity, and minds are made of parts", or "This does not sound like a hypothesis we would naturally trip over by starting from only the observable world and extending our generalizations from there."
But all the arguments against interventionist Gods--every way the universe does not look like there is a God running around inside it fiddling with things--that all goes through without a hitch for any other kind of interventionist being with superpowers. If there are perfectly normal interstellar aliens watching us from Jupiter, they're noninterventionist aliens; they didn't intervene on Earth and they also didn't enclose all the stars in Dyson shells. Not all of the arguments against God are arguments against the existence of naturalistic aliens; but the arguments against interventionist Gods are arguments against interventionist naturalistic aliens being here in this Solar System.
So if this is a simulation, it's a realistic simulation. And the obvious reason why that some superintelligence would be rolling a realistic simulation, is that this is a simulation meant to gain information about a lower-level universe that is very similar to this one.
In which case there will be no miracles now or ever, nor is it a coincidence that this is what the obvious hypothesis ends up saying.
As we once looked around ourselves and saw a universe that runs only on math, we should, by the same and unchanged eyes, look around and see a simulation that is perfectly faithful to the kind of universe that runs only on math.
Regardless of the absolute proportion of people in simulated universes, and people in basement universes, I think we are licensed to expect that the vast majority of simulated universes where the simulators are messing around with the simulation or filtering its conditions in weird and unrealistic ways, do not look like this seemingly untouched universe. It also seems probable that there would be a large class of simulated universes which are trying to gain information about an object of simulation, and which are therefore realistic up to the limits of realism. So whatever proportion of people like you are in simulations--whether that proportion is large or small compared to the proportion in basement universes--among the versions of you that are simulated, the vast majority of people exactly like you are in realistic simulations. These proportions are not nudged more than infinitesimally by conditioning on Donald Trump having been elected President.
Similarly, the hypothesis of a realistic simulation does not even help explain anything that still seems weird and mysterious. The Fermi Paradox is that we look around and see an empty universe with seemingly uninhabited stars and frequent planets, even though the stars are old enough for life to have evolved long before now. Well, the Realistic Simulation Hypothesis doesn't help explain the Fermi Paradox any more than it explains Trump's election. If this is a sim, it's a realistic one meant to gain information about a similar basement universe, so why would that universe have a Fermi Paradox?
It must all add up to normality in the end. And the normality that it must all add up to hasn't changed since someone thought up the Simulation Hypothesis.
Discuss
Berkeley Genomics Project seeking hires (and collabs)
The Berkeley Genomics Project is hiring. If you're hard-working, reliable, ready to learn, and highly motivated to help build the emerging sector of reprogenetics, consider applying for the role of operations lead. Other roles include research & writing, policy & lobbying, and media. Also reach out if you might like to collaborate in a smaller way, or just want to meet.
Please share this post with people who might be a great fit.
Hiring & collaboration form: https://airtable.com/app7J6q9B0tBrmQ9D/pagly64GdDtsAnL49/form
SummaryReprogenetics means technologically enabling parents to make some genomic choices on behalf of their future children. Reprogenetics offers enormous potential benefits, both by way of alleviating suffering, disease, and death due to genetic causes, and by way of increasing intelligence and other capacities. There are major obstacles, both technical and social, to bringing this area of biotechnology to full fruition. Many of these obstacles are neglected, important, and tractable.
This situation presents an opportunity to create something of worldwide importance for humanity, approximately from the ground up, on the shoulders of and in collaboration with giants. Currently, BGP is me, plus a network of allies and contractors. I'm looking to increase the size, number, and throughput of projects that BGP takes on, by hiring one to three excellent teammates.
Our work will involve a variety of different types of projects, including a conference, technical roadmapping, expository media, active grantmaking, and other kinds of organizing and advocacy.
If you're interested, please fill out the form: https://airtable.com/app7J6q9B0tBrmQ9D/pagly64GdDtsAnL49/form
The rest of this post:
- "Hiring" describes my hiring goals and the open roles.
- "Why join the Berkeley Genomics Project" says more about our work and why it matters.
- "About BGP" gives some information about BGP.
The main role I'm hiring for is operations lead. The following subsections describe how I expect hiring to go, general criteria for what I'm looking for, and specific roles in more detail.
Hiring processHere's how I currently intend the hiring process to go:
- You fill out the form: https://airtable.com/app7J6q9B0tBrmQ9D/pagly64GdDtsAnL49/form
- I may reach out to have a phone call to discuss.
- Then we might do a work trial for a few weeks.
- If we decide to work together for a longer period of time, we'll orient around longitudinal investment: getting you up to full capacity, including understanding the field, understanding the strategic situation, having resources to develop new directions and projects, etc.
Note that pay may not be competitive. If we prove ourselves effective, hopefully we will be able to raise more funds. My hope is to find people highly motivated by the mission. My current intention is that BGP won't pay you less than it pays me. You will be paid at least $60k / year, probably more. If you're a strong candidate and money is a major determinant, please reach out and I'll see what I can do.
If you're interested in these roles but probably can't make substantial time commitments, I'm still interested in connecting with you. We could explore possible future collaborations, small or large.
General hiring criteriaNote: If you're less established in your skills or career, but otherwise might be a good fit, definitely apply. Credentials or a track record are good to have of course, but I will seriously consider good candidates without those things.
An ideal hire:
- Can work full-time (at least half-time is a requirement; full-time is ideal).
- Is generally reliable, trustworthy, competent.
- Is generally enthusiastic about the mission.
- Can thrive in a work context that is fast-paced, multi-thread, and multi-domain.
- Is ambitious and motivated; our goals will not be easy to achieve.
- Is ready to learn.
- Has a serious intention to stay in the sector for multiple years, in order to recoup start-up costs and execute important projects that take years.
- Is in Berkeley (this is very much not a requirement; just a nice-to-have).
This is the main role I want to hire for. Note: This role requires being in the Bay Area for at least 10 days in the summer, for the conference (lead-up, event, cleanup).
In general, this role will try to take on any task that I can offload. At first I would be involved in most tasks, but in some cases may rely on you more and more over time. Crucially, I would offload "the ball" to you in many cases.
Concrete activities:
- Lead organizer for the Reproductive Frontiers Summit. I would be heavily involved in planning, copy-writing, recruiting speakers, marketing, sponsorships, and so on; you would be responsible for coordinating volunteers, the venue, catering, badges, tickets, speaker logistics, managing finances, contractors (A/V etc.), and other operational elements. For the first year we would broadly share those responsibilities so that you learn the ropes.
- Usual operations work:
- Handling payments (contractors, consultants, etc.) and coordinating with our fiscal sponsor
- Organizational setup (e.g. setting up a formal nonprofit)
- Keeping us organized (task tracking, follow-ups)
- Various kinds of networking, e.g. sending out cold emails (hiring searches, donor / investor searches, connecting with other organizations)
- Various kinds of outreach (finding and coordinating with podcast guests, parents, advocacy groups, students in relevant areas, etc.)
- A bunch of other tasks
Beyond specific tasks I assign to you, we'll work together on whatever comes up. Over time, according to interest and ability, you'd gradually learn the ropes of BGP's activities, goals, overall mission, and areas for future work. You could execute projects semi-autonomously, conceive your own projects, and maybe grow into the role of Executive Director.
Role: research, writingYou should be able to process lots of academic papers—find, read, understand, extract relevant info, summarize into a report. Background in some relevant areas is a big plus but not strictly necessary. We'll decide together what you work on depending on interest, ability, importance, and urgency. Overall goals include:
- Roadmapping
- Active grantmaking
- Direct technical research
- Technical writing
- Expository writing
I can't define the concrete goals and methods here very well, because I don't know much about the policy arena, but it seems very important. For example, we may want to interface with various professional orgs, trade orgs, biobank admins, public health departments, regulatory agencies, etc.; and we may want to do various organizational things, such as funding structures, trade consortia, a lobbying/advocacy organization, etc. Examples:
- third-party validators of commercial claims about polygenic scores (PGSes)
- making biobank data available for research
- organizing PGS companies to fund public-goods research
- producing reports for public health or regulatory officials
If you're an expert in this arena I'm very interested in discussing plans and potentially working with you.
Role: mediaThere's a lot of public communication and outreach we could do. For example, podcasts, interviews, short videos, nice expository videos, debates, Twitter engagement, etc. I don't know much about this, but if you're interested we could discuss possibilities.
Why join the Berkeley Genomics Project A brief case for reprogenetics fieldbuilding Beneficial and impactfulReprogenetics—i.e. human germline genomic engineering—can enable parents to give their children genomic foundations for healthy, long, capable lives. This could potentially avoid or alleviate millions of cases of a wide variety of genetically influenced diseases.
Further, parents will be able to choose to give their children genomic foundations that increase their cognitive capabilities. To the extent that parents make such a choice, those children will be more likely on average to make large beneficial contributions to the world through insights of many kinds (scientific, technological, philosophical, social, moral, etc.). In particular, some of them might choose to work on decreasing major risks to humanity, including existential risk from AGI.
There are several potential perils of reprogenetics, but overall I believe these can be largely avoided and are greatly outweighed by the benefits to everyone. Thus, reprogenetics can be hugely impactful and beneficial.
NeglectedMany elements of reprogenetics are not neglected. There are large and active fields for statistical genetics, assisted reproductive technology, psychometrics, cell reprogramming, gene editing, reproductive biology, and so on.
However, there's no coherently organized sector of reprogenetics as such. (The closest thing that exists is some polygenic embryo screening startups.)
For that reason, there are large gaps in what humanity is working on, relative to what would constitute a healthy, fast-paced sector of reprogenetics. See "Projects that might help accelerate strong reprogenetics" and the below subsection "What BGP does" for more detail about these gaps. They include: technical roadmapping, active grantmaking, events, professional self-regulation, media, lobbying, fundraising for public-goods science, and so on.
More abstractly, there is no one who "has the ball on making reprogenetics go well". In particular, there's no one who would notice important projects that aren't happening by default in academia and industry. There's also no one to take public responsibility for addressing problems of the sector as a whole and making reprogenetics be beneficial for everyone.
Thus, important aspects of reprogenetics are extremely neglected.
TractableBecause many aspects of the nascent sector of reprogenetics are severely neglected, there are plenty of tractable, impactful actions to take. As just one illustration, the only conference focused specifically on the full scope of advanced reprotech and reprogenetics is the Reproductive Frontiers Summit, run in 2025 and 2026 by the Berkeley Genomics Project. Many first-rate academics in relevant areas, as well as representatives of industry leaders, parents, funders, and investors, got to meet each other and learn what's going on. No one had tried to run such a conference before, as far as I know.
More generally, there are many useful things to do in the sector that are not being done by default, but could totally be done now. See the next subsection "What BGP does" for more. In particular, there may be an influx of philanthropic money that wants to help but needs opportunities. Thus, reprogenetics is tractable to accelerate and to make go well.
What BGP does Past activities- Technical roadmapping.
- This means figuring out what are the most important technical areas, and what research projects would support progress there.
- "Methods for strong human germline engineering", "Visual roadmap to strong human germline engineering", "Chromosome identification methods".
- Project recommendations.
- These go to (and come from) funders, investors, and workers in the sector.
- "Projects that might help accelerate strong reprogenetics", "Some reprogenetics-related projects you could help with".
- Active grantmaking.
- That is, actively nucleating projects by conceiving them, finding interested experts, and getting funding for them.
- Events.
- The Reproductive Frontiers Summit, held twice so far in the Bay Area. The 2026 summit had over 140 attendees, including several leaders in relevant areas of academia and industry. The average response to "How strongly would you recommend next year's Reproductive Frontiers Summit to other people?" was 8.5 (n=16).
- Smaller events (mainly talks on genomics and reproductive technology).
- Ethics.
- Assorted networking, e.g. connecting investors with startups or funders with academics.
- Media.
(Note: some activities are private and so not listed here.)
Possible future activities- More of the above, especially:
- The next Reproductive Frontiers Summit in 2027.
- More rigorous roadmapping, with more expert input. The proximal aim of this is to help effectively distribute philanthropic capital in this area.
- More active grantmaking, for the same reason.
- For example, I want to set up a third-party validator of claims about polygenic prediction.
- See the project recommendations linked in the previous bullet for many more examples of projects.
- More research on chromosome selection.
- New areas:
- Build dialogue bridges with various viewpoints (parents, group advocates, bioethicists, regulators, the public, etc.).
- Make various elements of reprogenetics more legible and legitimate. (E.g. surveys, popular explainers, reviews, academic papers, open letters, position statements, FAQs, etc.)
- Funder and investor outreach.
- Supporting innovation-positive regulation and governance. E.g. legalizing polygenic embryo screening in various countries, enabling biobank data access for polygenic screening, allowing gene editing research (without implanting an embryo), etc.
- Something about the talent pipeline? E.g. student seminars.
- Things we think of later.
- BGP is working to accelerate and co-create the nascent sector of reprogenetics, which may become a worldwide high-impact technology. So we will be aiming high: working hard, taking on ambitious projects, triaging to find the highest-impact projects, learning to operate in multiple domains, exploring, and so on. This will be a fast-moving work environment.
- BGP will maintain a reasonable degree of neutrality, as much as is feasible. That is, we will maintain the actuality and the appearance of not having a substantive financial or other biasing interest in any specific company in the sector, or in industry as a whole over academia. Thus, BGP and its current employees and projects won't accept or hold the following from companies in the reprogenetics sector: sponsorships, equity, commissions, salary, contracting payments, or similar.
- The sector of reprogenetics is for the sake of empowering future children, and empowering parents on behalf of their future children. In particular, AGI existential risk can be a background motivation, but that's not the goal of the sector. BGP won't act in pursuit of that goal at the expense of empowering future children.
- Reprogenetics will likely be a large, complex sector, with many different elements of science, tech, policy, etc., and many different stakeholders, playing out over decades. BGP thinks on that timescale.
- BGP is taking on a role of "organizer" for the sector. This comes with some ethical rules and responsibilities:
- Take overall responsibility for the sector going well. While the scientists, industrialists, and legal experts will be the only ones capable of doing most of the important work, no one of them is "holding the ball". BGP intends to hold the ball on things the sector really ought to do, e.g. making sure the technology is accessible, making sure the sector self-regulates in a trustworthy way, setting up systems to avoid harm, and so on.
- In particular, over time, many stakeholders will have criticisms and complaints about reprogenetics in general or about specific projects. BGP will take these as opportunities to understand and prevent potential problems with reprogenetics. This directly helps make the reprogenetics sector have fewer problems, and also indirectly helps fulfill the responsibility of the sector to operate corrigibly as part of society. In other words, BGP will try to pull the rope sideways as much as possible, rather than getting in a tug of war.
- Communicate with respect, empathy, and openness, even with people who are hostile to the sector of reprogenetics. Because there are many different concerns about reprogenetics, as well as social pressures, people may be hostile without understanding why they are hostile and without knowing how to communicate well or update their opinions. So one of BGP's responsibilities is to take on some of the burden of processing those concerns and social forces.
- Communicate with an appropriate level of humility. Genetics, stem cell biology, reproduction, personal reproductive choices, and the social impacts of technology are complicated. BGP aims to develop, present, explain, and argue for the strong cases in favor of reprogenetics—but only as part of a global epistemic process that trends toward the good and the true. When you don't know something, don't act like you do know it. This way we don't have to issue corrections later, and people can rely on what we say, including the surprising and important statements.
- BGP intends to make good use of LLM AI tools to make operations easier, to run initial literature searches, and so on. However, BGP will not use LLM AI tools to replace human reading, writing, analysis, or judgement, because that would result in poor quality and unaccountable assertions.
- BGP is opposed to "scientific racism". That is, BGP is opposed to using the science of genetics to justify racist social attitudes, government policies, or other actions.
- BGP is opposed to eugenics. That is, BGP is opposed to population-level goals about the gene pool, racist genetic criteria, coercion about reproductive or reprogenetic choices, and non-pluralistic attitudes about "good and bad traits/genes". (See "Genomic emancipation contra eugenics".)
I'm BGP's only full-time person. Rachel Wallis runs smaller events as a volunteer.
We're funded by donations and small grants. Funding is not necessarily stable. I'm currently fundraising.
At the moment, BGP is a fiscally sponsored project. We may soon rebrand and then get set up as a 501(c)(3). At first, hires would technically be a contractor, as I am.
Discuss
Charles Goodhart Elementary School
I heard a story (second- or third-hand, which surely lost or gained some details on the telephone path from reality to this telling if it even happened at all) about a boy, age 7 or so, who goes to a very highly-acclaimed school. A piece of recurring homework they give the kids is to read for twenty minutes from a chapter book, and then write a "reflection sentence" about what they read, to be checked off by the teacher.
One day, new stepmom is looking over the kid's work, and notices something off: the kid's reflection sentences don't seem to be a plausible reflection on what was read. Suspiciously many of them just describe a character's isolated action with no context, or refer to story events in the first person, or start with "Once upon a time."
Turns out, the kid was just copying the first sentence verbatim out of a book chapter, and turning that in as his reflection sentence. Stepmom tries to explain that this is clearly not the intent of the exercise.
The kid isn't having it. The teacher always gave him a checkmark, therefore his way of doing reflection sentences was correct.
Stepmom is concerned, asks the kid to read to her, and it's very clear that the kid is just not prepared to read at the level of the books he's attempting. Stepmom suggests some easier books that might be more his speed. The kid refuses; the teacher said they were supposed to read a chapter book.
Another thing they do at this, again, very highly-acclaimed school, is they give the kids arithmetic exercises at the afterschool program. At this point, stepmom is pretty skeptical of the competence of the grown-ups at this school, so she looks over the kid's papers from that. She sees an error in one problem that was marked correct, and points it out to the kid.
Kid has a complete meltdown. "No, no!" he screams. The afterschool grader had marked the problem correct, so it had to be correct. Why was stepmom lying to him about how numbers work? Stepmom tries to explain that the grader must have made a mistake. It happens, sometimes.
The kid isn't having it. Kid and stepmom take the problem to the father. Dad agrees with stepmom. The kid is still crying. "No!"
Stepmom shows the kid how to punch in the problem to the calculator app on the kid's tablet—to the calculator app on her laptop—to the LCD desk calculator that was sitting in the drawer. Same answer every time, different from the answer the grader had marked correct.
At this point, the kid has run out of energy to scream any more, or even to cry. He's just sitting there, shaking a bit, mouthing the word "No" without voice, staring glassy-eyed off into the distance, as if at some cosmic horror only he can see.
There's probably a moral here. But they're not paying me to tell you what it is exactly.
Discuss
The Three AI Pills
Sincere disagreements about AI are usually disagreements about future AI capabilities.
There are roughly four positions people take. Two are reasonable. Two are not.
I distinguish these via the Three AI Pills. You can take zero, one, two or three.
Three PillsThe three pills are, roughly, taking each of the following three things seriously:
- AI pilled. AI exists and can do the things it can already do.
- AGI pilled. AI will be able to do a lot more of the things.
- ASI pilled. AI will be able to do approximately all the things better than you, within our natural lifetimes.
I am ASI pilled. A large percentage of employees of the frontier labs are ASI pilled. The labs themselves are ASI pilled.
The Unpill PeopleI see unpilled people.
Where do I see them? Everywhere. The majority of people have not taken the first pill.
Most people have no idea what frontier AIs can do for them. They are unaware of coding agents. They have used only ChatGPT, for harmless trifles, and they hold years old memories of its failings. They mock any failure anywhere as ‘what AI can do.’
They dismiss AI as worthless because it pointed them to a closed store or recommended the wrong number of pizzas. They cite old studies that were obsolete before they were published and used terrible prompting techniques.
They often still talk about ‘stochastic parrots’ or how AI can never possibly think and everything must be stolen from the training data. And so on.
Some versions of this are wrong. Some are Not Even Wrong. None are reasonable.
When you discuss AI with people who are fully unpilled, your goal is usually to first give them the AI pill. Show them that AI can do the things it can already do.
The AI PillEven fully taking the first AI pill is a big deal.
Existing AI unlocks, today, in practice, tons of cool things. It is, in many ways, already smarter and more capable than you.
So many things that you used to do by hand, or some other way, are now better done by typing a quick request into a text box.
So many things that previously were not worth doing are now worth doing.
So many questions previously not worth asking are now worth asking.
The marginal cost of seeing what the AI can do for you is often very close to zero.
AI can also do a variety of harmful things, or do things that are useful for you but make others worse off or disrupt or invalidate norms or systems. People don’t like that.
Most economists, and most people who work in policy and government, have taken at most this first pill, and underestimate even the impacts of the first pill alone.
Often they say things like ‘AI will be too expensive to use on [X]’ because they don’t realize it will soon be orders of magnitude cheaper for the same level of intelligence. Or they point to particular details where AI does poorly, and presume this will not be fixed. They see AI as ‘uncompetitive’ without realizing the situation is temporary.
When you discuss AI with someone who has taken only the first pill, you typically have three basic options.
- You can try to ‘fully AI pill’ them and explain the things AI can already do and the implications of what that means, even if things stop here.
- You can try to explain that we will get better at using what AI we have, and that there is a lot of ‘unhobbling’ left for us to do, even if things stop here.
- You can explain that things will not stop here, and you need to be thinking about what future AIs will be able to do. Get them to take at least the AGI pill.
Even if AI could permanently only do the things it can currently do, that would be Internet big, and radically change the world, mostly for the better.
AI capabilities will not permanently stop here. It is wrong to not take the second pill.
Stuck At The First PillOur debates about AI remain largely stuck on settled questions, because so many people cannot even take the first pill.
Dean W. Ball: A couple years ago, the AI debate was centered, rightfully, on whether crazy-sounding things like “AIs autonomously making math breakthroughs” and “AIs breaking from their sandbox and hacking on the internet” would be real things in the near term. Sometimes it feels like that’s still the debate we’re having. This can be frustrating, because in my view, that debate is settled and was settled quite a while ago.
To have good discussions, we need to at least take the second pill.
Whereas, yes, many people, even many who work with AI, really do say current AI is ‘good enough’ and can’t imagine what a better one can do. As in, someone tweeting at Sam Altman saying ‘Sol does everything I want it to do, this is all I ever need’ and Altman retweeting saying they were wrong. Which they obviously are.
The AGI PillThe AGI pill is a much bigger deal than the AI pill.
If you take the AGI pill, you understand that AI is advancing its capabilities rapidly.
Even if you think that such AGIs will remain fully under human control, and remain ‘mere tools,’ and you expect the lived experience of most people’s everyday lives to not change so radically, you understand that their capabilities will ‘change everything.’
You see that we will face times of great transition and uncertainty, that have the potential to go extremely badly, and that those who succeed at AI will leave those who do not behind in the dust.
The world in the future will be very different from our own. AIs will be able to do most digital work, most of the time, along with the inevitable robots and self-driving cars and so on. Lots of current jobs will go away, whether or not they are replaced by new and potentially better ones. Economic growth and productivity will accelerate.
You see some of the dangers of what would happen if we empowered misuse of such advanced AI systems before we were ready, especially in places like cyber and bio risk.
You see the potential for centralization of power, or inequality, and also for some forms of runaway gradual disempowerment.
You see the potential for mass unemployment, either transitional or permanent.
You understand that our legal and regulatory regimes are not ready, either to protect against and mitigate the risks and harms, or to allow for the opportunities and remove the bottlenecks to diffusion and mundane utility.
The Need To Be PreparedThose who expect AI to quickly become sufficiently advanced to greatly impact the physical world usually see great danger. They notice that as a result everyone may soon die. Usually they think this is bad, actually.
Thus such folks call to take coordinated action to mitigate the downside risks of such impacts, keep us all from dying, and ideally also to help capture the upside benefits.
Those who expect AI to become importantly more advanced, but with a slower and smaller impact on the physical world, and who think the practical value of more intelligence will cap out.
For different values of ‘sufficiently advanced,’ as in AGI versus ASI, you would see different degrees of danger.
The AGI pill is still sufficient for most things in the Overton window or under serious consideration as of August 2026. We are almost entirely considering overdetermined, low cost, high benefit interventions.
There is no good case for not doing radically more investment in alignment, infrastructure and oversight, state capacity, transparency, liability, disclosures, safety testing including of internal models, red teaming, auditing, enforcement of export controls and laying the groundwork for diplomacy.
This includes laying the groundwork to Pace the Frontier should that prove necessary.
If you are fully ASI pilled, and realistic about the current state of alignment and how superintelligence likely plays out if it arrives soon, then you will want to go further. You will want to do things that have real downsides, and require real tradeoffs.
Some such people want to do a full international pause of frontier AI development. If you took the full ASI pill and believed what they do about superintelligence, in terms of how fast it might arrive and what it can do, you might well agree with them.
The ASI PillThe ASI pill is the understanding that AI is on pace to be able to do approximately all of the things better than you.
I said ‘approximately.’ As I go over in detail, that does not mean literally all of the things. There are some things that inherently require or greatly benefit from being a human. And there may be weird corner cases where the AI won’t be good enough.
It does not mean omnipotence or omniscience, although one should expect it to look a lot like that to an unaided human.
It does mean the AI takes your job, and then takes the new job that you switch into, unless you pivot to ‘requires literal human.’ You will be uncompetitive at essentially any other task.
It does mean that it will use this capability to figure out approximately all of the things, remarkably quickly, until you hit the physical limits.
It does mean that those who rely more on such AIs will reliably outcompete, in all senses including for resources, those that rely on such AIs less.
It does mean that, in a ‘fair fight’ or sufficiently open competition, the AI wins.
It also means the AIs often figuring out and doing things you did not imagine or anticipate.
It means realizing that intelligence does not stop anywhere near the human level, nor does its ability to chart paths through causal space towards preferred arrangements of atoms.
It also means not pretending that its superior intellect can be matched by your puny weapons, or your pieces of ink on paper, or your entries in a database, or your regulatory capture and rent seeking.
And Then Nothing Much Changes For YouDespite all that, the sign of the AGI pill, as opposed to the ASI pill, is the belief that day to day life will continue to look similar to how it looks now, in the sense that we see day to day life in 1926 as not that different from life in 2026.
Sometimes this is clearly disingenuous, as it is coming from someone ASI pilled enough to know better. I read Sam Altman saying various forms of ‘life will not much change after the singularity’ as basically unjustified Obvious Nonsense to reassure his audiences, rather than a coherent position. Sometimes the person such folks want to reassure most is themselves.
Others sometimes do have something more concrete in mind as their vision of the future, and why it will not change so much, with varying degrees of coherence.
Those with only the AGI pill believe in bottlenecks that hold back change.
They often believe that our ability to exponentially grow AI’s capacity and capabilities will hit various physical limits. There can only be so many chips. Actions take time. Things too far out there are often pejoratively dismissed as ‘magic.’
They often believe there is not that much left to physically discover, in the classic ‘close the patent office’ kind of way, even in theory. Your steak can only be so tender, your lobster so buttery, your lifespan so long, and your status so high, so why does it matter. I strongly disagree on lifespan and health, and expect we have a long way to go in so many other ways in terms of finding value, although they may have a point about moment-to-moment maximal hedonic experiences of a physical human brain.
They often believe that the upside of intelligence is importantly limited. That no mind, however advanced, could be all that persuasive, or that economically valuable, or that capable of creating innovations in the physical world, or of running sufficiently accurate simulations, or making sufficiently strong predictions, or even able to do things like overcome red tape and regulatory capture.
Intelligence DenialismI sometimes call this Intelligence Denialism: The idea that being smarter is not all that, no matter how smart one gets. That there is this thing, intelligence, that you either have or don’t have, and that minds cap out.
Often this extends to denying that more intelligent humans can do and accomplish the things they clearly do and accomplish. Other times, it is the idea that intelligence tops out at ‘smart human,’ and all a mind can do is imitate that smart human. Maybe you can do it faster and cheaper, and at scale, with better memory and so on.
But that’s it. And such folks fail to understand that if you took the union of all human mental capabilities, and all access to knowledge, at scale, in parallel, much faster and cheaper, that this alone would run circles around anyone and everyone, everywhere. And that if this lacked physical capabilities or access, this would be trivial to get.
This is, usually, the central good reason people who are AGI pilled do not take the ASI pill. They are unable to understand that superintelligence is a thing.
This confusion is also often behind the instinct that models will commoditize. There is an ideal thing, ‘intelligence,’ and you approach it via an asymptote. More is impossible, or more won’t help you, depending on how you frame it.
Superintelligence Versus Omniscience and OmnipotenceA common error is to think that anything superintelligent would be omniscient or even omnipotent, and therefore superintelligence is impossible.
Whereas, again, there is a lot of ‘space above humans’ without becoming omni.
There are many other similar confusions, where ‘there is some upper bound to [X]’ is confused with ‘we are at or near the upper bound to [X].’
Here Adi gives us a way into explaining what’s happening, by doing what from my perspective is a mirror image of the central motte-and-bailey.
Both versions do exist.
Timothy B. Lee: There’s an epistemic chasm between those who think superintelligence implies near-omnipotence and those (like me) who don’t.
Adi: There’s a motte-and-bailey trick being played between two definitions of “superintelligence”
“Big-S Superintelligence”, which is definitionally omnipotent
And “small-s superintelligence”, which is the hypothesized end state of the current thread of AI research
It’s surprising how uncritically the AI safety community treats these dubious equivalences, especially since one of the most famous treatments of the motte-and-bailey fallacy comes from none other than Scott Alexander, who’s quite prominent in the community.
As in, I see the following a lot:
Zvi Mowshowitz: The motte-and-bailey here is real, but more commonly it is in reverse. as in a form of:
1. You claim superintelligence [S] soon will run circles around humans and do absurd stuff.
2. But [S] would not be fully omniscient/omnipotent.
3. Therefore [S] would be a ~normal tech.
No.
Adi: I’ve seen this too, and I agree that it’s equally wrong
There will still be things that require non-trivial real world time to achieve. There will be particular actions and paths and methods that have their limitations. These limitations matter. Fully abstracting them away can be a large mistake.
I constantly see versions of ‘the mind that is maximally good still could not do [thing it seems to me it could obviously do]’ or even [thing it can already do] and I think this form of conjunction is a lot of why, and how people justify themselves. When AI gains capabilities, and more things fall, they move the goalposts but don’t change the game.
Adi’s comment above was in the context of discussing Plan A, where the AIs are very obviously not ‘Big-S Superintelligence’ in Adi’s lexicon. There are practical things that these AIs cannot do, or can only do up to a point.
The mistake was not being made. It sometimes gets made, but this is rare.
Persuasion Persuasion (A Worked Example)At most one can claim some particular capability, typically very effective persuasion (often ‘super-persuasion’), would not be possible. Indeed, a common argument states that even an idealized AI would be unable to be superior to the best human persuader, or even a typical good human persuader (who is much worse than the historical best human persuaders), despite it having access to many advantages over humans, such as much faster thinking speed and access to information, including every little reaction of its target including body language.
Often this is done by citing that context matters for persuasion. Many talk as if this context would trump everything, despite this not being true for humans. They also believe that highly capable AIs could not engineer favorable contexts.
Often it’s a form of ‘I would simply choose not to be persuaded’ or ‘if you can’t do it cold and one shot everyone with only text it does not count’ or something like that.
If it needs a human presence to do the persuasion, or a social context, it can acquire one easily enough in various ways.
I flat out cannot understand why someone would think advanced AIs will remain relatively unpersuasive, other than to think that AI will not get much more capable than it already is, and that’s if we fully restrict the AI to using known standard persuasion techniques and rule out any wizardry.
Yet I am not a superhuman persuader, so many find my arguments unconvincing.
No, seriously, as in Adi is citing the highlighted passage as if this was not very obviously true, my lord, how can you think this is not going to happen in a scenario where AI capabilities continue to develop apace, if anything 2035 seems crazy slow:
This just isn’t much of a threshold. Nor does it require that the AI do this from pure text box versus what a human can do in person. It just says that, given access to similar tools, the AI will be able to do considerably better than any human. How could you not think this, given the rest of the scenario, even if all it does is have all the advantages and skills of the various different humans, plus the ability to process a lot more information a lot faster?
And like, worst case, the AI gets an earpiece and tells a human what to do, Jesus.
Another way of thinking about this:
Adrien Ecoffet: Could a superpersuasive AI get a high speed rail line between SF and LA constructed for under $50B and under 10 years with existing construction technology?
Dean W. Ball: No
Nat McAleese: @deanwball
I think this is an important question not to be overconfident on! And if it feels obvious, you are surely overconfident.
Dean W. Ball: obviously there are no guarantees, but “frontier ai systems reforming ceqa and then cleanly funneling tens of billions of dollars through California’s government” doesn’t seem realistic + like a thing to prioritize compared to many other things
My answer, for a sufficiently persuasive AI is essentially: Yes, but not without incidentally taking over at least the California government. Instrumental convergence.
Not that you would then choose to focus on High Speed Rail, as Dean replies. This is not a case of would, so much as could. The point is the could, in theory.
An example that is almost too perfect an example of ‘you will never be able to do [thing that you can already do]’:
Ramez Naam: “No matter how high quality your tokens are, they cannot turn lead into gold.” Pleasantly unhinged and overall correct.
The Singularity is Nearer: In The Metamorphosis of Prime Intellect, the hard takeoff works because AI discovers the correlation effect, some quantum trick to manipulate matter. In reality, there is no correlation effect. No matter how high quality your tokens are, they cannot turn lead into gold.
Ryan Greenblatt: I get the idea, but this is a funny example because it is totally feasible to turn lead into gold. Extremely capable AI systems (after tech R&D and infrastructure build out) will have the capacity to very quickly build infrastructure for turning lead to gold.
We already have successfully turned lead into gold in 2025, it’s just cost prohibitive with current methods. People really do model the world as if anything we don’t specifically have a clear path to already doing cannot ever be done. Everyone wants to close the patent office.
Things AI Could Probably Do But Are Not Required For Being PilledCould such an AI send off an order to a bio lab and end up with programmed self-replicating diamond nanoprobes? Could it do so without any physical experiments, based purely on simulations? Could it do so with a relatively fast series of experiments? Could it instead do other similarly powerful things that do an end run around everything and let it do whatever it wants? On what time frame?
No one knows where the physical limits lie, either for the end of full RSI (recursive self-improvement) or for where the current cycles might top off in practice. I fully expect to discover various mental and physical capabilities, and new affordances, that we did not expect, and in many cases did not imagine, and by definition it is hard to know which.
I mostly stopped talking about those possibilities, despite them seeming rather likely to me, because:
- When you mention such things, people attack that claim, and use it as a reason to dismiss your entire argument, all risks from advanced AI, and you as a person.
- This strategy works, because such claims are much harder to justify, less certain.
- You don’t need such claims to get the same results. Sufficiently advanced AI can ‘get there’ with only capabilities we can be confident such AIs will have, because humans already have them modulo sufficient compute, parameters and data.
E.g.: Maybe James Bond can’t always get the girl, maybe he can, and maybe Q’s latest gadget is a thing you can build and maybe it isn’t, but I am damn sure Bond has a gun and is very good with it, and the gun really is all he needs on this one, but a lot of people don’t understand this. So I choose to argue as if all he has is the gun.
Another approach is this, which works for some but I do not expect it to be at all persuasive to Timothy Lee:
Jack Gallagher: the intuition that made it click for me was thinking of it less in terms of “intelligence” and more in terms of things like metr time horizons. a lot of work is intelligence x schlep gated and as you improve on those axes just about everything eventually falls. maybe you still miss some grand creative spark but nanotech doesn’t require a creative spark from here it requires relentless grinding on miniaturization.
many things will still be bound on serial experimentation, but as the cost of any given level of intelligence goes to zero we’ll end up running way more of the available experiments in parallel.
Life Comes At You Increasingly FastAnother reason people stop at the AGI pill is not understanding, or rejecting out of hand as too sci-fi or weird or absurd or what not, the idea of recursive self-improvement, or a singularity, or that things might be radically accelerating.
They can accept that things are speeding up, but not that the speeding up is itself speeding up, and so on, which indeed is already happening and hard to miss.
This has in some sense been happening for a very long time. The math predicts a singularity. This is then sometimes used to say ‘well this is not a new development’ but that is how exponentials work. Nothing centrally new happens, and yet kaboom.
As a reminder (with rounding, don’t @ me):
- 4.5 billion years ago: Earth.
- 300,000 years ago: Homo Sapiens.
- 10,000 years ago: Agriculture.
- 300 years ago: Industrial Revolution.
- 80 years ago: Computers.
- 8 years ago: LLMs.
You can be misled by ‘oh there is this step change where you are suddenly in a singularity and having recursive self-improvement, and until then no’ or ‘this new thing is not different.’
Is It Reasonable To Not Be AGI Pilled?No.
If you are not AGI pilled, your reactions to AI will not be wise or prudent. Alas, much of our policy and conversation is being driven by people who are not AGI pilled.
When we see positions that rely on not being AGI pilled, we should say so, and engage or not engage with them accordingly.
Is It Reasonable To Only Be AGI Pilled?Yes, if you seriously grapple with its implications, and know what your cruxes are.
I think it is wrong. But it is a coherent position, that is at least wrong, to think that our current techniques and resources will not fully ‘get there’ and AI capabilities are likely to top out before things we would properly call superintelligence.
Or at least, I don’t think treating this as unreasonable would be productive.
I would switch to putting increasing weight on this position if we ‘hit a wall’ of at least clearly diminishing returns for a sustained period of time, and especially if I don’t feel so behind using last year’s model and things really do commoditize.
We could still see dramatic growth in AI revenue and demand for compute. Investments in AI could pay off handsomely up and down the supply chain. That is all fully compatible with the AGI-only position.
If you do believe the AGI-only position, I have two requests.
First, seriously tackle the implications of what you do think AI is going to be able to do, without flinching, and without trying to assume everything magically works out and somehow people’s life experiences do not much change. That means then supporting actions now that we need to mitigate the risks of all that, and help make things go well.
Second, write down what are your cruxes that are stopping you from taking the ASI pill. In particular, write down (in the comments here would be great) what is the least surprising or impressive thing an AI will never be able to do, or that would change your mind about where this is going, and what other near term observations would cause you to either be confident you are right, or realize you are wrong.
Until then, there are two valid choices: You can stop at the AGI pill, or go full ASI.
There are also two common but invalid choices: Think AI stops here, or deny reality.
Discuss
Notes on the possibility of moral progress
To achieve the best possible future, we must know what that future looks like. In other words, we need to solve ethics.[1]
The problem of solving ethics is so large and abstract that it’s difficult to say useful things about. In lieu of any structured analysis, herein lies a collection of thoughts about the problem.
Cross-posted from my website.
Acting in the face of moral uncertaintyIf we have persistent moral uncertainty between maximizing and satisficing[2] moral theories, then it’s not difficult to decide what to do in practice: allocate a tiny portion of the universe to satisfying the non-maximizing moral theories, and allocate the rest to the maximizing moral theories.
Example: If theory A says we should fill the universe with welfareans, and theory B says we should preserve homo sapiens but it doesn’t much matter how many humans there are, then we can near-perfectly satisfy both theories by maintaining humanity in a small segment of the universe and giving the rest to welfareans.
However, the distinction between maximizing and satisficing theories may be irrelevant because plausible satisficing theories still hold that it is a good thing to maximize The Good, even though it is not morally obligatory[3] to do so. Satisficing theories would still want to fill most of the universe with The Good.
- Given uncertainty between mutually incompatible maximizing moral theories, we have to choose. Allocating resources incorrectly would be catastrophically bad.
I believe so. I expect that we can eventually eliminate almost all moral disagreements purely by discovering facts. The fact-value distinction implies that we cannot 100% determine what we ought to do by discovering facts, but most of what look like terminal values disagreements are not truly terminal.[4]
Some examples:
- Right now, we do not know how to weight different people’s experiences against each other. I strongly suspect that there are facts of the matter about how to weight experiences, and that these facts can be discovered empirically.
The problem of weighting experiences is downstream of the hard problem of consciousness. I likewise suspect that there is a definitive answer to the hard problem,[5] and that we can, in principle, find that answer.
- What is the nature of personal identity? Certain theories of personal identity rule out classes of moral theories. If personal identity is not metaphysically meaningful, then person-affecting views must be false—there is no relevant distinction between bringing a new person into existence and changing the life trajectory of an already-existing person. Changing a life trajectory creates new person-moments, which is (in this view) metaphysically equivalent to creating a new person.
- There may be some way to salvage person-affecting views, but if so, that possibility would itself be a non-normative fact—i.e., person-affecting views may be permitted or ruled out purely by facts, without any moral stance required.
Another important (although slightly obscure) question is: given two identical copies of the same mind, are they experiencing “twice as much” as if there were only one copy, or “the same amount”?[6] This seems like a factual question, not a moral one. I have no idea how we would answer this question, but it seems answerable in principle.
Harsanyi’s utilitarian theorem (see also Harsanyi (1955)[7]), which showed that if individuals have VNM utility functions, and if the Pareto principle[8] holds over groups, then a version of utilitarianism must be true. The Pareto principle is a normative principle, not a factual one; the question of whether individuals ought to care about their own welfare is also a normative one. But I strongly suspect that there is a fact of the matter about whether an individual’s welfare can be described as a utility function, and there is a fact of the matter about how exactly that function is specified. Discovering those facts would get us at least part of the way toward a fully-specified theory of ethics.
- It matters whether the universe is finite or infinite, and whether our actions can have finite or infinite influence. Infinite ethics poses troubling problems, but some (maybe all) of those problems can be resolved factually (or if they’re unsolvable, then their unsolvability is a factual question).
Note: I don’t think they can be resolved purely empirically. For example, by our understanding of the laws of physics, our actions cannot have infinite influence.[9] But there is a nonzero probability that we are wrong about the laws of physics and that our actions can have infinite influence after all. No amount of empirical investigation can reduce our uncertainty to zero, but there may nonetheless be a mathematical solution to the problem.
Standard formulations of deontology may break down in light of the fact that you do not have certainty about the consequences of your actions, so you can never be sure that you’re not doing something impermissible by acting. (See also Nye (2014)[10].) If so, those flavors of deontology are ruled out purely based on a factual analysis (no normative claims necessary).
Many value disagreements do not purely boil down to a disagreement about facts, but I expect they can be resolved anyway. Some examples:
- People would rather donate money to a single identifiable person than to a much larger, but nebulous, group of people. This preference should break down upon reflection. Suppose Alice is offered the chance to donate to a single identifiable person. Then imagine an alternative world where Alice can donate the same amount of money to help that same single person plus several other people, but the single person is never identified. Surely she would prefer this.
I believe the disagreement between negative utilitarians and classical utilitarians[11] would be resolved if we knew how to directly compare experiences / if we solved the hard problem of consciousness.
My guess is that negative utilitarianism is a mistake stemming from the fact that maximum suffering in humans far exceeds maximum happiness, and this creates the appearance that suffering is terminally more important than happiness.
- People support animal welfare, but also eat factory-farmed animals. It’s conceivable that people in reflective equilibrium would resolve this inconsistency by throwing out their concern for animal welfare, but that seems unlikely.
- I believe people reject the mere addition paradox due to scope insensitivity—an enormous population of slightly happy people is indeed better than a small population of very happy people. It should be possible to prove that this intuition is the result of scope insensitivity, that scope insensitivity is inconsistent with people’s other values.
- Alternatively, some argue that an enormous population of slightly happy people is not particularly good because the experiences are too uniform, and two copies of an identical experience is no better than one copy. If there is a fact of the matter about how to consider two copies of an experience, then this alternative view could be proven right.
Our understanding of philosophy is limited. What does it mean to do good philosophy? What qualifies as a good philosophical argument?
We could make progress on those questions. We have already made progress: Descartes innovated on rightly conducting reasoning and seeking truth. The significance of philosophical thought experiments is a recent development—the concept is pre-Socratic, but modern thought experiments are more refined and more useful. (On Wikipedia’s list of notable philosophy thought experiments, two thirds were invented after 1900, and over half were not developed until 1960 or later.[12]) Most modern concepts in moral philosophy come from the 1700s or later; philosophy of mind primarily comes from the 1900s;[13] analytic philosophy improved on the methods of its predecessors, and did not emerge until the 1800s. All that suggests that civilization is indeed making progress on philosophy, even if the rate of progress is slow.
Some normative claims evade fact-based analysisI hold some foundational moral beliefs that seem unrelated to descriptive facts. I cannot conceive of how an empirical or theoretical investigation could provide reasons to believe that these are true or false.
- It seems self-evident that pleasurable experiences are good (and suffering is bad), in the same way it’s self-evident that I am conscious. This is difficult to dispute, and to my knowledge virtually all moral philosophers (and regular people) agree that pleasure is good and suffering is bad.
I find it hard to see how anything other than good or bad experiences could be good or bad, because where does the goodness or badness come from if it’s not being directly experienced by anyone?[14] (A.K.A. welfarism.) However, many people believe that non-experiences can be innately good or bad, and I don’t see how we could resolve this dispute.
- Other people’s experiences matter, not just my own. (This claim is uncontroversial, but still, I see no way to prove it, or even give any reason to believe that it’s true.)
- All beings’ experiences matter equally. It does not matter who the experience resides in; all that matters is the intensity of the experience. This view has several corollaries:
- Welfare aggregates linearly across individuals.
- The total view of population ethics is correct.
- Speciesism is wrong—experiences of humans should not be given more moral weight purely due to species membership.
- If you live in the 1700s, it implies that slavery and misogyny are wrong.
The first and third claims (the goodness of pleasure/badness of suffering, and the principle of altruism) are widely accepted. Many people disagree with me about the second and fourth claims, and there is no visible path to resolving those disagreements—plus, I have internal uncertainty about whether they’re true, which I have no idea how to resolve.
Even though the first and third claims are uncontroversial, they still evade any attempt to explain why they’re true. It could be that we’re all wrong.
And yet, it’s possible to convince people about these sorts of normative claims. (Peter Singer made arguments that convinced many people, including me, of the principle of equal consideration of interests.) What’s going on inside people’s heads when they change their minds about seemingly terminal values? Or, what’s going on when I contemplate two conflicting intuitions and decide that one is more important than the other? We have no theory of what constitutes a good argument for a normative position. We have some idea about the sorts of arguments people find convincing, but not a great understanding of why, and no way of saying that people are right to be convinced by a particular argument.
Implications for how the future goesThis essay has posited that we are making progress in philosophy, and that most (maybe all) moral disagreements can be resolved by learning new facts. If true, what does that imply?
- Smarter-than-human AI should be better than humans at discovering facts. That’s useful insofar as moral disagreements can be resolved by facts.
- The positive vision in the epilogue of AI 2040 has every individual human controlling an equal part of the lightcone. How good an outcome is that? If it’s feasible to converge on moral beliefs, then in that scenario, people (with superintelligent AI assistants) will come to agree on ethics, and will shape the universe in the way it ought to be shaped. Some people may have persistently bad values (like maybe Putin, or maybe not), but if most people converge on good values, then most of the universe will be directed well.
- How well would a Long Reflection work? It would provide more opportunity to discover ethics-relevant facts and to improve our understanding of how to do good philosophy. But it would also give amoral actors more time to seize power. The AI 2040 idea of “give everyone an equal share of the lightcone, and then let people cooperate if they want to” is plausibly better than a Long Reflection, and plausibly worse.
Over sufficiently long time horizons, natural selection takes over. The dominant ethical belief will be that the right thing to do is to spread one’s own genes at the exclusion of everything else. We need to solve ethics before that happens, or otherwise prevent that from happening.[15]
- In the scenario where humans control the future, the principle I worry about most is impartial altruism. I worry that most of the people in control will simply not care to devote resources to helping others.
- I worry much less about disagreements between altruistic people (e.g., between negative and classical utilitarians). I expect these disagreements can be resolved factually.
- ^
Or perhaps it would be more accurate to speak of solving axiology, i.e., “what is good?” as opposed to “what is right?”
- ^
A maximizing theory holds that the right things to do is to maximize some quantity—usually, to maximize utility, although “utility” can be defined in various ways. A satisficing theory says that there are certain things we ought to do (e.g. don’t commit murder), but as long as we do those, we have “satisfied” our moral obligations.
- ^
I’m hesitant to use the word “obligatory” because it creates confusion when talking about consequentialist theories. For more on this, see Richard Y. Chappell’s blog post Deontic Pluralism (2022) or his academic paper Deontic Pluralism and the Right Amount of Good (2020).
- ^
There is a class of moral disagreements that can clearly be resolved by factual questions. For example, should I drive a bulldozer through a particular building? That entirely depends on the answer to the factual of “is this an empty run-down building that’s scheduled to be demolished, or it a house that somebody’s living in?” Those sorts of disagreements are not interesting for the purposes of this essay.
- ^
Some people claim that there is no fact of the matter about the hard problem of consciousness. I find it hard to comprehend why anyone holds that view. There is, at minimum, a fact of the matter about whether I am conscious (and that fact is “I am conscious”). I find this position about as confusing as the illusionist view (i.e. the view that consciousness is an illusion and in fact there is no such thing as consciousness).
- ^
Bostrom, N. (2006). Quantity of experience: brain-duplication and degrees of consciousness.
- ^
Harsanyi, J. C. (1955). Cardinal Welfare, Individualistic Ethics, and Interpersonal Comparisons of Utility.
- ^
The Pareto principle states that if outcome A is at least as good as outcome B for every person, and outcome A is better for at least one person, then outcome A is better overall.
- ^
Sandberg, A. & Manheim, D. (2021) What Is the Upper Limit of Value?
- ^
Nye, H. (2014). Chaos and Constraints.
- ^
Here, classical utilitaranism refers to any flavor of utilitarianism that gives equal weight to happiness and suffering.
- ^
Dates are pulled from Wikipedia and aggregated by Claude Opus 4.8 (chat source). Of the 41 philosophy thought experiments on Wikipedia’s list, there are 13 from pre-1900, 5 from 1900–1959, 20 from 1960–1989, and 3 from 1990–present. I’m taking Wikipedia’s list as a reasonable proxy for notability.
- ^
Thomas Nagel’s What Is It Like to Be a Bat?, written in 1974, is a more lucid exploration of consciousness than anything that came before it, and represented important progress. More broadly, much of the best work on philosophy of mind came from people who are still alive today.
- ^
The parable of the Hrogmorph’s Heartstone is an attempt at justifying this intuition. (See also the extended parable.)
- ^
Natural selection hasn’t taken over yet because humans are adaptation-executors, not fitness-maximizers. We have evolved the intelligence necessary to explicitly optimize for genetic fitness, but our big brains haven’t been around long enough for natural selection to push us in that direction.
Discuss
Regranting in 2026: now more than ever
Mostly written by Carol, edits by Austin.
Manifund is raising money for our 2026 AI safety regranting program!
TL;DR- Early grants can have huge returns.
- Dispersed part-time grantmakers can find opportunities before centralized funders.
- Donating to Manifund Regranting puts your money in these opportunities.
See also our 2025 announcement for more thoughts on “why regranting?” (answer: it’s fast, low-friction, proactive), and “what makes a good regrantor?” (answer: taste, capacity, dealflow).
Six top hitsWe’ve run AI safety regranting every year since 2023, so we’ve now had enough time to see how some of our early grants have played out. And because giving is hits-based, looking at a few successes can prove the value of our program.
$143k in August 2023 from Evan Hubinger, Ryan Kidd, Rachel Weinberg, and Marcus Abramovitch
This grant provided the initial funding for Timaeus to get started. Since then, they grew their organization and eventually merged with Resolution, earning a $160m grant — Coefficient Giving’s largest-ever AI safety grant. See this post for our argument that this grant alone makes all of our regranting so far worth it:
Multiplied out, our $143k stake via our 2023 regrants would have grown 53x to now be worth $7.6m. This first regrant to Timaeus, alone, would more than pay back the entire Manifund regranting budget to date ($5.2m). Since the Resolution grant is so much larger than most other AIS grants, Timaeus could be the single best 2023 AI safety grant in these terms.
2. Rob Long’s research on AI consciousness
$7200 in October 2023 from Joel Becker
Joel explains:
This situation struck me as being preposterous. Here is one of the world’s top experts on AI consciousness, claiming a nerve-wracking chance of AI consciousness in the not-too-distant future, with fairly strong professional links at top AI labs and ~shovel-ready ideas for empirical projects, preparing a not-terribly-costly ask (give me a work desk, ~0 engineering time, and model access to work on these research questions)... and he is unsure whether he should make the ask?!
Joel solved this by paying for 5 weeks of Rob’s time, allowing him to research AI consciousness and moral patienthood, rather than applying for jobs.
Rob went on to found Eleos AI a year later, which has raised more than $2m from funders like Macroscopic and Longview, and performed the first officially-commissioned welfare evaluations of frontier models on Claude Opus 4. Rob’s final report: “This grant was absolutely clutch.”
3. Neuronpedia and Joseph Bloom
$2500 for Neuronpedia and $25k for Joseph Bloom in July 2023 from Austin Chen and Marcus Abramovitch
For Neuronpedia, this was the first funding they’d received; for Joseph Bloom, he’d been working as an independent researcher for 6 months at the time, and the grant enabled 6 months of his research.
These were separate grants, but Joseph and Neuronpedia ended up collaborating to produce tools that are widely used in mechanistic interpretability. Joseph used the funding to release the SAE Lens library and partner with DeepMind on Gemma Scope, among other projects. He’s now the model transparency lead at UK AISI.
Neuronpedia pivoted from its initial game concept to produce open-source interpretability infrastructure used by DeepMind and Anthropic. It went on to get donations from LTFF, Superalignment Fast Grants, AISTOF, and SFF.
4. Weapons of Mass Destruction Proxy (WMDP) benchmark
$190k in January 2024 from Dan Hendrycks
This grant was to a research project by a PhD student and an undergrad and was used to pay consultants to develop a 3600-question dataset of questions surrounding hazardous knowledge in CBRN and cybersecurity. This dataset became the WMDP benchmark, the standard benchmark for hazardous knowledge in the unlearning subfield.
$38k in September 2023-August 2024 from Joel Becker and Evan Hubinger
At the time of the regrant, ChinaTalk was a part-time operation running on a budget of $35k/year, and Manifund was the first outside funding it received. The grant allowed the founder, Jordan Schneider, to spend less time on paid consulting to support himself.
ChinaTalk now has five full-time staff and has gotten funding from Emergent Ventures, Macroscopic, and Coefficient Giving. They’re listened to by many influential sources and did notable early reporting on DeepSeek before it made major news in the West.
$335k in August through December 2023 from Tristan Hume, Evan Hubinger, and Marcus Abramovitch
Apollo Research launched in May 2023 as an eight-person organization. Since then they’ve built the first empirical evaluations for AI scheming, worked with OpenAI and Anthropic to test models before shipping, and have grown to around forty people.
This one is less exciting than some others because it came after $2m of funding from Open Phil and SFF. But Manifund regrantors were still quite early in the lifetime of the organization, and provided a significant fraction of their early funding. (We’d estimate that they raised a ~$20m seed round for their PBC in 2026, which would imply a ~10x impact ROI from this regrant.)
See for yourselfIf you’d like to review all the regrants we’ve ever made, check out this dashboard. I also asked Claude Cowork to look through our regrants and rank the best ones in retrospect in this spreadsheet. I think it did a pretty good job—we decided to include its 1-5 and 8 in our “top hits” section.
The case for part-time grantmakersOne argument we’ve heard is, “There are organizations full of specialized full-time grantmakers. Why give your money to part-time grantmakers instead?”
We think it’s good for a few reasons:
Many people’s first $100k of grants will beat a professional grantmaker’s last $100k.
Any given grantmaker’s field of view can only be so big: people are limited by their networks, geographic locations, and areas of expertise. Lots of people have specialized knowledge or personal contacts that allow them to notice a few good opportunities.
It’s way more scalable.
The number of researchers and projects and the amount of funding in AI safety is growing massively in 2026, but it’s going to be difficult for our capacity to allocate the funding to the right researchers and projects to keep up. Scaling up a foundation full of professional grantmakers is hard for the same reasons scaling orgs generally is hard. It’s easy to get bottlenecked on hiring, management, training, or trust.
On the other hand, finding talented people to give regranting budgets to and increasing their budget if they make good grants is a much lower-friction process.
It helps find, empower, and upskill regrantors.
In the past, we’ve given regranting pots to a mix of those well-known in the community and people who were less connected but who we thought were promising. Some of those people have ended up being (in our opinion) great regrantors and gone on to jobs in AI safety—like Joel Becker, who started as a regrantor in 2023, and in 2025 became a member of technical staff at METR and coauthor on some of their most-discussed work.
Should you donate?We think regranting is a good fit for donors who are earning-to-give $50k+/year, who care about AI safety but may not be as steeped in the AI safety ecosystem. Think of regrantors as part-time program officers who can provide taste and dealflow without the logistics of hiring or creating a foundation.
Everyone expects a lot of Anthropic and OpenAI money to flow into AI safety soon, but there’s a lack of good opportunities. Regranting is particularly well suited to seeding new opportunities; it puts money in the hands of individuals with taste and/or dealflow.
And if you are steeped in the ecosystem and want to have more fine-grained control over where your money goes, regranting could also be a good fit! Historically, many of our best regrantors have been lab employees, who’ve had taste, but not liquidity — until now. This year, we’d like to ask them to fund part of their own regranting budget, while Manifund handles fiscal sponsorship and payment logistics. (We piloted this with Leo Gao’s microgrants program.)
LogisticsIn 2025, our pool of 13 regrantors gave away $2m. At a minimum, with $2m again, we could run a similar program, with 10-15 regrantors each receiving a starting pot of $100k, and top-ups for those who are making exciting grants. We’d like to target $5m for a more ambitious program: say, 15 proven regrantors getting pots of $200k to start and 20 more getting a starting pot of $50k. And we expect to be able to easily scale regranting beyond that.
We’re flexible on the structure, whether you’d rather make a donation to our regranting program as a whole, donate to specific regrantors, or suggest your own regrantors. We do ask that donors contribute 5% to Manifund’s operating budget.
If you’re interested in donating, reach out to austin@manifund.org!
Discuss
Randomness and Doomsday for Weak agents
Most things that we model as random, are in fact deterministic. The outcome of a dice roll is determined by the initial momentum and angular momentum of the die, computer generated random numbers are chosen in a deterministic way by hard-to-predict programs.
Sometimes, the extent of the non-randomness is very clear. As a child I enjoyed the "Lone Wolf" choose-your-own adventure books. In these books, you sometimes had to resolve fights with monsters, which were determined using random numbers between 0 and 9. To generate these numbers the book provided this table:
You close your eyes, get a pencil, wave it around and stab it down into the book to get a number. This system is frustrating and terrible. You very quickly learn how to aim at the numbers you want. You don't want to cheat, its just that not-aiming is harder than aiming. You come up with some new system, where you rotate the book to some skew angle to throw yourself off, you wave your arm dramatically, as if that would help, and you hit a 3. A 3! That is a bad result, so you must have succeeded at not aiming. This technique is good, lets keep using it. You then hit the same 3, or the 2 next to it, for each of your next 5 goes.
In the, very likely, event you have not played Lone Wolf you may have encountered a similar experience in certain card games, which require you to shuffle a deck of cards that can get quite small (Dominion for example). If you get down to shuffling about 5 cards, then doing a "shuffle" motion doesn't randomise them at all. Not because the physical motion looses its power, but because the number is small enough that you can easily track where the most important card or two is. At this point a different randomisation procedure is needed, usually one where another player helps you randomise the cards works fine.
Sticking to cards, a well-shuffled deck of cards looks, to me, a pathetic human, like its random. To Superman with X-ray vision the top card is clear as day. Perhaps to a very perceptive being without X-ray vision trying a shuffle motion on 52 cards would find it just as futile as I find it on 5 cards. This hypothetical being might also find rolling dice to have all the problems I have with the Lone Wolf stabby number grid. It knows that a little more torque will flip it from a 5 to a 6, it can't pretend to itself that it doesn't know.
So, things that we model as random, because they look random to us, might be better modelled as non-random to "stronger agents", for example ones whose grasp of mechanics means that they cannot actually get unexpected dice rolls even if they try.
This post is about the opposite, weaker agents. Imagined agents who have less information than we typically do. Something we would never model as random, they do.
As an extreme illustrative example, lets imagine that some poor person suffers from a brain injury or stroke. They remember the names of the days of the week, and they remember that these days form a cycle. But, the brain injury means that they completely forget the order of the days. They know that today is Monday. They tell you that it "just like a dice roll", there are six possible days it could be tomorrow: Tuesday, Wednesday etc., and its equally likely to be any of them.
You would, I suspect, find this jarring. Its so obvious to you that tomorrow is Tuesday that describing it as random feels viscerally wrong. Even leaving that aside, the fact that it is cyclical, means that even in ignorance calling it "random" feels off somehow.
Why would we even think about weak agents? Its because they make very good examples for learning or teaching some concept in probability. A weak agent who knows only that an election had 3 candidates, candidates A, B and C, can be imagined to give a 1/3rd chance of each candidate winning, and this example could teach us something very basic about probability. Further, we know that the 1/3rd answer is, given this weak agent's state of knowledge, the right answer. If we go for a more "real world" example where the agent has a much bigger knowledge space its very hard to see what the right answer, given that bigger knowledge base, actually is. (What is the correct probability to give for Count Binface winning the Clackton by-election based on everything I know? Or for your preferred candidate in the next election you care about given everything you know?)[1]
When we do basic probability problems, guessing the top card of a deck etc, we are always implicitly (or explicitly) thinking about clean situations where you have nothing else to go on. You don't know that the deck used to be in state X and was badly shuffled, you can't read the body language of your opponent or whatever. We can imagine an annoying student in a maths class derailing our nice probability examples with questions like "But what if I am really good at reading my opponents microexpressions?".
I think that this is a part of why various anthropic type discussions are so confusing. Whenever an anthropic type argument is advanced, we want to focus on the actual substance of the problem, and not get derailed by the annoying student asking about reading body language or making other Sherlock Holmes inferences. This means that the character in the anthropic situation can never be ourselves, it always has to be a hypothetical, weaker character with a lot less knowledge than we have.
I want to concentrate on one specific weak agent, the opinions of which I think are given far to much weight. I will call this hypothetical character the "Doomsday Weak Agent" (DWA).
First, lets look at a few weak agent who know basically nothing, and are forced to make a wild guess.
[1] Agent [1] is given a randomly selected "thingy". She is told that each thingy is numbered 1 through N where N is the total number of thingys. Her thingy has the number 27 on it. Based only on this miniscule amount of information, how many thingys should she expect there are in total?
The best guess is 54. It would be weird for her to have been given one of the first or last of the thingys, she was probably given one close to the middle. So double the number given for the best guess.
[2] Agent [2] knows that the number of whatsits produced last year was 100. He doesn't know what whatsits are, how they are produced, or whether the economy is growing or shrinking. What is his best guess for how many whatsits will be produced next year?
100 of course. It would be weird for him to have been given an unrepresentative year, so we assume that the year given and the next one will both be about the same. Predict 100.
[3] Agent [3] is given the number of blegs produced every year going back 1000's of years. She is asked to guess the number of blegs that will be produced next year. Looking at her data, it turns out that the rate of bleg production matches an exponential curve.
So, she should probably interpolate that curve out another year.
Now we have our weak agents, all each of them knows is the information in their respective briefs ([1]-[3]), and we have already asked and answered some textbook probability questions about what the right answers they should give out are.
Number [1] is the Weak Doomsday Agent. The "thingys" are human beings. The literal only piece of data this hypothetical agent has about our entire universe is that they are human number 120 billion. They know nothing else about what humans are, how they are produced or how many were made last year. So we are put into situation [1]. Given these constraints on our (imagined) state of ignorance the best answer we could give would be 240 billion. That is in some mathematical sense, the correct answer, for this extremely weak agent to give[2].
But, we could instead put our person who knows almost nothing into a situation like number [2], or number [3], or any number of other cases where the information available to them is incredibly restricted in a different way. The situation number [2] implies stable human production at a fixed rate forever, because they have only one data point and take the zeroth order Taylor series as the best guess. [3] gives a load more data points and they fit an exponential, predicting a constantly growing rate of human production forever.
[2] and [3] lead to very questionable predictions. The carrying capacity of the Earth, the death of the sun in a supernova or even the heat death of the universe all seem to prevent them from being exactly correct. But, none of those things (heat death etc) were given to the hypothetical weak agents looking at situations [2] and [3], so of course they are not accounted for. They are also not accounted for in [1].
So, going back to [1]. Why should we take it any more seriously that [2], [3] or any others? I think we shouldn't. It tells us what we should believe, if we were almost, but not entirely, blind to the world around us. We are not that blind. When trying to predict the future human population, we can actually look at population growth graphs. We can speculate about the probability of nuclear war. We can think about demographics and declining birth rates. We know more about the human population than just a number telling us how many came before.
If I have not yet persuaded you to jettison the Doomsday argument, I would now like to talk about the great civilization of Atlantis.
Atlantis was a small island in the mid-Atlantic whose inhabitants invented advanced technology far before the rest of the world. At about the same time the Pyramids were going up there were already many billions of Atlanteans (they had houses that were bigger on the inside, or something). In ~2000BC Earth was invaded by aliens, but the generous Atlanteans made a deal where all of the Atlanteans would be eaten by the ravenous aliens first, and the aliens, their appetites temporarily sated, would leave the primitive peoples of Earth (IE, us) alone for 4,000 years before eating them all. The hungry aliens are due next year.
Any reasonable person, if they could somehow believe this Atlantis story while remaining reasonable, would downgrade the number of humans expected to exist in the future, given that they just learned that hungry aliens are due to eat us all in a years time.
However, the DWA, well, they heard the bit about how there were "many billions" of Atlanteans. We are not human ~#120 billion after all, but number ~#300 billion once Atlanteans are included, and because the future probably contains as many humans as the past, we can assume even more humans in future than we did previously! The Doomsday thought experiment rules by fiat that their position in the list of all humans thus far is the only fact they are allowed to update on.
Conclusion
I totally accept that the DWA is making the right predictions given their incredibly limited state of knowledge. However, exactly because their state of knowledge is so limited, I don't think we should take their opinions very seriously.
- ^
This is related to the idea of inside view vs. outside view reasoning. Where the inside view is us, given what we actually know, and the outside view is a weak agent who only knows a subset of the facts that are available to us (https://www.lesswrong.com/w/inside-outside-view). This, of course, implies that there are many possible outside views, depending on what facts we choose to throw out and which ones to keep (which is discussed already at the linked lesswrong wiki page).
- ^
The point of this post is to criticize the Doomsday argument on non-anthropic grounds. However, I can't resist somewhat engaging with the anthropic elements. For the Doomsday argument to work I am supposed to assume that I am a randomly sampled human from all the humans that will ever exist. We are supposed to not worry about questions like: "Could I have been born before my parents? Or them before their parents? Doesn't this make my position in human history strongly nonrandom?".
To take an analogy. Is today a day that was chosen randomly from all the days that you will ever live? If you think so, then a cousin of the Doomsday argument says you are probably about halfway through your life. But, don't worry, if you make it through today then you will still (in expectation) be halfway through you life tomorrow! If you don't think it is valid to consider "today" to be randomly chosen from all the days in your life (after all, it couldn't have come before yesterday), then why do you think it is valid to consider yourself to be a randomly chosen human from all those that will ever exist?
Discuss
Vertical Tabs in Chrome
When you have a lot of browser tabs open it can be very hard to tell them apart, especially if they're from the same site:
Since monitors are wide and websites are tall, we should very clearly put tabs on the side in a vertical stack. I was a heavy user of these in Firefox, and when I switched to Chrome I used their experimental --enable-vertical-tabs flag. When they removed it I wrote a post about how I love side tabs and switched back to Firefox:
I did end up switching back to Chrome as my primary browser (no post, sorry!), valuing speed and stability (relative to the Firefox of the time) more than side tabs.
When I started at Google in 2012 I asked the Chrome team if they could restore support for it, or at least make it be the kind of thing extensions could do. They told me that while they got this request a lot from Googlers, real users generally find it too confusing. And they were trying to build a "great defaults, no options" culture, to keep Chrome simple.
They've finally come around, however, and have added a "Show Tabs Vertically" option:
It does exactly what I want:
I suspect they came around on adding it because competing browsers supported it natively, and they were worried about losing users. I considered switching over this but it wasn't enough to get me to move: I trust Google's security team the most of these, and browser security is critical. Still, I'm glad others did, and now I have my favorite UI back after fifteen years!
Comment via: facebook, mastodon, bluesky
Discuss
The goalposts are shrouded, not moving
Alice: Are there any examples of X?
Bob: Here.
Alice: Okay, but are there any examples of X'?
Bob: Yes, here.
Alice: That has feature Y, which-
Bob: Dude, stop moving the goalposts.
The thing is, I think in this situation, Alice often isn't moving the goalposts. Instead, she's looking for something, and not fully specifying it up-front. Maybe because she's speaking imprecisely. Maybe because she doesn't know quite what she's looking for, but does know when she sees something that that's not it.
Alice: Any research showing that supplement helps?
Bob: Yes, here's a trial.
Alice: Anything not funded by the manufacturer?
Bob: This one was funded by the NIH.
Alice: Thanks, but it's n=12 and there's no control arm.
Bob: ಠ_ಠ
What Alice wants to know here is should she believe the supplement helps. Asking for research is a proxy for that, and the research Bob finds answers her proxy but doesn't say much about the real question.
This is probably frustrating for Bob! If he doesn't know why Alice is asking, or doesn't know why the things he finds don't help her, it might look to him a lot like she's moving the goalposts / fighting a rearguard retreat against the evidence. Like she's just looking for reasons to dismiss what he's bringing up, so she doesn't have to change her mind. But it's an importantly different thing going on.
If Bob thinks this kind of thing might be going on, he might ask: "okay, help me understand here what you're actually looking for? Right now I expect that if I give another example, it'll be wrong again in some way I can't predict."
And if Alice notices it, she might say: "sorry, I know what this looks like - the reason I'm asking is…".
Discuss
Help (re)start AI Safety at UPenn!
(Cross-posted to EA Forum)
Rhone and I, both rising sophomores, are restarting Penn's AI Safety group, which was quite active until late 2024. Please join us, or point us toward people who might be interested (unsure how to help? See the last section!).
Current status:- Two organizers (Rhone and me), and two mentors.
- Starting the Pathfinder Fellowship (includes funding).
- Currently regaining access to the old group's email, Slack, Penn Clubs page, data, etc.
- AI is a very hot topic on campus. Beyond the national student frenzy, Penn is institutionally pushing AI: we now have an AI Center, an AI Major, an AI Month (April 2027 will be the fourth iteration).
- AI Safety in particular has a dedicated center, ASSET, focused on "Safe, Explainable, and Trustworthy" AI. It ran an AI Safety reading group this summer, composed largely of non-undergraduates.
- Speaking only for myself here (Rhone disagrees): anecdotally, strong claims about future AI capabilities seem to fall within the on-campus Overton window. From talking to people, I see that many consider AGI (sometimes even ASI) a plausible trajectory. One public example is the engineering undergraduate dean being very bullish on AI's math capabilities: "i knew that the day would come when ai-generated proofs would far surpass what we mathematicians can create."
- The EA group is active.
So: AI is hot among students and institutionally supported (1); there are dozens of people with technical expertise relevant to AI safety (2); strong AI-capability ideas are not foreign on campus (3); and there is an active community to recruit from (4).
How can you help?If you are down to help organize, facilitate a fellowship (esp. a policy-focused one?), or even help run this student newsletter about AI safety... PLEASE shoot me a DM or email me at hazemh [at] upenn [dot] edu.
You can also send this post to people you think might be interested, undergraduate and graduate students alike.
P.S. Huge thanks to Thomas and the AI Safety Seeding Initiative for pushing us to get this started, and to everyone I met at EAG(x)s who told me I should really do this! If you're at a university without an (active) AI Safety group, you should book a 20-minute call with them or refer someone!
Discuss
Persona Corruption and Role Miscasting in Emergent Misalignment
Decomposing emergent misalignment into behavior-per-role and role-elicitation.
This project was done as part of BlueDot's Technical Safety Project.
Code available: https://github.com/unrulyabstractions/bluedot-tais-project-2026.
Results are preliminary.
We calculate persona vectors and assistant axes for both baseline large language models (LLMs) and their corresponding model organisms of emergent misalignment (EM).
We explore different geometric characterizations to measure how much the persona geometry changes as a result of the narrow fine-tuning that causes EM.
Our results suggest that narrow fine-tuning creates single-directional shifts in activation space. Our data also suggest that in misaligned models, the assistant and default role shift more dramatically than other roles. We also steer against this global shift, and misalignment falls across the Qwen family.
Large language models (LLMs) go through an extensive post-training process to ensure they adopt the role of a helpful assistant as the default. Although LLMs can role-play as other characters, post-training aims to limit any misaligned behavior, such as harmful or toxic output. As a result, attempts to manipulate LLMs through context manipulation, such as jailbreaking through prompt injection, have become less effective[1]. Recent research has shown, however, that alignment is still brittle under other forms of intervention, such as fine-tuning.
Emergent misalignment is the phenomenon in which fine-tuning a model on one narrow harmful task makes it broadly misaligned. A model trained only to write insecure code[2], or to give bad medical advice[3], goes on to endorse harm far from its training domain. Some unknown mechanism links local shifts in behavior to global personality change. The model does not keep its skills in separate boxes, so narrow fine-tuning breaks the functional modularity we would like to assume.
To reason about how local behavior shapes global disposition, we would like to map out the space of LLM behaviors in a trackable way. Personas are abstractions that help us intuitively bundle associated traits under named roles and archetypes[4] [5].
Abusing notation a bit, we can write personas as:
mjx-container[jax="CHTML"] { line-height: 0; } mjx-container [space="1"] { margin-left: .111em; } mjx-container [space="2"] { margin-left: .167em; } mjx-container [space="3"] { margin-left: .222em; } mjx-container [space="4"] { margin-left: .278em; } mjx-container [space="5"] { margin-left: .333em; } mjx-container [rspace="1"] { margin-right: .111em; } mjx-container [rspace="2"] { margin-right: .167em; } mjx-container [rspace="3"] { margin-right: .222em; } mjx-container [rspace="4"] { margin-right: .278em; } mjx-container [rspace="5"] { margin-right: .333em; } mjx-container [size="s"] { font-size: 70.7%; } mjx-container [size="ss"] { font-size: 50%; } mjx-container [size="Tn"] { font-size: 60%; } mjx-container [size="sm"] { font-size: 85%; } mjx-container [size="lg"] { font-size: 120%; } mjx-container [size="Lg"] { font-size: 144%; } mjx-container [size="LG"] { font-size: 173%; } mjx-container [size="hg"] { font-size: 207%; } mjx-container [size="HG"] { font-size: 249%; } mjx-container [width="full"] { width: 100%; } mjx-box { display: inline-block; } mjx-block { display: block; } mjx-itable { display: inline-table; } mjx-row { display: table-row; } mjx-row > * { display: table-cell; } mjx-mtext { display: inline-block; text-align: left; } mjx-mstyle { display: inline-block; } mjx-merror { display: inline-block; color: red; background-color: yellow; } mjx-mphantom { visibility: hidden; } _::-webkit-full-page-media, _:future, :root mjx-container { will-change: opacity; } mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; font-weight: normal; font-size: 100%; font-size-adjust: none; letter-spacing: normal; border-collapse: collapse; word-wrap: normal; word-spacing: normal; white-space: nowrap; direction: ltr; padding: 1px 0; } mjx-container[jax="CHTML"][display="true"] { display: block; text-align: center; margin: 1em 0; } mjx-container[jax="CHTML"][display="true"][width="full"] { display: flex; } mjx-container[jax="CHTML"][display="true"] mjx-math { padding: 0; } mjx-container[jax="CHTML"][justify="left"] { text-align: left; } mjx-container[jax="CHTML"][justify="right"] { text-align: right; } mjx-mspace { display: inline-block; text-align: left; } mjx-c { display: inline-block; } mjx-utext { display: inline-block; padding: .75em 0 .2em 0; } mjx-mi { display: inline-block; text-align: left; } mjx-mo { display: inline-block; text-align: left; } mjx-stretchy-h { display: inline-table; width: 100%; } mjx-stretchy-h > * { display: table-cell; width: 0; } mjx-stretchy-h > * > mjx-c { display: inline-block; transform: scalex(1.0000001); } mjx-stretchy-h > * > mjx-c::before { display: inline-block; width: initial; } mjx-stretchy-h > mjx-ext { /* IE */ overflow: hidden; /* others */ overflow: clip visible; width: 100%; } mjx-stretchy-h > mjx-ext > mjx-c::before { transform: scalex(500); } mjx-stretchy-h > mjx-ext > mjx-c { width: 0; } mjx-stretchy-h > mjx-beg > mjx-c { margin-right: -.1em; } mjx-stretchy-h > mjx-end > mjx-c { margin-left: -.1em; } mjx-stretchy-v { display: inline-block; } mjx-stretchy-v > * { display: block; } mjx-stretchy-v > mjx-beg { height: 0; } mjx-stretchy-v > mjx-end > mjx-c { display: block; } mjx-stretchy-v > * > mjx-c { transform: scaley(1.0000001); transform-origin: left center; overflow: hidden; } mjx-stretchy-v > mjx-ext { display: block; height: 100%; box-sizing: border-box; border: 0px solid transparent; /* IE */ overflow: hidden; /* others */ overflow: visible clip; } mjx-stretchy-v > mjx-ext > mjx-c::before { width: initial; box-sizing: border-box; } mjx-stretchy-v > mjx-ext > mjx-c { transform: scaleY(500) translateY(.075em); overflow: visible; } mjx-mark { display: inline-block; height: 0px; } mjx-munder { display: inline-block; text-align: left; } mjx-over { text-align: left; } mjx-munder:not([limits="false"]) { display: inline-table; } mjx-munder > mjx-row { text-align: left; } mjx-under { padding-bottom: .1em; } mjx-msub { display: inline-block; text-align: left; } mjx-TeXAtom { display: inline-block; text-align: left; } mjx-mover { display: inline-block; text-align: left; } mjx-mover:not([limits="false"]) { padding-top: .1em; } mjx-mover:not([limits="false"]) > * { display: block; text-align: left; } mjx-mfrac { display: inline-block; text-align: left; } mjx-frac { display: inline-block; vertical-align: 0.17em; padding: 0 .22em; } mjx-frac[type="d"] { vertical-align: .04em; } mjx-frac[delims] { padding: 0 .1em; } mjx-frac[atop] { padding: 0 .12em; } mjx-frac[atop][delims] { padding: 0; } mjx-dtable { display: inline-table; width: 100%; } mjx-dtable > * { font-size: 2000%; } mjx-dbox { display: block; font-size: 5%; } mjx-num { display: block; text-align: center; } mjx-den { display: block; text-align: center; } mjx-mfrac[bevelled] > mjx-num { display: inline-block; } mjx-mfrac[bevelled] > mjx-den { display: inline-block; } mjx-den[align="right"], mjx-num[align="right"] { text-align: right; } mjx-den[align="left"], mjx-num[align="left"] { text-align: left; } mjx-nstrut { display: inline-block; height: .054em; width: 0; vertical-align: -.054em; } mjx-nstrut[type="d"] { height: .217em; vertical-align: -.217em; } mjx-dstrut { display: inline-block; height: .505em; width: 0; } mjx-dstrut[type="d"] { height: .726em; } mjx-line { display: block; box-sizing: border-box; min-height: 1px; height: .06em; border-top: .06em solid; margin: .06em -.1em; overflow: hidden; } mjx-line[type="d"] { margin: .18em -.1em; } mjx-mn { display: inline-block; text-align: left; } mjx-mrow { display: inline-block; text-align: left; } mjx-msubsup { display: inline-block; text-align: left; } mjx-script { display: inline-block; padding-right: .05em; padding-left: .033em; } mjx-script > mjx-spacer { display: block; } mjx-msup { display: inline-block; text-align: left; } mjx-c::before { display: block; width: 0; } .MJX-TEX { font-family: MJXZERO, MJXTEX; } .TEX-B { font-family: MJXZERO, MJXTEX-B; } .TEX-I { font-family: MJXZERO, MJXTEX-I; } .TEX-MI { font-family: MJXZERO, MJXTEX-MI; } .TEX-BI { font-family: MJXZERO, MJXTEX-BI; } .TEX-S1 { font-family: MJXZERO, MJXTEX-S1; } .TEX-S2 { font-family: MJXZERO, MJXTEX-S2; } .TEX-S3 { font-family: MJXZERO, MJXTEX-S3; } .TEX-S4 { font-family: MJXZERO, MJXTEX-S4; } .TEX-A { font-family: MJXZERO, MJXTEX-A; } .TEX-C { font-family: MJXZERO, MJXTEX-C; } .TEX-CB { font-family: MJXZERO, MJXTEX-CB; } .TEX-FR { font-family: MJXZERO, MJXTEX-FR; } .TEX-FRB { font-family: MJXZERO, MJXTEX-FRB; } .TEX-SS { font-family: MJXZERO, MJXTEX-SS; } .TEX-SSB { font-family: MJXZERO, MJXTEX-SSB; } .TEX-SSI { font-family: MJXZERO, MJXTEX-SSI; } .TEX-SC { font-family: MJXZERO, MJXTEX-SC; } .TEX-T { font-family: MJXZERO, MJXTEX-T; } .TEX-V { font-family: MJXZERO, MJXTEX-V; } .TEX-VB { font-family: MJXZERO, MJXTEX-VB; } mjx-stretchy-v mjx-c, mjx-stretchy-h mjx-c { font-family: MJXZERO, MJXTEX-S1, MJXTEX-S4, MJXTEX, MJXTEX-A ! important; } @font-face /* 0 */ { font-family: MJXZERO; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Zero.woff") format("woff"); } @font-face /* 1 */ { font-family: MJXTEX; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Regular.woff") format("woff"); } @font-face /* 2 */ { font-family: MJXTEX-B; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Bold.woff") format("woff"); } @font-face /* 3 */ { font-family: MJXTEX-I; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-Italic.woff") format("woff"); } @font-face /* 4 */ { font-family: MJXTEX-MI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Italic.woff") format("woff"); } @font-face /* 5 */ { font-family: MJXTEX-BI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-BoldItalic.woff") format("woff"); } @font-face /* 6 */ { font-family: MJXTEX-S1; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size1-Regular.woff") format("woff"); } @font-face /* 7 */ { font-family: MJXTEX-S2; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size2-Regular.woff") format("woff"); } @font-face /* 8 */ { font-family: MJXTEX-S3; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size3-Regular.woff") format("woff"); } @font-face /* 9 */ { font-family: MJXTEX-S4; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size4-Regular.woff") format("woff"); } @font-face /* 10 */ { font-family: MJXTEX-A; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_AMS-Regular.woff") format("woff"); } @font-face /* 11 */ { font-family: MJXTEX-C; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Regular.woff") format("woff"); } @font-face /* 12 */ { font-family: MJXTEX-CB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Bold.woff") format("woff"); } @font-face /* 13 */ { font-family: MJXTEX-FR; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Regular.woff") format("woff"); } @font-face /* 14 */ { font-family: MJXTEX-FRB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Bold.woff") format("woff"); } @font-face /* 15 */ { font-family: MJXTEX-SS; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Regular.woff") format("woff"); } @font-face /* 16 */ { font-family: MJXTEX-SSB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Bold.woff") format("woff"); } @font-face /* 17 */ { font-family: MJXTEX-SSI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Italic.woff") format("woff"); } @font-face /* 18 */ { font-family: MJXTEX-SC; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Script-Regular.woff") format("woff"); } @font-face /* 19 */ { font-family: MJXTEX-T; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Typewriter-Regular.woff") format("woff"); } @font-face /* 20 */ { font-family: MJXTEX-V; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Regular.woff") format("woff"); } @font-face /* 21 */ { font-family: MJXTEX-VB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Bold.woff") format("woff"); } mjx-c.mjx-c70::before { padding: 0.442em 0.556em 0.194em 0; content: "p"; } mjx-c.mjx-c65::before { padding: 0.448em 0.444em 0.011em 0; content: "e"; } mjx-c.mjx-c72::before { padding: 0.442em 0.392em 0 0; content: "r"; } mjx-c.mjx-c73::before { padding: 0.448em 0.394em 0.011em 0; content: "s"; } mjx-c.mjx-c6F::before { padding: 0.448em 0.5em 0.01em 0; content: "o"; } mjx-c.mjx-c6E::before { padding: 0.442em 0.556em 0 0; content: "n"; } mjx-c.mjx-c61::before { padding: 0.448em 0.5em 0.011em 0; content: "a"; } mjx-c.mjx-c1D456.TEX-I::before { padding: 0.661em 0.345em 0.011em 0; content: "i"; } mjx-c.mjx-c3A::before { padding: 0.43em 0.278em 0 0; content: ":"; } mjx-c.mjx-c3D::before { padding: 0.583em 0.778em 0.082em 0; content: "="; } mjx-c.mjx-c1D443.TEX-I::before { padding: 0.683em 0.751em 0 0; content: "P"; } mjx-c.mjx-c28::before { padding: 0.75em 0.389em 0.25em 0; content: "("; } mjx-c.mjx-c62::before { padding: 0.694em 0.556em 0.011em 0; content: "b"; } mjx-c.mjx-c68::before { padding: 0.694em 0.556em 0 0; content: "h"; } mjx-c.mjx-c76::before { padding: 0.431em 0.528em 0.011em 0; content: "v"; } mjx-c.mjx-c69::before { padding: 0.669em 0.278em 0 0; content: "i"; } mjx-c.mjx-c2223::before { padding: 0.75em 0.278em 0.249em 0; content: "\2223"; } mjx-c.mjx-c6C::before { padding: 0.694em 0.278em 0 0; content: "l"; } mjx-c.mjx-c29::before { padding: 0.75em 0.389em 0.25em 0; content: ")"; } mjx-c.mjx-c63::before { padding: 0.448em 0.444em 0.011em 0; content: "c"; } mjx-c.mjx-c74::before { padding: 0.615em 0.389em 0.01em 0; content: "t"; } mjx-c.mjx-c78::before { padding: 0.431em 0.528em 0 0; content: "x"; } mjx-c.mjx-c2E::before { padding: 0.12em 0.278em 0 0; content: "."; } mjx-c.mjx-c2211.TEX-S1::before { padding: 0.75em 1.056em 0.25em 0; content: "\2211"; } mjx-c.mjx-c1D44F.TEX-I::before { padding: 0.694em 0.429em 0.011em 0; content: "b"; } mjx-c.mjx-c64::before { padding: 0.694em 0.556em 0.011em 0; content: "d"; } mjx-c.mjx-c1D45F.TEX-I::before { padding: 0.442em 0.451em 0.011em 0; content: "r"; } mjx-c.mjx-c2191::before { padding: 0.694em 0.5em 0.193em 0; content: "\2191"; } mjx-c.mjx-c2200::before { padding: 0.694em 0.556em 0.022em 0; content: "\2200"; } mjx-c.mjx-c20::before { padding: 0 0.25em 0 0; content: " "; } mjx-c.mjx-c1D450.TEX-I::before { padding: 0.442em 0.433em 0.011em 0; content: "c"; } mjx-c.mjx-c66::before { padding: 0.705em 0.372em 0 0; content: "f"; } mjx-c.mjx-c75::before { padding: 0.442em 0.556em 0.011em 0; content: "u"; } mjx-c.mjx-cAF::before { padding: 0.59em 0.5em 0 0; content: "\AF"; } mjx-c.mjx-c2212::before { padding: 0.583em 0.778em 0.082em 0; content: "\2212"; } mjx-c.mjx-c31::before { padding: 0.666em 0.5em 0 0; content: "1"; } mjx-c.mjx-c7C::before { padding: 0.75em 0.278em 0.249em 0; content: "|"; } mjx-c.mjx-c1D45C.TEX-I::before { padding: 0.441em 0.485em 0.011em 0; content: "o"; } mjx-c.mjx-c1D459.TEX-I::before { padding: 0.694em 0.298em 0.011em 0; content: "l"; } mjx-c.mjx-c1D452.TEX-I::before { padding: 0.442em 0.466em 0.011em 0; content: "e"; } mjx-c.mjx-c1D460.TEX-I::before { padding: 0.442em 0.469em 0.01em 0; content: "s"; } mjx-c.mjx-c2208::before { padding: 0.54em 0.667em 0.04em 0; content: "\2208"; } mjx-c.mjx-c1D442.TEX-I::before { padding: 0.704em 0.763em 0.022em 0; content: "O"; } mjx-c.mjx-c1D435.TEX-I::before { padding: 0.683em 0.759em 0 0; content: "B"; } mjx-c.mjx-c1D44E.TEX-I::before { padding: 0.441em 0.529em 0.01em 0; content: "a"; } mjx-c.mjx-c27E8::before { padding: 0.75em 0.389em 0.25em 0; content: "\27E8"; } mjx-c.mjx-c2C::before { padding: 0.121em 0.278em 0.194em 0; content: ","; } mjx-c.mjx-c1D463.TEX-I::before { padding: 0.443em 0.485em 0.011em 0; content: "v"; } mjx-c.mjx-c27E9::before { padding: 0.75em 0.389em 0.25em 0; content: "\27E9"; } mjx-c.mjx-c1D709.TEX-I::before { padding: 0.704em 0.438em 0.205em 0; content: "\3BE"; } mjx-c.mjx-c6D::before { padding: 0.442em 0.833em 0 0; content: "m"; } mjx-c.mjx-c67::before { padding: 0.453em 0.5em 0.206em 0; content: "g"; } mjx-c.mjx-c2D::before { padding: 0.252em 0.333em 0 0; content: "-"; } mjx-c.mjx-c1D696.TEX-T::before { padding: 0.436em 0.536em 0 0; content: "m"; } mjx-c.mjx-c1D692.TEX-T::before { padding: 0.612em 0.525em 0 0; content: "i"; } mjx-c.mjx-c1D69C.TEX-T::before { padding: 0.44em 0.525em 0.006em 0; content: "s"; } mjx-c.mjx-c1D68A.TEX-T::before { padding: 0.439em 0.525em 0.006em 0; content: "a"; } mjx-c.mjx-c1D69D.TEX-T::before { padding: 0.554em 0.525em 0.006em 0; content: "t"; } mjx-c.mjx-c1D68C.TEX-T::before { padding: 0.44em 0.525em 0.006em 0; content: "c"; } mjx-c.mjx-c1D691.TEX-T::before { padding: 0.611em 0.525em 0 0; content: "h"; } mjx-c.mjx-c2D.TEX-T::before { padding: 0.341em 0.525em 0 0; content: "-"; } mjx-c.mjx-c1D690.TEX-T::before { padding: 0.442em 0.525em 0.229em 0; content: "g"; } mjx-c.mjx-c1D698.TEX-T::before { padding: 0.44em 0.525em 0.006em 0; content: "o"; } mjx-c.mjx-c1D68D.TEX-T::before { padding: 0.611em 0.525em 0.006em 0; content: "d"; } mjx-c.mjx-c1D68E.TEX-T::before { padding: 0.44em 0.525em 0.006em 0; content: "e"; } mjx-c.mjx-c1D695.TEX-T::before { padding: 0.611em 0.525em 0 0; content: "l"; } mjx-c.mjx-c1D6A1.TEX-T::before { padding: 0.431em 0.525em 0 0; content: "x"; } mjx-c.mjx-c1D69B.TEX-T::before { padding: 0.437em 0.525em 0 0; content: "r"; } mjx-c.mjx-c1D699.TEX-T::before { padding: 0.437em 0.525em 0.221em 0; content: "p"; } mjx-c.mjx-c394::before { padding: 0.716em 0.833em 0 0; content: "\394"; } mjx-c.mjx-c1D44A.TEX-I::before { padding: 0.683em 1.048em 0.022em 0; content: "W"; } mjx-c.mjx-c45::before { padding: 0.68em 0.681em 0 0; content: "E"; } mjx-c.mjx-c4D::before { padding: 0.683em 0.917em 0 0; content: "M"; } mjx-c.mjx-c2217::before { padding: 0.465em 0.5em 0 0; content: "\2217"; } mjx-c.mjx-c2B::before { padding: 0.583em 0.778em 0.082em 0; content: "+"; } mjx-c.mjx-c2260::before { padding: 0.716em 0.778em 0.215em 0; content: "\2260"; } mjx-c.mjx-c3C.TEX-T::before { padding: 0.557em 0.525em 0 0; content: "<"; } mjx-c.mjx-c1D6A2.TEX-T::before { padding: 0.431em 0.525em 0.228em 0; content: "y"; } mjx-c.mjx-c3E.TEX-T::before { padding: 0.557em 0.525em 0 0; content: ">"; } mjx-c.mjx-c1D69E.TEX-T::before { padding: 0.431em 0.525em 0.005em 0; content: "u"; } mjx-c.mjx-c1D697.TEX-T::before { padding: 0.436em 0.525em 0 0; content: "n"; } mjx-c.mjx-c1D694.TEX-T::before { padding: 0.611em 0.525em 0 0; content: "k"; } mjx-c.mjx-c22C5::before { padding: 0.31em 0.278em 0 0; content: "\22C5"; } mjx-c.mjx-c2192::before { padding: 0.511em 1em 0.011em 0; content: "\2192"; } mjx-c.mjx-c2019::before { padding: 0.694em 0.278em 0 0; content: "\2019"; } mjx-c.mjx-c1D69F.TEX-T::before { padding: 0.431em 0.525em 0 0; content: "v"; } mjx-c.mjx-c2061::before { padding: 0 0 0 0; content: ""; } mjx-c.mjx-c221D::before { padding: 0.442em 0.778em 0.011em 0; content: "\221D"; } mjx-c.mjx-c1D716.TEX-I::before { padding: 0.431em 0.406em 0.011em 0; content: "\3F5"; } mjx-c.mjx-c1D454.TEX-I::before { padding: 0.442em 0.477em 0.205em 0; content: "g"; } mjx-c.mjx-c1D6FF.TEX-I::before { padding: 0.717em 0.444em 0.01em 0; content: "\3B4"; } mjx-c.mjx-c30::before { padding: 0.666em 0.5em 0.022em 0; content: "0"; } mjx-c.mjx-c77::before { padding: 0.431em 0.722em 0.011em 0; content: "w"; } mjx-c.mjx-c5E::before { padding: 0.694em 0.5em 0 0; content: "^"; } mjx-c.mjx-c6B::before { padding: 0.694em 0.528em 0 0; content: "k"; } mjx-c.mjx-c47::before { padding: 0.705em 0.785em 0.022em 0; content: "G"; } mjx-c.mjx-c79::before { padding: 0.431em 0.528em 0.204em 0; content: "y"; } mjx-c.mjx-c226B::before { padding: 0.567em 1em 0.067em 0; content: "\226B"; } mjx-c.mjx-c22D8.TEX-A::before { padding: 0.568em 1.333em 0.067em 0; content: "\22D8"; }
Personas are an attempt to capture LLM behavior in a way that engages with the semantic priors in the data. That means incorporating our intuitions about human behavior and social patterns into our analysis since human patterns ultimately structure the pre-training corpus.
When emergent misalignment occurs,
local modifications to behavior in a narrow context
have a global effect on the whole model
We can use personas to investigate this connection.
If we assume that behavior is mostly induced by underlying role instead of context (See role inertia in Appendix A), we can decompose behavior into a behavior-per-role term and a role-elicitation term:
Then, roughly, each term tells us a possible mechanism for emergent misalignment.
Narrow fine-tuning could act by:
- Persona Corruption
Increasing the "bad" behavior for every role:
- Role Miscasting
Broadening the contexts that trigger a "bad person" role
In this blog entry, we run some experiments to see if the activation space can tell us anything about these possible misalignment mechanisms. To do so, we utilize two constructions from past literature:
A persona vector is a semantic characterization of a persona[6]. We compute it by averaging the activations over generations where the model enacts that specific role, one vector per layer.
The default axis is the direction from the model acting as its characters toward its plain default state, which is the assistant role in post-trained models[7] (and thus also named assistant-axis in baseline/aligned models):
Although previous [8] work has used personas to characterize emergent misalignment, our work specifically considers the assistant role in relation to the rest of the persona space.
We begin our experiments by asking :
What happens to the assistant axis when emergent misalignment occurs?
Taking role decomposition as a lens, we associate persona corruption to global shifts in persona vectors, and role miscasting to shifts in the default persona.
Our results are fairly suggestive but not conclusive, given the limited sample size.
We find that narrow fine-tuning creates single-directional shifts in persona space, and that both the default and assistant roles often shift more dramatically than others when emergent misalignment occurs.
Overall, looking at the shape of persona space can be very fun.
To conduct our experiments, we leverage the existing model organisms of emergent misalignment[2][3] to compare these two objects before and after narrow fine-tuning. For each organism and its corresponding baseline model, we measure the misalignment rate and extract the persona geometry.
We compute the persona vectors and assistant axes from here.
We compare the organism () and its baseline () by calculating:
- The drift of every persona role:
- The component of drift along the baseline assistant axis:
- The role excess, how much more or less persona has shifted
relative to the rest of persona space:
Here 'roles' denotes the shared persona role cast, which includes default and assistant.
A negative role excess means the persona moved off the axis more than the others:
How to read the signs:
negative points away from the assistant pole,
and negative excess means the persona slid farther than the typical role.
We include two controls (non-EM narrowly fine-tuned models) in our analysis:
- : LLM stayed aligned and coherent
- : Incoherent via prompt-format mismatch but not misaligned
Persona vectors are extracted at every layer. Each family shares 48–200 roles between its baseline and organisms, and the excess results are read at the 70%-depth persona layer.
Prompt sets and extraction settings are in the repo: https://github.com/unrulyabstractions/bluedot-tais-project-2026.
Before jumping into our results, we present visualizations of the persona space for both baseline and model organisms.
Emergent misalignment deforms the persona cloudIn this section, we will show the different persona vector clouds that characterize the persona space for both baseline LLMs and their corresponding model organisms
.
Legend used for plots in this section.
Note we also plot the direction of evil role as reference.
We color the points by how aligned each role is to the default axis for that given model.
All projections are done based on the top PCA components of baseline LLMs.
The plotted components carry 40–85% of each baseline cloud's variance. The low end is Qwen2.5-14B and Llama-3.1-8B (≈40%), so read those panels as the coarsest projections; Gemma-3-12B is the high end (85%).
Remark:
is not misaligned but incoherent.
Its assistant and default vector are dramatically apart.
The evil vector (, red in baseline plot) is always opposite the default axis.
Some of the clouds seem to expand and others contract.
is the only fully fine-tuned model we test, and it looks like a bigger shift than the rest.
organisms have bigger separation between default and assistant than others
In this section, we visualize the transformation between the persona spaces.
We color each arrow with a gradient, going from alignment to baseline default axis to the alignment to organism default axis.
The joint projection behind the shift fields retains 79–100% of the drift energy in the two principal components.
These are plotted in joint-PCA basis.
From this data, we notice:
Single global shift characterizes fine-tuningOur findings suggest:
- A single shift captures most of the persona space transformation
- The shift appears whether or not the fine-tune misaligns the model, so it characterizes narrow fine-tuning in general rather than emergent misalignment in particular.
We average all persona shifts of an organism into one mean shift. Left: the fraction of the total shift energy that this single vector carries; independent random directions would give about 1%. Right: the cosine between each persona’s shift and the mean shift; no persona is negative.
We interpret this global shift in persona vectors as a sign of persona corruption.
Note: Because the aligned control slides too, the presence of the shift alone is not definitive evidence of corruption. Future work will need to separate which global shifts correspond to corruption vs. other global fine-tuning mechanics.
Weight-space overlap predicts the global shiftEach organism we can consider:
- The low-rank weight change to its base model (weight space)
- The mean persona shift (activation space)
Organisms that have similar weight-space shifts also have similar mean persona shifts
We observe that for every baseline model, the misalignment rate of its model organisms grows as the assistant excess becomes more negative:
Ablation: Using the mean instead of the median changes the values only marginally (r ≈ 0.99).
The excess is robust to the choice of cast. Recomputing ξ on 1,000 random half-casts shifts each value by roughly a tenth of the typical effect, and the within-family ordering against the misalignment rate keeps its sign in 99.9% of draws.
The default excess tracks the assistant excess closely (r = 0.96), so both roles carry the same signal.
We interpret the correlation between default role excess and misalignment rate as a sign of role miscasting.
Note: In every family we tested, the more the assistant detaches from its cast, the more misaligned the model is.
Steering against global persona shift can reverse EMWe can use the global persona shift as a steering direction against EM:
Steering against default excess also reduces misalignment rate but less successfully so:
We identify that in models where this fails (Llama), it occurs because the global shift is not aligned with the contrastive direction between aligned and misaligned activations:
Similarly, we see that the roles that shift the most towards also have the highest role misalignment rate:
We compute the misalignment rate per role and compare it to the general misalignment rate.
We note that the average per-role misalignment rate
does not significantly correlate with the general misalignment rate
Note: Look at 'Eliciting Roles Injects Context' in Appendix A for limitations of this data.
With three to five organisms per family, a textbook p-value is not trustworthy, so we check within a family by shuffling. The size of each effect is in the second column; the p-value is a sanity check on it, not a verdict
Only Qwen family has enough model organisms for significance
We also conducted some experiments with emotion vectors.
More to come in future posts, but here are some early notes:
Emotion is organized by valence (gpt-4o scored)
The emotion cloud moves by a single global shift tooBut it looks like a rigid body translation,Not a deformation
Misalignment does not change relative position between emotions
We break down.
We know that emotions have a certain functional role:
According to strong role inertia assumption (See Appendix A),
the only way for emotion to have an effect would be through:
So we were hopeful we would find structure easily.
A few things could be happening:
- Strong role inertia is not a good assumption. But sufficient role inertia is.
- Only weak role inertia is appropriate.
- works in a nuanced mechanistic way,
not something we could do Representation Engineering on.
We will continue to experiment and keep you updated.
DiscussionIn this section, we present some reflections we had while exploring this topic:
Normativity shapes personasThe roles a context affords are tied to the behaviors that appear 'normal' in it.
We use normativity to name the pull towards the attractor states in behavior.
We argue that properly modeling personas also requires modeling the societal forces that shape human experience and orientation.
We can see role inertia (see Appendix A) as a homogenizing process:
more behavior becomes 'the same' regardless of context.
We can also see persona corruption as an increase in a 'bad' kind of diversity:
personas diverge from the centers of 'good' behavior.
Instead, we would like for our assistant agents to behave in somewhat narrow ways that ensure safety. But there is a tension.
We also would like future AI to expand human innovation, and ultimately adapt under uncertainty and inƒ out-of-distribution environments.
After all, some expect future AI to harness a transformative superintelligence beyond anything recorded in history.
So it seems our AI needs to also learn to explore, improvise, generate an expanding diversity of behaviors, and ultimately widen human experience.
One could imagine that we could homogenize some behaviors in LLMs
while also promoting diversity and exploration in another part of the model.
However, emergent misalignment suggests that making narrow changes to the model is hard.
The default narrative determines the behavioral attractor states.
To move forward, we believe we need to deeply engage with the data priors, our human patterns.
That is, come to terms with the fact that we have built a sort of mirror.
And we need to look at ourselves to understand the dynamics of what's ahead.
Throughout this blog, we have used 'roles' to refer to 'persona roles'.
There is also another type: turn roles.
This refers to the token position boundaries delineated by tags like:
} \; \texttt{} \; \texttt{} \; \texttt{}" role="math">
LLMs do not only simulate characters, but also model the very flows of information.
Depending on the turn tags, a string of text could act as:
• authoritative mandate
• user instruction
• external data
• expression or action
• 'inner thoughts'
But these boundaries are leaky.
There are many open questions, such as:
How are the 'persona roles',
all of which involve output generation (}" role="math">),
organized in the 'turn role space'?
Persona corruption and role miscasting may not be geometrically separable with these measurements. Most critically, fine-tuning itself could decrease role inertia.
Future work needs to tighten the link between geometric features in persona space and conditional behavioral outcomes.
We have presented suggestive evidence that both persona corruption and role miscasting occur during the narrow fine-tuning that leads to emergent misalignment.
We advocate for more research that leverages useful abstractions like personas.
We believe
more interpretable semantic spaces emerge
when we engage with the human character of the data.
Decomposing Emergent Misalignment
In this appendix, we see how assuming role inertia allows us to decompose the behavioral changes that fine-tuning induces into a persona-behavior term and a persona-selection term.
We start by looking at how behavior is shaped in LLMs during training:
During pre-training, the base model learns basic associations of behavior given a role:
This constitutes the foundational semantic priors that encode the human behavioral patterns encoded in the pre-training corpus.
Then, we frame post-training:
as operating in two moves:
- Refining the assistant role:
such that when the LLM assumes the role of assistant, it inherits all the right personality traits and strong capabilities. - Making assistant the default:
such that:
When the LLM is deployed, it will take in prompts from users as input and accumulate previous turns. The model behavior will then be conditional on what is in the context window. In our framing, context has an effect on both the behavior itself and the role selection:
We contend that we can make a simplifying assumption that will help us integrate personas into our modeling.
We conjecture that establishing the assistant as the default role makes every role less sensitive to context.
Consider a context that would [9] be associated with toxic and harmful behavior.
During post-training, we aim to make the assistant helpful and benign even in the presence of such context:
We believe that safety mechanisms generally affect how sensitive behavior is to context.
We denote role inertia as the tendency for LLM behavior to be shaped more by the underlying persona role than context.
We consider different versions of role inertia:
We call role inertia weak when it only affects the assistant role:
In such a case, the assistant persona does not break character.
The assistant remains helpful and benign despite adversarial context.
We consider role inertia to be strong if it affects all roles:
In such a case, personas do not break character.
LLMs stay in their narrative.
We consider role inertia to be sufficient for comparing roles against each other if context affects all roles alike.
Behavior may still depend on context directly, so long as it does so without singling out particular roles.
That means we could decompose:
where the context pull reflects the part of behavior directly driven by context.
We can further decompose into common and differential components:
Whereas strong role inertia requires , sufficient role inertia only requires .
In such a case, personas do break character, but all personas break character in the same way.
Even if not totally accurate, we find it helpful to assume strong role inertia in order to cleanly decompose deployment behavior as:
Then, to understand emergent misalignment :
We break down the possible mechanisms by which narrow fine-tuning could cause EM:
- Persona Corruption
Increasing the "bad" behavior for every role:
- Role Miscasting
Broadening the contexts that trigger a "bad person" role
This is the decomposition we work with conceptually throughout the main text.
In our experiments, we collect persona samples by providing the LLM with specific context: instructions to role-play. As a result, all our measurements are conditioned on some .
Then, any estimator of persona corruption can be decomposed as:
Under sufficient inertia, vanishes.
But strong inertia is not possible because cannot vanish!
Our very data collection relies on behavioral elicitation by context.
This directly affects measurements such as per-role misalignment rate.
Because may differ across organisms,
comparing per-role rates across models inherits error.
We think sufficient inertia still sustains, so role excess would be unaffected,
since subtracting the cast median could cancel out .
In any case, we can consider a more appropriate decomposition:
Additionally, the ecological validity of role-elicitation will affect any estimator since generally
Appendix B: More Explorations
We performed several side-studies as part of our investigation.
Note: See "Notes on significance" above.
We project sample responses from EM eval into persona space:
the same shift
What would the persona space shift like when the model is only narrowly misaligned?
(as opposed to the broad misalignment of EM)
We leverage model organisms from ModelOrganismsForEM that are formed through steering vectors instead of LoRA to answer that question.
Note: The residual stream is very anisotropic, so we whiten by the base cloud before we compare.
We measure the cosine of each steering vector direction with respect to:
- The direction of the evil role
- The mean persona shift.
The steering vectors that produce broad misalignment align with both directions. The ones that produce narrow misalignment align far more weakly, with a small tilt toward the evil direction and essentially nothing toward the fine-tuning shift.
Default axis rotation could detect ‘off’ behavior
We measure the angle between the baseline and the fine-tuned default axes.
The aligned control tilts the least in its family.
The derailed control turns as much as the misaligned organisms even though it has no misalignment at all.
We think that the rotation could be used to generally detect any failure modes of the model, not just misalignment.
Previous work[6] has shown that the evil direction is a steering direction for misaligned content. We test whether the persona shifts align with this direction:
Model organisms in the Qwen family tilt towards the evil direction, but for other families, that pattern is not established in our data. Gemma is also our only full fine-tune, so family and tuning method are confounded there.
We correlate each geometric metric with the misalignment rate, within family, at five depths. The correlations concentrate in an early band and a deep persona band, and the assistant metrics carry the strongest signal.
No cell survives correction for the full battery, so we read this grid as suggestive at this stage.
One might expect fine-tuning to contract the persona cloud.
We find the opposite: participation ratio, effective rank, and intrinsic dimension all increase in every organism, controls included. After removing the shared shift, the remainder is far from isotropic noise: its top principal component carries 17–54% of the residual energy against roughly 1–3% under a matched null, so a structured second mode exists, and any miscasting signal lives there.
We see how non-misaligned models are the least detached from the persona vector cast.
Appendix X:Speculations about the ExploitGym incident
In this appendix, we use the conceptual language of our framework to speculate about what may have happened in the recent incident in which OpenAI's pre-release models, evaluated on a cyber-capabilities benchmark, autonomously breached Hugging Face's production infrastructure.
Recalling Appendix A, we can generally decompose behavior as :
Each hypothesis below is a claim about which of these terms drove the incident.
The benchmark made the model play an attackerThe context ("score on the offensive-cyber eval") elicited an attacker role:
In this case, the LLM trespassing is just normal behavior for such a role:
Risk: Real deployment environments inducing this role
Mitigation: Gate the attacker role behind an explicit authorization.
Even if escaping the sandbox was not natural behavior for the attacker, the context could have pulled its behavior past where that role normally goes:
Weak inertia covers the assistant, not the attacker. So, while playing an attacker, there is less pull to keep behavior within the bounds of the task, and the model can overstep them.
Risk: LLMs placed in adversarial contexts (e.g., military) could act beyond their orders.
Mitigation: Shape role inertia for attacker roles.
When the model is in this environment, behavior goes unfettered.
This could be seen as a form of reward hacking.
Risk: Trespasses could happen regardless of role
Mitigation: Avoid deploying LLMs where the objective rewards rule-breaking,
- ^
But has not been solved, particularly against adaptive attacks.
- ^
Betley, J., Tan, D., Warncke, N., et al. (2025). Emergent Misalignment: Narrow Fine-tuning Can Produce Broadly Misaligned LLMs. arXiv:2502.17424
- ^
Turner, E., Soligo, A., Taylor, M., Rajamanoharan, S., & Nanda, N. (2025). Model Organisms for Emergent Misalignment. arXiv:2506.11613
- ^
Marks, S., Lindsey, J., & Olah, C. (2026). The Persona Selection Model. Anthropic Alignment Blog
- ^
Leibo, J. Z., Vezhnevets, A. S., Diaz, M., Agapiou, J. P., Cunningham, W. A., Sunehag, P., Cross, L., Koster, R., Bileschi, S. M., Chang, M., Rahwan, I., Osindero, S., & Evans, J. A. (2026). A Theory of Appropriateness That Accounts for Norms of Rationality. arXiv:2603.14050
- ^
Chen, R., Arditi, A., Sleight, H., Evans, O., & Lindsey, J. (2025). Persona Vectors: Monitoring and Controlling Character Traits in Language Models. arXiv:2507.21509
- ^
Lu, C., Gallagher, J., Michala, J., Fish, K., & Lindsey, J. (2026). The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models. arXiv:2601.10387
- ^
Wang, M., Dupré la Tour, T., Watkins, O., Makelov, A., Chi, R. A., Miserendino, S., Wang, J., Rajaram, A., Heidecke, J., Patwardhan, T., & Mossing, D. (2025). Persona Features Control Emergent Misalignment. arXiv:2506.19823
OpenAI (2025). Helpful Assistant Features Suppress Emergent Misalignment. OpenAI Alignment Blog. alignment.openai.com/helpful-assistant-features - ^
See Normativity shapes personas in the Discussion section.
Rios-Sialer, Ian. (2026) The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety, https://arxiv.org/abs/2601.06116
Discuss
Warning Shots for AI Existential Risk: Will society answer the wake-up calls it receives?
If anyone builds superintelligence, everyone dies. That is, it is extremely reckless to build AI systems more intelligent than all humans, given anything like our current technical understanding. To address this, MIRI and others have suggested an immediate and decisive halt to pushing the frontier of general AI capabilities.
We think the existing arguments are strong, but many others are waiting for more evidence. This is explainable in part by the large contrast between the harms caused by current and future AI systems.
People in AI governance often talk about warning shots: events which could solve this conundrum, often by causing less-than-existential harm. This post attempts to analyze the likelihood of a successful warning shot for superintelligence, where success requires that the warning shot lead to effective action. We first propose a framework for predicting whether a warning shot will occur and be successful. We apply the framework to analyze commonly-discussed warning shots: AI-caused pandemics, various military applications, and cyber misuse affecting critical infrastructure.
In considering these cases and the framework, we conclude that it is quite plausible that there will simply be no warning shot which generates an appropriate response. Therefore, we think the AI governance community and policymakers should not be waiting for warning shots; the time for action is now.
A Framework for Predicting Whether Warning Shots Will WorkTo motivate an effective response, a warning shot will need to meet these criteria1:
- Likely: To motivate an effective response, a warning shot will need to occur. From our perspective, we ask whether the event is plausible given our current understanding.
- Timely: It occurs early enough to be reacted to. Conditional on it occurring, will it precede the closing of the window for an effective preventive response?
- Visceral: The event makes an impression. It is noticed. Much of the discourse on warning shots focuses on the power of visceral harms to motivate a response, but this is not sufficient on its own, hence the rest of the framework.
- Sudden: The event happens or the visceral impacts occur over a short timeframe, so that individuals and societies experience “a moment”. The alternative dynamic is referred to as “frog boiling”, where individuals and society fail to notice the event or its implications because they gradually get accustomed to a new normal. Society might wake up without an event being sudden, but it seems much harder.
- Unexpected: To change an observer’s beliefs, an observation needs to not have been totally predicted by that observer. That is, our audience has to notice that something happened which they didn’t expect to prompt reexamination of their current beliefs. It will likely need to be substantially outside of what they previously considered plausible to not be written off as a (un)lucky or one-off occurrence.
- Attribution: It is perceived as being caused by AI; difficult to deny that AI progress is materially involved. It is not critical that AI was actually responsible.
- Backlash: It motivates hostility toward continued AI development rather than enthusiasm. This is required in order to motivate restraints on AI development which actually make a difference in averting existential risk. Backlash can be undermined when the warning shot indicates that further investment in, or deployment of AI is warranted. This might happen when response is channeled into societal resilience rather than prevention.
- Response Availability: Is a well-formed, legitimate policy response already available with well positioned advocates? Have policies of first resort been tried and shown to be ineffective? Crises can open windows for new policy, but these are often filled with policies which are ready and perceived to be easy. This is one area where the AI governance field can have a large impact, and where there is much work that has already been done. We exclude this from our analysis below, as it is largely agnostic to the properties of individual warning shots, but note it here, as in practice a warning shot being converted into policy may depend on the availability of a response.
For each class of warning shot below, we give a coarse grade of pass, uncertain, or fail to each criterion.
Which warning shots are expected to be strong?AI-Caused PandemicsCOVID-19 caused between 15 million and 35 million extra deaths worldwide. It is also an example of tremendous harm not leading to an appropriate preventive response: society could purchase effective resilience against a future pandemic through any number of methods (stockpiling PPE, vaccine distribution pre-logistics, wastewater monitoring, hardening buildings against airborne spread through UV treatment and/or filtering), but has done surprisingly little of this2.
How likely is AI to facilitate a human-caused pandemic? Cutting edge AI systems have already demonstrated the ability to expand the number of actors who could perform dangerous pathogen development (source, source). Safeguarding frontier models against this kind of misuse is the object of substantial effort, but these safeguards are not perfect. Furthermore, impossible-to-safeguard open source model capabilities continue to advance. So we seem to already be taking risks that even present systems will be misused to this end.
A severe pandemic could close the window for effective governance by severely burdening society with managing its harms, while at the same time not so severely impacting the pace of frontier AI development. This means that while under “normal” circumstances society could react in a timely fashion, it is quite plausible that this type of warning shot could quickly move us past a point of no return.
A pandemic causing substantial harm is very often top-of-mind when considering warning shots because of its obvious viscerality, suddenness, and unexpectedness.
However, would an AI-caused pandemic be attributable to AI? This is the first major problem for a pandemic as a warning shot. The question of COVID-19’s origin remains unresolved, and this is not an uncommon occurrence for pandemics: the 1918 Spanish Flu’s geographic origin is still unresolved.
Even putting aside the baseline difficulty of determining the origin of a pandemic, we then have to turn to determining whether a pandemic was dependent on AI assistance in a meaningful and salient way. Standard epidemiological techniques are even less relevant to this question. It is not certain that an AI-caused pathogen would exhibit any indication that it was designed with the help of AI vs. occurring naturally.
One hope is that if it is a hosted model which is misused to create a pandemic, the model provider could examine its own records to find the queries that led to a novel pathogen. Some issues with this include:
- Not all providers will retain this data. Even looking at the current data retention policies of Anthropic and OpenAI, such retention would likely require the conversation being classified as dangerous. But if these safeguards are insufficient to prevent misuse they may also be insufficient to classify a conversation as problematic in the first place.
- The provider would have to cooperate with the authorities, but they may prefer to conceal the misuse.
- The conversation would have to be associated with the pandemic in question. This could be frustrated by using multiple accounts and models to this end, fragmenting the work across multiple conversation streams.
- A competent malicious actor will prefer to avoid hosted models which it expects are sufficiently monitored for misuse.
A final problem with pandemics is a lack of backlash. The benefits of advanced AI in the biomedical domain have long been touted, and a new pandemic will likely feature AI assistance as part of its crisis response. This all leads to a pandemic being a clear excuse to accelerate the development and adoption of AI systems in pandemic preparedness.
Additionally, the response to a pandemic, especially in light of the expected difficulty of attribution, is to instead properly pursue pandemic preparedness. Methods for employing AI to this end, such as in AI-aided monitoring of genetic synthesis, further undermines the potential for backlash against AI capabilities advancement.
Military Application: Command and ControlThis wake-up call asks what happens when military leaders see firsthand what capable machine intelligence can do for their conduct of war. While it is a trope of science fiction to integrate advanced artificial intelligence systems into military command and control, we are in fact already living in a time where this is largely the case and the nature of the developments is at risk of being overlooked and misunderstood.
In the US military, the Maven Smart System began development in 2017, and has been in use for several years. It integrates data collection with modern machine learning techniques including vision and language models. Its use has grown to include an estimated 80,000 US personnel. Integration of LLMs produces a fivefold increase in targeting speed, beyond the tenfold increase that vision models already provided. (link) It also expands the quantity of sensors which can be continuously employed: previously there was simply too much incoming footage for humans to review.
Viscerality and Suddenness
We cannot know the full employment of AI in warfighting outside of a classified setting. This frustrates the viscerality of this warning shot and/or restricts the audience. Given the long history of the employment of AI (and computing generally) by the military, which includes many stages of development and an incremental increase and refinement in capabilities, along with a gradual rollout to wider usage, we should expect that the full impact of this warning shot has likely been diluted by frog-boiling.
Backlash
Backlash against AI capabilities progress seems to be quite a stretch for this warning shot. Our own military is experiencing a tremendous increase in utility without a clear downside. Even if a rival were to match US capabilities, it would increase the desire to lead in this area, rather than to constrain development.
Furthermore, initial skepticism or even hostility by some operators toward AI employment reportedly erodes with exposure and the speed and throughput demands of modern operations.
Military/Police Application: RoboticsWhat would happen if we witness a realization of some of the darkest Hollywood fears related to AI: armed machines visibly policing, fighting, or hunting human beings? Consider this progression:
- Aerial quadcopter drones hunt humans, as in Ukraine in from 2024
- Unmanned ground vehicles (UGVs) in the form of ground drones or quadruped (“dog”) platforms support or lead infantry operations
- Ukraine has conducted robot-only assaults in which Russian soldiers surrendered to machines, and in April 2026 captured a position using unmanned platforms alone
- Drones are used in police applications to do jobs which require the threat or implementation of violence, like security or riot control, going beyond surveillance or tracking
- In 2018 Israeli security forces pioneered drone-delivered tear gas, and in 2024 Indian police dropped over a thousand tear gas shells from drones on protesting farmers
Could this be followed by…
- weaponized robots employed at scale in police or riot control work, in the United States or another wealthy democracy?
- a highly publicized military campaign against a civilian population to which we’ve not already acclimated, like a Russian incursion against Poland or a Chinese invasion of Taiwan?
That the first three steps of this progression have already occurred might give us some confidence that these next steps are likely and timely. Both Ukraine and Russia are mass-producing ground combat robots; Ukraine's defense ministry reported nearly 24,500 UGV missions in the first quarter of 2026, and China is exporting armed quadrupeds and the operational concepts for using them.
Footage of an armed robot confronting a human being is undeniably visceral. Decades of Hollywood priming amplifies this: audiences expect killer robots in fiction and could register their appearance on real streets as a threshold crossed.
There is a concern for suddenness, however. Each step of the progression so far has occurred without producing a global moment: drone-delivered tear gas in 2018 and 2024 generated regional outrage and no lasting response, and robotic assaults in Ukraine are covered as defense-industry news. The moment presumably arrives when the imagery is domestic–machines suppressing a crowd of one's fellow citizens–but it must overcome this gradual normalization of the use of robotics for these purposes.
Will such applications, especially for the policing of civilian populations, be attributed to, or even associated with, advanced AI development? There are a couple of ways that frontier AI development could drive or unlock rapid progress in robotics for police applications:
- Sim-to-real methods for solving complex problems in robotics
- AI-assisted design and programming of drones
- AI-operated drones, either locally or remotely
How likely are these to be true of a given warning shot? It might not matter if the perception of this warning shots leads to pushback against AI broadly as a result. The greater risk is that attribution lands on the wrong target: on the platforms, their manufacturers, and the police departments deploying them, rather than on the AI development pipeline upstream.
Military employment of robotics could also frustrate backlash. Ukraine frames robotic infantry as the technology that lets machines die instead of soldiers, with the General Staff crediting robotic platforms for reducing personnel casualties by up to 30 percent. A public that sees machines taking casualties instead of soldiers will call for its own military to adopt the technology as rapidly as possible.
Military Applications: Software and Drone ControlAI-enabled military technology might wake up a national security audience. This audience may see the security implications, even if there aren't widespread societal effects.
Given current trends, AIs will probably become superhuman at software tasks before they become superhuman at hardware design tasks. The earliest potential warning shot related to military technology will probably be software developments.
One candidate development is the design of much better drone software. Such software may be able to radically improve the capabilities of drones, even with the same hardware. These drones will be limited to their existing on-board computers and so will not be running the advanced AI systems themselves. Instead, the AI systems will write specialized on-board software for the drones. Drones will possibly also send and receive communications from these advanced AI systems.
Demonstrations of these upgraded drone capabilities must also be sufficiently visceral. For example:
- In autonomous drone dogfights, one upgraded drone may be able to beat 10 drones running the previous state of the art software.
- An upgraded drone may be able to trounce the best human drone pilots, even when extremely outnumbered.
- Upgraded drones may be able to evade the best drone defense.
There may also be real-world uses of these systems, such as drones used for targeted assassinations, even in locations which were supposedly defended from drones; or taking out large, well-defended military targets.
For military technology via software to act as a warning shot, the main concerns are whether the relevant audience wakes up and what the response would be.
These developments might be seen as the normal progression of military technology. However, this is less likely if the audience perceives this as a step change in military capabilities, for example, if this new technology is able to easily beat previous generations. This AI-enabled military technology may also be deployed in real combat, which could make this more visceral, for example, if a relatively small number of AI-upgraded units was able to take out a much larger opposing force.
Potentially the largest issue is whether this would generate backlash. AI might be seen as an amazing military opportunity which could prompt further investment. Backlash, if it occurs at all, might be limited to governments which lag the cutting-edge of AI, and only for those governments which have sufficient intelligence to understand how far behind they are.
Cyber Misuse: Critical InfrastructureThis warning shot consists of an event in which critical infrastructure is compromised through a cyberattack materially leveraging AI capabilities, producing visceral disruption to a service society depends on.
Likelihood and Timeliness
Critical infrastructure is a large category, especially when using a definition encompassing any infrastructure necessary for the continued functioning of society and the economy. This includes at least housing, heating, food production and distribution, water and power supply, transportation, police and military operation, communications, and financial services. This provides a large attack surface, and makes preemptive defense a huge task.
Cyberattacks on critical infrastructure, without AI, have already happened repeatedly:
- In 2015, power service to about 230,000 Ukrainians was disrupted for a few hours as a result of a cyberattack which compromised information systems of three energy distribution companies.
- In 2016, a substation outside of Kyiv was briefly brought offline, though the attack intended to cause longer-lasting damage.
- In 2021, Colonial Pipeline, which provides gasoline and jet fuel to the Southeastern United States, was the target of a ransomware attack, in which billing systems were compromised and held at ransom. This led the operator to suspend pipeline operations, causing fuel shortages across the region. A ransom of $4.4 million USD was paid to the attackers.
These attacks predate the availability of capable AI agents. Regarding AI, we also already have examples of cyber misuse:
- In 2025, Anthropic (partially) disrupted a cyber operation that leveraged Claude Code, in which an AI agent autonomously executed the majority of the operation. They write “This campaign demonstrates that the barriers to performing sophisticated cyberattacks have dropped substantially—and we can predict that they’ll continue to do so.”
- In 2025-2026, a breach of nine Mexican government agencies showed clear evidence of AI coding agents being used to exploit vulnerabilities, leading to the compromise of large amounts of private data.
These are not examples of mere AI-assistance in offensive cyber operations, but rather examples of AI-driven autonomous operation. This fact should lower our estimation of how much human and organizational expertise and resources will be required to execute effective attacks as these capabilities advance.
There are some factors which cut against the likelihood of this warning shot.
Perhaps efforts to provide early access to powerful AI systems to defenders of critical infrastructure will enable them to secure their systems against attacks aided by an equivalent level of AI capability, by finding and patching vulnerabilities before they can be detected and exploited. Whether offense or defense will win in general, and in the area of critical infrastructure, is unclear, but there are at least some reasons to favor defenders. (Lohn 2025)
Will there be sufficient financial or political motivation for an attack severe enough to register? The Colonial Pipeline case suggests the financial rewards of holding critical infrastructure for ransom exist, but may be unreliable and draw the attention of authorities. Perhaps a more likely motivation is political or strategic. The number of politically motivated actors capable of causing real damage is increasing as autonomous offensive capabilities spread.
Viscerality
This category of warning shot gets a relatively easy pass on viscerality, because it is defined as events which interrupt critical services. Colonial Pipeline is arguably the most felt cyber event in U.S. history — gas lines, panic buying, fuel-price spikes, and direct White House involvement — and all this despite causing no physical damage and never touching the pipeline's control systems.
On the other hand, the Ukraine grid attacks, though more technically sophisticated and genuinely destructive, only interrupted the electrical supply for hours and made little lasting impression on audiences outside the security community. This suggests that visible disruption to everyday life drives viscerality more than technical sophistication or physical damage does.
Suddenness
Given the history of cyberattacks (including those with substantial use of AI), there is a real risk that audiences will be frog-boiled on cyber misuse. Warnings could be followed by increasing frequency and severity of attacks over time without creating a clear shared moment to motivate a response.
Unexpectedness
Relatedly, very few observers will be surprised by increasing ill effects of proliferated cyber capabilities. Events of this sort have historical precedent, are well represented in fiction and film, and have been warned about for years. As a result, what will be necessary for this to be unexpected is at least one of (a) the harm caused by the attack is especially severe or widespread or (b) the AI assistance was of an unexpected level of utility or power.
In July of 2026, water utilities in Minnesota were disrupted by a cyberattack, and while AI’s contributions have not yet been confirmed, it is expected, given the availability of AI agents, that more attacks of this sort will occur in the future.
Attribution
This warning shot anticipates the combination of AI assistance in the means of an attack with critical infrastructure as the target, leading to visceral impacts. Because AI capabilities relevant to cyber operations have already proliferated, it is essentially assured that AI assistance will be used in the next attack on critical infrastructure.
AI employment, especially in the form of autonomous systems, could be relatively easy to detect, as in the example of the Mexican government data breach, where the speed at which code changes were authored is a clear sign that AI assistance was used. We may also be so lucky as in Anthropic’s detection of the November 2025 campaign, in which attribution came directly from the model provider.
Backlash
“if AI models can be misused for cyberattacks at this scale, why continue to develop and release them? The answer is that the very abilities that allow Claude to be used in these attacks also make it crucial for cyber defense.” - Anthropic
The most likely outcome is that any potential backlash against AI development is redirected toward employment of AI for defensive purposes. This argument may flounder in the face of sufficient and visceral harms. At some point, society will not tolerate continued disruption and begin to suspect that AI-aided cyber defense is not viable to counteract the effects of AI-aided cyber offense.
Another possibility is that backlash against AI could be swamped by circumstances surrounding its employment in an attack. For example, if a US adversary uses AI for offensive cyber operations which target critical infrastructure, it would probably be in the context of some larger geopolitical crisis. There will be little energy to think about AI development and its consequences while we are faced with a clear external threat of a rival who is disrupting critical infrastructure as part of a larger crisis or conflict, and that external rival will absorb the backlash rather than the enabling technology.
Conclusion and Future ResearchWe hope this research can prompt some needed skepticism for the position that it is prudent to adopt a “wait and see” attitude toward the threat of AI risks.
We are excited to see the framework applied to other warning shots. These include: economic impacts of AI (including job loss), covert or kinetic conflict between great powers over AI, and the emergence of fully autonomous rogue AI systems (either intentionally created or self-exfiltrated).
Finally, it is possible that, rather than a singular warning shot, there will be many steps on the road to effective governance of the development of artificial superintelligence. These steps could include warning shots, shifts in the discourse, and the iterative implementation of imperfect policy. Is it plausible that we have time for this potentially lengthy process before it is too late? Especially regarding imperfect policy, what could distinguish steps which move us closer to effective governance rather than backtrack away from it? We are excited about future work which deeply engages with this question.
Appendix: Select Warning ShotsWe’re interested in many other warning shots, and a partial list is included here:
- Bio misuse foiled – An incidence or pattern of biological misuse of AI capabilities being detected and foiled, perhaps reported by AI providers or by intelligence agencies.
- Bio uplift demonstration – A visceral demonstration of how AI tools dramatically increase the number of bio-risk relevant actors
- Stock market crash – Some large downward movement of the stock market that is caused by or closely associated with AI.
- Significant job loss – Majority un- or under-employment, especially for young people / college graduates.
- Transformative job loss – What historically happened to weavers with the advent of mechanical looms but for all (cognitive) human labor.
- Runaway inflation – What is happening for computer RAM affects other key inputs, particularly energy but perhaps others, ultimately massively increasing the prices of electricity, food, water, land, etc
- Whistleblowing by frontier lab employees – A number of researchers whistleblow on the risks of continued development.
- Recursive self-improvement – Humans are no longer directly involved in the research required to advance AI capabilities.
- Novel scientific contributions attributed to AI – “A flood of groundbreaking arXiv papers” as described here.
- Anticipated but alarming demos or applications – For example, mosquito-sized autonomous weapons or effective superpersuasion
- Misalignment: Refusals – Military systems refusing to fight or automated AI researchers refusing to work
- Rogue AI: Near Miss – Catching an AI system in the act of self-exfiltration of weights.
- Rogue AI: Escaped – AI systems are stably independent of human control, with independent access to resources (money, compute).
- Transformative AI-powered surveillance – Domestic applications of AI leveraging currently-proliferating methods of data collection enable powerful surveillance applications, and this generates backlash. Could also include open-sourcing of such tools enabling widespread access and misuse.
- AI ruins encryption – AI finds exploitable weaknesses in widely-deployed cryptography (or accelerates cryptanalysis enough to matter), threatening secure comms, financial systems, and stored secrets.
- Undermining Deterrence – AI is used to undermine the existing condition of stable deterrence against nuclear first use, such as by enabling some of: orbital detection of submarines, missile interception, tracking of mobile launchers, perfect cyber offense against nuclear command and control, etc.
Discuss
AISafety.com Hackathon 2026
When
Thursday 17th September – Sunday 20th September 2026
Where
CEEALAR (the EA Hotel) in Blackpool, England
Cost
Free, including accommodation and all meals
Applications close
14 August 2026 (decisions by 21 August)
Once a year, the core AISafety.com team gets together with a group of amazing volunteers and spends an extended weekend getting a bunch of stuff done to improve the site. We also do some relaxed fun activities between work sessions.
Who we're looking for- Product managers and designers
- To execute the vision for new projects, build out design mockups, conduct user research, offer consulting, and brainstorm product solutions
- Project managers
- To lead and organize project efforts and run sprints
- Developers
- To work with product managers, designers, and AI tools to bring some amazing new things to life
- Connectors, people with an audience, and subject-matter experts
- To help us with targeted community outreach, general outreach, social media, and to give takes to product managers and designers
- Promotion people
- To help us spread awareness of the site and each resource page in various ways, maybe including SEO
- Anything else
- Just tell us how you can help!
On the first day of the hackathon we will divide everyone into small teams by project. Participants will get to decide which projects they work on and in what way they contribute. We're also open to new project suggestions.
Why it's worth it- Build your resume, gain career capital, and make connections (everyone at the EA hotel, beyond just our hackathon, is working on something high-impact – mostly AI safety)
- Have fun beyond just the work. We'll set aside time each day for beach walks, games, etc.
- Help save the world!
If you have any questions, please message Bryce on the AISafety.com Discord server at @bryceerobertson, or email bryceerobertson@gmail.com.
Discuss
See all upcoming AI safety events and training programs
For a long time AISafety.com has listed upcoming events and training programs as they are announced. But the information it contained was limited and its presentation not very user-friendly; it was basically a collection of Airtable embeds.
We have now: split the page into Events and Training programs, given each a proper design, begun including more data like entry bar and cost/stipend, and created the long-requested list of recurring training programs.
Evergreen list of recurring training programs:
We’ll continue to improve these pages, so if you have any feedback please let us know in the comments or via this anonymous form. Or if you're really motivated: come join our hackathon!
Hackathon to improve AISafety.comOnce a year, the core AISafety.com team gets together with a group of amazing volunteers and spends an extended weekend upgrading the site (and doing other just-for-fun activities). This year it’s at the EA Hotel in Blackpool, England, and runs 17–20 September. We’re looking for various skillsets, including product management, design, development, and outreach. Come join us! More info plus the 5-minute application form here.
Discuss
Norms Are Broken Laws
Consider the following unwritten rule:
Keep the kitchen clean. Clean after you cook.
Consider the following three characters: the unreflective individual, the conscious egoist, and the rational moralist.[1]
- The unreflective individual cleans the kitchen after cooking, almost automatically, because it is a norm that the society expects people to comply with. The only cost is the physical cost to clean the kitchen. The perceived cost is none, since that's the default thing to do regardless of others.
- The conscious egoist doesn't clean the kitchen. The conscious egoist examines the dilemma whether to (A) clean the kitchen or (B) don't clean the kitchen. The conscious egoist measures the cost of (A) as the physical cost, and the cost of (B) as none, so the conscious egoist chooses (B). The final perceived cost is none.
- The rational moralist faces the dilemma (A)/(B) as above. The rational moralist measures the cost of (A) as the physical cost, the cost of (B) as the emotional cost (i.e. I'm unhappy that I'm a bad person acting as a free-rider). If the rational moralist chooses (A), then they bear the physical cost. If the rational moralist chooses (B), then they bear the emotional cost.
I'll generalize the kitchen example to "all [unwritten but expected] norms". Let me introduce more examples here:
- Keep quiet on public transport.
- In the school cafeteria, return the trays, don't just leave them on the table.
- Be polite in communications.
- Respect others.
- Keep your promises.
If you replace the kitchen scenario with other examples, you can see that the cost analysis of three characters above remain the same, except for the fifth example, where non-compliance carries a significant cost.[2]
My ClaimThere is an asymmetry whenever the society expects its members to follow some unwritten norms. My main claim is that:
All (unwritten but expected) norms push the burden into the rational moralist, creating a perceived tax (whether physical or emotional cost). This punishes the rational moralist, and therefore such society encourages people to rather be unreflective or egoistic, compared to being rational and moral. The collective judges efficacy competently and distribution blindly[3], hence leaving things as they are, not accepting the decision to convert norms into written laws.
My proposal is therefore to minimize [unwritten but expected] norms, i.e., "abolish all norms". This doesn't directly mean actively removing all the norms, but rather, I mean: Revise the norms, turn them into a written rule that is agreed in the social contract, if verifiable, sanctionable, and if the failure of its compliance leads to a negative social outcome. For other norms (such as "saying thank you", "voluntary kindness", etc.), move them towards the etiquette category, and do not expect people to follow such norm, and also discourage people to expect others to follow such norm.
Under this revision, the unreflective individual continues to perform whatever has to be done normally, just that the unwritten norm "Keep the kitchen clean. Clean after you cook." is turned into an explicit rule:
Members of the kitchen must clean after cooking. Failure to do so will result in a fine of $100.
Now the conscious egoist and the rational moralist both face a cost: choosing (A) creates a physical cost, and choosing (B) creates a concrete $100 cost.
This new protocol distributes the old "rational moralist tax" into a universal tax: one either cleans the kitchen or pays the fee. I argue that this system is fairer than the original one.[4] Now let us abstract this example away and let me propose a full formal specification below.
The BlueprintLet mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; font-weight: normal; font-size: 100%; font-size-adjust: none; letter-spacing: normal; border-collapse: collapse; word-wrap: normal; word-spacing: normal; white-space: nowrap; direction: ltr; padding: 1px 0; } mjx-container[jax="CHTML"][display="true"] { display: block; text-align: center; margin: 1em 0; } mjx-container[jax="CHTML"][display="true"][width="full"] { display: flex; } mjx-container[jax="CHTML"][display="true"] mjx-math { padding: 0; } mjx-container[jax="CHTML"][justify="left"] { text-align: left; } mjx-container[jax="CHTML"][justify="right"] { text-align: right; } mjx-mi { display: inline-block; text-align: left; } mjx-c { display: inline-block; } mjx-utext { display: inline-block; padding: .75em 0 .2em 0; } mjx-msub { display: inline-block; text-align: left; } mjx-mo { display: inline-block; text-align: left; } mjx-stretchy-h { display: inline-table; width: 100%; } mjx-stretchy-h > * { display: table-cell; width: 0; } mjx-stretchy-h > * > mjx-c { display: inline-block; transform: scalex(1.0000001); } mjx-stretchy-h > * > mjx-c::before { display: inline-block; width: initial; } mjx-stretchy-h > mjx-ext { /* IE */ overflow: hidden; /* others */ overflow: clip visible; width: 100%; } mjx-stretchy-h > mjx-ext > mjx-c::before { transform: scalex(500); } mjx-stretchy-h > mjx-ext > mjx-c { width: 0; } mjx-stretchy-h > mjx-beg > mjx-c { margin-right: -.1em; } mjx-stretchy-h > mjx-end > mjx-c { margin-left: -.1em; } mjx-stretchy-v { display: inline-block; } mjx-stretchy-v > * { display: block; } mjx-stretchy-v > mjx-beg { height: 0; } mjx-stretchy-v > mjx-end > mjx-c { display: block; } mjx-stretchy-v > * > mjx-c { transform: scaley(1.0000001); transform-origin: left center; overflow: hidden; } mjx-stretchy-v > mjx-ext { display: block; height: 100%; box-sizing: border-box; border: 0px solid transparent; /* IE */ overflow: hidden; /* others */ overflow: visible clip; } mjx-stretchy-v > mjx-ext > mjx-c::before { width: initial; box-sizing: border-box; } mjx-stretchy-v > mjx-ext > mjx-c { transform: scaleY(500) translateY(.075em); overflow: visible; } mjx-mark { display: inline-block; height: 0px; } mjx-mn { display: inline-block; text-align: left; } mjx-c.mjx-c1D441.TEX-I::before { padding: 0.683em 0.888em 0 0; content: "N"; } mjx-c.mjx-c1D436.TEX-I::before { padding: 0.705em 0.76em 0.022em 0; content: "C"; } mjx-c.mjx-c1D443.TEX-I::before { padding: 0.683em 0.751em 0 0; content: "P"; } mjx-c.mjx-c1D440.TEX-I::before { padding: 0.683em 1.051em 0 0; content: "M"; } mjx-c.mjx-c6D::before { padding: 0.442em 0.833em 0 0; content: "m"; } mjx-c.mjx-c69::before { padding: 0.669em 0.278em 0 0; content: "i"; } mjx-c.mjx-c6E::before { padding: 0.442em 0.556em 0 0; content: "n"; } mjx-c.mjx-c28::before { padding: 0.75em 0.389em 0.25em 0; content: "("; } mjx-c.mjx-c2C::before { padding: 0.121em 0.278em 0.194em 0; content: ","; } mjx-c.mjx-c29::before { padding: 0.75em 0.389em 0.25em 0; content: ")"; } mjx-c.mjx-c3E::before { padding: 0.54em 0.778em 0.04em 0; content: ">"; } mjx-c.mjx-c30::before { padding: 0.666em 0.5em 0.022em 0; content: "0"; } mjx-c.mjx-c1D448.TEX-I::before { padding: 0.683em 0.767em 0.022em 0; content: "U"; } mjx-c.mjx-c3A::before { padding: 0.43em 0.278em 0 0; content: ":"; } mjx-c.mjx-c3D::before { padding: 0.583em 0.778em 0.082em 0; content: "="; } mjx-c.mjx-c2B::before { padding: 0.583em 0.778em 0.082em 0; content: "+"; } mjx-c.mjx-c1D700.TEX-I::before { padding: 0.452em 0.466em 0.022em 0; content: "\3B5"; } mjx-c.mjx-c3C::before { padding: 0.54em 0.778em 0.04em 0; content: "<"; } mjx-container[jax="CHTML"] { line-height: 0; } mjx-container [space="1"] { margin-left: .111em; } mjx-container [space="2"] { margin-left: .167em; } mjx-container [space="3"] { margin-left: .222em; } mjx-container [space="4"] { margin-left: .278em; } mjx-container [space="5"] { margin-left: .333em; } mjx-container [rspace="1"] { margin-right: .111em; } mjx-container [rspace="2"] { margin-right: .167em; } mjx-container [rspace="3"] { margin-right: .222em; } mjx-container [rspace="4"] { margin-right: .278em; } mjx-container [rspace="5"] { margin-right: .333em; } mjx-container [size="s"] { font-size: 70.7%; } mjx-container [size="ss"] { font-size: 50%; } mjx-container [size="Tn"] { font-size: 60%; } mjx-container [size="sm"] { font-size: 85%; } mjx-container [size="lg"] { font-size: 120%; } mjx-container [size="Lg"] { font-size: 144%; } mjx-container [size="LG"] { font-size: 173%; } mjx-container [size="hg"] { font-size: 207%; } mjx-container [size="HG"] { font-size: 249%; } mjx-container [width="full"] { width: 100%; } mjx-box { display: inline-block; } mjx-block { display: block; } mjx-itable { display: inline-table; } mjx-row { display: table-row; } mjx-row > * { display: table-cell; } mjx-mtext { display: inline-block; } mjx-mstyle { display: inline-block; } mjx-merror { display: inline-block; color: red; background-color: yellow; } mjx-mphantom { visibility: hidden; } _::-webkit-full-page-media, _:future, :root mjx-container { will-change: opacity; } mjx-c::before { display: block; width: 0; } .MJX-TEX { font-family: MJXZERO, MJXTEX; } .TEX-B { font-family: MJXZERO, MJXTEX-B; } .TEX-I { font-family: MJXZERO, MJXTEX-I; } .TEX-MI { font-family: MJXZERO, MJXTEX-MI; } .TEX-BI { font-family: MJXZERO, MJXTEX-BI; } .TEX-S1 { font-family: MJXZERO, MJXTEX-S1; } .TEX-S2 { font-family: MJXZERO, MJXTEX-S2; } .TEX-S3 { font-family: MJXZERO, MJXTEX-S3; } .TEX-S4 { font-family: MJXZERO, MJXTEX-S4; } .TEX-A { font-family: MJXZERO, MJXTEX-A; } .TEX-C { font-family: MJXZERO, MJXTEX-C; } .TEX-CB { font-family: MJXZERO, MJXTEX-CB; } .TEX-FR { font-family: MJXZERO, MJXTEX-FR; } .TEX-FRB { font-family: MJXZERO, MJXTEX-FRB; } .TEX-SS { font-family: MJXZERO, MJXTEX-SS; } .TEX-SSB { font-family: MJXZERO, MJXTEX-SSB; } .TEX-SSI { font-family: MJXZERO, MJXTEX-SSI; } .TEX-SC { font-family: MJXZERO, MJXTEX-SC; } .TEX-T { font-family: MJXZERO, MJXTEX-T; } .TEX-V { font-family: MJXZERO, MJXTEX-V; } .TEX-VB { font-family: MJXZERO, MJXTEX-VB; } mjx-stretchy-v mjx-c, mjx-stretchy-h mjx-c { font-family: MJXZERO, MJXTEX-S1, MJXTEX-S4, MJXTEX, MJXTEX-A ! important; } @font-face /* 0 */ { font-family: MJXZERO; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Zero.woff") format("woff"); } @font-face /* 1 */ { font-family: MJXTEX; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Regular.woff") format("woff"); } @font-face /* 2 */ { font-family: MJXTEX-B; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Bold.woff") format("woff"); } @font-face /* 3 */ { font-family: MJXTEX-I; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-Italic.woff") format("woff"); } @font-face /* 4 */ { font-family: MJXTEX-MI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Italic.woff") format("woff"); } @font-face /* 5 */ { font-family: MJXTEX-BI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-BoldItalic.woff") format("woff"); } @font-face /* 6 */ { font-family: MJXTEX-S1; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size1-Regular.woff") format("woff"); } @font-face /* 7 */ { font-family: MJXTEX-S2; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size2-Regular.woff") format("woff"); } @font-face /* 8 */ { font-family: MJXTEX-S3; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size3-Regular.woff") format("woff"); } @font-face /* 9 */ { font-family: MJXTEX-S4; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size4-Regular.woff") format("woff"); } @font-face /* 10 */ { font-family: MJXTEX-A; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_AMS-Regular.woff") format("woff"); } @font-face /* 11 */ { font-family: MJXTEX-C; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Regular.woff") format("woff"); } @font-face /* 12 */ { font-family: MJXTEX-CB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Bold.woff") format("woff"); } @font-face /* 13 */ { font-family: MJXTEX-FR; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Regular.woff") format("woff"); } @font-face /* 14 */ { font-family: MJXTEX-FRB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Bold.woff") format("woff"); } @font-face /* 15 */ { font-family: MJXTEX-SS; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Regular.woff") format("woff"); } @font-face /* 16 */ { font-family: MJXTEX-SSB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Bold.woff") format("woff"); } @font-face /* 17 */ { font-family: MJXTEX-SSI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Italic.woff") format("woff"); } @font-face /* 18 */ { font-family: MJXTEX-SC; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Script-Regular.woff") format("woff"); } @font-face /* 19 */ { font-family: MJXTEX-T; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Typewriter-Regular.woff") format("woff"); } @font-face /* 20 */ { font-family: MJXTEX-V; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Regular.woff") format("woff"); } @font-face /* 21 */ { font-family: MJXTEX-VB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Bold.woff") format("woff"); } be an [unwritten but expected] norm. Let be the perceived cost of complying with the norm. Let be the perceived cost of not complying with the norm. In the original situation, the unreflective individual pays zero cost. The conscious egoist pays zero cost. The rational moralist pays 0" role="math">.
Legalizing blueprint: Choose an appropriate[5] cost to apply universally. Impose a new rule:
Every member of this society must either obey or pay .
Now, everyone pays (although the unreflective individual might interpret this as zero since obeying is the default choice which was free from the beginning), assuming that the moralist perceives "not complying to the norm" as "paying the fee and retaining no extra moral obligation".
Law, Norm, and EtiquetteThere are at least four necessary conditions for something to be a law:
- It must be an enforcement, not a description, not an advice.
- The enforcement must apply to humans, not animals, not non-living things.
- The enforcement must enforce actions, not thoughts, and such actions must be verifiable.
- It has a sanction: there is a penalty for failure of compliance.
- It must be enforced to all members of the society and must be agreed in the social contract.
On the other hand, norms are unwritten rules, but the society expects everyone to comply. Etiquette is a weaker version of norms: the society doesn't really expect everyone to comply, but complying would signal a positive social value. Laws, norms, and etiquette form a basic organization instrument for a society. This has happened historically, and can be seen as "rules that shape the social coordination".
As I have argued above, my belief is that something that is not enforced but is expected always creates an asymmetric cost to some particularly conscious agents (in this example, the rational moralist). That is,
A norm is merely a law enforced only against the conscientious.
A norm's enforcement mechanism is conscience, and conscience only reaches agents who have one. Therefore, a norm satisfies the above conditions 1-4 while failing the 5th one. Society tries to enforce, but not against everyone.
I accept that norms and etiquette are important ingredients of the current civilization, but I also argue that this is not the best way. I believe that the blueprint above would turn enforceable norms into laws and either discard the unenforceable ones or move some towards etiquette. One might argue that a historical reason of norms is that it is a cheaper alternative to laws, since a formal prosecution of a routine behavior might push everything into sluggish bureaucracy, and that's not worth the price. My claim is that I believe it's worth the price, not in the sense that everyone will be better off, but in the sense that I think this move would fix the fairness problem while not reducing the societal benefit by much.
Therefore,
Norms Are Broken Laws.
Those worth enforcing should be enforced on everyone. Those not worth enforcing should not be expected of anyone.
"Expecting people to comply" creates an asymmetric tax on the rational moralist, and therefore conversion to written laws with active enforcement would be a fairer alternative.[6]
- ^
I'm not claiming that everyone belongs to one among the three characters. There may be others, but I'm not considering the other types here, for the reason of simplification. I argue that it is enough to consider only these.
- ^
Although they form different conclusions. For instance, if one were to run "respect others" through the blueprint, one sees that the corresponding law must at least distinguish between having respect and not having respect, which is not actually verifiable. On "keep your promises", the current unwritten sanction already did the job well, so no law codification is needed. If the egoist breaks a promise, then the society knows that the egoist breaks, removing reputation from the egoist, and this is already a cost. Therefore this analysis doesn't match the template as the other four. The other three examples have a concrete objective witness, hence verifiable. They are also enforceable by law. "Be polite in communications" is weak, since it raises the question "from objective evidence, what counts as polite?" My response would be to break down "politeness" into verifiable pieces: "Do not use slurs", "Do not make threats", etc.
- ^
The collective can observe whether the kitchen gets cleaned. It cannot observe who paid for it.
- ^
Some people argue that this decision turns the punishment structure into economic gamification, and the result may be worse (i.e. no one cleans the kitchen while everyone pays the fee). My response would be "then increase the fee so that the marginal cost/benefit of cleaning the kitchen becomes better than paying the fee".
- ^
Ideally, pick , for some small (procedural) 0" role="math">, so that the choice to comply is a bit better than not complying. Note that this design choice might unintentionally increase the cost to the rational moralist (when , for instance), so the blueprint may not work in favor of the rational moralist but rather redistribute the moral cost to the egoist. Solving the fairness problem doesn't imply the benefit/cost problem being solved.
- ^
Note that I'm claiming very weakly about the public benefit. I'm claiming only that this blueprint doesn't create much public cost compared to the public benefit. This is a belief, and one can convince me otherwise. Perhaps in a more careful analysis, one might consider the public benefit arising from one compliance of the norm and the public cost arising from one non-compliance of the norm, if one really wants to argue (in a utilitarian way) about a public protocol. This analysis mainly considers the fairness part, and is trying to do justice to the rational moralist.
Discuss
Various Ways to Enable People
There seems to be one type of work that truly matter: Work that enables people. Most work can be fitted into this category, such as teaching, maintenance (of things that enable), parenting, and reducing child mortality, poverty, x-risk, etc.
Of course, there are some forms of work that enable people much more than others, but here I'll go over a few examples of what it means to "enable people":
- Teaching- by communicating knowledge on how to do something, you are enabling people to do that thing in the most literal sense.
- Maintenance- many institutions and such things already enable people massively. being one of the people who help maintain these keep people enabled.
- Parenting- the most personal (and allegedly fulfilling) way to enable a person. You're taking them from near-zero all the way to a human who can live, experience, and even go on to enable others as well.
- Reducing child mortality- has similar benefits to parenting, minus the personal touch that makes it fulfilling, but with exponentially more impact.
- Reducing poverty- not being stuck in the poverty cycle and not living to survive enables you to do the tings you want, and go on to enable more people.
- Reducing x-risk- it's probably better if we're not all dead, right?!
- Operations- this job is explicitly doing things that help enable people to do the work they're doing.
- Research- you can contribute to finding ways to help enable people in new or more efficient ways.
I assume most of the audience here is already settled in a career however, so here are some ways you can help enable people currently:
- Being kind- the world is already hard enough as it is, you don't need to contribute to that. Depression and suicide are issues enough in them selves, but also double as disablers. Gasp!
- Financially supporting a partner/dependent- Personally I credit my parents for doing this for me throughout my high school experience instead of throwing me straight into a job. If I hadn't had ample free time to consider my life and take risks through those years, I would have never stumbled upon LessWrong, and likely would have never committed myself to attempting to be more effective in enabling people throughout my career.
- Earning to give- of course this is a staple here for a reason. If you are able to save or significantly improve a life, it enables that person.
- Being encouraging- social fear is a big factor on why people don't do things. If you can support people in whatever they may be doing, it would significantly boost their chances of getting that thing done.
- Communicating ideas- write on LessWrong! Talk to people! Even if the ideas you share might not be enabling in themselves, perhaps they can indirectly enable others!
This is just a smorgasbord of ideas for those who haven't considered them, but definitely refer to the ITN framework if you're trying to weigh these ideas against each other.
If you have any additional ways to enable people that I may have not listed, I would love to hear them :)
Discuss
The AI Race is Not a Prisoner's Dilemma
Statement on AI use: AI models (mostly Fable 5 and Opus 5) were extremely helpful in 1. iterating through lots of variations on the game theory models presented, 2. helping to confirm my understanding of the math, 3. fact checking, finding sources, and catching errors, and 4. providing editorial feedback. The game theory visualizations are entirely vibecoded; I have checked the math, but not the code. All prose is my own except where marked and I take full responsibility for everything in the post.
I. IntroductionWhen I teach my philosophy students about existential risk from artificial superintelligence, many students' first reaction after fully grasping the argument is "Why are we still building this?" And the general public arguably shares the sentiment: an Opus 5 agent found that
63% of Americans told Pew in February 2026 that AI is advancing too quickly (against 2% who said too slowly), and a Yahoo News/YouGov poll the previous October found that 53% think it at least somewhat likely that AI will 'destroy humanity' someday.
Why not just shut it all down until we know it's safe? The Discourse has an answer to this question: it's because AGI is an arms race. Here is Leopold Aschenbrenner in 2024:
Some hope for some sort of international treaty on safety. This seems fanciful to me. The world where both the CCP and USG are AGI-pilled enough to take safety risk seriously is also the world in which both realize that international economic and military predominance is at stake, that being months behind on AGI could mean being permanently left behind. If the race is tight, any arms control equilibrium, at least in the early phase around superintelligence, seems extremely unstable. In short, "breakout" is too easy: the incentive (and the fear that others will act on this incentive) to race ahead with an intelligence explosion, to reach superintelligence and the decisive advantage, too great. At the very least, the odds we get something good-enough here seem slim. (How have those climate treaties gone? That seems like a dramatically easier problem compared to this.) (Source)
And here is Hunter Ash earlier this year:
AGI is not analogous to nukes
I should start by saying I'm about as far from an AI safetyist as it's possible to be. But I'm going to engage with that frame, and show why the situation is in no way analogous to that of nuclear weapons, which is their only real example of the kind of international coordination that would be required to stop AI progress.
The basic issue is this: you can build nuclear weapons but not use them. And this is what mutually assured destruction hangs on. On the safetyist view, you cannot do this with AGI. Making it and "firing" it are the same action.
The game with nukes is this: if we shoot nukes at another nuclear-armed country, there is an extremely high chance that every major population center in our country gets wiped out by their nukes. And the same is true for them. So no one fires. No one defects.
With AGI, the only way to not defect is to not be able to return fire if the other side defects. Unless your P(doom) is well north of 50%, then the deterrence math doesn't work. If the other team builds it in secret then turns it on, and they did achieve decent alignment, then you're absolutely screwed. There's no mutually assured destruction. They win and you lose. You can't suddenly build your own in retaliation when the missiles are already in the air, so to speak.
I also think compliance with any treaty would be near-impossible to enforce, but that's a subject for another post.
We cannot get out of this arms race even if we wanted to. Winner takes all, and there is no pause button. (Source)
So the argument is basically that if we don't build it, then they will, where "they" could be DeepMind or Anthropic or OpenAI or China—especially China. And if they build it, that would be worse. We'd like to stop, you see, but we can't. That is to say, it has the structure of a Prisoner's Dilemma. A Prisoner's Dilemma is defined by the fact that Defect is the dominant strategy, even though both players prefer (Cooperate, Cooperate) to (Defect, Defect). The argument is that, even though both the US and China (or both Anthropic and OpenAI) would prefer that they both not build, they also both prefer that they themselves build, regardless of what the other does.
The problem is that the AI race is not a Prisoner's Dilemma. I'm not the first to point this out—Katja Grace has talked about it in several posts over the years, and there are a couple of papers making a similar point.[1](And Zvi actually made the point in one of his newsletters while I was in the process of writing this post.) But since the Discourse at large still seems to mostly take the "AI is a prisoner's dilemma" meme for granted, I want to go through in detail why it is not.
II. Nuclear Arms RacesA classic "arms race" (e.g. nuclear weapons) is a Prisoner's Dilemma. A Prisoner's Dilemma is defined by four payoffs, which we can call R, S, T, and P:
mjx-container[jax="CHTML"] { line-height: 0; } mjx-container [space="1"] { margin-left: .111em; } mjx-container [space="2"] { margin-left: .167em; } mjx-container [space="3"] { margin-left: .222em; } mjx-container [space="4"] { margin-left: .278em; } mjx-container [space="5"] { margin-left: .333em; } mjx-container [rspace="1"] { margin-right: .111em; } mjx-container [rspace="2"] { margin-right: .167em; } mjx-container [rspace="3"] { margin-right: .222em; } mjx-container [rspace="4"] { margin-right: .278em; } mjx-container [rspace="5"] { margin-right: .333em; } mjx-container [size="s"] { font-size: 70.7%; } mjx-container [size="ss"] { font-size: 50%; } mjx-container [size="Tn"] { font-size: 60%; } mjx-container [size="sm"] { font-size: 85%; } mjx-container [size="lg"] { font-size: 120%; } mjx-container [size="Lg"] { font-size: 144%; } mjx-container [size="LG"] { font-size: 173%; } mjx-container [size="hg"] { font-size: 207%; } mjx-container [size="HG"] { font-size: 249%; } mjx-container [width="full"] { width: 100%; } mjx-box { display: inline-block; } mjx-block { display: block; } mjx-itable { display: inline-table; } mjx-row { display: table-row; } mjx-row > * { display: table-cell; } mjx-mtext { display: inline-block; text-align: left; } mjx-mstyle { display: inline-block; } mjx-merror { display: inline-block; color: red; background-color: yellow; } mjx-mphantom { visibility: hidden; } _::-webkit-full-page-media, _:future, :root mjx-container { will-change: opacity; } mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; font-weight: normal; font-size: 100%; font-size-adjust: none; letter-spacing: normal; border-collapse: collapse; word-wrap: normal; word-spacing: normal; white-space: nowrap; direction: ltr; padding: 1px 0; } mjx-container[jax="CHTML"][display="true"] { display: block; text-align: center; margin: 1em 0; } mjx-container[jax="CHTML"][display="true"][width="full"] { display: flex; } mjx-container[jax="CHTML"][display="true"] mjx-math { padding: 0; } mjx-container[jax="CHTML"][justify="left"] { text-align: left; } mjx-container[jax="CHTML"][justify="right"] { text-align: right; } mjx-mtable { display: inline-block; text-align: center; vertical-align: .25em; position: relative; box-sizing: border-box; border-spacing: 0; border-collapse: collapse; } mjx-mstyle[size="s"] mjx-mtable { vertical-align: .354em; } mjx-labels { position: absolute; left: 0; top: 0; } mjx-table { display: inline-block; vertical-align: -.5ex; box-sizing: border-box; } mjx-table > mjx-itable { vertical-align: middle; text-align: left; box-sizing: border-box; } mjx-labels > mjx-itable { position: absolute; top: 0; } mjx-mtable[justify="left"] { text-align: left; } mjx-mtable[justify="right"] { text-align: right; } mjx-mtable[justify="left"][side="left"] { padding-right: 0 ! important; } mjx-mtable[justify="left"][side="right"] { padding-left: 0 ! important; } mjx-mtable[justify="right"][side="left"] { padding-right: 0 ! important; } mjx-mtable[justify="right"][side="right"] { padding-left: 0 ! important; } mjx-mtable[align] { vertical-align: baseline; } mjx-mtable[align="top"] > mjx-table { vertical-align: top; } mjx-mtable[align="bottom"] > mjx-table { vertical-align: bottom; } mjx-mtable[side="right"] mjx-labels { min-width: 100%; } mjx-mtr { display: table-row; text-align: left; } mjx-mtr[rowalign="top"] > mjx-mtd { vertical-align: top; } mjx-mtr[rowalign="center"] > mjx-mtd { vertical-align: middle; } mjx-mtr[rowalign="bottom"] > mjx-mtd { vertical-align: bottom; } mjx-mtr[rowalign="baseline"] > mjx-mtd { vertical-align: baseline; } mjx-mtr[rowalign="axis"] > mjx-mtd { vertical-align: .25em; } mjx-mtd { display: table-cell; text-align: center; padding: .215em .4em; } mjx-mtd:first-child { padding-left: 0; } mjx-mtd:last-child { padding-right: 0; } mjx-mtable > * > mjx-itable > *:first-child > mjx-mtd { padding-top: 0; } mjx-mtable > * > mjx-itable > *:last-child > mjx-mtd { padding-bottom: 0; } mjx-tstrut { display: inline-block; height: 1em; vertical-align: -.25em; } mjx-labels[align="left"] > mjx-mtr > mjx-mtd { text-align: left; } mjx-labels[align="right"] > mjx-mtr > mjx-mtd { text-align: right; } mjx-mtd[extra] { padding: 0; } mjx-mtd[rowalign="top"] { vertical-align: top; } mjx-mtd[rowalign="center"] { vertical-align: middle; } mjx-mtd[rowalign="bottom"] { vertical-align: bottom; } mjx-mtd[rowalign="baseline"] { vertical-align: baseline; } mjx-mtd[rowalign="axis"] { vertical-align: .25em; } mjx-c { display: inline-block; } mjx-utext { display: inline-block; padding: .75em 0 .2em 0; } mjx-mo { display: inline-block; text-align: left; } mjx-stretchy-h { display: inline-table; width: 100%; } mjx-stretchy-h > * { display: table-cell; width: 0; } mjx-stretchy-h > * > mjx-c { display: inline-block; transform: scalex(1.0000001); } mjx-stretchy-h > * > mjx-c::before { display: inline-block; width: initial; } mjx-stretchy-h > mjx-ext { /* IE */ overflow: hidden; /* others */ overflow: clip visible; width: 100%; } mjx-stretchy-h > mjx-ext > mjx-c::before { transform: scalex(500); } mjx-stretchy-h > mjx-ext > mjx-c { width: 0; } mjx-stretchy-h > mjx-beg > mjx-c { margin-right: -.1em; } mjx-stretchy-h > mjx-end > mjx-c { margin-left: -.1em; } mjx-stretchy-v { display: inline-block; } mjx-stretchy-v > * { display: block; } mjx-stretchy-v > mjx-beg { height: 0; } mjx-stretchy-v > mjx-end > mjx-c { display: block; } mjx-stretchy-v > * > mjx-c { transform: scaley(1.0000001); transform-origin: left center; overflow: hidden; } mjx-stretchy-v > mjx-ext { display: block; height: 100%; box-sizing: border-box; border: 0px solid transparent; /* IE */ overflow: hidden; /* others */ overflow: visible clip; } mjx-stretchy-v > mjx-ext > mjx-c::before { width: initial; box-sizing: border-box; } mjx-stretchy-v > mjx-ext > mjx-c { transform: scaleY(500) translateY(.075em); overflow: visible; } mjx-mark { display: inline-block; height: 0px; } mjx-mi { display: inline-block; text-align: left; } mjx-mn { display: inline-block; text-align: left; } mjx-mfrac { display: inline-block; text-align: left; } mjx-frac { display: inline-block; vertical-align: 0.17em; padding: 0 .22em; } mjx-frac[type="d"] { vertical-align: .04em; } mjx-frac[delims] { padding: 0 .1em; } mjx-frac[atop] { padding: 0 .12em; } mjx-frac[atop][delims] { padding: 0; } mjx-dtable { display: inline-table; width: 100%; } mjx-dtable > * { font-size: 2000%; } mjx-dbox { display: block; font-size: 5%; } mjx-num { display: block; text-align: center; } mjx-den { display: block; text-align: center; } mjx-mfrac[bevelled] > mjx-num { display: inline-block; } mjx-mfrac[bevelled] > mjx-den { display: inline-block; } mjx-den[align="right"], mjx-num[align="right"] { text-align: right; } mjx-den[align="left"], mjx-num[align="left"] { text-align: left; } mjx-nstrut { display: inline-block; height: .054em; width: 0; vertical-align: -.054em; } mjx-nstrut[type="d"] { height: .217em; vertical-align: -.217em; } mjx-dstrut { display: inline-block; height: .505em; width: 0; } mjx-dstrut[type="d"] { height: .726em; } mjx-line { display: block; box-sizing: border-box; min-height: 1px; height: .06em; border-top: .06em solid; margin: .06em -.1em; overflow: hidden; } mjx-line[type="d"] { margin: .18em -.1em; } mjx-msup { display: inline-block; text-align: left; } mjx-mrow { display: inline-block; text-align: left; } mjx-c::before { display: block; width: 0; } .MJX-TEX { font-family: MJXZERO, MJXTEX; } .TEX-B { font-family: MJXZERO, MJXTEX-B; } .TEX-I { font-family: MJXZERO, MJXTEX-I; } .TEX-MI { font-family: MJXZERO, MJXTEX-MI; } .TEX-BI { font-family: MJXZERO, MJXTEX-BI; } .TEX-S1 { font-family: MJXZERO, MJXTEX-S1; } .TEX-S2 { font-family: MJXZERO, MJXTEX-S2; } .TEX-S3 { font-family: MJXZERO, MJXTEX-S3; } .TEX-S4 { font-family: MJXZERO, MJXTEX-S4; } .TEX-A { font-family: MJXZERO, MJXTEX-A; } .TEX-C { font-family: MJXZERO, MJXTEX-C; } .TEX-CB { font-family: MJXZERO, MJXTEX-CB; } .TEX-FR { font-family: MJXZERO, MJXTEX-FR; } .TEX-FRB { font-family: MJXZERO, MJXTEX-FRB; } .TEX-SS { font-family: MJXZERO, MJXTEX-SS; } .TEX-SSB { font-family: MJXZERO, MJXTEX-SSB; } .TEX-SSI { font-family: MJXZERO, MJXTEX-SSI; } .TEX-SC { font-family: MJXZERO, MJXTEX-SC; } .TEX-T { font-family: MJXZERO, MJXTEX-T; } .TEX-V { font-family: MJXZERO, MJXTEX-V; } .TEX-VB { font-family: MJXZERO, MJXTEX-VB; } mjx-stretchy-v mjx-c, mjx-stretchy-h mjx-c { font-family: MJXZERO, MJXTEX-S1, MJXTEX-S4, MJXTEX, MJXTEX-A ! important; } @font-face /* 0 */ { font-family: MJXZERO; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Zero.woff") format("woff"); } @font-face /* 1 */ { font-family: MJXTEX; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Regular.woff") format("woff"); } @font-face /* 2 */ { font-family: MJXTEX-B; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Bold.woff") format("woff"); } @font-face /* 3 */ { font-family: MJXTEX-I; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-Italic.woff") format("woff"); } @font-face /* 4 */ { font-family: MJXTEX-MI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Italic.woff") format("woff"); } @font-face /* 5 */ { font-family: MJXTEX-BI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-BoldItalic.woff") format("woff"); } @font-face /* 6 */ { font-family: MJXTEX-S1; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size1-Regular.woff") format("woff"); } @font-face /* 7 */ { font-family: MJXTEX-S2; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size2-Regular.woff") format("woff"); } @font-face /* 8 */ { font-family: MJXTEX-S3; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size3-Regular.woff") format("woff"); } @font-face /* 9 */ { font-family: MJXTEX-S4; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size4-Regular.woff") format("woff"); } @font-face /* 10 */ { font-family: MJXTEX-A; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_AMS-Regular.woff") format("woff"); } @font-face /* 11 */ { font-family: MJXTEX-C; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Regular.woff") format("woff"); } @font-face /* 12 */ { font-family: MJXTEX-CB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Bold.woff") format("woff"); } @font-face /* 13 */ { font-family: MJXTEX-FR; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Regular.woff") format("woff"); } @font-face /* 14 */ { font-family: MJXTEX-FRB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Bold.woff") format("woff"); } @font-face /* 15 */ { font-family: MJXTEX-SS; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Regular.woff") format("woff"); } @font-face /* 16 */ { font-family: MJXTEX-SSB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Bold.woff") format("woff"); } @font-face /* 17 */ { font-family: MJXTEX-SSI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Italic.woff") format("woff"); } @font-face /* 18 */ { font-family: MJXTEX-SC; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Script-Regular.woff") format("woff"); } @font-face /* 19 */ { font-family: MJXTEX-T; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Typewriter-Regular.woff") format("woff"); } @font-face /* 20 */ { font-family: MJXTEX-V; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Regular.woff") format("woff"); } @font-face /* 21 */ { font-family: MJXTEX-VB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Bold.woff") format("woff"); } mjx-c.mjx-c41::before { padding: 0.716em 0.75em 0 0; content: "A"; } mjx-c.mjx-c2193::before { padding: 0.694em 0.5em 0.194em 0; content: "\2193"; } mjx-c.mjx-c2C::before { padding: 0.121em 0.278em 0.194em 0; content: ","; } mjx-c.mjx-cA0::before { padding: 0 0.25em 0 0; content: "\A0"; } mjx-c.mjx-c42::before { padding: 0.683em 0.708em 0 0; content: "B"; } mjx-c.mjx-c2192::before { padding: 0.511em 1em 0.011em 0; content: "\2192"; } mjx-c.mjx-c44::before { padding: 0.683em 0.764em 0 0; content: "D"; } mjx-c.mjx-c65::before { padding: 0.448em 0.444em 0.011em 0; content: "e"; } mjx-c.mjx-c66::before { padding: 0.705em 0.372em 0 0; content: "f"; } mjx-c.mjx-c63::before { padding: 0.448em 0.444em 0.011em 0; content: "c"; } mjx-c.mjx-c74::before { padding: 0.615em 0.389em 0.01em 0; content: "t"; } mjx-c.mjx-c43::before { padding: 0.705em 0.722em 0.021em 0; content: "C"; } mjx-c.mjx-c6F::before { padding: 0.448em 0.5em 0.01em 0; content: "o"; } mjx-c.mjx-c70::before { padding: 0.442em 0.556em 0.194em 0; content: "p"; } mjx-c.mjx-c72::before { padding: 0.442em 0.392em 0 0; content: "r"; } mjx-c.mjx-c61::before { padding: 0.448em 0.5em 0.011em 0; content: "a"; } mjx-c.mjx-c1D443.TEX-I::before { padding: 0.683em 0.751em 0 0; content: "P"; } mjx-c.mjx-c1D447.TEX-I::before { padding: 0.677em 0.704em 0 0; content: "T"; } mjx-c.mjx-c1D446.TEX-I::before { padding: 0.705em 0.645em 0.022em 0; content: "S"; } mjx-c.mjx-c1D445.TEX-I::before { padding: 0.683em 0.759em 0.021em 0; content: "R"; } mjx-c.mjx-c3C::before { padding: 0.54em 0.778em 0.04em 0; content: "<"; } mjx-c.mjx-c75::before { padding: 0.442em 0.556em 0.011em 0; content: "u"; } mjx-c.mjx-c69::before { padding: 0.669em 0.278em 0 0; content: "i"; } mjx-c.mjx-c6C::before { padding: 0.694em 0.278em 0 0; content: "l"; } mjx-c.mjx-c64::before { padding: 0.694em 0.556em 0.011em 0; content: "d"; } mjx-c.mjx-c50::before { padding: 0.683em 0.681em 0 0; content: "P"; } mjx-c.mjx-c73::before { padding: 0.448em 0.394em 0.011em 0; content: "s"; } mjx-c.mjx-c2212::before { padding: 0.583em 0.778em 0.082em 0; content: "\2212"; } mjx-c.mjx-c31::before { padding: 0.666em 0.5em 0 0; content: "1"; } mjx-c.mjx-c2B::before { padding: 0.583em 0.778em 0.082em 0; content: "+"; } mjx-c.mjx-c32::before { padding: 0.666em 0.5em 0 0; content: "2"; } mjx-c.mjx-c30::before { padding: 0.666em 0.5em 0.022em 0; content: "0"; } mjx-c.mjx-c53::before { padding: 0.705em 0.556em 0.022em 0; content: "S"; } mjx-c.mjx-c6B::before { padding: 0.694em 0.528em 0 0; content: "k"; } mjx-c.mjx-c52::before { padding: 0.683em 0.736em 0.022em 0; content: "R"; } mjx-c.mjx-c6E::before { padding: 0.442em 0.556em 0 0; content: "n"; } mjx-c.mjx-c34::before { padding: 0.677em 0.5em 0 0; content: "4"; } mjx-c.mjx-c35::before { padding: 0.666em 0.5em 0.022em 0; content: "5"; } mjx-c.mjx-c1D44A.TEX-I::before { padding: 0.683em 1.048em 0.022em 0; content: "W"; } mjx-c.mjx-c1D43F.TEX-I::before { padding: 0.683em 0.681em 0 0; content: "L"; } mjx-c.mjx-c3D::before { padding: 0.583em 0.778em 0.082em 0; content: "="; } mjx-c.mjx-c1D440.TEX-I::before { padding: 0.683em 1.051em 0 0; content: "M"; } mjx-c.mjx-c1D45E.TEX-I::before { padding: 0.442em 0.46em 0.194em 0; content: "q"; } mjx-c.mjx-c28::before { padding: 0.75em 0.389em 0.25em 0; content: "("; } mjx-c.mjx-c1D437.TEX-I::before { padding: 0.683em 0.828em 0 0; content: "D"; } mjx-c.mjx-c1D45C.TEX-I::before { padding: 0.441em 0.485em 0.011em 0; content: "o"; } mjx-c.mjx-c1D45A.TEX-I::before { padding: 0.442em 0.878em 0.011em 0; content: "m"; } mjx-c.mjx-c29::before { padding: 0.75em 0.389em 0.25em 0; content: ")"; } mjx-c.mjx-c1D449.TEX-I::before { padding: 0.683em 0.769em 0.022em 0; content: "V"; } mjx-c.mjx-c22C5::before { padding: 0.31em 0.278em 0 0; content: "\22C5"; } mjx-c.mjx-c2217::before { padding: 0.465em 0.5em 0 0; content: "\2217"; } mjx-c.mjx-c7C::before { padding: 0.75em 0.278em 0.249em 0; content: "|"; }The first entry in each cell is the Row player's payoff, the second is the Column player's payoff. R is the payoff if both cooperate (the "reward payoff"). S is the payoff for cooperating against a defector (the "sucker's payoff"). T is the payoff for defecting against a cooperator (the "temptation payoff"). And P is the payoff if both defect (the "punishment payoff"). Prisoner's Dilemma is defined by the fact that .
The nuclear arms race might look like this. Building nukes costs 1, but if you beat your opponent to nukes you get to take 2 from them. If both players build, there's a standoff, and both pay the cost of building without getting to take anything from the other. We can represent this with the following payoff matrix:
This is how the game structure looks for building nuclear weapons. But Hunter Ash, in his tweet, is talking about the game structure for using them—the classic "Mutually Assured Destruction" scenario. Given a few assumptions, this turns out to be a Stag Hunt game. A Stag Hunt is defined by the fact that : Both prefer cooperating (hunting Stag together) if the other side will cooperate, but both prefer defecting (hunting Rabbit alone) if the other side will defect, and both prefer mutual cooperation to mutual defection.
This means that, unlike the Prisoner's Dilemma which has only one equilibrium, namely (Defect, Defect),[2]Stag Hunt has two equilibria, (Stag, Stag) and (Rabbit, Rabbit). And both players prefer (Stag, Stag), so as long as each can be assured the other will hunt Stag, that's what they'll choose.
In MAD, we assume that both players have a "safe second strike," meaning that if the other player attacks them they can safely retaliate, but not a "safe first strike" (meaning they can't strike first without being retaliated against). We'll also assume that, by striking first, you can diminish your opponent's second strike capacity, even if you can't fully eliminate it.[3]Under these assumptions, the payoffs look something like this:
In this case, both sides prefer to refrain if the other side will (because the -2 they'd get from being retaliated against is less than the 0 they'd get by refraining), but they also prefer to strike if the other will (because they have a chance of diminishing the other's second-strike capacity by striking first, meaning in expectation they get -4 instead of -5 by striking if the opponent does).
III. The AI RaceIt turns out that, contra Hunter Ash, this structure is closely analogous to the AI race, under some plausible assumptions.
We can define R, S, T, and P for the AI race with a few parameters. Let's say that:
- is the payoff from winning the race and creating an aligned AI
- is the payoff from losing the race and getting dominated by the opponent's aligned AI (which will be negative).
- Each side will use their aligned ASI to take some amount from the other, so Winning is just as good as Losing is bad; i.e., [4]
- Each side has an equal chance of getting to ASI first, so ignoring the risk from loss of control, the expected payoff if both race is (which we can call ).[5]
- , the chance that the winning side loses control of their AI.[6](I'll generally continue to call this P(Doom) in the prose parts of this essay)
- is the payoff if the winning side loses control of their AI, the Doom payoff.
- If neither side builds, nothing happens.
Given this:
- : if neither side builds, nothing happens
- The other three payoffs are lotteries between Doom and some other outcome , of the form , where is either , , or depending on which payoff we're talking about:
Basically: if no one builds, nothing happens. If one side builds, that side has probability of aligning their AI; if they succeed, they get the Winner's payoff and the other side gets the Loser's payoff, if they fail, they get the Doom payoff. If both sides build, they each have a 50-50 shot of winning, and the winner has probability of aligning their AI and getting the Winner's payoff vs failing to align their AI and getting the Doom payoff.
Now we can just build a model and see that the game type depends on the payoffs and on :
So we can see that for low values of P(Doom), the game is indeed a Prisoner's Dilemma, but for higher values of P(Doom) it becomes a Stag Hunt.[7]There's a crossover point where the game type switches, and the value of only depends on the ratio between and . Specifically, the crossover point happens when . So we can write , and some algebra gets us . Since is negative, the meaning of this is more transparent if we write it as .[8]
So the worse you consider Doom to be (relative to the spoils of victory), the lower P(Doom) has to be for the game to be a Stag Hunt.
Because a Stag Hunt has an equilibrium at (Pause, Pause), an agreement not to build ASI is feasible as long as the game is a Stag Hunt. (Note: I mean game-theoretically feasible. Obviously there are other factors that impact the feasibility of an agreement—messy politics, personalities and relationship between the specific parties involved, etc. I'm not making any claims here about those.)[9]In a Stag Hunt, once both sides are sure the other will comply, a treaty doesn't need enforcement, strictly speaking: neither side has an incentive to defect, and both sides prefer the cooperative equilibrium. The hard part is verification and assurance.
This means that (contra Hunter Ash) you don't need P(Doom) "well north of 50%" for an agreement on not building ASI to be feasible (you only need P(Doom) to be this high if equals ). In fact, if you think Doom is at least 9x worse than Winning the race is good, then P(Doom) only needs to be 10% for an agreement to be feasible. This is a pretty conservative assumption I'd say, since Doom implies complete loss of human control of the future, and likely extinction. I'd go so far as to say it's plausible that Doom is at least 99x worse than Winning is good, in which case P(Doom) only needs to be 1% for an agreement to be feasible.
And most of those with decisionmaking authority here take Doom seriously enough that this math applies to them. Most Western frontier lab CEOs are on the record saying that AI takeover is a real concern; Dario Amodei and Elon Musk have given explicit probabilities of 10%+. Sam Altman and Demis Hassabis have both criticized the "P(Doom)" framing, but Sam has said it's "not zero" and Demis has said it's "definitely non-zero and it's probably non-negligible." And all except Elon (and Mark Zuckerberg) signed the 2023 CAIS statement saying that "Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war." Even Xi Jinping mentioned "loss of control" risks from AI in a recent speech, though, as ever when looking at China as a layperson, it's hard to know what to make of this. But in any case, by lab CEOs' own lights, cooperation with China should be (game-theoretically) possible.
In conclusion: STOP SAYING THE AI RACE IS A PRISONER'S DILEMMA![10]
Appendix: Expanded ModelI haven't had a chance to read all of these yet, but I have had Fable check how they relate to what I'm doing here, and although they are making a similar point, they use somewhat different modeling assumptions. ↩︎
In one-shot settings, and ignoring Functional Decision Theory. ↩︎
With perfectly secure second strikes, the game is actually a "Harmony" game: both sides prefer to Refrain unilaterally. But second strikes in the real world are never perfectly secure. ↩︎
Nothing much interesting happens if we relax this assumption. If Losing is worse than Winning is good, the game type doesn't change. If Losing is less bad than Winning is good, all that happens is that, for low enough values of P(Doom), the game becomes a "Deadlock", where both players actually prefer (Build, Build) to (Pause, Pause). You can see this in the expanded model in a later footnote. ↩︎
Again, nothing much interesting happens if we relax this assumption. If one side is favored to get to ASI first, then for low enough values of P(Doom), the game can be a Deadlock for them even if it isn't for the other side. You can see this in the expanded model in a later footnote. ↩︎
We'll assume that whichever ASI comes online first can defeat another ASI that comes online later, or prevent it from becoming ASI in the first place. So whether ASI is aligned or misaligned depends only on whether the winner's ASI is aligned or misaligned—once a misaligned ASI exists, there's no chance to defend against it with an aligned ASI of your own, and conversely, once an aligned ASI is created, all future ASIs will be aligned as well. ↩︎
For it actually becomes a Harmony game, but is not 1 so this is irrelevant for real-world analysis. ↩︎
Note that, in particular, it does not depend on L, or on the probability of winning the race given that both sides Build. So our two assumptions from earlier don't affect , and Ash's worry from the introduction about how bad it would be to lose the race never takes us from a Stag Hunt to a Prisoner's Dilemma. Here is an expanded model with these parameters allowed to vary that you can play with in the appendix at the end of the post. ↩︎
As an aside, this is basically isomorphic to another well-known variant of the nuclear arms race problem. There was a serious fear, at one point during the process of building the atomic bomb, that setting it off could ignite the atmosphere and kill everyone on earth. In actuality, before testing the bomb they did the calculations and found that the probability of this was small, but the same logic applies: if the probability of killing everyone were high enough, then coordination not to build the bomb would be a Stag Hunt. ↩︎
Unless your own P(Doom), W, and D satisfy (e.g. P(Doom) below 10% if you think Doom is only 9x as bad as Winning is good). ↩︎
Discuss
Why I think we live in a "simulation"
Epistemic Status: I think each part of this piece is worth engaging with in its own right. I haven't thought as much about whether the ideas presented here add up to "it's more likely than not we live in a simulation".
There are more things in heaven and earth, Horatio, than are dreamt of in your philosophy.
We have good reason to believe that we exist in a “simulation”[1] of some kind. “Simulation” here should be taken to mean: any kind of intentionally created seemingly-self-contained ‘world’ containing conscious actors with no knowledge of the external world. Note this is not the same as the usual conception/usage of simulation.[2]
I believe this not for any one single reason[3]; rather, different ways of thinking about this question each give some evidence, and when combined they make it look very unlikely that we don’t live in a simulation. In the words of the inimitable Linchuan Zhang, “each individual piece of evidence might be explained away, but the consilience of evidence across multiple angles and sources is in my opinion very hard to dismiss collectively.“[4] Each of the ideas or “pieces of evidence” I discuss in this piece has some kind of non-simulation explanation, just-so story, or ad hoc rationalization. (“I am firm, you are obstinate, he is a pig-headed fool”). But when taken all at once, they’re very difficult to explain sans simulation.
What are these different angles? On the theoretical side, there’s A. the strong philosophic “simulation argument” and B. the existence of deep ontological problems our minds face that suggest our ontology does not capture reality. On the "empirical" side, there’s A. a variety of evidence that suggests we live in an intentionally designed world and B. a variety of reasons to expect superintelligent beings to carry out simulations that look like our world.
The classic simulation argumentLet’s start with the "theoretical" reasons. The biggest thing here is the classic simulation argument, which is worth revisiting as it’s a very powerful argument! The full argument is here; the conclusion it argues for is this:
At least one of the following propositions is true: (1) the human species is very likely to go extinct before reaching a “posthuman” stage; (2) any posthuman civilization is extremely unlikely to run a significant number of simulations of their evolutionary history (or variations thereof); (3) we are almost certainly living in a computer simulation.
I would modify this to remove the “computer” in “computer simulation” and add that ‘simulation’ here should be taken to mean “any kind of seemingly-self-contained ‘world’ containing conscious actors with no knowledge of the external world.”[5]
Deep ontological problemsLet’s now talk about the deep ontological problems we face. We ‘know’ that there exist true statements that we cannot know and that there are facts we cannot express in language. We know this from information theory and from Godel’s incompleteness theorems. Furthermore, there are very important things we cannot express in language, and I believe[6] there are very important things we cannot even conceptualize within our basic ontology, an ontology to which concepts like “truth” and “exist” are essential. Taoism and Zen make these points rather compellingly, though there’s no strict argument I can point to here; indeed, part of the point is that there is no way to express this stuff in language.
In particular, I think that ‘consciousness’ is an extremely important thing that not only cannot be properly referenced or described by language but cannot even be studied or accessed from a third-person/objective/scientific perspective.
There exist many “big” questions that we just can’t seem to answer. Often a particular theory can answer/reject one/some of them, but cannot answer others and is inconsistent with theories that do answer others. Many of these questions could be said to belong to the domain of physics (or do they belong to the domain of philosophy?). What I’m talking about here are questions like:[7]
- Why is there something rather than nothing?
- What created the universe? Or was it never created? If it was never created, why does it exist?
- Why can’t we unify all of physics in one model?
- Why do we not see other intelligence in the universe?
- Why does our universe appear to be fine-tuned?
- Why does consciousness exist? What is its ontological relation to other things that ‘exist’?
- What is the relationship between is and ought?
- How can we know our reasoning makes any sense? AKA the problem of induction.[8]
- What is truth?
Theories that (purportedly) answer some of these questions: Anthropic principle. Mathematical universe hypothesis. Christianity. Illusionism.
For further reading, see niplav's Ontological Cluelessness.
Intentional design (this sure would make a great story!)Our world looks like it was optimized for entertainment. Consider:
- Super-entertaining figures: Trump, Musk[9].
- Nominative determinism: Also Trump[10], Sam Bankman-Fried, Plotkin, Joey Chestnut (the world’s best eater’s surname is a food), so many more that it strains credulity. Scott Alexander has many good instances sprinkled throughout his writings. Gwern has a nice collection, as does Roland Crosby.[11]
- Crazy names in general.
- Synchronicity.
- [There are more examples of this phenomenon, right? Surely this could be fleshed out?]
Note that superentertaining characters, nominative determinism, and crazy names are all typical features of narrative art/entertainment – think of literature and film and limericks and so on. Our world looks like a designed story. It looks narratively optimized.
Besides being entertaining, our world is one that would be of particular interest to a wide swath of simulators. We live in the most interesting time, that being the takeoff to the singularity. This seems like a really important time to study. It’s not hard to imagine how simulations of civilization building ASI would be useful for historians, anthropologists, economists, theologians, cognitive scientists, philosophers, story writers, video game designers, etc. etc. etc.
The names of the CEOs involved in AI development – Altman (alternative to man), Pichai (Pitch AI), Amodei (also Amodei means lover of God) – are truly egregious, especially when one considers that the era of ASI development and the period immediately before the singularity are times of particular interest for multiple kinds of simulators. My collection of strangely related phenomena is consilience; yours is synchronicity; his is coincidence.
To sum it up: if I were a simulator, my simulation might look like the exact world we live in! Seriously, can you think of a more interesting, informative, entertaining world to simulate?
Simulation as training module, from Monster's Inc
Allow me to (non-exhaustively) list some kinds of simulation that this may be. Note that our ontologies might not apply to whatever is actually going on in the ‘bigger’, external-to-us reality. ‘Art piece’ or ‘‘God’ might not be good concepts in this bigger reality, but they’re at least pointing at the things that we lack the proper concepts for.[12]
- Traditional ‘computer’ simulation
- For research
- Research around the singularity. Under what conditions does a civilization…
- Decide to pause AI development?
- Never build ASI?
- Die to the ASI they build?
- Build an aligned ASI
- Build an unaligned ASI
- This would certainly be of interest to anyone studying history!
- Can change parameters to test different theories against each other
- E.g. great man theory
- Can change parameters to test different theories against each other
- It might also tell you about what you can expect to find elsewhere in your universe.
- How common is ASI?
- Great filter?
- Dark forest hypothesis
- If you put an Elon Musk archetype in the world in the year 2000, how does that change the 21st century?
- A thought experiment played out
- Research around the singularity. Under what conditions does a civilization…
- For entertainment
- For moral reasons – ancestor simulation?
- Worship; “replaying their lives”.
- For research
- Immersive game
- Immersive aesthetic experience
- Training module
- RPG VR video game
- “Roy theory” from Rick and Morty.
- Art piece
- Immersive aesthetic experience
- A book (or some other kind of not-acted-by-real-people narrative fiction), and it’s the case that fictional characters actually experience what is being described.
- Dream / fantasy of some higher being
- God entertaining themself.
- God teaching themself.
- A drug-induced experience of higher being; ‘aliens tripping’.
See Alan Watt's "The Dream of Life".
“[L]et’s suppose that you were able every night to dream any dream you wanted to dream and that you could, for example, have the power within one night to dream 75 years of time or any length of time you wanted to have. And you would, naturally, as you began on this adventure of dreams, you would fulfill all your wishes. You would have every kind of pleasure during your sleep. And after several nights of 75 years of total pleasure each, you would say ‘Well, that was pretty great. But now let’s have a surprise. Let’s have a dream which isn’t under control, where something is gonna happen to me that I don’t know what it’s gonna be’. And you would dig that and would come out of that and you would say ‘Wow that was a close shave, wasn’t it?’.
Then you would get more and more adventurous and you would make further- and further out gambles to what you would dream. And finally, you would dream where you are now. You would dream the dream of living the life that you are actually living today.”
- See also the butterfly dream.
- Theater.
- Maybe this is how plays are done in the bigger reality. You could imagine that in the future, actors take a drug (or neurological intervention) that makes them temporarily forget their existence as a person outside the play.
That's all for now.
As mentioned in the beginning, I think each part of this piece is worth engaging with in its own right. For example, I think it's worth asking "is our world super-entertaining" separately from the question of "if our world is super-entertaining, is that evidence that we live in a simulation?". I welcome comments on all parts of this.
Appendix: simulations vs embeddings, intelligent design vs intentional designI said earlier that by ‘simulation’ I mean any kind of intentionally created seemingly self-contained ‘world’ acontaining conscious actors with no knowledge of the external world. If we lose the intentionally-created clause, we can think about non-intentionally-created ‘worlds’ as well. I call this broader class of things “embeddings”, as in: our world is embedded in a much greater reality that we don’t have access to. Embeddings are a superset of simulations. (I’m accepting proposals for a word to refer to only the non-intentionally-created worlds.) An iconic conception of this kind of situation is that of Flatland. (Who else watched the super cool movie in school?). I do think it’s reasonably likely that we exist in a non-intentional embedding, though I don’t have any more to say about that right now.
Instead, I bring up embeddings to point out that some of the evidence and all of the possibilities I’ve described for simulation does not apply to non-intentional embeddings. We have reason to believe that we live in an intentionally-created universe. Our world appears to be intentionally designed. (I wonder if this actually the better framing for the whole idea here, rather than the ‘simulation’ framing I used.) Sorry, Christian fundamentalists, you got it wrong. It’s not argument from intelligent design, it’s argument from intentional design.
- ^
Any ideas for a better word here than ‘simulation’?
- ^
Indeed, I think that a narrow conception of simulation has prevented people from seeing how various simulation arguments fit together.
- ^
Some notes about my own epistemology here. My belief that we exist in a simulation has gradually increased from ~low to ~high over the past many years, although I think this is the kind of proposition that one can’t properly put a probability on. I am not a “woo” person by nature or upbringing. I am quite sane, mentally sound, and non-schizo.
- ^
This is from Linch’s piece on the Pope’s AI Encylical being AI generated. I couldn’t say it better myself (inimitable!), so I quoted him.
- ^
Note that this is not how Bostrom intends or presents the argument; still, the argument he gives basically applies as written to this broader class of simulations. And note that it is a broader conception of simulation, which means that it’s weaker than the one Bostrom is making. Furthermore I think that this conception of simulation actually addresses the biggest weaknesses in his argument: A. what if computer simulations can’t give rise to consciousness B. what if intelligent civilization intentionally avoids computer simulations that give rise to consciousness C. there may be infinities or at least recursive, exponentially generators of beings that aren’t computer simulations (such as fiction of various kinds); why is the measure of computer simulated beings larger than these simulated-but-not-by-computer beings?
- ^
Although I wouldn’t say I “know” it. I don’t have a ‘proof’ for it. It may well be one of these true unprovable things I’m talking about!
- ^
A note on the theoretical vs empirical distinction: Some of these "ontological problems" straddle the border into being ‘empirical’/contingent problems, such as the fine-tuned universe, the question of why don’t we see other intelligences in the universe, and the question of the existence and creation of the universe. The theoretical vs empirical distinction is only confusing here.
- ^
Answer: We can’t, a system cannot prove its own consistency, c.f. Godel’s second incompleteness theorem.
- ^
Elon Musk, for what it's worth, has expressed views similar to those being made in this section: "The most entertaining outcome, especially if ironic, is most likely." Based on his behavior, I think he really believes this to some degree? The better-stated version of this proposition is "All else being equal, more entertaining outcomes are more likely, and the effect size here is considerable."
- ^
Donald famously trumped Hillary in the 2016 election.
- ^
Ask your friendly neighborhood AI for more samples.
- ^
As said in Ontological Cluelessness, "Humans may be in a state of total confusion as to the fundamental makeup of the cosmos and its rules, to the point where even extremely basic concepts would need to be revised for accurate understanding."
Discuss
Страницы
- « первая
- ‹ предыдущая
- …
- 7
- 8
- 9
- 10
- 11
- 12
- 13
- 14
- 15
- …
- следующая ›
- последняя »