Вы здесь
Сборщик RSS-лент
Value generalisation Theory of Change: putting it into practice
In the previous post, I presented my theory of change for why value generalisation is vital for AI alignment. Here I'll add the practical part of the argument: given those facts, why explicitly try to do value generalisation, what are the dangers of the approach, how should it be done, and how do we mitigate the risks?
The formal theory of change is down below, but I'll put a collapsible version here, to make references easier:
Theory of Change
- Inputs / activities: investment/grants, a small research team, small-medium compute resources.
- Outputs:
- Solutions to the three components of value generalisation (see here) in academically rigorous forms, demonstrations of these solutions on toy problems, public benchmarks, and new benchmarks. If going commercial, applications of this value generalisation to saleable products.
- The first component is recognising that the AI is off-distribution in a value-relevant way.
- The second component is establishing what features should be used to reach a decision while off-distribution this way.
- The third component is to reach a good decision while off distribution, and to loop through the process again to critique and improve the decision, including using outside feedback.
- Combination of these approaches in a successful prealigned learning model.
- If going commercial and it is safe, sale of the model to the public.
- Solutions to the three components of value generalisation (see here) in academically rigorous forms, demonstrations of these solutions on toy problems, public benchmarks, and new benchmarks. If going commercial, applications of this value generalisation to saleable products.
- Intermediate outcomes:
- Development of value generalisation alignment methods for AIs.
- Making these methods available to AI safety researchers.
- Prealigned models being generally available to safely empower users and keep humanity on track for a positive outcome.
- If the approach fails, the failed attempts and the partial successes will be made available to others.
- Impact:
- If failed:
- A seemingly promising approach crossed off the list of possible paths to alignment.
- A better understanding of the issues around the "value generalisation" framing.
- If intermediate outcome:
- A tool or collection of useful tools or models that can improve some AI alignment techniques.
- If successful:
- AI alignment, or a big step forwards towards it.
- An increase in AI capabilities.
- If failed:
- Claim mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; font-weight: normal; font-size: 100%; font-size-adjust: none; letter-spacing: normal; border-collapse: collapse; word-wrap: normal; word-spacing: normal; white-space: nowrap; direction: ltr; padding: 1px 0; } mjx-container[jax="CHTML"][display="true"] { display: block; text-align: center; margin: 1em 0; } mjx-container[jax="CHTML"][display="true"][width="full"] { display: flex; } mjx-container[jax="CHTML"][display="true"] mjx-math { padding: 0; } mjx-container[jax="CHTML"][justify="left"] { text-align: left; } mjx-container[jax="CHTML"][justify="right"] { text-align: right; } mjx-mi { display: inline-block; text-align: left; } mjx-c { display: inline-block; } mjx-utext { display: inline-block; padding: .75em 0 .2em 0; } mjx-TeXAtom { display: inline-block; text-align: left; } mjx-msub { display: inline-block; text-align: left; } mjx-mo { display: inline-block; text-align: left; } mjx-stretchy-h { display: inline-table; width: 100%; } mjx-stretchy-h > * { display: table-cell; width: 0; } mjx-stretchy-h > * > mjx-c { display: inline-block; transform: scalex(1.0000001); } mjx-stretchy-h > * > mjx-c::before { display: inline-block; width: initial; } mjx-stretchy-h > mjx-ext { /* IE */ overflow: hidden; /* others */ overflow: clip visible; width: 100%; } mjx-stretchy-h > mjx-ext > mjx-c::before { transform: scalex(500); } mjx-stretchy-h > mjx-ext > mjx-c { width: 0; } mjx-stretchy-h > mjx-beg > mjx-c { margin-right: -.1em; } mjx-stretchy-h > mjx-end > mjx-c { margin-left: -.1em; } mjx-stretchy-v { display: inline-block; } mjx-stretchy-v > * { display: block; } mjx-stretchy-v > mjx-beg { height: 0; } mjx-stretchy-v > mjx-end > mjx-c { display: block; } mjx-stretchy-v > * > mjx-c { transform: scaley(1.0000001); transform-origin: left center; overflow: hidden; } mjx-stretchy-v > mjx-ext { display: block; height: 100%; box-sizing: border-box; border: 0px solid transparent; /* IE */ overflow: hidden; /* others */ overflow: visible clip; } mjx-stretchy-v > mjx-ext > mjx-c::before { width: initial; box-sizing: border-box; } mjx-stretchy-v > mjx-ext > mjx-c { transform: scaleY(500) translateY(.075em); overflow: visible; } mjx-mark { display: inline-block; height: 0px; } mjx-mn { display: inline-block; text-align: left; } mjx-msup { display: inline-block; text-align: left; } mjx-mtext { display: inline-block; text-align: left; } mjx-c.mjx-c1D43D.TEX-I::before { padding: 0.683em 0.633em 0.022em 0; content: "J"; } mjx-c.mjx-c1D43E.TEX-I::before { padding: 0.683em 0.889em 0 0; content: "K"; } mjx-c.mjx-c45.TEX-C::before { padding: 0.705em 0.564em 0.022em 0; content: "E"; } mjx-c.mjx-c1D444.TEX-I::before { padding: 0.704em 0.791em 0.194em 0; content: "Q"; } mjx-c.mjx-c1D44B.TEX-I::before { padding: 0.683em 0.852em 0 0; content: "X"; } mjx-c.mjx-c46.TEX-C::before { padding: 0.683em 0.829em 0.032em 0; content: "F"; } mjx-c.mjx-c1D452.TEX-I::before { padding: 0.442em 0.466em 0.011em 0; content: "e"; } mjx-c.mjx-c2208::before { padding: 0.54em 0.667em 0.04em 0; content: "\2208"; } mjx-c.mjx-c7B::before { padding: 0.75em 0.5em 0.25em 0; content: "{"; } mjx-c.mjx-c1D453.TEX-I::before { padding: 0.705em 0.55em 0.205em 0; content: "f"; } mjx-c.mjx-c28::before { padding: 0.75em 0.389em 0.25em 0; content: "("; } mjx-c.mjx-c29::before { padding: 0.75em 0.389em 0.25em 0; content: ")"; } mjx-c.mjx-c7D::before { padding: 0.75em 0.5em 0.25em 0; content: "}"; } mjx-c.mjx-c30::before { padding: 0.666em 0.5em 0.022em 0; content: "0"; } mjx-c.mjx-c2C::before { padding: 0.121em 0.278em 0.194em 0; content: ","; } mjx-c.mjx-c31::before { padding: 0.666em 0.5em 0 0; content: "1"; } mjx-c.mjx-c7C::before { padding: 0.75em 0.278em 0.249em 0; content: "|"; } mjx-c.mjx-c32::before { padding: 0.666em 0.5em 0 0; content: "2"; } mjx-c.mjx-c1D448.TEX-I::before { padding: 0.683em 0.767em 0.022em 0; content: "U"; } mjx-c.mjx-c3A::before { padding: 0.43em 0.278em 0 0; content: ":"; } mjx-c.mjx-c2192::before { padding: 0.511em 1em 0.011em 0; content: "\2192"; } mjx-c.mjx-c211D.TEX-A::before { padding: 0.683em 0.722em 0 0; content: "R"; } mjx-c.mjx-c2032::before { padding: 0.56em 0.275em 0 0; content: "\2032"; } mjx-c.mjx-c3A6::before { padding: 0.683em 0.722em 0 0; content: "\3A6"; } mjx-c.mjx-c1D703.TEX-I::before { padding: 0.705em 0.469em 0.01em 0; content: "\3B8"; } mjx-c.mjx-c210E.TEX-I::before { padding: 0.694em 0.576em 0.011em 0; content: "h"; } mjx-c.mjx-c2217::before { padding: 0.465em 0.5em 0 0; content: "\2217"; } mjx-c.mjx-c39E::before { padding: 0.677em 0.667em 0 0; content: "\39E"; } mjx-c.mjx-c1D449.TEX-I::before { padding: 0.683em 0.769em 0.022em 0; content: "V"; } mjx-c.mjx-c1D43F.TEX-I::before { padding: 0.683em 0.681em 0 0; content: "L"; } mjx-c.mjx-c1D440.TEX-I::before { padding: 0.683em 1.051em 0 0; content: "M"; } mjx-c.mjx-c1D441.TEX-I::before { padding: 0.683em 0.888em 0 0; content: "N"; } mjx-c.mjx-c1D442.TEX-I::before { padding: 0.704em 0.763em 0.022em 0; content: "O"; } mjx-c.mjx-c1D443.TEX-I::before { padding: 0.683em 0.751em 0 0; content: "P"; } mjx-c.mjx-c49::before { padding: 0.683em 0.361em 0 0; content: "I"; } mjx-c.mjx-c1D45B.TEX-I::before { padding: 0.442em 0.6em 0.011em 0; content: "n"; } mjx-c.mjx-c2308::before { padding: 0.75em 0.444em 0.25em 0; content: "\2308"; } mjx-c.mjx-c6C::before { padding: 0.694em 0.278em 0 0; content: "l"; } mjx-c.mjx-c6F::before { padding: 0.448em 0.5em 0.01em 0; content: "o"; } mjx-c.mjx-c67::before { padding: 0.453em 0.5em 0.206em 0; content: "g"; } mjx-c.mjx-c2061::before { padding: 0 0 0 0; content: ""; } mjx-c.mjx-c2309::before { padding: 0.75em 0.444em 0.25em 0; content: "\2309"; } mjx-container[jax="CHTML"] { line-height: 0; } mjx-container [space="1"] { margin-left: .111em; } mjx-container [space="2"] { margin-left: .167em; } mjx-container [space="3"] { margin-left: .222em; } mjx-container [space="4"] { margin-left: .278em; } mjx-container [space="5"] { margin-left: .333em; } mjx-container [rspace="1"] { margin-right: .111em; } mjx-container [rspace="2"] { margin-right: .167em; } mjx-container [rspace="3"] { margin-right: .222em; } mjx-container [rspace="4"] { margin-right: .278em; } mjx-container [rspace="5"] { margin-right: .333em; } mjx-container [size="s"] { font-size: 70.7%; } mjx-container [size="ss"] { font-size: 50%; } mjx-container [size="Tn"] { font-size: 60%; } mjx-container [size="sm"] { font-size: 85%; } mjx-container [size="lg"] { font-size: 120%; } mjx-container [size="Lg"] { font-size: 144%; } mjx-container [size="LG"] { font-size: 173%; } mjx-container [size="hg"] { font-size: 207%; } mjx-container [size="HG"] { font-size: 249%; } mjx-container [width="full"] { width: 100%; } mjx-box { display: inline-block; } mjx-block { display: block; } mjx-itable { display: inline-table; } mjx-row { display: table-row; } mjx-row > * { display: table-cell; } mjx-mtext { display: inline-block; } mjx-mstyle { display: inline-block; } mjx-merror { display: inline-block; color: red; background-color: yellow; } mjx-mphantom { visibility: hidden; } _::-webkit-full-page-media, _:future, :root mjx-container { will-change: opacity; } mjx-c::before { display: block; width: 0; } .MJX-TEX { font-family: MJXZERO, MJXTEX; } .TEX-B { font-family: MJXZERO, MJXTEX-B; } .TEX-I { font-family: MJXZERO, MJXTEX-I; } .TEX-MI { font-family: MJXZERO, MJXTEX-MI; } .TEX-BI { font-family: MJXZERO, MJXTEX-BI; } .TEX-S1 { font-family: MJXZERO, MJXTEX-S1; } .TEX-S2 { font-family: MJXZERO, MJXTEX-S2; } .TEX-S3 { font-family: MJXZERO, MJXTEX-S3; } .TEX-S4 { font-family: MJXZERO, MJXTEX-S4; } .TEX-A { font-family: MJXZERO, MJXTEX-A; } .TEX-C { font-family: MJXZERO, MJXTEX-C; } .TEX-CB { font-family: MJXZERO, MJXTEX-CB; } .TEX-FR { font-family: MJXZERO, MJXTEX-FR; } .TEX-FRB { font-family: MJXZERO, MJXTEX-FRB; } .TEX-SS { font-family: MJXZERO, MJXTEX-SS; } .TEX-SSB { font-family: MJXZERO, MJXTEX-SSB; } .TEX-SSI { font-family: MJXZERO, MJXTEX-SSI; } .TEX-SC { font-family: MJXZERO, MJXTEX-SC; } .TEX-T { font-family: MJXZERO, MJXTEX-T; } .TEX-V { font-family: MJXZERO, MJXTEX-V; } .TEX-VB { font-family: MJXZERO, MJXTEX-VB; } mjx-stretchy-v mjx-c, mjx-stretchy-h mjx-c { font-family: MJXZERO, MJXTEX-S1, MJXTEX-S4, MJXTEX, MJXTEX-A ! important; } @font-face /* 0 */ { font-family: MJXZERO; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Zero.woff") format("woff"); } @font-face /* 1 */ { font-family: MJXTEX; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Regular.woff") format("woff"); } @font-face /* 2 */ { font-family: MJXTEX-B; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Bold.woff") format("woff"); } @font-face /* 3 */ { font-family: MJXTEX-I; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-Italic.woff") format("woff"); } @font-face /* 4 */ { font-family: MJXTEX-MI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Italic.woff") format("woff"); } @font-face /* 5 */ { font-family: MJXTEX-BI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-BoldItalic.woff") format("woff"); } @font-face /* 6 */ { font-family: MJXTEX-S1; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size1-Regular.woff") format("woff"); } @font-face /* 7 */ { font-family: MJXTEX-S2; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size2-Regular.woff") format("woff"); } @font-face /* 8 */ { font-family: MJXTEX-S3; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size3-Regular.woff") format("woff"); } @font-face /* 9 */ { font-family: MJXTEX-S4; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size4-Regular.woff") format("woff"); } @font-face /* 10 */ { font-family: MJXTEX-A; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_AMS-Regular.woff") format("woff"); } @font-face /* 11 */ { font-family: MJXTEX-C; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Regular.woff") format("woff"); } @font-face /* 12 */ { font-family: MJXTEX-CB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Bold.woff") format("woff"); } @font-face /* 13 */ { font-family: MJXTEX-FR; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Regular.woff") format("woff"); } @font-face /* 14 */ { font-family: MJXTEX-FRB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Bold.woff") format("woff"); } @font-face /* 15 */ { font-family: MJXTEX-SS; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Regular.woff") format("woff"); } @font-face /* 16 */ { font-family: MJXTEX-SSB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Bold.woff") format("woff"); } @font-face /* 17 */ { font-family: MJXTEX-SSI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Italic.woff") format("woff"); } @font-face /* 18 */ { font-family: MJXTEX-SC; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Script-Regular.woff") format("woff"); } @font-face /* 19 */ { font-family: MJXTEX-T; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Typewriter-Regular.woff") format("woff"); } @font-face /* 20 */ { font-family: MJXTEX-V; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Regular.woff") format("woff"); } @font-face /* 21 */ { font-family: MJXTEX-VB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Bold.woff") format("woff"); } : explicit value generalisation is much more useful than implicit value generalisation.
Fundamentally, there's a difference between an AI that knows what course humans would consider the best one, and one that actually follows that course. Having the AI capable of value generalisation inside its own mind is not useful, unless we can incorporate that generalisation into its goals. And explicit generalisation allows that.
This leads to the crucial and unfortunate claim:
- Claim : value generalisation will aid in empirical generalisation. But empirical generalisation will likely not aid value generalisation.
Empirical generalisation is the ability to update features usefully across model splinterings, in ways that preserve or improve the ability of the modeller to effectively influence the world.
Claim derives in part from claim : since useful value generalisation is the explicit kind, implicit empirical generalisation is not likely to usefully help. It also derives from the fact that empirical generalisation can afford to discard its features if necessary: we refined concepts like heat while discarding vitalism. But we can't just discard the "suffering" feature in "avoid human suffering"; it has to be redefined and extended. So value generalisation is strictly harder.
Somewhat connected to that is the fact that "in the limit" of infinite computation and infinite observation, empirical generalisation is doable, but value generalisation involves moral choices that don't come free from mere observations.
More details on the dis-equivalence between implicit or explicit empirical generalisation, and explicit value generalisation
Assume we have , a set of environments, with a probability distribution over . An agent uses features to model the set of environments; these can be seen as numerical functions on the set of environments, so environment is modelled by the value , and there is a probability distribution over the values of the features. These features will include the agent's actions, allowing it to plot a policy.
For simplicity, assume the features are Boolean valued[1], so a world is mapped, via the , to an element of , which is also written as .
Then give a utility function , we can say that is useful for -maximising if there exists a utility function such that policies that optimise or fail to optimise , will also optimise or fail to optimise . Thus agent can stay in their modelling space and plan and act there.
Typically, the agent doesn't start with access to ground reality, so they don't start with , but with a . Note that just because a is defined over features, doesn't mean that the features are useful for maximising that function. The score in Pacman is enough of a feature to define an objective, but not enough to play the game.
A model splintering is a change in and/or , replacing with (a model splintering might be as simple as the agent realising they have extra options - such as in reward tampering or wire-heading - or it might be a complete change of their view of reality). Note that model splinterings are not things the agent directly observes - the need for a model splintering is inferred from anomalous-seeming observations.
Empirical feature generalisation is an algorithm that takes in the agent's internal state (which corresponds to ) and the agent's history . The then moves the agent to an internal state (which corresponds to new features ), such that the new features are useful for maximising the original (or ), given the evidence has provided about model splintering.
Explicit feature generalisation does the process in terms of the features themselves: maps and to the new . Technically, neither nor define each other; but it is generally much easier to construct a black-box method from an explicit method than the opposite. And, in the limit, might be constructible by brute force analysis.
Value generalisation instead uses an that plays the same role as , except that it also has to map to a new . Empirical generalisation (or just exploration) might establish that spinning the boat endlessly gives maximal points; value generalisation also has to figure out that this is not a desired generalisation of the original goal. For a similar reason, while can discard a feature as no longer useful, or simplify it to the extreme, can't discard or over-simplify a value-relevant feature; indeed, these features are often expected to grow more complex over time (e.g. the definition of sentient being).
But why "also"? Could , or the explicit version , not simply focus on the value generalisation and ignore the empirical changes? The first problem is that, in order to generalise , the agent will have to figure out a lot about the features of the environment that are correlated or not correlated with and how they relate to each other (and how that relationship might be changeable).
For instance, human smiling is correlated with the feeling of happiness, the release of certain hormones, physiological changes in the brain, reported happiness, and so on. So to generalise "smiling" to "true happiness" and beyond, the agent has to learn a lot about the structure of the world and what features best describe it. It also has to learn how to change the various potential : if a given candidate is suddenly very easy to optimise, that's a sign that it might be a Goodhart proxy.
The second problem is that could be replaced with any instrumental goal , and any change of .
So it seems that , which we need for effective and checkable value generalisation, must be able to do a deep analysis of any instrumental goal over any model splintering, including figuring out ways to optimise the original and the new instrumental goals. This seems to feed straight into empirical generalisation.
I will ruthlessly search for that don't give powerful , but I am not optimistic.
This asymmetry leads to the unfortunate result that:
- Claim : a value generalisation project will increase AI capabilities.
I've previously made the claim that:
- Claim : strong generalisation is a uniquely human ability, that current LLMs don't have and will likely not develop; nor will any similar model develop strong generalisation either.
However, this claim is not crucial to the approach. If is wrong, claim becomes less worrying (AIs will develop the ability anyway) but the research becomes more urgent: we need to get a decent start on explicit value generalisation before AIs get good at empirical generalisation.
The weaker claim is:
- Claim : studying how humans do strong generalisation is a fruitful way of analysing the skill, at least initially.
The arguments in the second value generalisation post apply also for claim . Even if one is of the claim that "LLMs will always fail at strong generalisation", there is certainly evidence that humans are doing something very different (and much more data-efficient) when we generalise.
The question of timingGiven that value generalisation is essential to alignment, but also could lead to capability increase, the question is: should we research it now (early) or later (late)?
I believe:
- Claim : value generalisation should be researched early.
The arguments for this are that it's better to have a capability increase while the models are weaker and more controllable, and early value generalisation can more easily be integrated into models from the get-go rather than retro-fitting at a later date. We want a minimal capability overhang that empirical generalisation can unleash. And, of course, we want to avoid the scenarios where the research arrives too late.
The corporate argumentGiven the above, I propose creating either a research program or a commercial entity to find useable solutions to value generalisation. But I will claim:
- Claim : if value generalisation is indeed revolutionary but does lead to dangerous capability increases, then the commercial route weakly dominates the research route.
The main argument for claim comes from the question: assume that value generalisation is solved or partially solved, then what? We have a powerful alignment technology that is essential for alignment, but also a powerful capability technology, in a way that can't be separated. What do we do with it?
Well, we'd probably want to hand it over to some trustworthy entity (some have suggested a "CERN for AI") to implement the alignment part at some point.
But that is easier to achieve for a corporation than for a research program. A corporation is much better placed to keep its research private or patent-protected and concealed from the world. It can sell or license the research to a trustworthy entity, or sell it to a semi-trustworthy entity with conditions. And it can also choose to be the trustworthy entity and start implementing alignment itself. Indeed:
- Claim : if the capability boost from generalisation is inevitable, a corporation can mix the capability increase with the alignment increase (such as prealigned AIs) so that the first generalising AIs are aligned, and the initial income from generalising AIs flows to value aligned entities.
I've mentioned the advantages of the corporate route, but what of the advantages of the academic or research route - such as openness allowing more scrutiny and feedback, getting more trust from the safety community? Well, the research is potentially dangerous, so we can't expect to have it publicly or semi-publicly available. Thus:
- Claim : the potential need for secrecy reduces the standard advantages from the academic or research institution route.
Of course, the corporate path has its weaknesses, mainly revolving around the profit motive. Investors, the legal system, management, and employees will all want a company to cash in on legal innovative ideas, even if the ideas are potentially dangerous. So a first step would be:
- Design : the corporation should have an independent AI ethics board, with the power to block the use of IP it deems dangerous. To facilitate this, the AI ethics board should be the owner of the IP.
That solves the problem in the formal sense, which means that it doesn't really solve it. Additional measures would be:
- Design : the employees and management should be value aligned with AI safety.
- Design : the investors should be value-aligned with AI safety.
That also helps; also not enough. The situation will never be "do I kill everyone with certainty to make $10,000 more this year"? The situation will be more like "when I feel the ethics board is being fussy and unreasonable and overly cautious, do I nevertheless bow down to their irrational decrees that will cost me a lot of my expected income that I have worked so hard on and so earned, and also give up my possibility of improving the world for the better"? And value-aligned investors remain investors: they expect to make a profit (a not-unreasonable demand, which the legal system backs up).
I don't expect myself to be immune to that pressure and those arguments. And it's not just a question of holding firm; as OpenAI's experience demonstrates:
- Claim : if the employees are ready to jump ship to a partner organisation that can continue the research with a minimum of fuss, the AI ethics board's power is theoretical rather than real.
So I've been considering ways to remove the stark tension. One design is:
- Design : the AI ethics board will have the power to implement a "pivot to immediate profit". The long-range dangerous IP will be removed from the company, and the company will turn to making money from all the partial ideas and designs that it has made and rejected (as not being alignment relevant) along the way. The board will, of course, vet these ideas for danger, but immediate short-term profitability is less likely to overlap with truly dangerous IP.
Arguably, the pivot to immediate profit would be more profitable, in expectation, than speculative long term IP. This would relieve much of the commercial pressure. And if the long term IP was truly world-improving, the ethics board would find a way to get it carefully deployed, so world-improving impulses are preserved (if the research is merely dangerous, the ethics board would bury it, and good riddance).
Plan, milestones, and assumption checksLet's gather all the previous together to put it all in one plan:
- Inputs / activities: investment/grants, a small research team, small-medium compute resources.
- Outputs:
- Solutions to the three components of value generalisation (see here) in academically rigorous forms, demonstrations of these solutions on toy problems, public benchmarks, and new benchmarks. If going commercial, applications of this value generalisation to saleable products.
- The first component is recognising that the AI is off-distribution in a value-relevant way.
- The second component is establishing what features should be used to reach a decision while off-distribution this way.
- The third component is to reach a good decision while off distribution, and to loop through the process again to critique and improve the decision, including using outside feedback.
- Combination of these approaches in a successful prealigned learning model.
- If going commercial and it is safe, sale of the model to the public.
- Solutions to the three components of value generalisation (see here) in academically rigorous forms, demonstrations of these solutions on toy problems, public benchmarks, and new benchmarks. If going commercial, applications of this value generalisation to saleable products.
- Intermediate outcomes:
- Development of value generalisation alignment methods for AIs.
- Making these methods available to AI safety researchers.
- Prealigned models being generally available to safely empower users and keep humanity on track for a positive outcome.
- If the approach fails, the failed attempts and the partial successes will be made available to others.
- Impact:
- If failed:
- A seemingly promising approach crossed off the list of possible paths to alignment.
- A better understanding of the issues around the "value generalisation" framing.
- If intermediate outcome:
- A tool or collection of useful tools or models that can improve some AI alignment techniques.
- If successful:
- AI alignment, or a big step forwards towards it.
- An increase in AI capabilities.
- If failed:
This theory of change has been a bit light on specific probability estimates for different assumptions, and for the overall program.
The reason is that, for this approach, the proof of the pudding is in the eating. Now, I've presented solid theoretical arguments for the necessity and usefulness of value generalisation, I can point to my own research track record and intermediate successful results and a plan for the R&D path and how it builds on human generalisation abilities.
I feel the case is compelling, but it rests ultimately on a judgement call: that it makes sense to group all these problems under the heading "value generalisation", to see it all as a single specific goal, and to aim to tackle it directly.
That judgement call may be valid (I certainly feel it is!) and is the true crux in this theory of change. And the best way of figuring out its truth or falsity is to... attempt the project and see what happens, see whether the framework gives swift success or falls apart into disparate disconnected problems.
To check on that, we'll need intermediate benchmarks and assessments. The initial phase of the program (carried out in part while the setup is happening) will include establishing key benchmarks for each subsequent step.
We'll be using public benchmarks for these purposes, though they need to be used with care[2]. A lot of benchmarks get saturated quite easily by models that don't show the true performance the benchmark was supposed to measure. We need algorithms that solve benchmarks via value generalisation, not via any intermediate incomplete shortcuts. The benchmark design can do part of the work, but controlling the information and methods the algorithm can use is also crucial.
Conclusion: value generalisation is a useful path for AI alignmentTo summarise:
- Most AI alignment failures are value generalisation failures.
- Non-decomposability: alignment likely can't be decomposed into smaller simpler parts.
- Thus value generalisation is necessary for alignment, and will make alignment a lot easier.
- Explicitly targeting value generalisation is necessary.
- Unfortunately, value generalisation aids empirical generalisation (a capability) while the converse is not true.
- A corporate structure weakly dominates a research structure for solving value generalisation and for using the results as safely as possible.
- Good design and a good AI ethics board can mitigate the vulnerabilities of the corporate route.
- There is a plan to solve value alignment step by step, focusing first on how agents identify they are out of distribution, then what features to use for deciding in those circumstances, and finally what decision to make and how to assess and improve that decision.
So, let's go and solve this little alignment problem, aye? ^_^
- ^
Any function taking values can be represented by Boolean functions.
- ^
Take the "wolf vs husky" image classification problem, where the wolf images were all taken on snow, causing the classifier to misclassify any light-background image as a wolf.
We had a similar benchmark, "tanks vs forests" where the tank images were taken on a cloudy day (a recreation of an apocryphal but traditional tale in machine learning). The ACE algorithm successfully disambiguated "darkness" from "tankiness".
But a well-trained image recognition model could likely solve both of these benchmarks, just by having enough wolf/dog/tank/forest data that it can resolve these images anyway. That's not out-of-distribution learning, that's increasing the training data until the benchmarks are in-distribution.
Discuss
The separation principle: where do beliefs and desires come from?
TLDR: Psychology, economics, and other disciplines describe agents as systems driven by beliefs and desires. This post argues that the belief-desire view can be derived from classic theorems from optimal control and reinforcement learning. This suggests seeing beliefs and desires as properties of optimal policies rather than as assumptions from folk psychology.
IntroductionOne way to think about agents is as "systems that act for reasons".[1] This compact statement can be interpreted as encapsulating two key implications:
- The notion of action assumes a boundary between agent and environment, so that the former can act on the latter.
- The term reason captures two kinds of internal activity: motivations associated with how to achieve specific goals or outcomes, and beliefs regarding what the agent infers to be the current state of affairs.
In other words, an agent is a well-differentiated system that acts based on beliefs and desires. This view is compatible with perspectives that have been developed by various disciplines:
- Behavioural science, which sees agency as goal-directed behaviour.
- Economics, which treats agency as the ability to select policies to achieve an objective.
- Cybernetics, which conceptualises agency as the ability to regulate the environment and keep it within a desired subset of states.
Instead of taking the working definition[2]
agency = beliefs + desires
at face value, in this post I will discuss conceptual tools that can allow us to derive this construction from formal considerations. In particular, I will discuss the separation principle, which I believe has potential to provide a rigorous foundation for this conceptualisation. These ideas take inspiration from cybernetics and computational mechanics, but distinctly leverage classic results in optimal control theory and reinforcement learning.
Motivation — cybernetics.
A related line of thinking pertains to the (in)famous Good Regulator Theorem. This result, put forward by cyberneticists Conant and Ashby in 1970, suggests that an effective controller requires an internal model of what is being controlled. This makes intuitive sense: it is hard to predict a system that follows an intricate internal mechanism, and perhaps the only way to effectively control it is by understanding it.
Unfortunately, the original paper has been both a source of inspiration and confusion. Lots have been written trying to clarify it; to read more, see this post and this post, and also this manuscript. This post will follow these ideas in spirit, but not the formalisation as proposed there.
Motivation — computational mechanics.
Another related line of work is computational mechanics, which combines principles of information theory, theoretical computer science, and some sprinkles of statistical mechanics to describe the dynamics of prediction. Computational mechanics derives minimal structures required for optimal prediction, which take form in the mjx-container[jax="CHTML"] { line-height: 0; } mjx-container [space="1"] { margin-left: .111em; } mjx-container [space="2"] { margin-left: .167em; } mjx-container [space="3"] { margin-left: .222em; } mjx-container [space="4"] { margin-left: .278em; } mjx-container [space="5"] { margin-left: .333em; } mjx-container [rspace="1"] { margin-right: .111em; } mjx-container [rspace="2"] { margin-right: .167em; } mjx-container [rspace="3"] { margin-right: .222em; } mjx-container [rspace="4"] { margin-right: .278em; } mjx-container [rspace="5"] { margin-right: .333em; } mjx-container [size="s"] { font-size: 70.7%; } mjx-container [size="ss"] { font-size: 50%; } mjx-container [size="Tn"] { font-size: 60%; } mjx-container [size="sm"] { font-size: 85%; } mjx-container [size="lg"] { font-size: 120%; } mjx-container [size="Lg"] { font-size: 144%; } mjx-container [size="LG"] { font-size: 173%; } mjx-container [size="hg"] { font-size: 207%; } mjx-container [size="HG"] { font-size: 249%; } mjx-container [width="full"] { width: 100%; } mjx-box { display: inline-block; } mjx-block { display: block; } mjx-itable { display: inline-table; } mjx-row { display: table-row; } mjx-row > * { display: table-cell; } mjx-mtext { display: inline-block; } mjx-mstyle { display: inline-block; } mjx-merror { display: inline-block; color: red; background-color: yellow; } mjx-mphantom { visibility: hidden; } _::-webkit-full-page-media, _:future, :root mjx-container { will-change: opacity; } mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; font-weight: normal; font-size: 100%; font-size-adjust: none; letter-spacing: normal; border-collapse: collapse; word-wrap: normal; word-spacing: normal; white-space: nowrap; direction: ltr; padding: 1px 0; } mjx-container[jax="CHTML"][display="true"] { display: block; text-align: center; margin: 1em 0; } mjx-container[jax="CHTML"][display="true"][width="full"] { display: flex; } mjx-container[jax="CHTML"][display="true"] mjx-math { padding: 0; } mjx-container[jax="CHTML"][justify="left"] { text-align: left; } mjx-container[jax="CHTML"][justify="right"] { text-align: right; } mjx-mi { display: inline-block; text-align: left; } mjx-c { display: inline-block; } mjx-utext { display: inline-block; padding: .75em 0 .2em 0; } mjx-mo { display: inline-block; text-align: left; } mjx-stretchy-h { display: inline-table; width: 100%; } mjx-stretchy-h > * { display: table-cell; width: 0; } mjx-stretchy-h > * > mjx-c { display: inline-block; transform: scalex(1.0000001); } mjx-stretchy-h > * > mjx-c::before { display: inline-block; width: initial; } mjx-stretchy-h > mjx-ext { /* IE */ overflow: hidden; /* others */ overflow: clip visible; width: 100%; } mjx-stretchy-h > mjx-ext > mjx-c::before { transform: scalex(500); } mjx-stretchy-h > mjx-ext > mjx-c { width: 0; } mjx-stretchy-h > mjx-beg > mjx-c { margin-right: -.1em; } mjx-stretchy-h > mjx-end > mjx-c { margin-left: -.1em; } mjx-stretchy-v { display: inline-block; } mjx-stretchy-v > * { display: block; } mjx-stretchy-v > mjx-beg { height: 0; } mjx-stretchy-v > mjx-end > mjx-c { display: block; } mjx-stretchy-v > * > mjx-c { transform: scaley(1.0000001); transform-origin: left center; overflow: hidden; } mjx-stretchy-v > mjx-ext { display: block; height: 100%; box-sizing: border-box; border: 0px solid transparent; /* IE */ overflow: hidden; /* others */ overflow: visible clip; } mjx-stretchy-v > mjx-ext > mjx-c::before { width: initial; box-sizing: border-box; } mjx-stretchy-v > mjx-ext > mjx-c { transform: scaleY(500) translateY(.075em); overflow: visible; } mjx-mark { display: inline-block; height: 0px; } mjx-munder { display: inline-block; text-align: left; } mjx-over { text-align: left; } mjx-munder:not([limits="false"]) { display: inline-table; } mjx-munder > mjx-row { text-align: left; } mjx-under { padding-bottom: .1em; } mjx-TeXAtom { display: inline-block; text-align: left; } mjx-msub { display: inline-block; text-align: left; } mjx-msup { display: inline-block; text-align: left; } mjx-mn { display: inline-block; text-align: left; } mjx-mspace { display: inline-block; text-align: left; } mjx-mover { display: inline-block; text-align: left; } mjx-mover:not([limits="false"]) { padding-top: .1em; } mjx-mover:not([limits="false"]) > * { display: block; text-align: left; } mjx-munderover { display: inline-block; text-align: left; } mjx-munderover:not([limits="false"]) { padding-top: .1em; } mjx-munderover:not([limits="false"]) > * { display: block; } mjx-msubsup { display: inline-block; text-align: left; } mjx-script { display: inline-block; padding-right: .05em; padding-left: .033em; } mjx-script > mjx-spacer { display: block; } mjx-c::before { display: block; width: 0; } .MJX-TEX { font-family: MJXZERO, MJXTEX; } .TEX-B { font-family: MJXZERO, MJXTEX-B; } .TEX-I { font-family: MJXZERO, MJXTEX-I; } .TEX-MI { font-family: MJXZERO, MJXTEX-MI; } .TEX-BI { font-family: MJXZERO, MJXTEX-BI; } .TEX-S1 { font-family: MJXZERO, MJXTEX-S1; } .TEX-S2 { font-family: MJXZERO, MJXTEX-S2; } .TEX-S3 { font-family: MJXZERO, MJXTEX-S3; } .TEX-S4 { font-family: MJXZERO, MJXTEX-S4; } .TEX-A { font-family: MJXZERO, MJXTEX-A; } .TEX-C { font-family: MJXZERO, MJXTEX-C; } .TEX-CB { font-family: MJXZERO, MJXTEX-CB; } .TEX-FR { font-family: MJXZERO, MJXTEX-FR; } .TEX-FRB { font-family: MJXZERO, MJXTEX-FRB; } .TEX-SS { font-family: MJXZERO, MJXTEX-SS; } .TEX-SSB { font-family: MJXZERO, MJXTEX-SSB; } .TEX-SSI { font-family: MJXZERO, MJXTEX-SSI; } .TEX-SC { font-family: MJXZERO, MJXTEX-SC; } .TEX-T { font-family: MJXZERO, MJXTEX-T; } .TEX-V { font-family: MJXZERO, MJXTEX-V; } .TEX-VB { font-family: MJXZERO, MJXTEX-VB; } mjx-stretchy-v mjx-c, mjx-stretchy-h mjx-c { font-family: MJXZERO, MJXTEX-S1, MJXTEX-S4, MJXTEX, MJXTEX-A ! important; } @font-face /* 0 */ { font-family: MJXZERO; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Zero.woff") format("woff"); } @font-face /* 1 */ { font-family: MJXTEX; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Regular.woff") format("woff"); } @font-face /* 2 */ { font-family: MJXTEX-B; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Bold.woff") format("woff"); } @font-face /* 3 */ { font-family: MJXTEX-I; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-Italic.woff") format("woff"); } @font-face /* 4 */ { font-family: MJXTEX-MI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Italic.woff") format("woff"); } @font-face /* 5 */ { font-family: MJXTEX-BI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-BoldItalic.woff") format("woff"); } @font-face /* 6 */ { font-family: MJXTEX-S1; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size1-Regular.woff") format("woff"); } @font-face /* 7 */ { font-family: MJXTEX-S2; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size2-Regular.woff") format("woff"); } @font-face /* 8 */ { font-family: MJXTEX-S3; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size3-Regular.woff") format("woff"); } @font-face /* 9 */ { font-family: MJXTEX-S4; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size4-Regular.woff") format("woff"); } @font-face /* 10 */ { font-family: MJXTEX-A; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_AMS-Regular.woff") format("woff"); } @font-face /* 11 */ { font-family: MJXTEX-C; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Regular.woff") format("woff"); } @font-face /* 12 */ { font-family: MJXTEX-CB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Bold.woff") format("woff"); } @font-face /* 13 */ { font-family: MJXTEX-FR; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Regular.woff") format("woff"); } @font-face /* 14 */ { font-family: MJXTEX-FRB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Bold.woff") format("woff"); } @font-face /* 15 */ { font-family: MJXTEX-SS; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Regular.woff") format("woff"); } @font-face /* 16 */ { font-family: MJXTEX-SSB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Bold.woff") format("woff"); } @font-face /* 17 */ { font-family: MJXTEX-SSI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Italic.woff") format("woff"); } @font-face /* 18 */ { font-family: MJXTEX-SC; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Script-Regular.woff") format("woff"); } @font-face /* 19 */ { font-family: MJXTEX-T; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Typewriter-Regular.woff") format("woff"); } @font-face /* 20 */ { font-family: MJXTEX-V; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Regular.woff") format("woff"); } @font-face /* 21 */ { font-family: MJXTEX-VB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Bold.woff") format("woff"); } mjx-c.mjx-c1D716.TEX-I::before { padding: 0.431em 0.406em 0.011em 0; content: "\3F5"; } mjx-c.mjx-c1D453.TEX-I::before { padding: 0.705em 0.55em 0.205em 0; content: "f"; } mjx-c.mjx-c28::before { padding: 0.75em 0.389em 0.25em 0; content: "("; } mjx-c.mjx-c1D465.TEX-I::before { padding: 0.442em 0.572em 0.011em 0; content: "x"; } mjx-c.mjx-c29::before { padding: 0.75em 0.389em 0.25em 0; content: ")"; } mjx-c.mjx-c3D::before { padding: 0.583em 0.778em 0.082em 0; content: "="; } mjx-c.mjx-c1D454.TEX-I::before { padding: 0.442em 0.477em 0.205em 0; content: "g"; } mjx-c.mjx-c2B::before { padding: 0.583em 0.778em 0.082em 0; content: "+"; } mjx-c.mjx-c210E.TEX-I::before { padding: 0.694em 0.576em 0.011em 0; content: "h"; } mjx-c.mjx-c6D::before { padding: 0.442em 0.833em 0 0; content: "m"; } mjx-c.mjx-c69::before { padding: 0.669em 0.278em 0 0; content: "i"; } mjx-c.mjx-c6E::before { padding: 0.442em 0.556em 0 0; content: "n"; } mjx-c.mjx-c7B.TEX-S1::before { padding: 0.85em 0.583em 0.349em 0; content: "{"; } mjx-c.mjx-c7D.TEX-S1::before { padding: 0.85em 0.583em 0.349em 0; content: "}"; } mjx-c.mjx-c2265::before { padding: 0.636em 0.778em 0.138em 0; content: "\2265"; } mjx-c.mjx-c1D461.TEX-I::before { padding: 0.626em 0.361em 0.011em 0; content: "t"; } mjx-c.mjx-c2208::before { padding: 0.54em 0.667em 0.04em 0; content: "\2208"; } mjx-c.mjx-c211D.TEX-A::before { padding: 0.683em 0.722em 0 0; content: "R"; } mjx-c.mjx-c1D45B.TEX-I::before { padding: 0.442em 0.6em 0.011em 0; content: "n"; } mjx-c.mjx-c1D466.TEX-I::before { padding: 0.442em 0.49em 0.205em 0; content: "y"; } mjx-c.mjx-c1D45A.TEX-I::before { padding: 0.442em 0.878em 0.011em 0; content: "m"; } mjx-c.mjx-c1D462.TEX-I::before { padding: 0.442em 0.572em 0.011em 0; content: "u"; } mjx-c.mjx-c1D451.TEX-I::before { padding: 0.694em 0.52em 0.01em 0; content: "d"; } mjx-c.mjx-c2115.TEX-A::before { padding: 0.683em 0.722em 0.02em 0; content: "N"; } mjx-c.mjx-c31::before { padding: 0.666em 0.5em 0 0; content: "1"; } mjx-c.mjx-c1D434.TEX-I::before { padding: 0.716em 0.75em 0 0; content: "A"; } mjx-c.mjx-c1D435.TEX-I::before { padding: 0.683em 0.759em 0 0; content: "B"; } mjx-c.mjx-c1D464.TEX-I::before { padding: 0.443em 0.716em 0.011em 0; content: "w"; } mjx-c.mjx-c1D436.TEX-I::before { padding: 0.705em 0.76em 0.022em 0; content: "C"; } mjx-c.mjx-c1D463.TEX-I::before { padding: 0.443em 0.485em 0.011em 0; content: "v"; } mjx-c.mjx-c2C::before { padding: 0.121em 0.278em 0.194em 0; content: ","; } mjx-c.mjx-c30::before { padding: 0.666em 0.5em 0.022em 0; content: "0"; } mjx-c.mjx-c223C::before { padding: 0.367em 0.778em 0 0; content: "\223C"; } mjx-c.mjx-c4E.TEX-C::before { padding: 0.789em 0.979em 0.05em 0; content: "N"; } mjx-c.mjx-cAF::before { padding: 0.59em 0.5em 0 0; content: "\AF"; } mjx-c.mjx-c3A3::before { padding: 0.683em 0.722em 0 0; content: "\3A3"; } mjx-c.mjx-c1D44A.TEX-I::before { padding: 0.683em 1.048em 0.022em 0; content: "W"; } mjx-c.mjx-c1D449.TEX-I::before { padding: 0.683em 0.769em 0.022em 0; content: "V"; } mjx-c.mjx-c2026::before { padding: 0.12em 1.172em 0 0; content: "\2026"; } mjx-c.mjx-c1D70B.TEX-I::before { padding: 0.431em 0.57em 0.011em 0; content: "\3C0"; } mjx-c.mjx-c2212::before { padding: 0.583em 0.778em 0.082em 0; content: "\2212"; } mjx-c.mjx-c1D43D.TEX-I::before { padding: 0.683em 0.633em 0.022em 0; content: "J"; } mjx-c.mjx-c1D53C.TEX-A::before { padding: 0.683em 0.667em 0 0; content: "E"; } mjx-c.mjx-c5B.TEX-S2::before { padding: 1.15em 0.472em 0.649em 0; content: "["; } mjx-c.mjx-c2211.TEX-S1::before { padding: 0.75em 1.056em 0.25em 0; content: "\2211"; } mjx-c.mjx-c1D447.TEX-I::before { padding: 0.677em 0.704em 0 0; content: "T"; } mjx-c.mjx-c1D70F.TEX-I::before { padding: 0.431em 0.517em 0.013em 0; content: "\3C4"; } mjx-c.mjx-c28.TEX-S1::before { padding: 0.85em 0.458em 0.349em 0; content: "("; } mjx-c.mjx-c22A4::before { padding: 0.668em 0.778em 0 0; content: "\22A4"; } mjx-c.mjx-c1D444.TEX-I::before { padding: 0.704em 0.791em 0.194em 0; content: "Q"; } mjx-c.mjx-c1D445.TEX-I::before { padding: 0.683em 0.759em 0.021em 0; content: "R"; } mjx-c.mjx-c29.TEX-S1::before { padding: 0.85em 0.458em 0.349em 0; content: ")"; } mjx-c.mjx-c1D446.TEX-I::before { padding: 0.705em 0.645em 0.022em 0; content: "S"; } mjx-c.mjx-c5D.TEX-S2::before { padding: 1.15em 0.472em 0.649em 0; content: "]"; } mjx-c.mjx-c2217::before { padding: 0.465em 0.5em 0 0; content: "\2217"; } mjx-c.mjx-c1D43E.TEX-I::before { padding: 0.683em 0.889em 0 0; content: "K"; } mjx-c.mjx-c5E::before { padding: 0.694em 0.5em 0 0; content: "^"; } mjx-c.mjx-c5B.TEX-S1::before { padding: 0.85em 0.417em 0.349em 0; content: "["; } mjx-c.mjx-c7C::before { padding: 0.75em 0.278em 0.249em 0; content: "|"; } mjx-c.mjx-c5D.TEX-S1::before { padding: 0.85em 0.417em 0.349em 0; content: "]"; } mjx-c.mjx-c1D460.TEX-I::before { padding: 0.442em 0.469em 0.01em 0; content: "s"; } mjx-c.mjx-c53.TEX-C::before { padding: 0.705em 0.642em 0.022em 0; content: "S"; } mjx-c.mjx-c1D45C.TEX-I::before { padding: 0.441em 0.485em 0.011em 0; content: "o"; } mjx-c.mjx-c4F.TEX-C::before { padding: 0.705em 0.796em 0.022em 0; content: "O"; } mjx-c.mjx-c1D44E.TEX-I::before { padding: 0.441em 0.529em 0.01em 0; content: "a"; } mjx-c.mjx-c41.TEX-C::before { padding: 0.728em 0.819em 0.05em 0; content: "A"; } mjx-c.mjx-c1D705.TEX-I::before { padding: 0.442em 0.576em 0.011em 0; content: "\3BA"; } mjx-c.mjx-c3A::before { padding: 0.43em 0.278em 0 0; content: ":"; } mjx-c.mjx-cD7::before { padding: 0.491em 0.778em 0 0; content: "\D7"; } mjx-c.mjx-c2192::before { padding: 0.511em 1em 0.011em 0; content: "\2192"; } mjx-c.mjx-c394::before { padding: 0.716em 0.833em 0 0; content: "\394"; } mjx-c.mjx-c1D708.TEX-I::before { padding: 0.442em 0.53em 0 0; content: "\3BD"; } mjx-c.mjx-c22C5::before { padding: 0.31em 0.278em 0 0; content: "\22C5"; } mjx-c.mjx-c1D45D.TEX-I::before { padding: 0.442em 0.503em 0.194em 0; content: "p"; } mjx-c.mjx-c220F.TEX-S1::before { padding: 0.75em 0.944em 0.25em 0; content: "\220F"; } mjx-c.mjx-c1D45F.TEX-I::before { padding: 0.442em 0.451em 0.011em 0; content: "r"; } mjx-c.mjx-c221E::before { padding: 0.442em 1em 0.011em 0; content: "\221E"; } mjx-c.mjx-c1D6FE.TEX-I::before { padding: 0.441em 0.543em 0.216em 0; content: "\3B3"; } mjx-c.mjx-c5B::before { padding: 0.75em 0.278em 0.25em 0; content: "["; } mjx-c.mjx-c5D::before { padding: 0.75em 0.278em 0.25em 0; content: "]"; } mjx-c.mjx-c1D44F.TEX-I::before { padding: 0.694em 0.429em 0.011em 0; content: "b"; } mjx-c.mjx-c1D719.TEX-I::before { padding: 0.694em 0.596em 0.205em 0; content: "\3D5"; } mjx-c.mjx-c58.TEX-C::before { padding: 0.683em 0.807em 0 0; content: "X"; } -machine and -transducer. These ideas offer a rigorous explanation to (i) why optimal prediction requires specific computations, and (ii) what properties those computations need to satisfy. To read more, see this thesis or this paper.
Figure adapted from Crutchfield & Feldman, Regularities unseen, randomness observed: levels of entropy convergence, arXiv:cond-mat/0102181
More recently, computational mechanics is providing effective tools to study the internal representations in deep learning. Moreover, while originally formulated to study prediction, computational mechanics also provides a promising foundation for understanding perception-action loops. The ideas developed here are heavily influenced by both the concepts and formalism of computational mechanics.
What is a separation principle?Let's start by discussing what separation principles are, in general.
As a first approximation, consider a situation where we want to minimise the function . A direct calculation shows that
.
Hence, solving the problem for and separately does not yield the solution for . Put simply, one does not solve the problem by solving its parts separately.
A separation principle takes place when the divide-and-conquer strategy actually works, so one can solve a complicated optimisation problem by optimising various subcomponents separately. Concretely, a separation theorem guarantees that optimal performance can be achieved by solving sub-tasks in isolation.
Let me further illustrate this idea by presenting an important separation theorem at the core of information theory.
The communication problem studied by information theory is about how to send bits from source to destination without error. The tools that a communication engineer employs to face this problem are two types of coding techniques:
- Source coding, which compresses the original signal by removing redundancy in order to transmit as few bits as necessary.[3]
- Channel coding corresponds to error-correcting codes, which adds redundancy to protect the information bits from noise.[4]
Most people have heard of the following results by Claude Shannon:
- The source coding theorem, which states that the shortest achievable average code length is given by the entropy of the source.
- The noisy-channel coding theorem, which states that the maximal transmission rate that allows arbitrarily low decoding error is given by the mutual information between input and output symbols (maximised over input distributions).
However, these results would not be as useful in the absence of the third, less known result known as the source-channel separation theorem.[5] This result states that, under relatively mild conditions,[6] an optimal solution for the overall communication problem can be achieved by breaking the problem into two separate parts: compression and error correction. This means that an optimal communication strategy can be designed by a team separately working on compression and another working on error-correction, without requiring coordination between them. Needless to say, this makes the problem much more tractable!
The inference-control separation principleLet us now review results from optimal control theory and reinforcement learning that provide conditions for separating inference and control. In doing this, I will use the non-standard term inference-control separation principle to group together results that are often presented independently of each other, despite having a lot in common.
Separation principle in optimal control theoryLet us consider the classic instantiation of separation in the context of optimal control theory. For this, consider a control setting made by three components:
- a system to be controlled, whose state is described by ,
- the measurements an agent has access to, described by , and
- the actions .
For simplicity, we assume that the system evolves dynamically over discrete time via linear dynamics and noisy observations that can be described as
,
,
where and are matrices, is the system's initial condition, and and are independent white Gaussian noise terms.
The action is determined according to a policy that has no direct access to the state , but only to the sequence of measurements . Thus, we take to be a deterministic function of the sequence of past measurements and actions . The controller aims to minimise the quadratic cost
,
where , , and are matrices and is the length of the time window over which the optimisation takes place. This is known as the discrete-time LQG problem, as it involves Linear dynamics, a Quadratic loss function, and Gaussian variables — being the one setting where everything is exactly solvable.
Separation principle — To understand the solution of this problem, let us first consider the case of full observability, where . In this setting, the optimal controller can be shown to only depend on the current state of the system, and the optimal solution can be expressed as
,
where is the "optimal gain" matrix (whose formula can be found in standard textbooks). So, if the controller can directly see the state of the system, this directly tells what the optimal action should be.
What happens with partial observability? In this setting, the optimal controller can be written as
,
where is the same as above and is an estimation of the state of the system given by
.
In fact, is the result of Kalman filtering, which is how recursive Bayesian filtering looks like under Gaussian dynamics.[7]
In summary, the optimal policy can be described as "estimate first, and then act as if your estimate were the actual state of the system". Thus, a separation theorem is taking place separating inference and control.
Separation principle in reinforcement learningTo understand what the separation principle looks like in reinforcement learning, let us slightly modify the nomenclature and notation:
- the state of the system to be controlled is described by a discrete variable ,
- the measurements that the agent has access to are described by , and
- the actions are denoted by .
Let's assume that the variables follow a partially observed Markov decision process (POMDP), so that the environment dynamics are Markovian conditioned on the agent's action via a Markov kernel ,[8] and the current observation depends on the current state of environment via a Markov kernel . The agent takes actions following a policy that maps sequences of previous actions and observations into distribution of next actions.
In this setting, the probability of observing a sequence of observations and actions is
,
where is a prior distribution over the environment's initial state. This is roughly the same as the control scenario studied above, but considering general discrete variables instead of Gaussian ones.
Instead of a loss function, consider a reward that the agent tries to maximise. More specifically, let us define the discounted future reward
,
where is the so-called discount factor. Then, the goal of the agent is to find a policy that maximises .
Separation principle — To solve this POMDP, let us first consider the Bayesian beliefs about the latent state . There are two ways to understand Bayesian beliefs:
- As random variables given by the conditional expectation , where the trajectory is regarded as a random variable.
- As probability distributions given by , which correspond to the above expectation when conditioned on a specific realisation of .
Bayesian beliefs have two key properties. First, as random variables, they contain all the information in that is useful for predicting (i.e., they are sufficient statistics). Second, as distributions, they can be efficiently updated via
,
where is a deterministic function, and hence the dynamics of are Markovian conditioned on (as they only depend on ). These facts can be used to build a new Markov decision process (MDP), known as belief MDP, defined by the following variables:
- the states are the beliefs of the latent state of the POMDP,[9]
- actions are (same as in the POMDP),
- the reward function is .
In contrast with the original problem, this MDP is fully observable (as the agent knows its own beliefs), and therefore can be solved using standard RL techniques. The crucial point is that combining the estimation of Bayesian beliefs with the optimal solution for the belief MDP gives rise to a solution of the full POMDP.[10] Thus, a separation principle is at play separating inference (calculating beliefs) and control (solving the belief MDP).
Interim summaryBoth results in LQG control and POMDPs give rise to the same solution: they show that the optimal controller for a partially observed system can be built by
- first performing optimal prediction over the latent state of the system, and then
- solving the control problem using that prediction as a state.
To further understand how the LQG and POMDP settings relate, let's discuss the shape of beliefs and the effect of actions (not essential for the rest of the post).
Beliefs are distributions, not point estimates.
In the LQG setting, the optimal controller does not need the full distribution of the prediction — the prediction is a Gaussian posterior, but the controller only uses the mean and not the variance. This property is known as certainty equivalence, which suggests to act as if a point estimate were the true state.
It would be tempting to describe beliefs as "best guesses" (e.g. maximum a posteriori), but certainty equivalence only holds because of the friendly properties of the LQG setup — a more general setting would not allow for this simplification.
In the more general POMDP case, the agent must carry the full belief distribution, not a point estimate. Nonetheless, separation survives in the Bayesian filter + belief-MDP form. Thus, "beliefs" in the agency sense are not mere best guesses but the full posterior, and only under special conditions do they collapse to a point estimate. More generally, the separation principle suggests to think of an agent's epistemic state not as what it thinks is true, but how its uncertainty is shaped.
The dual effect.
Actions do two things at once: they change the state of the system (control effect) and also change what the agent will subsequently be able to observe and thus learn (informational effect). Think of a robot with a noisy range sensor deciding whether to drive straight toward a goal or take a detour past a landmark: the detour is suboptimal for the control objective in isolation, but it sharpens the robot's position estimate, which improves every subsequent decision. This second channel — action shaping future belief — is what control theorists call dual effect.
Part of why LQG separation is so clean is that while the evolution of the Kalman mean estimate depends on past actions and observations, its covariance does not. This means that the agent's uncertainty about the state at any future time is fixed the moment the problem is specified, as no action-observation sequence can make it larger or smaller. This implies that actions have no informational effect: there is no possible epistemic bonus that can reduce uncertainty, no matter what the agent does. This also explains why certainty equivalence holds: given that the covariance of the posterior evolves in a predetermined manner, it carries no useful information for the controller.
This simplification does not hold in more general POMDP settings, where actions can influence future information as well as the environment. In the belief MDP construction, the "state" is the belief , and the transition kernel captures both effects at once: how the action moves the underlying environment and how it reshapes the posterior.
ImplicationsLet us now return to the question we started with: can the working definition
agency = beliefs + desires
be derived, rather than assumed? The separation theorems reviewed above suggest that, at least in a specific sense, the answer is yes.
Beliefs and desires as properties of solutionsLet's recapitulate what the separation principle delivers. We posed a single, monolithic optimisation problem: find a policy mapping histories into actions that minimise/maximise loss/reward. Nothing in this problem statement mentions beliefs, inference, or internal representations — the problem is stated in terms of pure behaviour. Yet, the solution naturally factorises into sub-components. Indeed, the above results state that the optimal policy can always be realised as the composition of two blocks:
- An inference module, which compresses the history into a Bayesian belief about the latent state that can be recursively updated. This module knows about dynamics but is oblivious about loss/rewards.
- A control module, which selects actions as a function of that belief alone. This module knows about loss/rewards but not about raw observations.
The interface between them is exactly the belief state, which is the output of the inference module and the input of the control module. Thus, the inference module can be interpreted as "where beliefs are made"; similarly, the control module can be interpreted as "where desires take place", as it is the only part of the policy that is directed by the reward/loss function. In this way, beliefs and desires are recovered not as assumptions about agents but as properties of solutions.
This is worth pausing on. The folk-psychological notions of belief and desire have been with us at least since Aristotle, and they are usually treated as primitives — as the ingredients one starts with when theorising about minds. What the separation principle offers is something different: not a definition of these terms but a derivation of them. Beliefs and desires appear here not because we built them in, but because the optimal solution to a problem stated in purely behavioural terms turns out to have that shape. Whether or not one takes folk psychology seriously as a theory of mind, this is at least a reason to think its central distinction is not arbitrary.
It is remarkable that this modularity arises from solving a single, monolithic optimisation problem. We have seen this shape before: it is precisely the payoff that made the source-channel separation theorem so valuable! There, an optimal communication system could be designed by a compression team and an error-correction team working independently, without coordination — the source coder needed to know nothing about the channel, and the channel coder needed to know nothing about the source.
The inference-control separation delivers the exact analogue for agency, conferring a form of compositional generalisation. Indeed, if the reward changes but the environment does not, only the control module needs to be updated; if the environment shifts but the goals remain, only the inference module needs revising. Thus, the two components can be developed, improved, and swapped independently.
Agents as cognitive light-conesThese results may also provide some formal ground to conceiving of agents as cognitive light-cones.[11] The idea goes as follows: an agent is a system whose behaviour is both sensitive to what it can remember about the past (the past component of the cone) and what it can foresee (the future component of the cone).
Figure from Levin (2019), Frontiers in psychology, 10, 2688.
This idea suggests an operational way to assess the capabilities of an agent: measure the size of its light-cone. This could be done using tools from computational mechanics — for example, studying how deep into the past causal states go, and doing the same for retrodictive causal states (which build sufficient statistics from future to past).[12]
The separation principle is normative, not descriptivePlease note that the separation principle says that an optimal agent can be built as inference + control, but it does not say that any given agent is built this way. In particular, a policy trained end-to-end by gradient descent is under no obligation to organise itself into a filter feeding a planner.
But the separation principle gives us something subtler: it tells us that the belief state is a sufficient statistic for optimal behaviour. This suggests that any agent approaching optimality must be computing something as informative as the Bayesian posterior (whether or not it is legible as such). In this way, the separation principle turns into a lens for interpretability: we should expect belief-like structure inside competent agents, and we can go looking for it. Recent work from colleagues at Simplex in finding belief-state geometry in the residual streams of transformers is, I think, an early vindication of this expectation. At XOR Labs we are currently investigating to what degree these results transfer to deep RL agents trained to do control tasks.
At this point, it is useful to distinguish two separate concerns:
- Control problems / POMDPs can have multiple optimal policies. In these settings, the separation principle states that there exists one optimal policy that can be factored, but doesn't guarantee all do.[13]
- When only one optimal solution exists, the separation principle says this unique behaviour admits a realisation as filter + belief-controller — but it does not say that's the only realisation. Indeed, the same mapping from histories to next action can be implemented in different ways; for instance, in a finite setting the mapping can always be implemented as a large look-up table.
Regarding the second concern, there are various pieces of evidence suggesting that deep learning architectures have an implicit bias for simplicity, which makes them find simpler implementations of algorithms whenever they exist.[14] Therefore, while it is often possible to implement a policy via a gigantic look-up table, there are reasons to believe that policies obtained via deep learning techniques will tend towards factorised ones.[15]
Where does separation fail? As Shannon's source-channel separation theorem, the inference-control separation theorems hold under fairly general conditions — not relying on finite alphabets or stationarity. That said, for it to be useful one needs to assume that the agents under consideration have enough compute to build (quasi-)optimal policies. Under bounded rationality, model misspecification, or multi-agent interaction, the clean factorisation can break: what to represent starts depending on what you want, and inference becomes goal-driven.[16] Many of the most interesting aspects of agency may come from the ways in which real agents deviate from the separation principle — but that is a topic for a future post.
CodaBefore concluding, let me ask: does the separation principle derive beliefs and desires from first principles?
Concretely, the separation principle offers the following:
- It reveals that optimal policies can be factored into two separate modules.
- The signal shared between the modules can be interpreted as beliefs, so that the first (inference) module can be interpreted as "where beliefs are made".
- The second (control) module can be interpreted as "where desires take place", as it is the only part of the policy that is directed by the reward/loss function.
Thus, the separation principle provides formal ground for articulating the view of agents as "systems that act for reasons", displaying behaviour that is sensitive to both the past (what it remembers as part of its beliefs) and the future (due to foreseeing and planning to achieve goals).
However, one has to acknowledge that the separation principle assumes an optimisation problem which already contains a reward or cost function. Following this reasoning, one could argue that the desire or preference has not been really derived, as it has been baked in externally. More precisely, one can interpret the separation principle not as deriving why desires exist in the first place, but as explaining why they can be treated as separate entities from beliefs.[17]
- ^
There is a huge literature discussing various views on what agency is; if you want to read more, perhaps take a look into this and this paper and references therein.
- ^
This post will take the boundary between agent and environment as given to focus on their internal activity. To read about boundaries, see these posts.
- ^
Examples include Huffman and arithmetic coding.
- ^
Examples include Hamming, Reed-Solomon, or convolutional codes.
- ^
In contrast with the other two, this result doesn't have a wikipedia page...
- ^
For technical details, see Cover & Thomas Chapter 7.13.
- ^
For more details about the separation principle in LQR, see this paper. For a good introduction to Kalman and Bayesian filtering, see this book.
- ^
denotes the space of probability distributions over the set .
- ^
The transition probabilities can be computed directly; for formulas see this paper or wikipedia.
- ^
For proofs, see this and this paper.
- ^
I first saw this idea in this inspiring paper of Michael Levin.
- ^
I'll expand on this in a future post.
- ^
This is a real issue, but can be disregarded in some settings. For instance, "almost every" finite POMDP has a single optimal solution. Multiple optimal solutions in POMDPs are caused, for example, by ties in the optimal Q-function, which are broken by arbitrarily small perturbations on the reward.
- ^
There have been numerous discussions about the simplicity biases of deep learning, see for example this post.
- ^
For example, we have evidence that transformers find factored representations whenever they are available — see this paper.
- ^
This is arguably where phenomena like motivated reasoning and wishful thinking become rational responses to bounded resources rather than mere biases. To read more, see this and this.
- ^
For answering the why question, one may need to consider arguments based on natural selection, or perhaps approaches to agency based on the enactive tradition — as developed, for example, in this paper. See related discussions questioning utility functions in this post.
Discuss