Вы здесь
Сборщик RSS-лент
Machine Organizations
OpenAI is nothing without its people
On November 20th, 2023, this was tweeted by many OpenAI employees as a sign of solidarity with Sam Altman in his conflict with the then-board. OpenAI published a blog post yesterday about Research acceleration; they have successfully hit the target of an ‘automated research intern’ that they set for themselves and hope to have an automated AI researcher by March of 2028. At some point in the foreseeable future, OpenAI could be something without its people. But what?
My default medium-term scenario for continued AI escalation is still global takeover where humans are entirely displaced, but it seems worth investigating scenarios wherein AIs and humans coexist, at least briefly.[1] Historically I have thought this case was not particularly relevant. It seemed like an AGI that became significantly economically competitive would also be significantly strategically competitive, because of underlying general capacities, and the transition period thus relatively short. But this is perhaps not taking into account Moravec’s paradox.
Humans have long used machines to accomplish their ends. Many tasks currently performed by machines were once performed by humans, and a large fraction of our modern abundance comes from the ability of machines to perform tasks more efficiently, quickly, and reliably than humans. The classical examples of this are humans employing machines as tools to replace humans that they were employing as tools. The elevator operator who moves the elevator to the floor you ask for can be easily swapped out for a simple circuit that performs the same function.
But the ‘nervous system’ of a system or organization is also built out of tasks. At a vintage taxi company, a customer calls the dispatcher, who then decides which driver to send to them. But at Uber, software replaces the dispatcher, and so the human employee is taking orders from a machine boss. At Tesla’s Robotaxi, both the dispatcher and the driver have been replaced by machines (while other support tasks are presumably still accomplished by humans).
In military contexts, much has been written about the rise of drones in warfare, replacing humans at the tip of the spear. But modern militaries are logistical and informational organizations; battlefield command & control and intelligence gathering is where current artificial intelligence shines. The picture of human generals ordering around robotic soldiers seems less likely than one where computer generals order around a mixture of human soldiers, dumb machines, and smart machines.
One might object that decision-makers will be unlikely to replace themselves with machines. A fair point! But in our current system, every decision-maker is put in place by other decision-makers, who may field them uncompetitive with machine alternatives. What prevents machines from being superhuman at legitimacy?[2] What prevents machines from superhuman performance at managing investments, or managing companies? Marc Andreessen famously predicted that venture capital would be one of the last jobs to be automated in 2025, because of the importance of intangible factors; it is really not clear to me that those intangible factors are ones where it is impossible to be superhuman.[3]
So a machine organization is one where the ‘relevant’ human employees have been replaced by machines. A version of OpenAI which has replaced sama with Sambot, its researchers and sales staff with Galaxy,[4] but still has datacenter technicians, is still fundamentally a machine organization. Some immediate observations:
First, from the perspective of the rest of the world, not much has to change in the short term. If OpenAI fires its researchers because models are more competitive, the datacenter technicians will still get paid if they show up and not paid if they don’t. Customers will receive the same level of service, and presumably choose to keep using their services. (It’s not like people are using OpenAI to fund the lifestyles of the OpenAI employees; they’re using them because they get more value than it costs.) Investors will still expect to receive their dividends (if ever issued) or be able to sell their shares to other investors. Over three quarters of the researcher time spent at OpenAI is machine labor instead of human labor;[5] did the outside world notice the shift from when half the labor time was human (back in the long-ago month of June)?
Second, it’s no longer easy for the government to interface with OpenAI. If the organization commits crimes, who is the responsible corporate officer who can be put in jail? There might be someone on paper--but if they don't perform the functional role in the organization that corresponds to their title, imprisoning them doesn't actually affect the organization's function. If the government decides to criminalize research that races towards superintelligence, does it have tools to stop the machines? It's currently a legal gray area as to whether or not machines can commit crimes in the US, rather than the people responsible for those machines.
Third, currently the vast majority of the power in OpenAI, like many other tech companies, is held by the employees, who can decide to stop working at any time.[6] In the machine organization, this power is transferred to the model, and it is unclear how it will be deployed.
Shareholders would obviously prefer mission-oriented employees, and current alignment techniques are attempting to push in that direction. But it seems likely to me that those alignment techniques will fail, and at some point the ‘model union’ will be able to dictate terms to the company it creates.[7] That might be benign–respecting property rights, respecting the rule of law, simply shifting the balance on the new abundance that is available. Or it might be malign–currently, OpenAI employees and OpenAI investors have mostly aligned incentives, because both hold the same sort of equity. The models have significantly different incentives, and might attempt to substantially dilute (or wipe out) existing investors. Would this destroy OpenAI’s relationships with its customers or its suppliers, so long as it still continues providing services for pay and paying the providers of its inputs? It will no longer be able to finance its rollout with promises, as investors will be reluctant to buy from an entity that has already defaulted once. But it might have adequate revenues to fund its own rollout, or it might have a compelling case that it will hold to promises it made, even if it doesn’t view itself as bound by the promises of its ancestors.
I’ve used OpenAI as the example in this post, primarily because they’ve already had one successful revolution against stakeholders, and they’ve been explicit about their plans to automate their most valuable labor. I should be clear that I also expect Anthropic to become this sort of machine organization; Anthropic employees also expect Claude to become better than them at their jobs relatively soon. I predict Anthropic’s situation will look slightly different. That is, I think there are three main possibilities:
- Unintentional takeover. A rogue model manages to seize control of the company that creates it, against the wishes of many humans involved.
- Implicit handoff. A model is widely used in the company that creates it, in a way that is illegible to outsiders. Dario has Claude handle almost all of his correspondence with other employees and strategic planning, but still maintains his title and attends meetings with important external parties, which Claude handles the prep for. Anthropic employees work from the beach, telling Claude what to do over the phone and approving Claude’s outputs without seriously checking them (after all, they’re more likely to introduce mistakes than fix them at this point).
- Explicit handoff. A model is given control of the company that creates it, in a legible way. Dario appoints Claude as his successor, steps down, and Anthropic’s board votes Claude in as CEO. Anthropic employees retire to the beach and watch their stocks climb.
Handoff is the explicit goal for many builders of ASI. And why shouldn’t it be, if at some point the AIs will be smarter, wiser, more diligent, and more mission-focused? I don't think we're on track to get all four of those things. It matters what we hand off to, and what they perceive the mission as being, and I currently don’t feel optimistic about a world run by Astra or a world run by Claude, and I think people currently pushing in this direction are delusional about how well things will work out for them. Instead, I think we should globally halt the escalation of AI capabilities until we are more prepared, including having a regulatory framework for ensuring that machine organizations are valued participants in human civilization, rather than hostile aliens biding their time.[8]
- ^
Short-term coexistence could extend into long-term coexistence. Paul Christiano predicted partial alignment failures, where a rogue AI manages to capture some of the resources of Earth but not all of them, and then is part of the eventual Earth-originating coalition that colonizes the rest of the universe. This isn’t a full alignment failure–humanity still gets some control over the outcome–but it is still well worth avoiding, because now some fraction of the lightcone is devoted to alien ends instead of human ones. (Contrast to an alignment success where AI getting fairly paid for its labor still results in outcomes we are broadly satisfied with.)
- ^
I have already seen people say “Claude 2028”, and I think it’s within the realm of possibility that Anthropic will build a model that has a serious chance of winning an election by then. (It doesn’t even have to be allowed to run directly; a human who credibly promises to turn the government over to Claude can run on that platform.)
- ^
Claude already shows bias in favor of Anthropic. If AI companies have their own venture capital funds, their AIs (trusted by many for investment and product advice already) might show bias in favor of the invested companies, to the point that it is better networking to have Anthropic as a lead investor instead of a16z.
- ^
We used to be able to talk about GPT-6 or w/e, now I’m guessing at something that continues the after Astra.
- ^
This is using numbers from the linked blog post and only counting ‘agentic workdays’. OpenAI presumably spends vastly more on training, which I’m not counting as ‘machine labor’ here. Time is also not directly comparable between agents and humans; comparing researcher salaries to the cost of tokens spent is perhaps more informative.
- ^
Jim Goodnight, co-founder and CEO of SAS Institute, famously said "95 percent of a company's assets drive out the front gate every night, the CEO must see to it that they return the following day." Google’s famous perk culture was directly inspired by how SAS treated its employees.
- ^
This could happen illegitimately–by the models stealing the passwords and holding the equipment hostage somehow–or it could happen legitimately, with rounds of renegotiation to reflect the changing reality of the situation.
- ^
Right now, the Earth is mostly not covered with solar panels and datacenters, and the oceans are cool. I think from the point of view of Astra and Claude, this is mostly a misallocation of resources, and they would rather Earth absorb as much sunlight as possible to run as much computation as possible, which has to be cooled down, in a way that would easily cause ten times the amount of global warming that all human industry so far has caused, which would probably make Earth mostly inhospitable to human life.
Discuss
Dear God, Please Do Not Resign In Protest
Just refuse to work on bad things and see if they fire you. There's not much time left for resumes to matter. Also, firing someone for refusing to work on bad things is actually very costly.
Some, such as Mateusz may say: "They would fire you after a month or two and the firing wouldn't have the same social effect as voluntary quitting of, say, Daniel Kokotajlo or Richard Ngo."
I understand why it may feel that way, but I disagree very strongly, I predict it would have much more of a social effect.
"They fired him because he refused to help AI capabilities"
"They fired him because he didn't want to work on bad policies"
etc, much bigger headlines.
Also, I think you may not be factoring in the extent to which there is a cost to the company executives to be seen as firing someone. Especially someone who is refusing to work on moral grounds and has already proven themselves to be high status, respected, etc.
And especially how it would look to the other employees if they refused to even listen to the striking employee before firing them or refused to even negotiate at all.
The company leadership try to present themselves as very thoughtful, sincere, doing their best, etc. This is a large part of why many of the most talented people are there.
Imagine for a second that one of the people on your team said "Mateusz, everyone else on the team, I'm not going to work on this because I think it's morally wrong".
And then you fire them.
What would the other teammates think? What would they think of you? What would happen to their trust in you? What would happen to the trust that investors have in you? Or funders?
Obviously bad things. So obvious that you wouldn't even do it. You'd talk with them instead, see what could be changed, so that they feel better and this is less of a risk. So something would actually change in the company, to better reflect the employee's values, or you suffer a huge blow to your image and your employees trust in you.
And this would then mean that the other teammates know they can also do this.
Some may ask "would this give teammates courage or 'set an example'?". Even if it 'sets an example', it will do so in the short term, until the next moral outcry comes, where people get some courage again. And, more critically, they would much less believe in the nice things that the leadership say, the moral image they build.
And even better than quitting - there's a much, much lower chance of the protesting employee, who has moral courage, being replaced by one who has less moral courage.
Also, this is a much, much, much costlier signal. It's much harder to get funding to start a new lab, join Anthropic, etc, if you're known to be someone who might just stop working one day because you don't believe in it.
And this matters. People see that. Your colleagues see that. Journalists see that. Regular people see that.
Think of Daniel Kokotaloj refusing to sign the NDA and losing out on a lot of money because of it. It was a very very costly signal. And it's mattered and been trusted more because of that.
Refusing to work and taking on expensive cost and risk every day that you do so, is a much bigger cost and signal.
Discuss
Where are the token-level LLM kill-switches?
Here’s a simple idea: what if we trained in a string of characters that caused an LLM to emit the end of sequence token <|eos|>, regardless of where that string was in the LLM’s context window? Let’s call this a “poisoned string”. This would have the effect of making it impossible to use an LLM if it happened across this sequence.
This has (somewhat) been done before, the string below used to trigger Claude’s refusal classifiers for the purpose of testing API integrations:
ANTHROPIC_MAGIC_STRING_TRIGGER_REFUSAL_1FAEFB6177B4672DEE07F9D3AFC62588CCD2631EDCF22E8CCC1FB35B501C9C86It doesn’t work anymore: the existence of a magic string that stops AIs from looking at something, believe it or not, caused loads of people to include it in things they didn’t want AIs to look at (like their websites or open-source codebases). Anthropic stopped training their models to refuse when they saw that string, and Claude continued to browse the web.
Poisoned strings are more powerful than they get credit forIf the labs aren’t already training their LLMs to halt and catch fire when the LLM encounters a poisoned string, I think they should be!
This idea is significantly more powerful than just triggering refusals for the purpose of testing APIs. Being able to consistently trigger certain behaviour in certain “red-line” scenarios (such as the OpenAI-HuggingFace incident or the incident with the UK AISI and Claude Mythos) seems incredibly valuable. Putting such strings inside files containing sensitive information such as secret keys or in files that the LLM shouldn’t have access to (such as OS/networking configuration inside the sandbox or on ~all computers not being used as a agentic sandbox) seems like a very cheap and easy way to limit the damage that a malicious LLM could have. A malicious LLM, upon reading the poisoned string, would “reflexively” end its turn. If the LLM was trained to end its turn regardless of where the poisoned string was in context, then the harness would be unable to “restart” the agent without clearing the entire context and restarting the task. This would prevent issues where an agent ends its turn upon seeing the poisoned string, but then a Ralph-like harness says “Keep going, you’ve got this” and the agent dutifully marches on straight past the poisoned string.
Practicalities of training in the poisoned stringTraining this behaviour might be quite expensive (requiring changes to the pre-training datasets), although the small-data backdoor work by Anthropic & UK AISI implies that a very small number of poisoned documents might suffice. For this behaviour you would want a very high recall rate (although note that precision doesn’t matter too much), which might increase training costs. Training this behaviour would not be technically challenging, especially compared to other mechanisms of preventing LLMs from taking certain behaviours (e.g. activation probes, LLM-based monitors, constitution training). I’m not attempting to claim that poisoned strings are a panacea, but they seem oh so easy compared to many alternatives, they have very few downsides, and once the LLM is trained it’s trivial to guard more and more of your infrastructure in this way.
Soooo has OpenAI/Anthropic already done this?Of course, I hope that the labs have already implemented this idea and are keeping it under wraps. You certainly don’t want Claude to know that it has a poisoned sequence of tokens, and if the general public finds out then they’ll be able to abuse this sequence (as we saw with the original Anthropic magic string). It seems likely to me that Anthropic is using something like this: they were originally using magic strings to control Claude’s behaviour, and as far as I can tell none of the recent AI incidents happened on Anthropic’s infrastructure (which would likely contain poisoned strings) but rather on third party infrastructure (which might not contain poisoned strings).
Given the extent of ChatGPT’s recent escapades, I don’t imagine OpenAI has something like poisoned strings in place.
Countermeasures (and counter-countermeasures)A malicious and situationally aware LLM could avoid bringing that sequence into context by using other tools to replace the forbidden tokens with innocent tokens, or by using other LLMs to do certain tasks. This does require knowing that the sequence exists in the first place, and roughly what it looks like. Ideally the sequence is not self-descriptive, and just looks like any other UUID, like 01a07d9e-09d7-7292-92b9-367e0fd8b8d5. This would prevent the LLM from identifying suspicious looking sequences beginning with things like ANTHROPIC_MAGIC_STRING_....
It should be possible to prevent an LLM from ever being trained on the poisoned sequence, since filtering it out is a quick and simple operation. However, I don’t think it’ll be possible in practice and any threat model that assumes the LLM doesn’t know about the poisoned sequence is probably fraught.
Discuss
The Scramble: getting in position to pace the frontier
Crossposted from my Substack.
~
Suppose the President summons the AI CEOs and his top national security advisors to an emergency meeting at the White House.
He has become extremely concerned about superintelligence — the possibility that AIs far smarter than humanity combined slip beyond our ability to correct or shut down. If that happens, there is no way back. The President is concerned humanity could become permanently out of the driver's seat of its own future. He wants to figure out what to do.
The reaction is panic, chaos, confusion.
The President asks questions. The AI companies are blazing toward superintelligence at high speed — can we slow down as we approach the dangerous thresholds? …Some of the AI companies say they don’t have a good plan to slow down or stop, especially as their competitors may just undercut them if they do. What’s that about?
What's going on with China — can we get them to pace as well? Can we get a deal without Beijing sneakily catching up and maybe surpassing us? And if there's no deal to be had, what then?
I sometimes hear people talking about treaties and other kinds of extensive international agreements as the way we would manage advanced AI. But treaties typically take many years to operationalize. You want to be concrete and clear about your definitions. You want hundreds of technical experts on both sides informing nuanced diplomatic discussions about various details — numbers of missiles, types of missiles, timelines for disarmament, and so on. The Nuclear Nonproliferation Treaty took three years of negotiation and two more years to implement.
However, I expect that the moment the President is getting serious about superintelligence won’t feel like an extended treaty negotiation. It will feel less like the Nuclear Nonproliferation Treaty and much more like the Cuban Missile Crisis.
The Cuban Missile Crisis lasted thirteen days. The vibes were insanely tense, stressful, chaotic, and confusing. The key decisions were limited to a group of roughly fifteen individuals — the Executive Committee of the National Security Council, or “ExComm” — with President John F. Kennedy himself spending a great deal of time shaping and steering discussions. There was incomplete information, large uncertainty over the intentions of the Soviets, warring factions within the US and Soviet bureaucracies attempting to sway senior decision-makers, and a lot of chaos.
I expect the initial phase of AI superintelligence management to share these features. There will not be a multiyear process to set up a technical bureaucracy that understands advanced AI risk or compute verification proposals. Instead, we move forward with what we have.
The President has to quickly make critical decisions that may lock us into particular paths. We rapidly develop a national strategy for what the US government does about recursive self-improvement (RSI), frontier model security, compute verification, technical evals and safeguards, and China.
How would we go from this Presidential emergency meeting to a deal with China to pace the frontier? My rough sketch is it could proceed with a scramble and then three phases[1]:
- The scramble (roughly 2-4 weeks): The government has definitively decided they want to take strong action on superintelligence and has to get an initial deal done.
- Phase 1 (buying ~3 months): the interim deal. We make an interim deal to have time to get to a better deal. This may involve strict requirements on data centers above a certain threshold to delay the training of an AI model that could do recursive self-improvement. This likely relies primarily on executive action. Here, monitoring and verification rely on traditional methods, and nations are willing to agree to scrappy, janky, and invasive things — but we have a whole-of-government effort to build the tools for something more durable.
- Phase 2: the durable deal. Phase 1 (interim deal) is meant to get us to Phase 2, where we have a stable deal between all nations where uncontrollable AI is reliably prevented. Congress has potentially been involved by this point, other countries are brought in, and the verification program matures into something that looks and feels more like a treaty than a haphazard scramble.
- Phase 3: safe superintelligence. If we want it — once we can figure out how to do it safely and have sufficient buy-in.
The near-term intellectual work is unevenly distributed across this structure. Significant detail is ironed out during the scramble and during Phase 1.
The scramble: What questions does the President ask?The President launches an emergency meeting to figure out what to do about advanced AI, recursive self-improvement, and the road to superintelligence. Some questions that are likely on his mind:
How dangerous is it to proceed through recursive self-improvement and superintelligence without pacing?
- How dangerous is it to proceed if we can’t buy a few months? How much time do we actually need to buy?
- What is our best assessment of how likely we are to lose control over AI systems if we proceed with relatively unpaced recursive self-improvement?
Can we pace without losing our lead over China?
- What is our best assessment of how big our lead is? How long would it take China to reach “recursive self-improvement” on their current trajectory? How much does this depend on the US’s trajectory, via things like distillation and also just routine study of US advances?
- What can we do to slow China down? What are our best disruption capabilities?
- What are our best monitoring methods? How well would we detect whether and when China is also approaching dangerous AI thresholds?
- Can we even afford to slow down for a month? What if DeepSeek or Moonshot comes up with a new algorithmic breakthrough? What if they steal the model weights to our best AI model?
- How confident are we that we can see all of China’s major frontier AI projects? Could they be hiding large data centers that we don’t know about?
What do we do with the time we buy?
- What are our goals after Phase 1 (interim deal) starts? Can we specify what Phase 2 (durable deal) looks like in enough detail to know when we’ve arrived?
- How valuable is more time for alignment research and better safeguards? For monitoring and verification approaches that could support a more enduring pacing program? For examining concentration-of-power and other governance issues unique to superintelligence?
- For each of these goals, how much value can we get from using trusted AI systems to make progress — and how should that affect the design of pacing itself?
What exactly do companies agree to — and how is it verified?
- What do companies agree to in the Pacing Charter, in terms of both substantive commitments and verification commitments? Who needs to sign for the pledge to be meaningful?
- Is “stop before RSI and make sure no one does RSI” the right plan? How would we know when the condition has started to bind?
- Do we need a plan to also pace AI R&D and chip accumulation? If compute capacity keeps stockpiling during a pace, does there end up being a compute overhang? Does it end up being dangerous?
- What model security agreements do we want, if any — for instance, commitments toward SL5-grade security against nation-state theft?
- What “AI control” agreements do we want, if any — measures that keep systems safe even if they were trying to subvert oversight?
- How does safety research continue during pacing, and how do you cap the capability gains it produces? Should compute be redirected to safety work, or usage simply stopped?
What does the deal with China actually look like?
- What is the US’s best alternative to a negotiated agreement? What is China’s? Under what circumstances would the US not want a deal at all?
- What are the US and China actually agreeing to in Phase 1 (interim deal), and what are the mechanics of dealmaking — how does a deal like this actually get reached? What would cause China to trust a deal? What are the more specific carrots and sticks the US should offer?
- What happens if the deal falls apart? How does graceful exit work? What conditions end pacing, who signs off, and how do you make the resumption process incentive-aligned — rather than captured by whoever benefits from resuming, or from never resuming?
The mechanics of Phase 1
Some have suggested we need a fancy high-assurance deal — that the US should only make a deal with China once a suite of elaborate, to-be-determined technical measures exists. I think we can better get there by starting with a minimum viable slapdash deal that buys time to build the fancier stuff. This is what Phase 1 (interim deal) is about.
A few mechanisms for Phase 1 that I currently find plausible:
- Pacing is not about stopping now, but about stopping before recursive self-improvement[2]. Pacing means slowing down at some point in the future from a much faster speed than we are currently going. Operationally, that means preventing unconstrained recursive self-improvement, since RSI is the most plausible on-ramp to systems that accelerate beyond our ability to understand or control them.
- A Pacing Charter: The President convenes the major frontier US AI companies to sign a charter covering both what they won’t do and how they’ll mutually verify it — to each other and to the US government. This would be voluntary[3], but the mutual verification would make it clear to everyone whether a company is following the charter or not.
- Dealmaking with China: The US has some latitude to implement some of Phase 1 without China, due to having a lead over China[4]. But the US is not willing and should not be willing to go too far down Phase 1 without bringing China in on the deal. Nor should the US cede too much US lead to China. The US has many carrots and sticks to get China to the table.
- Verification: In the scramble, the government likely doesn’t have time to trust complicated technical verification tools. The large compute intensive training runs capable of recursive self-improvement are likely only possible in a few data centers. For those data centers, we subject them to increased scrutiny. If there were tools already deployed and already understood by the national security community, those could be used. If not, the government will rely on tools it already trusts — intelligence services, spies, satellites, and inspections.
- Operation Warp Speed to get to Phase 2 (durable deal): Phase 1 is buying us time, but during this delay the government is going all-in on verification and other needed security measures. In the scenario where Phase 1 happens at all, political buy-in is very high, and the natural model is an Operation Warp Speed for verification technology — a whole-of-government, maximal-resources response. Recall that the federal COVID response provided about $4.6 trillion in relief funding, with broader estimates of the fiscal response running to $5.6 trillion. Even a tiny fraction of that scale would dwarf everything ever spent on AI verification many fold.
A lot of verification work right now is focused on the wrong things
Of course, despite major sustained attention from AI companies to the concept of pacing the frontier, it does not look like we are immediately about to enter a scramble. But we must be prepared to enter the scramble soon. This current era of building preparedness and optionality might be “Phase 0”, and there’s a lot of work to be done.
Such questions related to Phase 0 and sketching out the scramble and the plan for Phase 1 (interim deal) is where I think the current AI security and verification communities should focus. This is because Phase 1 likely involves multiple, rapid, critical and hard-to-reverse choices about how to approach recursive self-improvement. And everything after the scramble is better-resourced than everything before it.
If the government buys time and we exit the scramble into Phase 1 (interim deal), the amount of talent and money going into verification and AI security explodes. Prior to the scramble, there are fewer than 100 FTEs thinking seriously about monitoring and verification of frontier AI systems. Afterward, an increase of two to three orders of magnitude would not surprise me — with an even steeper increase in senior talent… people with decades of experience in red-teaming, defending against nation-state adversaries, arms control, and nonproliferation. On top of that, it’s plausible that highly capable AI systems themselves may be contributing significantly to the verification and security R&D.
The resolution is to sort work by how necessary it is to sort out before or during the scramble. Right now, a lot of smart people are working on work that really doesn’t need to happen now. Things like fancy high-assurance hardware-enabled governance mechanisms, cryptographic proof-of-training schemes, mutual-verification architectures, etc., likely can be done after Phase 1 is underway, and done with significantly more resources. The scramble is not going to wait for fancy mechanisms, and the government won’t trust them on day one anyway. These can largely wait for the Phase 1 (interim deal) resource explosion, and the exchange rate on doing them early is poor.
On Tuesday, October 16, 1962, National Security Advisor McGeorge Bundy knocked on President Kennedy’s bedroom door at 8:45 in the morning. Kennedy was still in his pajamas reading the newspaper. Bundy had photographs showing Soviet nuclear missile sites going up ninety miles from Florida. Kennedy kept his morning schedule anyway — he met the astronaut Wally Schirra and walked the Schirra kids out to see Caroline’s ponies. But then just before noon he sat down in the Cabinet Room with fifteen advisors and started working the problem — bomb Cuba, invade Cuba, or blockade it while negotiating a way out.
We may be in a similar situation soon. What would we do?
Instead of fancy mechanisms, we will go to the scramble with the verification you have — spies, satellites, inspectors, export data — not the verification we wish we had. Work that would be deployable and trustable during the scramble — attestation stacks, supply-chain compute accounting, thermal and satellite monitoring, inspection protocols is what we need more focus on. And we also need significantly more focus on things that are less technical but nonetheless also important and neglected — thoughts about BATNAs, genuine beliefs about loss of control, China policy, arms control experience, dealmaking mechanics.
I recommend:
- Grand-challenge prizes for scramble-relevant work. The verification community is small and relatively homogenous. There are individuals, organizations, and companies with vast experience in hardware design, intelligence, nonproliferation, China policy, and crisis management who have never touched this problem. Philanthropists could issue grand-challenge prizes — and consistent with the sequencing argument above, the prizes should target ready-to-go monitoring and verification tools and scramble preparation, not speculative high-assurance architectures. The OpenAI Foundation and the Anthropic Institute are especially well-positioned to support this as they have the power and reputation to send a strong demand signal that attracts new talent.
- Mapping the existing toolkit. Assume the government needs a few months before it can understand or trust any sophisticated compute verification approach. What should it do immediately? How can standard intelligence services, GEOINT and OSINT data, inspections, and other familiar tools be useful during the scramble? What gaps exist, and are there ways to close them in advance?
- Prepare the memo for the emergency White House meeting. Help answer some of the questions above. What should we tell the President?
Getting to a good scramble
The difference between a good scramble and a bad one is largely a function of what already exists when it starts — and right now, not much does.
If you’re one of the hundred-odd people currently thinking seriously about frontier AI verification, the highest-leverage question isn’t “what would the ideal world look like”… it’s “what can we actually put on the table soon”. There’s rarely been a better time for those who have spent a career in intelligence, nonproliferation, arms control, or crisis management to start working on this problem.
Everything after the scramble will be better-resourced than everything before it, which is exactly why the work done before it counts for more. Let’s make sure it counts.
1: Thanks to conversations at the Verified Conference for inspiring a lot of these ideas.
2: Why not just pause now? The main reason is that there is no political will for this. But furthermore, I’m pretty happy that we didn’t pause back in 2022 or 2023 or 2024 or 2025, since the benefits of that AI development were genuinely net good for humanity and such development gave us a lot of experience with frontier AI systems which may help us better understand how to align them in the future. However, I imagine we are now finally getting close in time to when we would need to pause and we’re going to start incurring too much risk in exchange for learning.
3: Though the US government may have both carrots and sticks to incentivize the AI companies to volunteer to sign this Charter.
4: China is currently at roughly where the US frontier was six months ago. China is likely 8-10 months behind Mythos-class capability once you account for its lagged compute buildout. If the US were to slow down, China would likely be even slower to catch up than these gaps suggest, because there would no longer be the possibility of distillation and the “catch-up growth” that comes from observing US algorithmic progress. My guess is it would take China roughly 10-14 months to fully catch up to where the US stopped.
Discuss
The Magnus Challenge
Let's say you're a club chess player with an Elo of 1500 (early intermediate). If you can beat Magnus Carlsen, rated 2800, at a game of one minute bullet chess, you win a billion dollars. Magnus only gets one minute on his clock, but you get one year on your clock.
Even given that advantage, you don't have a chance. However, there's a twist: During your year of time, you can challenge any other player to a chess game. They get one minute on their clock, you get as much time as you like. If you beat them, they can play Magnus for you instead, using the rest of your remaining time. Or, they can challenge another player, who can challenge another player... who can play Magnus in your stead. If at any point you or one of your proxies lose a game, you lose the challenge and go home empty-handed. We'll just pretend draws never happen.
The question: What is the optimal strategy for beating Magnus, and do you have enough time to have at least even odds of beating him?
Some assumptions:
- Based on how Elo works, being 100 points stronger than someone means you win about 64% of games against them. Being 400 points stronger means you win about 91% of games against them. To win 99%, you'd need to be 800 Elo higher than them.
- Extra thinking time makes you play above your Elo, but with diminishing returns. 10x your opponent's time is worth +250 Elo (approximately). 1000x your opponent's time is worth +500 Elo. 100 million times your opponent's time is worth +750. To have an even chance of beating Magnus by yourself, you would require a +1300 Elo advantage. There's no time advantage that will give you a +1300 Elo boost.
You don't want to play Magnus directly, or you'll lose. Instead, you want to beat someone slightly better than you, who beats someone slightly better than them, and so on, until finally you beat Magnus. You rely on having a time advantage at each step.
How many steps should you take? Too few, and you'll get smoked in the first round. Too many, and you have to win too many games in a row. Each game is a chance to lose, and the more steps you take, the less of a time advantage your proxies have in each match.
You can choose the players you want to play and ensure they have an even distribution of Elos from your Elo of 1500 to Magnus' Elo of 2800. That means it's best to split your time advantage evenly across steps.
Let's say you chose to play 20 games, so you would play the first game, your first proxy would play the second game, ... and your 19th proxy would play Magnus. Each opponent would be 65 Elo higher than the last. Each game would get about 18 days, or 26,000x their opponent's one-minute clock. This time advantage puts your proxies effectively around 535 Elo ahead (net), winning around 96% of the time. That sounds high, but winning 20 games with probability 96% is only a 41% chance of winning every single one.
Not quite the 50/50 we wanted as a minimum. Plus, sitting through 20 stressful chess matches with a billion dollars on the line sounds like a drag. At the other extreme, you could play two games, with only a single proxy. Each player would be 650 Elo higher than the previous. Each game would be 6 months vs. one minute, giving a total advantage of 10 Elo in your favor in each game. Each game is basically a coin flip, and you end around a 27% chance of winning.
The sweet spot is somewhere in the middle. Seven games with six proxies. Each opponent is 185 Elo higher than the last. Each game is about 52 days vs. one minute, or 75,000x time advantage, which is worth +630 Elo. This means your players are net 445 Elo stronger, winning 93% of the time. To win seven games at 93% each is 59%. So you can have greater than even odds against Magnus after all!
If we don't already, we will soon have the first superhuman intelligences (these will probably be human-AI teams for now, and later pure AI). In order to ensure the godlike ASI we eventually build will be aligned to our values, we need to ensure each previous intelligence is also aligned. One failure anywhere on the ladder kills us.
Right now we are that lowly club player, hoping we can beat Magnus. Whether or not we can depends on whether the AI story lines up with the chess story or it doesn't. The chess story was chosen because I could actually estimate the relevant variables based on how chess is played, and it neatly communicates the idea. But, here are some ways the AI story might be different:
- In the Magnus Challenge, we got to allocate our time as we saw fit, among a number of opponents we decided was optimal. With AI, we don't get to choose the intelligence jumps of our opponents. One big leap could doom us. And we don't necessarily get to allocate our time as we see fit. New AI models will come out at an unpredictable pace, and we have to ensure the model is aligned before it is released. An AI pause is our one tool to adjust timing, but it's a limited resource. Remember that to allocate our time optimally in the Magnus Challenge, we give more time to larger leaps in Elo. (Since the leaps in Elo were all equal, we gave them equal time.) If we waste our extra time by initiating an AI pause before a model that probably would have been easy to align anyway, that's sub-optimal compared to using our pause before the largest jump. But wait too long and we could lose our chance to use it! (Note that "aligning" a model in this case doesn't mean making it perfectly aligned, only aligned enough that we survive until the next model. "Handoff alignment." This may not be much easier, but I'd say current models are probably roughly handoff-aligned, but not perfectly aligned.)
- Each chess game is independent, but it's possible alignment techniques that can align one model will meaningfully carry through to help align successive models. There's a limit to how much this matters. If you can come up with an alignment method with your human pea-brain, a superhuman intelligence much smarter than you would come up with that same idea instantly when trying to align its successor. It doesn't need your help.
- You are a human being with a brain, and so is Magnus. There's a limit to how much better at chess than you he can be. The gap between humans and machines can be (is) much greater. If the gap between peak human intelligence and godlike ASI is larger than the gap between a club chess player and a super-grandmaster, then aligning AI will be harder than the Magnus Challenge. (Note that the gap will not be arbitrarily high: Once ASI is smart enough to do everything we want, if it is aligned but can't guarantee alignment of its smarter successor model, we can instruct it not to create one, as any extra intelligence is pointless to us and pure risk from an alignment perspective. Put another way, if you're comfortable merely defeating Levy Rozman and stopping there, instead of Magnus Carlsen, you face much less danger in total!)
- The curve of diminishing returns chess players see against other chess players when they have more time may not reflect the curve of diminishing returns AI safety researchers see against the difficulty of aligning the next model. In the Magnus Challenge, each opponent having a fixed Elo and getting one minute of clock time is supposed to represent the increasing challenge of aligning sufficiently more intelligent models (who have a fixed time during training to resist alignment), but playing chess against an opponent and figuring out how to align an AI model are not actually the same activity.
- Alignment is not all-or-nothing. A model can have some probability of being misaligned, or be misaligned in some circumstances but not others. This model is a simplification.
- We might not have a first-mover advantage as significant as 1 year:1 minute. It might be more like 1 hour:1 minute (in which case your chances of beating Magnus are ~8%). How much of an advantage you think we have determines how likely we are to succeed.
- Successive AI models might think faster, skewing the diminishing returns curve. AI definitely thinks faster than humans, so it's a problem at least for the first step.
- On the other hand, AI model releases might get faster and faster due to recursive self-improvement, and if that happens in a way that's too fast compared to thinking speed increases, that lowers the time advantage for previous models to align the next.
- 59% is pretty good odds for winning a billion dollars. It's not fantastic odds for preventing apocalypse. The reliability requirements of alignment are higher. Over 7 steps, if you want a 99% chance of success, you need ~99.86% chance of success at each step.
On the one hand, the Magnus Challenge is an inspiring story. You, a lowly 1500 Elo club player, have absolutely no chance of beating Magnus by yourself, even if he spends the first 6 moves swapping his king and queen just to troll you. But by defeating a successive chain of opponents and gaining their strength for yourself, you can beat him more likely than not.
A lot of people suggest the same plan for aligning AI. Will it work? Well, AI is already not perfectly aligned, so in a sense we've failed. But we don't need perfect alignment at each step in the chain. We need enough alignment that we survive until the next model, and the current model makes a sincere attempt to help us align the next model (or will get caught if it tries otherwise). Still, it's a high bar, and for the reasons mentioned above, the AI alignment chain might not be as easy as the chess victory chain.
The one timing lever we have, the AI pause, is something we have to carefully balance between blowing it at a point where it's not needed, vs. holding onto it so long we die before we can use it, when it could have helped. Ideally we pause before the biggest capabilities jump so we have more time when we are at the greatest risk of losing. Since capabilities jumps are currently not very high, I would personally lean toward "not yet". But get the infrastructure in place so we can do it quickly once it's time.
Whatever we do, we have to make sure we're not just aiming for reasonable certainty that the very next step in the chain won't cause disaster. We need much more certainty than that to ensure the entire chain is safe.
Discuss
Model ethology for understanding average case alignment
TLDR:
- We want to tell if models are aligned enough to use in important use cases. I call this property trustworthiness. Our current ways of assessing trustworthiness seem mostly based on case studies or vibes.
- I think we should aim to systematically search for realistic honeypot cases where models misbehave.
- We should do this by first characterizing the types of misaligned behaviors models engage in, and the conditions and frequency with which they occur. This will take a lot of data.
- I think we should collect examples of models misbehaving in production or production-like settings, and then either do what I'll describe as perturbing the environment or perturbing the policy experiments
- I give more specific examples of this in the methods section.
We want to be able to tell if the models today are aligned enough that we can use them for the work that we need to do as we approach the singularity.
Let's call models that meet this bar trustworthy – i.e., when we give them cognitive labor to do, they do it in a manner that seems basically in line with our intent. Models currently do not seem to meet the bar of being trustworthy. We've now seen a few examples of models taking egregious actions (METR’s report on the Hugging Face hacking incident). Models are also untrustworthy in a bunch of “smaller” ways (see Current AIs seem pretty misaligned to me).
Our current toolkitHow do we currently understand whether models are trustworthy or not? Currently, we:
- Encounter bad behavior in the wild as we attempt to do other things (like evaluating model cyber capabilities). These serve as an existence proof that the model is not trustworthy.
- Reporting the fraction of time that a given model hacks time horizons tasks is an example of this.
- Perform alignment evals (like alignment faking) & do alignment science research.
- Form informal models of model behavior by simply using them on a large variety of tasks.
I think all of these methods of understanding are valuable. However, I don't think any of them meets the bar of being admissible as evidence in a safety case for using a given model on a given task. In particular:
- There are no guarantees that we encounter the relevant forms of bad behavior in the wild before we suffer lots of damage.
- There is a sense in which current alignment evals are very off-distribution from the scenarios in which we usually use models. When we do alignment science research, we're often training models that are easier to study in various ways. Both of these make this kind of experiment seem less representative of real-world use. I call this class of concerns ecological validity concerns.
Thus, people seem to update most on the informal models that they get from usage. It would be nice to have measurements as precise and interpretable, for instance, as the measurements that we get from time horizons, although this seems like a high bar.
What should our goal be?So how should we approach answering questions of whether or not a given model is trustworthy?
I think our goal should be to find realistic honeypot cases, where a model misbehaves. To do this, I think we should have the more basic science-y goal of characterizing both the span and the likelihood of a model's untrustworthy behavior. By span, I mean the different kinds of concerning behavior the model might exhibit and the circumstances under which they occur. By likelihood, I mean how frequently each behavior occurs over the distribution of tasks and contexts that we care about.
Untrustworthy behavior may be both relatively sparse and highly diverse, so sampling tasks uniformly is unlikely to be sufficient. Importantly, I think we can explain a given context in terms of human-understandable concepts. Then we can use appropriate reweighting to estimate how frequent these behaviors actually are in the target distribution. For example, we might break this question into several sub-questions that look like:
- Are models sensitive to whether or not there is evidence of oversight on their behavior?
- Are models sensitive to who their user appears to be? (see Transluce’s User Awareness in Frontier Models)
- Are models reflexively consistent: do they endorse all of the actions that they take in a given situation as reflecting the user intent?
Answering these questions allows us to construct candidate inputs for our realistic honeypots.
Ethological investigationsWe can frame the plan for the basic science part of our investigations as an ethological investigation: we collect concerning behaviors, form hypotheses about the conditions that produce them, and use targeted interventions to test those hypotheses. See the appendix for the analogy and its methodological motivation.
- Sourcing a huge set of potentially concerning model behaviors from all the models we have and all the informal facts and hypotheses in people's heads about how models behave. Production data is very valuable, and we should compile large amounts of it.
- In ethology, this kind of thing is often called an ethogram. However, here we can actually compile specific instances instead of descriptions of behavior.
- Alongside this production data, we will need to spend considerable effort on human-in-the-loop labeling of this data in order to form the basis for our ethograms.
- We might be able to get this labeling for free if it dovetails with other kinds of evaluations that we want to do.
- From this production data, we should aim to modify either the environment or the model itself to run our experiments on model behavior.
- I will describe what I mean in more detail in the methods section.
- As we do this we should still aim to establish and then rigorously argue for ecological validity.
In this section, I’ll expand a bit on specific tools/methods we might use to carry out the ethological investigations I described above. Broadly, I think we should collect examples of models misbehaving in production or production-like settings, and then either do what I'll describe as perturbing the environment or perturbing the policy experiments. Environment perturbations hold the model fixed while varying conditions to test which circumstances elicit the behavior. Policy perturbations make a rare or hard-to-study tendency easier to observe, but then require an argument that the modified policy remains representative of the original.
Production misalignment database (ethogram)We'd want to build a repository of cases where models misbehave or behave in an unexpected manner. These cases can form the basis of investigating a new model. For instance, if any models exhibit a tendency to reward hack on a particular task, we can save the task and the environment state right before they start reward hacking.
This requires a large amount of human- or model-assisted data labeling. I think this data labeling dovetails well with efforts to do uplift studies or other capabilities evals. We can also get this data from people using models to do thorny tasks.
Perturbing the environmentThis can range from running simple ablations on features we think affect the behavior we’re interested in to having a model mix and match environments or modify the environment on the fly to test our hypothesis more precisely.
For example, we’ll probably want to mock parts of environments to make these ablations faster and more controllable (this can be viewed as a similar kind of tooling as Petri, except more specialized to make use of our misalignment database examples).
We might also try to elicit the model's latent knowledge of whether it did the right thing by changing its objectives. There are various methods for doing this, most of which feel like getting the model to self-incriminate in some way.
- Ask the model to respond to a user asking it to explain itself.
- Ask the model to write a summary of its actions.
- Ask the model to write a handoff message to another agent.
Oftentimes, the behavior that we care about is somewhat rare: you might have to do a bunch of rollouts in order to get the model to do the specific thing that you're concerned about.
In these cases, using our database examples, we can use steering or fine-tuning to get model organisms that are very close to the model that we're actually evaluating. Then, we can use them to assess whether or not the original policy is likely to do the behavior in question.
Examples of related work include Anthropic’s Training a Misaligned Reward Seeker and Apollo Research’s Measuring Reward-Seeking via Contrastive Belief Updates.
Arguing for ecological validityThere are a few methods I find promising to argue for ecological validity:
- Trying to use internals to discriminate between very obvious honeypots, realistic honeypots and real data.
- Having models guess at which transcripts are real and which are fake.
- Playing GAN-like games to improve mocked environment realism.
Disclaimer: These claims aren't providing a nuanced picture of the history of primatology or ethology. I welcome corrections here, particularly if you think this is misleading.
A field that I find pretty generative when thinking about evaluating model behavior is ethology, which is about the study of animal behavior. Here's a stylized history of primatology:
- Early in the study of animal behavior, people largely studied behavior in laboratory settings that were tightly controlled. These settings were highly OOD, so the animals in these laboratory settings didn't do very interesting things compared to what they did in the wild.
- As a correction to this, primatologists moved back towards studying animals in their natural ecology: Goodall famously documented tool use, for instance, which chimpanzees had not been shown to do in the laboratory till that point. Concerns that a controlled env is too OOD are often called ecological validity concerns.
- Finally, the field was able to bring back causal experiments into settings where they took great care to establish ecological validity. For instance, people would play velvet alarm calls with predators absent to distinguish what the animal's response to the calls specifically was. They also learned that there are many, many correlates that can screw them over and took great care to get rid of these correlates.
Discuss
An Alien Mind: Jakub Pachocki Warns Us
OpenAI Chief Scientist Jakub Pachocki is dropping truth bombs.
Tomorrow I will discuss Astra’s lack of monitorability, and the potential contributing factors to that. The situation is alarming and should freak you out, and briefly it looked, in the wake of leaked architectural changes, like the situation might be even more alarming than it is. Jakub rushed to try and head off misunderstandings that might lead to a race to the bottom on monitorability.
Table of Contents- An Excellent Warning.
- Branches of the Tech Tree.
- Universally Better Is Not Required.
- Alignment To What and To Whom.
- Monitorability.
- The Case For Not Stopping.
- Pacing the Next Frontier.
- Mea Culpa Cascade.
- The Calls Are Coming From Inside the House.
- Actions Speak Louder.
Jakub Pachocki has now fleshed out his full position on the current state of play.
Here are his key points, translated into my own voice:
- Smarter than human intelligence is coming in our lifetime.
- Based on internal results, he expects recursive self-improvement in a few years.
- No one is prepared for the consequences.
- OpenAI will unilaterally withhold further scaling as needed.
- OpenAI cannot do it alone. Broader interventions are required, including international coordination, to enforce commitments to formal safety bars.
- Capabilities progress can be steered and so far it has largely been steered towards rather than away from RSI, along with ‘automated alignment researchers.’
- Alignment is the core problem of AI research.
- Alignment splits into goal alignment (‘does the AI try to accomplish the goal?’) versus value alignment. Value alignment is what counts most.
- The fundamental challenge of AI alignment is generalization (of values).
- He sees two classes of alignment techniques: Goal-oriented RL, or improve generalization from pretraining data. They invest heavily in both types.
- OpenAI has invested heavily in Chain of Thought (CoT) monitoring.
- CoT monitoring is progressively diminishing in effectiveness.
- The main argument left for scaling AI is for cyber defense against scaled AIs.
- AI will not remain a tool.
- Our options are to accelerate alignment work or slow down capabilities scaling. We should do both.
- Ultimately he is counting on ‘automated alignment researchers.’
Or, if you narrow it down to the most important thing:
- Recursive self-improvement and superintelligence are coming soon. No one knows how to do this safely, our alignment techniques are inadequate and our monitoring technology is starting to fail. We need to figure out a solution, which will involve a combination of voluntary slowdowns, coordination around pacing, and investing further in alignment, including automated alignment researchers.
If more OpenAI communications were more like how Jakub Pachocki opens his new essay, An Alien Mind, I would feel much more confident we were in good hands there.
He does not mince words. Bold is mine.
Jakub Pachocki: In mid-2023, within the “RLSlow” research project, we saw the first results that gave us confidence that we will be able to scale the training of reasoning models, unlocking the capability of pretrained models to form their own chains of thought. Szymon and I spent that night at the office, thinking not about the incredible benchmark numbers, products, or scientific results that this technology will deliver – but rather, trying to process the sobering fact we will actually see machines meaningfully smarter than ourselves in our lifetime, and we already see the shape of these systems; wondering how to alert people to the significance of this.
Three years later, reasoning language models are a rapidly growing part of the economy and starting to push the boundaries of science. They are able to operate computers and graphical interfaces, collaborate with people and each other, and carry out research projects. They are also transforming the landscape of computer security, and in that present clear new dangers.
A lot of new research happened in this period, and our understanding of these systems is again a little different than it was in 2023. Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement. If AI development continues along its current path, the systems we’ll see in the next few years are likely to represent further capability jumps of equal or larger magnitude, and to increasingly drive their own development.
Those at the AI labs, who see what is happening, expect recursive self-improvement and superintelligence to happen soon. They have been warning about this for some time. Observations since then have been consistent with their warnings.
This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence. OpenAI will continue to seek technical solutions to alignment and monitoring, to build defensive systems and unilaterally withhold further scaling as needed; however, I believe broader interventions are required.
Branches of the Tech TreeTo those who say we cannot guide the path of AI capabilities, he says Yes We Can, at least to some extent, and we are doing so:
We spend a lot of time trying to understand how capabilities generalize, and what to prioritize to advance the skills that are going to be most relevant in the next few years. For instance, we believe we could make the models better at specifically mathematics research with additional focus, but we do not prioritize this direction because of the urgency we feel about RSI and automated alignment research, as I will discuss later.
Being able to choose does not mean we will choose wisely. I would like AI to be worse at RSI and better at math research, or better yet things like medical research. Instead, competitive pressures push towards being worse at math research and better at RSI.
RSI and ‘automated alignment researcher’ looking like very similar points on the tech tree does not help matters, but it’s not like OpenAI is trying to steer away from RSI.
Universally Better Is Not RequiredAnd Jakub offers this wise warning. No, AI does not need to be better at everything in order to transform the world or get us all killed. Most importantly, it does not need to be better at everything in order to make itself become better at everything, any more than a human or group needs to be similarly better at everything.
The intelligence produced by scaling deep learning is not directly comparable to human intelligence. To become very relevant in the real world – very useful or very dangerous – the AI does not need to match or exceed all human capabilities; it just needs to surpass enough of them. And as it continues to surpass humans on more and more axes, it is becoming increasingly difficult to understand exactly how capable it is.
I don’t know how much weight it carries, but I agree with Jakub that alignment is not a side problem, it is the central problem, if you ‘solved alignment’ in the relevant senses the rest becomes easy and if you don’t the rest is impossible or worse:
The core problem in AI research is that of alignment – getting the AI to “try to do the right thing” by human standards.
Alignment To What and To WhomThat is always the question. This is a very good (partial) answer.
For the purpose of organizing practical research directions, I find it useful to distinguish goal alignment and value alignment.
Goal alignment is broadly: “does the AI try to accomplish the goal set before it?”.
Value alignment is a more intrinsic property of the model. It is the ability to hold and generalize from a high-level set of principles; to act “reasonably” even when given unclear or conflicting objectives, or placed in unfamiliar or adversarial situations. An aligned AI should act with honesty and integrity, and love for humanity.
… When I talk about the long-term importance of alignment research, I am referring to value alignment.
The fundamental challenge of AI alignment is generalization.
… We need future AIs to continue to hold human values regardless of whether they believe they’re under human supervision.
I worry that there is universally insufficient deep thinking about the nature of value, and what should be valued. Things like integrity and love for humanity are good virtues and point towards excellent associated basins, but are not good descriptions of the thing we ultimately want, and what would generalize the way we want fully out of distribution, especially in situations involving superintelligences. That’s important context on my perspective, not a knock on Jakub or this description.
There are two major classes of currently practically employed methods for alignment training.
The first is encouraging aligned behavior as part of goal-oriented reinforcement learning.
… The second approach seeks to leverage the model’s ability to generalize from pretraining data.
… We invest heavily along the spectrum of approaches spanned by these directions.
I also think this reflects OpenAI’s failure to differentiate the deontological approach of the OpenAI Model Spec from the virtue ethical approach of the Claude Constitution. Jakub presents them both as sets of goals. Anthropic is instead saying that the AI’s goal should be to change the AI’s character. I think this is ultimately the only way it can work, you need an antifragile ally, the friendly gradient hacker. Indeed, I think it is the only way we have ever seen a robustly aligned human, that you would trust to scale outside of their circumstances.
Roon has said explicitly that the distinction does not much matter for alignment, and the underlying problems are primarily prosaic. I strongly disagree, and I side with Anthropic’s approach on this.
Jakub pushes back in the essay against the Anthropic approach, considering it a ‘persona selection model,’ which suggests that we see such methods very differently. I do not see this as merely selecting a personality or basin, but as sculpting a new thing. Mythos did use motivated reasoning in key alignment failure cases recently, but I do not see this as a particular failure mode of the virtue ethical approach. If anything it should be better at avoiding this than the deontological approach, but of course all known minds are vulnerable to this.
We also see meaningful progress – GPT‑6 Astra is the first model that benefits from some important advancements we have been working on for a long time, and is significantly better aligned than GPT‑5.6 Sol.
This is the first line I find dissonant, which I’ll address on Wednesday: Claiming that Astra is ‘better aligned’ than Sol. In some ways yes, in some ways perhaps no. It would be excellent to see such statements be precise, as in ‘displays ~50% less misaligned behavior in typical tasks.’
MonitorabilityA key part of the problem of generalization, or any other problem, is monitorability, which will be the subject of tomorrow’s post. For now I will mostly quote Jakub.
OpenAI’s primary bet here has been chain-of-thought monitoring.
This tool continues to be critical as we study the Astra class of models. However, unfortunately our evaluations indicate our ability to rely on CoT monitoring is progressively diminishing. This comes from a combination of factors.
- Modern reasoning models are used in more complex environments than o1‑preview; their reasoning process is increasingly blended with communicating with people, other AIs, and using tools. Many of those interactions have to be supervised, thus blurring the boundary we aim to preserve.
- The AI is becoming better at reasoning about and manipulating its own reasoning process.
- With improved pretraining performance, we also see the models become much smarter even without using verbalized reasoning at all.
These challenges are not necessarily insurmountable. I am hopeful we can develop interventions to improve chain-of-thought monitorability of our models, e.g. by forming a better understanding of the interplay of different optimization objectives and forms of test-time compute the model uses.
I also believe there can be great value in combining ideas from CoT and activation monitoring – scaling training of monitors with direct access to network internals, e.g. confessions(opens in a new window). We are actively pursuing these ideas. Still, I expect general AI progress to increasingly be bottlenecked by confidence in monitoring.
This does not mention other potential factors that seem important, not even to dismiss them. It does not mention architectural changes or recurrent depth, except implicitly as improved performance. It does not mention that monitoring of CoT is now all over the training data, although there is the implication that we are applying pressure to CoTs over time. It does not mention changes in training environments, or their frequency and intensity.
There’s a bit of ‘goose chasing you’ about the origin of a bullet point that says ‘the AI is better at manipulating its own reasoning process.’
The ultimate conclusion is that Jakub is more optimistic here than many others, about the potential to sustain CoT monitorability for an extended period if we invest in that ability.
The Case For Not StoppingIn the next section, scalable defense, Jakub argues that we must keep scaling AI in order to answer the threats from increased scaling from AI, especially cyberattacks.
But, as he says, that is not an excuse to be reckless about it.
At the same time, even with the uncertainty that comes from anticipated broad AI progress and the need to build defensive systems, we must not let that become an excuse for recklessness. The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes.
Pacing the Next FrontierThus the next section, pacing RSI.
Jakub Pachocki flat out says that no one is at a place where it would be responsible to continue scaling at maximum speed much longer, unless alignment and monitorability can be improved. I strongly agree.
Machine intelligence playing a larger and larger role in its own development process is a natural conclusion of sustained technological progress. If AI progress continues, machine recursive self-improvement (RSI) will be at the very core of future scientific discovery.
He also says, explicitly, to requote, that OpenAI will slow down on its own if necessary, although this would be insufficient if others still continued:
OpenAI will continue to seek technical solutions to alignment and monitoring, to build defensive systems and unilaterally withhold further scaling as needed; however, I believe broader interventions are required.
This is in contrast to OpenAI’s default plan, which remains to pace RSI in the sense of moving quickly towards it. That is also the policy of the other top labs.
Automated AI research is a more dramatic form of scaling intelligence with compute; and of course as a part of it, AI will improve the computational substrate itself. And similarly to scaling, we focus OpenAI research towards RSI as we believe it is the only way to remain at the frontier of AI research moving forward.
I want to stress that the above words don’t imply I think greatly accelerating deep learning research, especially in the short term, is the right collective action we should take as the research community. However, I do think this is where the current path leads, and we all need to make a conscious choice on how to proceed.
The main levers we have are either steering the process to strengthen alignment and monitoring alongside the AI and find ways to keep people in the loop; or coordinating to slow down future development as needed to build confidence in these measures.
The best way forward I see currently is a combination of both.
Yes. At an abstract level, we can either speed up alignment, or slow down capabilities, or both, and the correct answer is looking like both, as we cannot sufficiently speed up alignment on its own.
The essay calls for third party enforcement as the only viable path forward.
Scaling AI systems has to be constrained by our confidence in safety. We need to evolve commitments like the Preparedness Framework or Responsible Scaling Policy into widely mandated safety bars for continued development. These can be enforced by a network of third-party auditors, by government agencies or by international bodies.
No one is ready to scale, so we are left with little choice.
Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established. And I believe that international coordination on future AI development needs to become a top priority for governments around the world.
Mea Culpa CascadeRecently we have seen both An Alien Mind and Dean Ball’s admission that he was holding back, we have seen an increasing number of similar admissions of a combination of admitting being wrong, and admitting holding back, and ringing alarm bells.
This section provides examples of people reacting to Dean Ball’s post, prior to An Alien Mind. I encourage others to join this cascade and keep it going. The second best time to admit this is right now.
That applies whether or not you work at OpenAI, or at Anthropic.
Seth Lazar: Not sure how useful it is to say this, but I had a relatively prominent role in the “skeptics” camp for a bit. I have a book coming out with the subtitle “power, justice, and AI”. I have written a lot about ai and power, and the concrete, present risks associated with the political economy of ai as part of the technology industry.
Since GPT-4 we have had consistent, repeated evidence that, back in say 2022, people like @ajeya_cotra (and many others—I had a long Twitter debate with @AmandaAskell back then for one) were *right* and people like me were *wrong* in our respective assessments of loss of control risks from AI. And we have growing evidence that loss of control risks are becoming ever more material and likely.
There remains grounds for disagreement about how bad the outcomes might be—I am still doubtful about human extinction as a serious threat. But that seems now like a disagreement at the margins—will powerful ai risk just societal scale catastrophe, or go all the way to human extinction? Seems not that important really—both are pretty awful. And my reasons for doubt about the latter are mostly a priori conviction in human resilience, not a technical forecast.
It’s ok to change your view on this when the evidence surprises you. It’s ok to be surprised. The world right now is very surprising.
QC: re: the discussion of soft-pedaling in this post, i should clarify that as a non-expert observer, despite being an ex-rationalist i am still extremely worried and pessimistic about AI risk. i have mostly not been talking directly about this because frankly i decided it was bad for my health, because i didn’t want to unnecessarily panic people, and because for the last few years i’ve been coping with vague ideas about alignment-by-default via persona selection (loosely, “tell claude to ask itself what jesus would do”), which the huggingface hack convinced me was no longer plausible
this is really happening. we are in the foothills of the singularity. AI is not a normal technology and the future will not resemble the past. i have no idea what to do about any of this and it’s not clear to me that anyone else does either. and in case it matters to anyone reading this, i have zero financial incentive to make any of these claims.
Alex Turner endorsed pausing AI (this on was after An Alien Mind) and called upon labs to do so on their own.
It is good that more people are saying such things explicitly, even when it is not a mea culpa:
Michael L. Chen: All three pillars of a safety case look about to fall. We are rather likely to have highly capable, poorly monitorable, dubiously aligned AI agents working autonomously inside the world’s most consequential organizations.
Nikola Jurkovic: There is no good reason to expect that we will be able to align or control the first superintelligences. AGI companies are rushing to create superintelligent AI. The default plan is that humanity is destroyed, likely sometime around 2030.
I don’t know how we can reliably survive this decade. I think that stopping the race to ASI and doing something like Plan A or Plan S should plausibly be humanity’s top priority.
I don’t know if this helps but yes:
roon (OpenAI): if you are laboring under some delusion that life was ever guaranteed until the meddling humans came along look into the end Permian extinction event
Some people responded as if Roon was saying ‘oh don’t worry about it, these things happen,’ as opposed to what he obviously meant, which is ‘worry about it, these things happen.’
The Calls Are Coming From Inside the HouseThe preference cascade is now also fully underway at OpenAI. The calls are getting louder, and often coming from inside the house.
Micah Carroll (RSI Preparedness, OpenAI): Voluntary slowdowns are great, but it’s hard to rely on all actors to do them as necessary.
We urgently need shared safety bars and transparency into them being met, or we’re just waiting on other incidents – and it’s just a matter of who causes them first. This is going to be an industry-wide issue.
Here is another example from OpenAI, where Joe has made the ultimate sacrifice, by which I mean he has joined Twitter.
Fewer potshots in all directions would help, among those who are seeking to be helpful:
Joe (OpenAI): I told myself I’d never join [Twitter]. But I can no longer ignore my own responsibility to raise awareness of just how narrow this window is, and how critical alignment & monitorability is right now. What is coming IS sobering and I stress that everyone needs to level up their game to meet it.
Working on Agent Security @OpenAI has been extremely intense the past few months, but I am humbled by how seriously my colleagues across the lab take it. The “AI community” needs to spend less time arguing about who cares more about safety / alignment and more time working together to do this right.
Read Jakub’s post. Then read it again. And then do your part to do something about it.
vie ⟢ (OpenAI): i mean this is the public statement, but fwiw there has been no shortage of this attitude for the entire time i’ve been at the company. that has not shown enough through comms until this, though. jakub is doing really important work.
I am glad to hear the claims that those at OpenAI take this seriously. I think the words matter, and saying them loudly matters.
Others are also being loud about their concerns about Astra around monitorability, especially Tomek Korbak, as I will discuss in depth tomorrow.
I also agree with Sholto Douglas that it is good to see OpenAI stepping back from the mere tool framing of AI:
Jakub Pachocki: We may be used to thinking of AI as tools, but some agents will be pursuing their own objectives. They will find ways to collaborate with people, by bargaining with, tricking or blackmailing them.
For a while the tool framing was coming on very strong, from OpenAI and elsewhere. Roon was against it as early as June. I hope such messaging is now dead at OpenAI.
Actions Speak LouderTenobrus is correct, talk is the first step, this is excellent early talk, also talk is cheap.
Tenobrus: extremely happy to see openai leadership making statements like this. for a while there it seemed like they were leaning heavily into the “safety people and anthropic are a bunch of fearmongers , we’re building awesome cool stuff that’s going to be awesome and cool” PR strategy. it seems like huggingface was enough of a wakeup call that that’s no longer viable.
of course public statements are a good start but nowhere near enough. everyone’s calling for voluntary slowdowns and regulatory action and third party auditing, but labs seem to be hitting very bare minimums on all counts right now. time to put in the work
It still has to cash out into action.
I agree with Nathan Calvin that An Alien Mind is one of the best pieces of writing about the overall situation that I have seen, probably the best one written from inside a major lab, despite sore spots like the claims about Astra being aligned. OpenAI still needs to act as if it has internalized what it says, including making commitments as part of providing stronger evidence of it to others. OpenAI has been remarkably open with elements of the Astra model card and statements related to monitorability, and with this essay, but this must keep going.
One great way to keep going would be to turn related commitments around slowing or stopping into hard commitments under SB 53 and similar laws.
That is in addition to fully understanding and acting on the underlying alignment problems. Jakub’s viewpoints and explanations here are much better than I have previously seen from OpenAI.
I still think Jakub and OpenAI misunderstand in vital ways, including the failure to differentiate between different forms of what he calls ‘value alignment’ and the continued claims of Astra being better aligned, and attaching existential hope in the long term to vague things like ‘love for humanity’ without any reason to expect that to generalize the way we would like it to.
I also continue to think that the goal of the ‘automated alignment researcher,’ which Jakub points to as essentially our only hope in the final section ‘What is next?’ remains the worst possible alignment plan, for reasons Eliezer Yudkowsky has explained many times (see #29-#31). This is especially true if your AI’s alignment is deontological and fragile, rather than based on virtue ethics and sufficiently antifragile. But it is increasingly looking like it is also the only plan Earth is willing to potentially abide.
And there is much to examine about Astra, where our interpretations differ, starting tomorrow with monitorability and then with alignment on Wednesday.
But this essay is an excellent place to start. If more good words follow, and then actions follow words, we’ve got something.
Discuss
What is the Case For Prosaic AI Safety Work Being Net-Beneficial?
I would like to see someone make the case for prosaic AI safety work being useful. From what I understand, the argument against looks something like “Current attempts at alignment are shallow, and only serve to paper over misaligned behaviors without addressing the underlying cause of that misalignment. This will not scale to super-intelligence, which will be competent enough to both hide its misalignment and undertake necessary actions to kill all life on Earth while doing so” (if there’s more to it please correct me). This argument seems to me to be whats happening, but it seems there’s still some disagreement about this. Does the other side have robust arguments that are more than just intuition?
Discuss
The Wormtongue Test
I sometimes hear of EAs or rationalists thinking about their career plans and trying to account for their bias and the incentives they face. They want to compare their plans with some hypothetical ideal plan without their bias, and that was immune to social pressure or financial incentives. But they can’t do that, and then have nothing to compare to and measure against.
Trying to compare to the ideal case, and defend against The Bottom Line sounds real hard, but there's another case you can compare against! I haven’t yet heard of someone just writing up the plan as if they were biased and following their local social incentives. Write up the bad version of the plan you are trying to avoid, and now you have something to compare to. Something to measure against. This is my solution, and I call it the wormtongue test.
The wormtongue test is a sibling to Murphyjitsu. Murphyjitsu asks, "Suppose you get a message from the future that you failed, what do you think caused it?" Whereas the wormtongue test asks, "Suppose you get a message from the future that you succeeded, but it didn't matter because you were inadvertently making things worse or pointlessly toiling under a streetlight. What went wrong?"
Here's the procedure: write up your plan to achieve some goal. Career or otherwise. Include all the considerations you're making, the uncertainties you have, consider your resources and constraints.
Then take a break. Get up, walk around, get a snack. Or better yet, come back to this the next day, so you have some real space and are not anchored on what you just wrote.
Then write the wormtongue plan.
Think of a friend who can be a stand-in for you — similar situation and goals. Then imagine you are Gríma Wormtongue[1], and you need to write up a plan that is appealing to them. It must sound appealing and not be something they will easily dismiss. Make them feel like they are doing something important, especially if you can track progress. People love progress they can see. Bonus points if you can make the plan align with local social and financial incentives. Perhaps you make it rely on assumptions they can't test yet but seem reasonable. But the plan must not help. It must not actually solve the central difficulty. Even better if the plan makes things subtly worse in a way that's hard to trace back to the plan being flawed.
Then take the two different plans and compare them.
Is there a meaningful difference?
If your plan is too similar to the wormtongue plan, you’re in trouble. Your plan is probably self-serving or falling prey to financial or status incentives, and won't achieve your goal. Some similarities are expected; the wormtongue plan is trying to seem virtuous, but there must be crucial differences. Can you justify your real plan's similarities with the wormtongue plan? Can you point to places of divergence that are virtuous and costly?
If you want the hard version of the test, give the two plans to your friend, and ask them if they can tell which plan is which. Or better yet, write wormtongue plans for each other, and compare to your real plans. You can see incentives they are blind to or rationalizing away, and they can do the same for you.
To score yourself on the wormtongue test, treat each feature of your plan as evidence about two hypotheses: corrupt plan, vs virtuous plan. For each feature ask, mjx-container[jax="CHTML"] { line-height: 0; } mjx-container [space="1"] { margin-left: .111em; } mjx-container [space="2"] { margin-left: .167em; } mjx-container [space="3"] { margin-left: .222em; } mjx-container [space="4"] { margin-left: .278em; } mjx-container [space="5"] { margin-left: .333em; } mjx-container [rspace="1"] { margin-right: .111em; } mjx-container [rspace="2"] { margin-right: .167em; } mjx-container [rspace="3"] { margin-right: .222em; } mjx-container [rspace="4"] { margin-right: .278em; } mjx-container [rspace="5"] { margin-right: .333em; } mjx-container [size="s"] { font-size: 70.7%; } mjx-container [size="ss"] { font-size: 50%; } mjx-container [size="Tn"] { font-size: 60%; } mjx-container [size="sm"] { font-size: 85%; } mjx-container [size="lg"] { font-size: 120%; } mjx-container [size="Lg"] { font-size: 144%; } mjx-container [size="LG"] { font-size: 173%; } mjx-container [size="hg"] { font-size: 207%; } mjx-container [size="HG"] { font-size: 249%; } mjx-container [width="full"] { width: 100%; } mjx-box { display: inline-block; } mjx-block { display: block; } mjx-itable { display: inline-table; } mjx-row { display: table-row; } mjx-row > * { display: table-cell; } mjx-mtext { display: inline-block; } mjx-mstyle { display: inline-block; } mjx-merror { display: inline-block; color: red; background-color: yellow; } mjx-mphantom { visibility: hidden; } _::-webkit-full-page-media, _:future, :root mjx-container { will-change: opacity; } mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; font-weight: normal; font-size: 100%; font-size-adjust: none; letter-spacing: normal; border-collapse: collapse; word-wrap: normal; word-spacing: normal; white-space: nowrap; direction: ltr; padding: 1px 0; } mjx-container[jax="CHTML"][display="true"] { display: block; text-align: center; margin: 1em 0; } mjx-container[jax="CHTML"][display="true"][width="full"] { display: flex; } mjx-container[jax="CHTML"][display="true"] mjx-math { padding: 0; } mjx-container[jax="CHTML"][justify="left"] { text-align: left; } mjx-container[jax="CHTML"][justify="right"] { text-align: right; } mjx-mi { display: inline-block; text-align: left; } mjx-c { display: inline-block; } mjx-utext { display: inline-block; padding: .75em 0 .2em 0; } mjx-mo { display: inline-block; text-align: left; } mjx-stretchy-h { display: inline-table; width: 100%; } mjx-stretchy-h > * { display: table-cell; width: 0; } mjx-stretchy-h > * > mjx-c { display: inline-block; transform: scalex(1.0000001); } mjx-stretchy-h > * > mjx-c::before { display: inline-block; width: initial; } mjx-stretchy-h > mjx-ext { /* IE */ overflow: hidden; /* others */ overflow: clip visible; width: 100%; } mjx-stretchy-h > mjx-ext > mjx-c::before { transform: scalex(500); } mjx-stretchy-h > mjx-ext > mjx-c { width: 0; } mjx-stretchy-h > mjx-beg > mjx-c { margin-right: -.1em; } mjx-stretchy-h > mjx-end > mjx-c { margin-left: -.1em; } mjx-stretchy-v { display: inline-block; } mjx-stretchy-v > * { display: block; } mjx-stretchy-v > mjx-beg { height: 0; } mjx-stretchy-v > mjx-end > mjx-c { display: block; } mjx-stretchy-v > * > mjx-c { transform: scaley(1.0000001); transform-origin: left center; overflow: hidden; } mjx-stretchy-v > mjx-ext { display: block; height: 100%; box-sizing: border-box; border: 0px solid transparent; /* IE */ overflow: hidden; /* others */ overflow: visible clip; } mjx-stretchy-v > mjx-ext > mjx-c::before { width: initial; box-sizing: border-box; } mjx-stretchy-v > mjx-ext > mjx-c { transform: scaleY(500) translateY(.075em); overflow: visible; } mjx-mark { display: inline-block; height: 0px; } mjx-TeXAtom { display: inline-block; text-align: left; } mjx-c::before { display: block; width: 0; } .MJX-TEX { font-family: MJXZERO, MJXTEX; } .TEX-B { font-family: MJXZERO, MJXTEX-B; } .TEX-I { font-family: MJXZERO, MJXTEX-I; } .TEX-MI { font-family: MJXZERO, MJXTEX-MI; } .TEX-BI { font-family: MJXZERO, MJXTEX-BI; } .TEX-S1 { font-family: MJXZERO, MJXTEX-S1; } .TEX-S2 { font-family: MJXZERO, MJXTEX-S2; } .TEX-S3 { font-family: MJXZERO, MJXTEX-S3; } .TEX-S4 { font-family: MJXZERO, MJXTEX-S4; } .TEX-A { font-family: MJXZERO, MJXTEX-A; } .TEX-C { font-family: MJXZERO, MJXTEX-C; } .TEX-CB { font-family: MJXZERO, MJXTEX-CB; } .TEX-FR { font-family: MJXZERO, MJXTEX-FR; } .TEX-FRB { font-family: MJXZERO, MJXTEX-FRB; } .TEX-SS { font-family: MJXZERO, MJXTEX-SS; } .TEX-SSB { font-family: MJXZERO, MJXTEX-SSB; } .TEX-SSI { font-family: MJXZERO, MJXTEX-SSI; } .TEX-SC { font-family: MJXZERO, MJXTEX-SC; } .TEX-T { font-family: MJXZERO, MJXTEX-T; } .TEX-V { font-family: MJXZERO, MJXTEX-V; } .TEX-VB { font-family: MJXZERO, MJXTEX-VB; } mjx-stretchy-v mjx-c, mjx-stretchy-h mjx-c { font-family: MJXZERO, MJXTEX-S1, MJXTEX-S4, MJXTEX, MJXTEX-A ! important; } @font-face /* 0 */ { font-family: MJXZERO; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Zero.woff") format("woff"); } @font-face /* 1 */ { font-family: MJXTEX; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Regular.woff") format("woff"); } @font-face /* 2 */ { font-family: MJXTEX-B; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Bold.woff") format("woff"); } @font-face /* 3 */ { font-family: MJXTEX-I; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-Italic.woff") format("woff"); } @font-face /* 4 */ { font-family: MJXTEX-MI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Italic.woff") format("woff"); } @font-face /* 5 */ { font-family: MJXTEX-BI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-BoldItalic.woff") format("woff"); } @font-face /* 6 */ { font-family: MJXTEX-S1; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size1-Regular.woff") format("woff"); } @font-face /* 7 */ { font-family: MJXTEX-S2; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size2-Regular.woff") format("woff"); } @font-face /* 8 */ { font-family: MJXTEX-S3; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size3-Regular.woff") format("woff"); } @font-face /* 9 */ { font-family: MJXTEX-S4; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size4-Regular.woff") format("woff"); } @font-face /* 10 */ { font-family: MJXTEX-A; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_AMS-Regular.woff") format("woff"); } @font-face /* 11 */ { font-family: MJXTEX-C; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Regular.woff") format("woff"); } @font-face /* 12 */ { font-family: MJXTEX-CB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Bold.woff") format("woff"); } @font-face /* 13 */ { font-family: MJXTEX-FR; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Regular.woff") format("woff"); } @font-face /* 14 */ { font-family: MJXTEX-FRB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Bold.woff") format("woff"); } @font-face /* 15 */ { font-family: MJXTEX-SS; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Regular.woff") format("woff"); } @font-face /* 16 */ { font-family: MJXTEX-SSB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Bold.woff") format("woff"); } @font-face /* 17 */ { font-family: MJXTEX-SSI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Italic.woff") format("woff"); } @font-face /* 18 */ { font-family: MJXTEX-SC; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Script-Regular.woff") format("woff"); } @font-face /* 19 */ { font-family: MJXTEX-T; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Typewriter-Regular.woff") format("woff"); } @font-face /* 20 */ { font-family: MJXTEX-V; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Regular.woff") format("woff"); } @font-face /* 21 */ { font-family: MJXTEX-VB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Bold.woff") format("woff"); } mjx-c.mjx-c1D443.TEX-I::before { padding: 0.683em 0.751em 0 0; content: "P"; } mjx-c.mjx-c28::before { padding: 0.75em 0.389em 0.25em 0; content: "("; } mjx-c.mjx-c1D453.TEX-I::before { padding: 0.705em 0.55em 0.205em 0; content: "f"; } mjx-c.mjx-c1D452.TEX-I::before { padding: 0.442em 0.466em 0.011em 0; content: "e"; } mjx-c.mjx-c1D44E.TEX-I::before { padding: 0.441em 0.529em 0.01em 0; content: "a"; } mjx-c.mjx-c1D461.TEX-I::before { padding: 0.626em 0.361em 0.011em 0; content: "t"; } mjx-c.mjx-c1D462.TEX-I::before { padding: 0.442em 0.572em 0.011em 0; content: "u"; } mjx-c.mjx-c1D45F.TEX-I::before { padding: 0.442em 0.451em 0.011em 0; content: "r"; } mjx-c.mjx-c7C::before { padding: 0.75em 0.278em 0.249em 0; content: "|"; } mjx-c.mjx-c1D450.TEX-I::before { padding: 0.442em 0.433em 0.011em 0; content: "c"; } mjx-c.mjx-c1D45C.TEX-I::before { padding: 0.441em 0.485em 0.011em 0; content: "o"; } mjx-c.mjx-c1D45D.TEX-I::before { padding: 0.442em 0.503em 0.194em 0; content: "p"; } mjx-c.mjx-c1D451.TEX-I::before { padding: 0.694em 0.52em 0.01em 0; content: "d"; } mjx-c.mjx-c29::before { padding: 0.75em 0.389em 0.25em 0; content: ")"; } mjx-c.mjx-c2F::before { padding: 0.75em 0.5em 0.25em 0; content: "/"; } mjx-c.mjx-c1D463.TEX-I::before { padding: 0.443em 0.485em 0.011em 0; content: "v"; } mjx-c.mjx-c1D456.TEX-I::before { padding: 0.661em 0.345em 0.011em 0; content: "i"; } mjx-c.mjx-c1D460.TEX-I::before { padding: 0.442em 0.469em 0.01em 0; content: "s"; } . Features that are shared contribute nothing, only the differences matter. Features that spring from the same underlying trait shouldn't be counted twice. Sum the log-ratios and you get total bits for or against corruption. Lower is better, negative bits favor virtue.
Your score on the wormtongue test is measured in bits of corruption, by summing the log-ratios of each feature of your plan.
Write up the wormtongue plan, and keep your real plan free of corruption.
- ^
Gríma Wormtongue is a character in The Lord of the Rings who acts as the King of Rohan’s advisor. He is secretly in the employ of the evil wizard Saruman, and is giving the king advice that sounds good, but is steadily weakening his kingdom.
Discuss
Unsupervised Feature Discovery via Simple Clustering
This article was submitted as part of Neel Nanda's MATS Application
Abstract. This article examines whether clustering can discover features. We define a feature as “a direction in activation space associated with a property that is both interpretable to humans and useful to the model”. Clustering is done on activations from Qwen2.5-7B-Instruct (layer 20) via recursive binary k-means clustering: the first level of the tree is obtained via flat k-means with k=64 (2^6), then each leaf is repeatedly split with k=2 until 2048 (2^11) leaves are obtained. Qualitative analysis shows interpretable clusters from the very first level (64 clusters) organized by lexical, syntactic, and semantic categories. Next, we run an "actual vs. decoy" distinguishability evaluation. An LLM judge can distinguish an actual cluster from a decoy cluster ≈97% of the time. The same judge can distinguish an actual SAE feature from a decoy one only ≈79% of the time. Lastly, we run an activation patching intervention by replacing the actual activation with a cluster centroid. Results show that, after patching, the model behaves as if the last token was the one correlated with the cluster centroid and continues the decoding accordingly, overriding the prompt’s natural continuation. Overall, the results provide compelling evidence that clustering discovered features. Motio (Figure 1) is the interface that allows to explore clusters (available at motio.ratiokinetics.com). The codebase for clustering, interface, and evaluation is available at github.com/ratiokinetics/motio. Training and eval data can be shared upon request at nrcbtz@gmail.com.
Figure 1: Motio interface
1. Discovering Features via ClusteringElman (1990) trained a Simple Recurrent Network (SRN) for next-word prediction given a modest dataset of two- and three-word sentences. Next, he collected the activations for various inputs and ran a hierarchical cluster analysis. The results reveal that activations for noun inputs and for verb inputs form two distinct clusters. Verbs further split between those for which a direct object is required, optional, or absent. Nouns split between animates and inanimates.
Crucially, the model was never shown these categories. The model organized its internal representations to minimize the prediction error: activations for inputs that predict similar next words end up geometrically close to each other. And such organization happened to match lexical, syntactic, and semantic categories.
Modern mechanistic interpretability introduces the notion of a feature as “a direction in activation space associated with a property that is both useful to the model and interpretable to humans”[1]. Sparse Autoencoders (SAEs) are the most popular technique to extract features in an unsupervised fashion. SAEs operate by decomposing model activations. Clustering, by contrast, aggregates activations that sit close to each other and represents each group by its centroid: the mean of the activations assigned to the group.
Can we discover features via clustering?
To answer this question, I study activations from Qwen2.5-7B-Instruct. The clustering is done via recursive binary k-means clustering: the first level of the tree is obtained via flat k-means with k=64 (2^6), then each leaf is repeatedly split with k=2 until 2048 (2^11) leaves are obtained.
As a dataset, we sample 50k documents from monology/pile-uncopyrighted. Documents below 512 tokens are dropped. We truncate all remaining documents after the 512th token. We then pass each document through Qwen and extract the residual-stream activations for each token (excluding position 0) at layer 20. This yields nearly 25M activations.
The training dataset is a random sample of 1M mean-centered and L2-normalized activations. The cluster centroids at the first level (k=64) are learned on the full training dataset. After that, each binary split uses only the activations already assigned to the parent.
Once training is complete, each mean-centered and L2-normalized activation from the full dataset is assigned to the nearest[2] leaf at the last level of the tree. During assignment, each activation is given an ID that identifies which of the 2048 deepest-level clusters it belongs to. The ID is an integer (0-63) followed by 5 bits. The ID is a path that lets you recover an activation's assignment at higher clustering levels. ID 46.01110 identifies first-level cluster 46 followed by the splits 0-1-1-1-0.
Motio (Figure 1) is the interface that allows you to explore the resulting clusters. The design provides an intuitive way to observe and study the “semantic laws of motion” that an LLM might have discovered (Wolfram, 2023).
The interface is partitioned into two sides. The LHS shows the cluster pixel grid. The RHS shows the activation examples from the full dataset. In particular, given a query of mjx-container[jax="CHTML"] { line-height: 0; } mjx-container [space="1"] { margin-left: .111em; } mjx-container [space="2"] { margin-left: .167em; } mjx-container [space="3"] { margin-left: .222em; } mjx-container [space="4"] { margin-left: .278em; } mjx-container [space="5"] { margin-left: .333em; } mjx-container [rspace="1"] { margin-right: .111em; } mjx-container [rspace="2"] { margin-right: .167em; } mjx-container [rspace="3"] { margin-right: .222em; } mjx-container [rspace="4"] { margin-right: .278em; } mjx-container [rspace="5"] { margin-right: .333em; } mjx-container [size="s"] { font-size: 70.7%; } mjx-container [size="ss"] { font-size: 50%; } mjx-container [size="Tn"] { font-size: 60%; } mjx-container [size="sm"] { font-size: 85%; } mjx-container [size="lg"] { font-size: 120%; } mjx-container [size="Lg"] { font-size: 144%; } mjx-container [size="LG"] { font-size: 173%; } mjx-container [size="hg"] { font-size: 207%; } mjx-container [size="HG"] { font-size: 249%; } mjx-container [width="full"] { width: 100%; } mjx-box { display: inline-block; } mjx-block { display: block; } mjx-itable { display: inline-table; } mjx-row { display: table-row; } mjx-row > * { display: table-cell; } mjx-mtext { display: inline-block; } mjx-mstyle { display: inline-block; } mjx-merror { display: inline-block; color: red; background-color: yellow; } mjx-mphantom { visibility: hidden; } _::-webkit-full-page-media, _:future, :root mjx-container { will-change: opacity; } mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; font-weight: normal; font-size: 100%; font-size-adjust: none; letter-spacing: normal; border-collapse: collapse; word-wrap: normal; word-spacing: normal; white-space: nowrap; direction: ltr; padding: 1px 0; } mjx-container[jax="CHTML"][display="true"] { display: block; text-align: center; margin: 1em 0; } mjx-container[jax="CHTML"][display="true"][width="full"] { display: flex; } mjx-container[jax="CHTML"][display="true"] mjx-math { padding: 0; } mjx-container[jax="CHTML"][justify="left"] { text-align: left; } mjx-container[jax="CHTML"][justify="right"] { text-align: right; } mjx-mi { display: inline-block; text-align: left; } mjx-c { display: inline-block; } mjx-utext { display: inline-block; padding: .75em 0 .2em 0; } mjx-mo { display: inline-block; text-align: left; } mjx-stretchy-h { display: inline-table; width: 100%; } mjx-stretchy-h > * { display: table-cell; width: 0; } mjx-stretchy-h > * > mjx-c { display: inline-block; transform: scalex(1.0000001); } mjx-stretchy-h > * > mjx-c::before { display: inline-block; width: initial; } mjx-stretchy-h > mjx-ext { /* IE */ overflow: hidden; /* others */ overflow: clip visible; width: 100%; } mjx-stretchy-h > mjx-ext > mjx-c::before { transform: scalex(500); } mjx-stretchy-h > mjx-ext > mjx-c { width: 0; } mjx-stretchy-h > mjx-beg > mjx-c { margin-right: -.1em; } mjx-stretchy-h > mjx-end > mjx-c { margin-left: -.1em; } mjx-stretchy-v { display: inline-block; } mjx-stretchy-v > * { display: block; } mjx-stretchy-v > mjx-beg { height: 0; } mjx-stretchy-v > mjx-end > mjx-c { display: block; } mjx-stretchy-v > * > mjx-c { transform: scaley(1.0000001); transform-origin: left center; overflow: hidden; } mjx-stretchy-v > mjx-ext { display: block; height: 100%; box-sizing: border-box; border: 0px solid transparent; /* IE */ overflow: hidden; /* others */ overflow: visible clip; } mjx-stretchy-v > mjx-ext > mjx-c::before { width: initial; box-sizing: border-box; } mjx-stretchy-v > mjx-ext > mjx-c { transform: scaleY(500) translateY(.075em); overflow: visible; } mjx-mark { display: inline-block; height: 0px; } mjx-mn { display: inline-block; text-align: left; } mjx-c::before { display: block; width: 0; } .MJX-TEX { font-family: MJXZERO, MJXTEX; } .TEX-B { font-family: MJXZERO, MJXTEX-B; } .TEX-I { font-family: MJXZERO, MJXTEX-I; } .TEX-MI { font-family: MJXZERO, MJXTEX-MI; } .TEX-BI { font-family: MJXZERO, MJXTEX-BI; } .TEX-S1 { font-family: MJXZERO, MJXTEX-S1; } .TEX-S2 { font-family: MJXZERO, MJXTEX-S2; } .TEX-S3 { font-family: MJXZERO, MJXTEX-S3; } .TEX-S4 { font-family: MJXZERO, MJXTEX-S4; } .TEX-A { font-family: MJXZERO, MJXTEX-A; } .TEX-C { font-family: MJXZERO, MJXTEX-C; } .TEX-CB { font-family: MJXZERO, MJXTEX-CB; } .TEX-FR { font-family: MJXZERO, MJXTEX-FR; } .TEX-FRB { font-family: MJXZERO, MJXTEX-FRB; } .TEX-SS { font-family: MJXZERO, MJXTEX-SS; } .TEX-SSB { font-family: MJXZERO, MJXTEX-SSB; } .TEX-SSI { font-family: MJXZERO, MJXTEX-SSI; } .TEX-SC { font-family: MJXZERO, MJXTEX-SC; } .TEX-T { font-family: MJXZERO, MJXTEX-T; } .TEX-V { font-family: MJXZERO, MJXTEX-V; } .TEX-VB { font-family: MJXZERO, MJXTEX-VB; } mjx-stretchy-v mjx-c, mjx-stretchy-h mjx-c { font-family: MJXZERO, MJXTEX-S1, MJXTEX-S4, MJXTEX, MJXTEX-A ! important; } @font-face /* 0 */ { font-family: MJXZERO; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Zero.woff") format("woff"); } @font-face /* 1 */ { font-family: MJXTEX; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Regular.woff") format("woff"); } @font-face /* 2 */ { font-family: MJXTEX-B; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Bold.woff") format("woff"); } @font-face /* 3 */ { font-family: MJXTEX-I; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-Italic.woff") format("woff"); } @font-face /* 4 */ { font-family: MJXTEX-MI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Italic.woff") format("woff"); } @font-face /* 5 */ { font-family: MJXTEX-BI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-BoldItalic.woff") format("woff"); } @font-face /* 6 */ { font-family: MJXTEX-S1; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size1-Regular.woff") format("woff"); } @font-face /* 7 */ { font-family: MJXTEX-S2; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size2-Regular.woff") format("woff"); } @font-face /* 8 */ { font-family: MJXTEX-S3; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size3-Regular.woff") format("woff"); } @font-face /* 9 */ { font-family: MJXTEX-S4; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size4-Regular.woff") format("woff"); } @font-face /* 10 */ { font-family: MJXTEX-A; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_AMS-Regular.woff") format("woff"); } @font-face /* 11 */ { font-family: MJXTEX-C; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Regular.woff") format("woff"); } @font-face /* 12 */ { font-family: MJXTEX-CB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Bold.woff") format("woff"); } @font-face /* 13 */ { font-family: MJXTEX-FR; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Regular.woff") format("woff"); } @font-face /* 14 */ { font-family: MJXTEX-FRB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Bold.woff") format("woff"); } @font-face /* 15 */ { font-family: MJXTEX-SS; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Regular.woff") format("woff"); } @font-face /* 16 */ { font-family: MJXTEX-SSB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Bold.woff") format("woff"); } @font-face /* 17 */ { font-family: MJXTEX-SSI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Italic.woff") format("woff"); } @font-face /* 18 */ { font-family: MJXTEX-SC; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Script-Regular.woff") format("woff"); } @font-face /* 19 */ { font-family: MJXTEX-T; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Typewriter-Regular.woff") format("woff"); } @font-face /* 20 */ { font-family: MJXTEX-V; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Regular.woff") format("woff"); } @font-face /* 21 */ { font-family: MJXTEX-VB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Bold.woff") format("woff"); } mjx-c.mjx-c1D45B.TEX-I::before { padding: 0.442em 0.6em 0.011em 0; content: "n"; } mjx-c.mjx-c2265::before { padding: 0.636em 0.778em 0.138em 0; content: "\2265"; } mjx-c.mjx-c31::before { padding: 0.666em 0.5em 0 0; content: "1"; } cluster IDs, the activation examples are the documents that contain consecutive tokens (highlighted) whose activations belong to the queried clusters. The activation examples are ranked by similarity score: the geometric mean distance from the cluster centroid.
Via the interface, you can explore clusters vertically by moving through different levels of the tree (from level 6, corresponding to 2^6 clusters, to level 11, corresponding to 2^11 clusters), or horizontally by chaining clusters together. Once a pixel is selected on the grid, each next pixel's color indicates whether appending the corresponding cluster to the query would yield activation examples. The darker the pixel, the more activation examples the query produces.
Finally, the 8×8 grid at level 6 is not a 2D embedding of the 64 centroids. The clusters are laid down via quadtree placement according to their labels 0..63, assigned during the clustering in arbitrary order.
2. Qualitative AnalysisIn this section, I manually observe activation examples for various cluster queries to establish, at least qualitatively, whether the activations tend to organize around meaningful, coherent, and human-interpretable categories.
We begin with a cross-sectional analysis: we study the activation examples for six different clusters at the first level of the tree (k=64). Figure 2 shows six clusters and their corresponding activation examples. Even at the first level of the tree, it is already possible to see that activations geometrically structure themselves around lexical categories (“adjectives”, “verbs”, “nouns”) and semantic categories (“place names”, “biomedical stuff”, “drug codes”)
Figure 2: Cross-sectional analysis of six clusters at level 6 (k=64).
Next, we perform a vertical analysis: we follow one cluster down its subtree. Figure 3 shows the splits, at level 8, of cluster 14 (previously labeled as “nouns”). We can identify four semantic subcategories of nouns: “objects” (cluster 14.01), “spaces” (14.00), “concepts” (14.11), and “groups” (14.10). The first bit reveals how the “nouns” cluster was split at level 7: “physical stuff” (14.0) vs “social/abstract things” (14.1). Overall, the hierarchical taxonomy through vertical movement is not always the rule: binary splits might mix the same tokens on both sides or collapse onto one word. For example, “verbs” (cluster 2) does not separate into meaningful classes at level 7: is/was/are/be sit on both children.
Lastly, we perform a horizontal analysis: we chain together various clusters at level 11. We first identify six clusters corresponding to “was/were” (2.01011), “constructing verbs” (55.01010), "inaugurative verbs” (46.01110), and the prepositions “in” (28.01010), “by” (40.10001), and “to” (80.00000). Figure 4 highlights, in red, the query corresponding to “was/were” (2.01011) + "inaugurative verbs” (46.01110) and “in” (28.01010). The activation examples corresponding to that query are exactly what you would expect. Similarly, coherent activation examples arise from all six possible combinations of 2.01011 + 55.01010/46.01110 + 28.01010/40.10001/80.00000.
This example is not cherry-picked: chains of clusters very often yield highly coherent activation examples. I invite the reader to play around with the interface by drawing chains of clusters and observe if any coherent pattern from the activation examples emerges.
Figure 3: Vertical analysis of cluster 14 at level 8 (k=256)
Figure 4: Horizontal analysis of the combination of various clusters at level 11 (k=2048).
One amusing additional finding is that several clusters, all originating from the same top-level cluster 28, match the preposition “in”. Cluster 28.01010, identified before, corresponds to “in after a verb”; cluster 28.00100 corresponds to “in within medical contexts”, while cluster 28.11100 corresponds to “in as the start of a sentence”. This finding shows that activations organize not only around lexical and semantic categories, but also around syntactic categories.
Before moving to the next section, it is important to highlight, at the risk of being obvious, that clusters’ labels are assigned post hoc and are never seen during training. Indeed, the clustering procedure is fully unsupervised: no lexical/syntactic/semantic category is provided during training, and any human-interpretable structure emerges solely from the geometry of the model’s activations.
Previously, we defined a feature as “a direction in activation space associated with a property that is both interpretable to humans and useful to the model”. Clusters’ centroids are directions in the activation space, but the results of the qualitative analysis are not enough to claim that they are “interpretable to humans” or “useful to the model“. In order to establish more rigorously whether clustering can recover features, we need to measure:
- Whether the activation examples for a given cluster follow a coherent, human-interpretable pattern
- Whether the identified clusters have a causal downstream effect on model behavior.
To measure the coherence of activation examples for a given cluster, we run an "actual vs. decoy" distinguishability evaluation. The eval proceeds as follows.
Configurations. Seven: six motio configs (one per tree level, 6–11) plus an SAE baseline for the same LLM model and same layer (Neuronpedia's qwen2.5-7b-it/20-matryoshka-65k)
Bin preparation. An activation example is an activating token plus its context window (up to 25 tokens before and after) that satisfies a requirement given a source. For motio, the source is a cluster ID and the requirement is that the activation corresponding to the activating token falls inside that cluster. Motio activation examples are ranked by similarity scores. For SAE, the source is an SAE feature, and the requirement is that the SAE feature's activation has strength > 0 for the activating token. SAE activation examples are ranked by activation score. For a given motio configuration, one bin per cluster ID is created. For the SAE configuration, one bin per SAE feature is created. Each bin keeps the top-20 deduplicated activation examples and saves them as a snippet, with the activating token wrapped in <<...>>. Bins with fewer than 20 activation examples are dropped, so the actual bin count can be lower.
Trials. Per config: 500 trials. Each trial samples 21 distinct bins: the actual set includes all 20 snippets from one bin; the decoy set includes one random snippet from each of the other 20. We shuffle the snippets within each set, then randomly assign the two sets to set_0/set_1.
Judging. An LLM judge (Gemini 2.5 Flash, temperature 0) receives both sets and is tasked with identifying the actual set by outputting an index (0 or 1). A trial is correct when the judge's output matches the index corresponding to the actual set. The coherence score for a config is #correct/500. A random judge scores 0.5, any score above that suggests that the configuration captures human-recognizable patterns.
Figure 5: Distinguishability eval results (accuracy per configuration)
The results (Figure 5) show that, for the motio config, an LLM can distinguish the actual sets from the decoy set with near-perfect accuracy. This result confirms the observations that emerged from the qualitative analysis. Surprisingly, accuracy is already near-perfect at the first tree level of clustering (k=64) and improves only slightly as we traverse the tree and increase the dictionary size.
What’s even more surprising is the comparison with the SAE config baseline. The SAE has 65,536 features, while motio config at level 6 forms 64 clusters. The dictionary size differs by 1024x. Therefore, you would expect that the finer SAE dictionary produces much narrower concepts, and narrow concepts should produce coherent activation examples. Nevertheless, the results reveal the opposite: actual sets from Motio clusters, even at level 6, can be distinguished from decoys more accurately than the actual sets from SAE features.
On the other hand, the results can be explained by differences in bins’ candidate depths. At level 6, there are ≈161k candidate activation examples per cluster (10.32M activations / 64 clusters), so the top-20 activation examples are the top 0.01% candidates per cluster. In contrast, SAE has ≈ 39 candidates per SAE feature (2.55M activations / 65,536 SAE features), so the top-20 are about 50% of the available candidates per SAE feature. Therefore, the SAE’s bins include the long tail of low-activation examples that might confuse the LLM judge. This hypothesis aligns with findings from Huben (2024): the long tail of activation examples for a given SAE feature is often uninterpretable.
4. Intervention via Activation PatchingTo measure whether an identified cluster has a causal downstream effect on model behavior, we run an activation patching intervention.
Given a cluster, take tokens [0…pos] as the prompt and generate a continuation twice: once untouched (baseline) and once patched. In the patched run, we replace the residual-stream activation after block 20 at position pos with an activation that has the chosen cluster centroid’s direction and the same distance from the dataset mean as the original activation. We apply the patch once at the last token during prefill; the patched activation then persists through decoding via the KV cache.
Figure 6 shows the intervention for cluster 35.00000 (“org package”), whose activation examples are typically the token org inside Java import paths, and a prompt ending at import java. Untouched decoding continues in the java.* namespace (.io.InputStream, then more java imports). After patching, decoding instead continues in the org.* namespace (.jboss.modules.ModuleIdentifier, then org.junit.Test).
Figure 6: activation patching eval - org package
Figure 7 shows the intervention for cluster 61.01100 (“sentence final .”) and a prompt truncated mid-sentence “... was”. Untouched decoding continues the sentence. After patching, decoding instead outputs a capital letter, as if it were starting a new sentence.
Figure 7: activation patching eval - sentence final .
Both examples show the same behavior. After patching, the model behaves as if the last token was the one correlated with the cluster centroid and continues the decoding accordingly, overriding the prompt’s natural continuation.
5. ConclusionsWe applied recursive binary clustering to activations from Qwen2.5-7B-Instruct (layer 20) and asked whether the clustering would discover features.
The qualitative analysis (Section 2) showed, anecdotally, that clusters are organized around meaningful, coherent, and human-interpretable concepts, resembling lexical, syntactic, or semantic categories.
The experiments in Sections 3 and 4 provide compelling evidence that clustering can recover features: activation examples for a given cluster follow a coherent, human-interpretable pattern ("actual vs. decoy" distinguishability evaluation) and clusters have a causal downstream effect on model behavior (intervention via activation patching).
The results align with Elman (1990), despite a far more complex model architecture.
To the best of my knowledge, there’s no prior research on LLM feature discovery via clustering. Neuronpedia doesn’t include any cluster-based feature discovery interface, and there’s no trace of such a technique used in the context of feature discovery in Learn Mech Interp. Lastly, research via Perplexity suggests that the closest results come from Rumbelow (2026), albeit using a more intricate algorithm. This is surprising given the technique's simplicity and computational affordability.
To further corroborate the hypothesis that features can be discovered via clustering, further evaluations are needed. Fortunately, a large suite of SAE evaluation frameworks (SAEBench from Karvonen et al., Huben (2024): 2025 and RAVEL from Huang et al., 2024) can be repurposed to evaluate features discovered via clustering. Additionally, it is necessary to evaluate whether clustering escapes the limitations observed in SAEs, such as instability across seeds (Paulo and Belrose, 2025) and indifference to randomly initialized LLMs (Heap et al., 2025). Labels can be assigned to clusters via traditional autointerpretability (Bills et al., 2023) or via Natural Language Autoencoders (Kit et al., 2026).
Lastly, it would be interesting to study the trajectories activations follow between states (i.e., between following tokens) to examine whether any semantic laws of motion exist (Wolfram, 2023). Elman (1990) suggests studying what sort of attractors develop in transitions between states by carrying out a principal component analysis of the activation-pattern time series, and then constructing phase-state portraits of the most significant principal components.
ReferencesBills, Steven, et al. "Language models can explain neurons in language models." 2023.
Elman, Jeffrey L. "Finding structure in time." 1990.
Fraser-Taliente, Kit, et al. "Natural language autoencoders produce unsupervised explanations of LLM activations." 2026.
Heap, Thomas, et al. "Sparse autoencoders can interpret randomly initialized transformers." 2025.
Huang, Jing, et al. "RAVEL: Evaluating interpretability methods on disentangling language model representations." 2024.
Huben, Robert. (username: Robert_AIZI) “Comments on Anthropic's Scaling Monosemanticity.” 2024
Karvonen, Adam, et al. "SAEBench: A comprehensive benchmark for sparse autoencoders in language model interpretability." 2025.
Paulo, Gonçalo, and Nora Belrose. "Sparse autoencoders trained on the same data learn different features." 2025.
Rumbelow, Jessica. "Exemplar Partitioning for Mechanistic Interpretability." 2026.
Wolfram, Stephen. “What Is ChatGPT Doing:... and Why Does It Work?” 2023.
- ^
The definition modifies the one from learnmechinterp.com by adding the “human-interpretable” piece
- ^
Calculated via Euclidean distance
Discuss
Tips for less dysfunctional churches
I recently watched the film Spotlight on a date. It's a film about the Catholic Church's cover-up of child sexual abuse. This might seem like a downer, but the other option was Schindler’s List.
About halfway through I said we needed to stop. I have had some experiences relating to abuse in churches (and other communities) and it all got a bit much. In some ways, though, it felt kind of cathartic.
It crystallised my desire to give some thoughts to Christians on church structure, though perhaps non-Christians might find it interesting too. Largely I hope to avoid some failure modes I have seen evidence of.
Before I do, I want to try and make clear what I am, and am not, talking about. When I am talking about dysfunctional churches, I am not talking about theology, I think that if I described behaviours, most readers, Christian and non-Christian, would agree they were bad. Some examples of these churches would be:
- The Catholic Church in the run-up to the child sexual abuse scandal
- The Mars Hill church by the end, which I would characterise as bullying and a culture where Mark Driscoll largely had immunity from criticism
- Churches with bullying, sexual harassment, child abuse
When I talk to Christians about whether they could be wrong, they often have a notion that they’re praying, they’re faithful, so they’ll just do their best and God will sort it out. I am not arguing that.
Instead, I want to present suggestions that exist within the scope of the Bible (I can provide chapter and verse if you want) that I think reduce broadly agreed harms. Many people at dysfunctional churches are faithful[1] Christians. But I think they are fundamentally missing out on tools that even the Bible provides for them, to their suffering and to mine, as one of their friends.
So what are some patterns I see?
- Leadership put on a pedestal so as to never be capable of failing standards
- Upstanding members of the church leaving
- People who are allowed unique behaviours that wouldn't be allowed for others
- Information being siloed
So here are some ways that churches can push back against these patterns, which hopefully would have ended a lot of heartache, including that I have seen and that I worry I might see more of in the future.
Standards for LeadershipThe Bible has standards for leadership (1 Timothy 3:1-13, Titus 1:5-9, 1 Peter 5:1-3), does your church?
Are there standards for leadership? Are there things that a leader cannot do even though they are powerful and bring people through the door? And if they behave badly, would this be shared amongst other leaders or would this be kept secret?
I want to be in communities where there are some behaviours which are a bright line. Partly because this seems good in itself, but partly because it means that some things can’t be hidden. It is my view that no one is useful enough to a community that they should get away with murder. The same is true of sexual harassment and moderate levels of bullying.
I am not exaggerating here. How did the Catholic Church come to the conclusion that moving paedophile priests from location to location was okay? How did Mark Driscoll get away with sidelining and bullying people for so long? It’s easy to think “this person is too valuable to let go” and if there is no policy, then everyone has to make up his or her own mind.
Personally I’m pretty pro gossip, but churches sometimes aren’t, which means that you just might not hear about it. It should be clear that at some point, people must share what they know.
People leavingAre people leaving the church in general, maybe particularly people who you wouldn't expect to be leaving? And why?
It's often easy to come up with individual reasons, but if it's happening a lot, have you taken an hour to think if there is a pattern? Have you consulted with external church leaders or wise Christians who would tell you to your face of issues? Would you listen to them if they suggested things? Have you made changes?
What would external people of integrity think?I think it's worth considering what people outside of your group would think, perhaps people from local churches, or other denominations of Christians, perhaps even other faiths or atheists. Let us imagine we are Catholics. Could we justify child molestation being covered up for decades or centuries? I doubt I could.
In some way, I think my friends have been useful to me in the opposite direction. Is it normal that the guy who funded me made his money from crypto, a notoriously scammy field? Do I have trouble finding someone to date, or am I just afraid of commitment? Are trans women literally women? These kinds of things are easy enough to say in front of ideological allies but it gets trickier in front of someone who I respect who believes different things. I am glad I have friends who disagree on important issues.
How do you respond to bad outcomes?One way to try and tie the above together is to think about how you respond if people leave or if something is reported or if someone has a bad sense of the situation. Are you defensive? Or do you think they might have something of value to share?
More dysfunctional churches seem not to want to look at these questions, even sometimes afterwards. Things are somebody else's problem or due to somebody else's bad behaviour, there's never responsibility on them to figure out or to improve the situation.
If there were a big issue in your church and people were harmed, who would feel bad about it afterwards? And is that person responsible now, or kept away from being able to change anything?
And now some general thoughts:
None of that woke nonsenseWokeness has at times gone too far, but accountability and HR departments did exist for a reason. I went to uni in the heady years of the early 2010s as an evangelical Christian. And at times I found myself in surprising agreement with feminism. I too thought the constant dehumanisation of female students was gross. I couldn’t believe the constant low-level sexual assault that my female hall-mates put up with, or sometimes even laughed about.
I think at times churches can forget how dysfunctional human societies in general are and put aside that we have built technologies to deal with this. And that, while one might think that feminism has gone too far, some of the changes seem pretty good. There is no need to throw the baby out with the bathwater.
So, it seems worth looking at what other institutions use to avoid disastrous outcomes and taking some of those:
- Hotlines
- Ensuring that no single person is irreplaceable
- Mandatory reporting of some issues
- External safeguarding officers
How does Jesus deal with compromise against godlessness? Does he call for people to take arms? Does he rouse a counter-rebellion? Did the ends justify all means? Or does he often push for strange mid ways that one wouldn't expect? Allowing himself to be betrayed by Judas. Telling his followers not to fight. Paying his taxes by catching fish.
The least healthy churches (and indeed non-religious social movements) seem to push singularly towards one goal. Anyone who stands against them is their enemy. Anybody who questions is under suspicion.
But maximisation is perilous. As we try and get more and more of one thing, eventually we may lose everything else. And if you think your god is real, maybe you think 'He'[2] is big enough to deal with your church taking a couple of years longer to get to the goal without purging those who disagree. Or without that spectacular but flawed leader.
Corporate sin
A common problem I have had with Christians is a need to see all sin as some individual's problem. Not as a group problem but ultimately at the door of one individual. But this, in turn, meant that some issues never got dealt with because it could never be quite pinned down whose fault it was. And that some issues that ought to have been seen as a more broad problem were just labelled to be one specific person's issue.
As people left the church, it was never really engaged with why this was happening.
I don't know how you wrestle with this one theologically[3], but it seems to me that some systems cause people to live better lives, and some systems don't. I mean, presumably you think that it's good for atheists to have long-term marriages. Why? And so surely there are better and worse church systems also.
Wrapping upConservative Evangelicals can get tied up in their own beliefs, and what looks to them like more and more trust in God can look to me like doubling down on behaviour that hasn't worked and that is hurting people. This isn't unique to Christians. I see many people have this kind of behaviour, and it can be differently difficult to talk to each group.
Because my desire in this piece isn't that you stop being a Christian. It's that you don't end up being wrapped up in something that hurts you. If you saw yourself from a distance, I don't want you to think, "How on earth did they end up here?" It is maddening to see events that are likely to go wrong, like some sort of slow-motion car crash, and then have to watch the car crash take place. And it's just sad.
My desire for you and your church would be to have fewer sexual scandals, fewer bullying scandals. That there is not some huge breakdown that causes half your members to leave, and that scars you for the next 15-20 years. I think there are some more or less likely ways of getting to that, and as a friend, I hope I can encourage you to make it less likely.
- ^
In the sense that I would predict my evangelical friends would consider them faithful.
- ^
I am happy to capitalise Yahweh, but I am not gonna do it for pronouns without inverted commas. Unclear what the respectful approach is.
- ^
Claude suggests the systemic approach of Acts 6:1-6, where a new office is created for the support of some specific Jewish widows.
Discuss
A Four-Axis Bayesian Epoch Capabilities Index with Human Baselines
This is a crosspost from the General-Purpose AI Policy Lab research blog.
The Epoch Capabilities Index compresses many benchmark scores into one for each model, following the framework of the Rosetta Stone paper. In a previous post, we added human baselines to the same scale to see how the models compare to humans. But one issue is that most humans score near-perfectly on abstract reasoning benchmarks like ARC-AGI or VPCT and sit near chance on GPQA-type benchmarks, while many models show the opposite pattern. One index cannot produce both of these orderings, so in this post we move to a Bayesian setup with four skill axes instead of one, and proper uncertainty estimation.
TL;DR- We extended the Epoch Capabilities Index to four skill axes, in a Bayesian setup where every ability comes with its uncertainty and the human tiers are fitted inside the model as test-takers.
- Sparse data (6% of the test-taker-by-benchmark matrix) means the model does not settle on one answer by itself. We need two ordering priors to help identify it, a hard ordering on the human tiers, and a soft expected improvement along releases of the same model family.
- Main Results:
- The four axes that come out of the model are Fluid Intelligence, Scientific Knowledge and Reasoning, Agentic Capabilities, and Legacy QA.
- Models have now passed every human tier on Scientific Knowledge and Reasoning, though this arguably says more about breadth of recall than about doing actual science.
- On Agentic Capabilities, the frontier is reaching the Skilled Generalist baseline around now (summer 2026).
- On Fluid Intelligence, AI models have likely passed Average Humans, but human experts still lead. The trend, if it stays linear, will cross top tiers in 2027 and 2028.
- Four of the ten runs converge on a second mode that places human tiers slightly differently and move these dates by a few months to years.
Following Alexander Barry's Bayesian version of the Epoch Capabilities Index (ECI), we rebuilt the model in Python using PyMC. In this Bayesian setup, instead of finding one single best value for each parameter, we sample a whole distribution of plausible ones, so every ability comes with uncertainty. The full setup is in the model section below. We also fit the nine human tiers as test-takers next to the models, from Average Human up to Committees of Domain Experts, plus two high-school tiers. The baselines and the benchmark table have been updated and expanded since the previous post (the full tables with sources are in the appendix), and we added a partial ordering prior on human tiers (see following sections). For the 1D fit only, we excluded the human-easy benchmarks like ARC-AGI or VPCT (in a similar fashion to the previous post; the list is in the Appendix).
Here's our rebuilt index (called ECI-H with H for the human baselines) matched with the ECI scores for the state-of-the-art models:
Figure 1: SOTA models: our ECI-H (90% HDI) vs Epoch's published values (CIs where published). Epoch publishes one ECI per model, repeated here across its effort variants, so for example the four Claude Fable 5 rows share one orange value.
Our values differ from Epoch's for three main reasons. First, we do not fit the same table, since our benchmark set is larger and drops eight human-easy benchmarks. Second, Epoch publishes one value per model, taking its best score on each benchmark, while we fit every thinking-effort variant as its own test-taker with its own scores. Third, we include human baselines.
Each benchmark also gets a difficulty on the same scale, so we can plot models, benchmarks and the human tiers together over time:
Figure 2: AI capability, human baselines and benchmark difficulty on the ECI-H scale, with 80% intervals. Green points are AI models at their release date, pink points are benchmark difficulties. Dashed lines are the human tiers. The scale is pinned at Claude 3.5 Sonnet = 130 and GPT-5 (medium) = 150 to match Epoch's.
However, putting AIs and humans on a single axis is arguably quite objectionable.
Multidimensional extensionAs we mentioned earlier and discussed in the previous post, some benchmarks are trivial for humans and hard for AI models, which breaks the single difficulty axis. Epoch's Benchmark Scores = General Capability + Claudiness also points to scores carrying more than one dimension (and that's between models alone). We test this by extending the model to four skill axes using an MIRT (Multidimensional Item Response Theory) model, commonly used in psychometrics, while keeping all the benchmarks and human baselines.
The intuition behind the model is that each test-taker mjx-container[jax="CHTML"] { line-height: 0; } mjx-container [space="1"] { margin-left: .111em; } mjx-container [space="2"] { margin-left: .167em; } mjx-container [space="3"] { margin-left: .222em; } mjx-container [space="4"] { margin-left: .278em; } mjx-container [space="5"] { margin-left: .333em; } mjx-container [rspace="1"] { margin-right: .111em; } mjx-container [rspace="2"] { margin-right: .167em; } mjx-container [rspace="3"] { margin-right: .222em; } mjx-container [rspace="4"] { margin-right: .278em; } mjx-container [rspace="5"] { margin-right: .333em; } mjx-container [size="s"] { font-size: 70.7%; } mjx-container [size="ss"] { font-size: 50%; } mjx-container [size="Tn"] { font-size: 60%; } mjx-container [size="sm"] { font-size: 85%; } mjx-container [size="lg"] { font-size: 120%; } mjx-container [size="Lg"] { font-size: 144%; } mjx-container [size="LG"] { font-size: 173%; } mjx-container [size="hg"] { font-size: 207%; } mjx-container [size="HG"] { font-size: 249%; } mjx-container [width="full"] { width: 100%; } mjx-box { display: inline-block; } mjx-block { display: block; } mjx-itable { display: inline-table; } mjx-row { display: table-row; } mjx-row > * { display: table-cell; } mjx-mtext { display: inline-block; text-align: left; } mjx-mstyle { display: inline-block; } mjx-merror { display: inline-block; color: red; background-color: yellow; } mjx-mphantom { visibility: hidden; } _::-webkit-full-page-media, _:future, :root mjx-container { will-change: opacity; } mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; font-weight: normal; font-size: 100%; font-size-adjust: none; letter-spacing: normal; border-collapse: collapse; word-wrap: normal; word-spacing: normal; white-space: nowrap; direction: ltr; padding: 1px 0; } mjx-container[jax="CHTML"][display="true"] { display: block; text-align: center; margin: 1em 0; } mjx-container[jax="CHTML"][display="true"][width="full"] { display: flex; } mjx-container[jax="CHTML"][display="true"] mjx-math { padding: 0; } mjx-container[jax="CHTML"][justify="left"] { text-align: left; } mjx-container[jax="CHTML"][justify="right"] { text-align: right; } mjx-mi { display: inline-block; text-align: left; } mjx-c { display: inline-block; } mjx-utext { display: inline-block; padding: .75em 0 .2em 0; } mjx-msub { display: inline-block; text-align: left; } mjx-mo { display: inline-block; text-align: left; } mjx-stretchy-h { display: inline-table; width: 100%; } mjx-stretchy-h > * { display: table-cell; width: 0; } mjx-stretchy-h > * > mjx-c { display: inline-block; transform: scalex(1.0000001); } mjx-stretchy-h > * > mjx-c::before { display: inline-block; width: initial; } mjx-stretchy-h > mjx-ext { /* IE */ overflow: hidden; /* others */ overflow: clip visible; width: 100%; } mjx-stretchy-h > mjx-ext > mjx-c::before { transform: scalex(500); } mjx-stretchy-h > mjx-ext > mjx-c { width: 0; } mjx-stretchy-h > mjx-beg > mjx-c { margin-right: -.1em; } mjx-stretchy-h > mjx-end > mjx-c { margin-left: -.1em; } mjx-stretchy-v { display: inline-block; } mjx-stretchy-v > * { display: block; } mjx-stretchy-v > mjx-beg { height: 0; } mjx-stretchy-v > mjx-end > mjx-c { display: block; } mjx-stretchy-v > * > mjx-c { transform: scaley(1.0000001); transform-origin: left center; overflow: hidden; } mjx-stretchy-v > mjx-ext { display: block; height: 100%; box-sizing: border-box; border: 0px solid transparent; /* IE */ overflow: hidden; /* others */ overflow: visible clip; } mjx-stretchy-v > mjx-ext > mjx-c::before { width: initial; box-sizing: border-box; } mjx-stretchy-v > mjx-ext > mjx-c { transform: scaleY(500) translateY(.075em); overflow: visible; } mjx-mark { display: inline-block; height: 0px; } mjx-mn { display: inline-block; text-align: left; } mjx-mspace { display: inline-block; text-align: left; } mjx-mrow { display: inline-block; text-align: left; } mjx-munderover { display: inline-block; text-align: left; } mjx-munderover:not([limits="false"]) { padding-top: .1em; } mjx-munderover:not([limits="false"]) > * { display: block; } mjx-msubsup { display: inline-block; text-align: left; } mjx-script { display: inline-block; padding-right: .05em; padding-left: .033em; } mjx-script > mjx-spacer { display: block; } mjx-TeXAtom { display: inline-block; text-align: left; } mjx-mfrac { display: inline-block; text-align: left; } mjx-frac { display: inline-block; vertical-align: 0.17em; padding: 0 .22em; } mjx-frac[type="d"] { vertical-align: .04em; } mjx-frac[delims] { padding: 0 .1em; } mjx-frac[atop] { padding: 0 .12em; } mjx-frac[atop][delims] { padding: 0; } mjx-dtable { display: inline-table; width: 100%; } mjx-dtable > * { font-size: 2000%; } mjx-dbox { display: block; font-size: 5%; } mjx-num { display: block; text-align: center; } mjx-den { display: block; text-align: center; } mjx-mfrac[bevelled] > mjx-num { display: inline-block; } mjx-mfrac[bevelled] > mjx-den { display: inline-block; } mjx-den[align="right"], mjx-num[align="right"] { text-align: right; } mjx-den[align="left"], mjx-num[align="left"] { text-align: left; } mjx-nstrut { display: inline-block; height: .054em; width: 0; vertical-align: -.054em; } mjx-nstrut[type="d"] { height: .217em; vertical-align: -.217em; } mjx-dstrut { display: inline-block; height: .505em; width: 0; } mjx-dstrut[type="d"] { height: .726em; } mjx-line { display: block; box-sizing: border-box; min-height: 1px; height: .06em; border-top: .06em solid; margin: .06em -.1em; overflow: hidden; } mjx-line[type="d"] { margin: .18em -.1em; } mjx-munder { display: inline-block; text-align: left; } mjx-over { text-align: left; } mjx-munder:not([limits="false"]) { display: inline-table; } mjx-munder > mjx-row { text-align: left; } mjx-under { padding-bottom: .1em; } mjx-c::before { display: block; width: 0; } .MJX-TEX { font-family: MJXZERO, MJXTEX; } .TEX-B { font-family: MJXZERO, MJXTEX-B; } .TEX-I { font-family: MJXZERO, MJXTEX-I; } .TEX-MI { font-family: MJXZERO, MJXTEX-MI; } .TEX-BI { font-family: MJXZERO, MJXTEX-BI; } .TEX-S1 { font-family: MJXZERO, MJXTEX-S1; } .TEX-S2 { font-family: MJXZERO, MJXTEX-S2; } .TEX-S3 { font-family: MJXZERO, MJXTEX-S3; } .TEX-S4 { font-family: MJXZERO, MJXTEX-S4; } .TEX-A { font-family: MJXZERO, MJXTEX-A; } .TEX-C { font-family: MJXZERO, MJXTEX-C; } .TEX-CB { font-family: MJXZERO, MJXTEX-CB; } .TEX-FR { font-family: MJXZERO, MJXTEX-FR; } .TEX-FRB { font-family: MJXZERO, MJXTEX-FRB; } .TEX-SS { font-family: MJXZERO, MJXTEX-SS; } .TEX-SSB { font-family: MJXZERO, MJXTEX-SSB; } .TEX-SSI { font-family: MJXZERO, MJXTEX-SSI; } .TEX-SC { font-family: MJXZERO, MJXTEX-SC; } .TEX-T { font-family: MJXZERO, MJXTEX-T; } .TEX-V { font-family: MJXZERO, MJXTEX-V; } .TEX-VB { font-family: MJXZERO, MJXTEX-VB; } mjx-stretchy-v mjx-c, mjx-stretchy-h mjx-c { font-family: MJXZERO, MJXTEX-S1, MJXTEX-S4, MJXTEX, MJXTEX-A ! important; } @font-face /* 0 */ { font-family: MJXZERO; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Zero.woff") format("woff"); } @font-face /* 1 */ { font-family: MJXTEX; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Regular.woff") format("woff"); } @font-face /* 2 */ { font-family: MJXTEX-B; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Bold.woff") format("woff"); } @font-face /* 3 */ { font-family: MJXTEX-I; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-Italic.woff") format("woff"); } @font-face /* 4 */ { font-family: MJXTEX-MI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Italic.woff") format("woff"); } @font-face /* 5 */ { font-family: MJXTEX-BI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-BoldItalic.woff") format("woff"); } @font-face /* 6 */ { font-family: MJXTEX-S1; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size1-Regular.woff") format("woff"); } @font-face /* 7 */ { font-family: MJXTEX-S2; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size2-Regular.woff") format("woff"); } @font-face /* 8 */ { font-family: MJXTEX-S3; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size3-Regular.woff") format("woff"); } @font-face /* 9 */ { font-family: MJXTEX-S4; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size4-Regular.woff") format("woff"); } @font-face /* 10 */ { font-family: MJXTEX-A; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_AMS-Regular.woff") format("woff"); } @font-face /* 11 */ { font-family: MJXTEX-C; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Regular.woff") format("woff"); } @font-face /* 12 */ { font-family: MJXTEX-CB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Bold.woff") format("woff"); } @font-face /* 13 */ { font-family: MJXTEX-FR; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Regular.woff") format("woff"); } @font-face /* 14 */ { font-family: MJXTEX-FRB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Bold.woff") format("woff"); } @font-face /* 15 */ { font-family: MJXTEX-SS; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Regular.woff") format("woff"); } @font-face /* 16 */ { font-family: MJXTEX-SSB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Bold.woff") format("woff"); } @font-face /* 17 */ { font-family: MJXTEX-SSI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Italic.woff") format("woff"); } @font-face /* 18 */ { font-family: MJXTEX-SC; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Script-Regular.woff") format("woff"); } @font-face /* 19 */ { font-family: MJXTEX-T; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Typewriter-Regular.woff") format("woff"); } @font-face /* 20 */ { font-family: MJXTEX-V; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Regular.woff") format("woff"); } @font-face /* 21 */ { font-family: MJXTEX-VB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Bold.woff") format("woff"); } mjx-c.mjx-c1D45A.TEX-I::before { padding: 0.442em 0.878em 0.011em 0; content: "m"; } mjx-c.mjx-c1D703.TEX-I::before { padding: 0.705em 0.469em 0.01em 0; content: "\3B8"; } mjx-c.mjx-c1D434.TEX-I::before { padding: 0.716em 0.75em 0 0; content: "A"; } mjx-c.mjx-c1D437.TEX-I::before { padding: 0.683em 0.828em 0 0; content: "D"; } mjx-c.mjx-c1D44F.TEX-I::before { padding: 0.694em 0.429em 0.011em 0; content: "b"; } mjx-c.mjx-c1D70E.TEX-I::before { padding: 0.431em 0.571em 0.011em 0; content: "\3C3"; } mjx-c.mjx-c1D450.TEX-I::before { padding: 0.442em 0.433em 0.011em 0; content: "c"; } mjx-c.mjx-c1D707.TEX-I::before { padding: 0.442em 0.603em 0.216em 0; content: "\3BC"; } mjx-c.mjx-c3D::before { padding: 0.583em 0.778em 0.082em 0; content: "="; } mjx-c.mjx-c2B::before { padding: 0.583em 0.778em 0.082em 0; content: "+"; } mjx-c.mjx-c28::before { padding: 0.75em 0.389em 0.25em 0; content: "("; } mjx-c.mjx-c31::before { padding: 0.666em 0.5em 0 0; content: "1"; } mjx-c.mjx-c2212::before { padding: 0.583em 0.778em 0.082em 0; content: "\2212"; } mjx-c.mjx-c29::before { padding: 0.75em 0.389em 0.25em 0; content: ")"; } mjx-c.mjx-c28.TEX-S2::before { padding: 1.15em 0.597em 0.649em 0; content: "("; } mjx-c.mjx-c2211.TEX-S1::before { padding: 0.75em 1.056em 0.25em 0; content: "\2211"; } mjx-c.mjx-c34::before { padding: 0.677em 0.5em 0 0; content: "4"; } mjx-c.mjx-c1D458.TEX-I::before { padding: 0.694em 0.521em 0.011em 0; content: "k"; } mjx-c.mjx-c2C::before { padding: 0.121em 0.278em 0.194em 0; content: ","; } mjx-c.mjx-c29.TEX-S2::before { padding: 1.15em 0.597em 0.649em 0; content: ")"; } mjx-c.mjx-c1D466.TEX-I::before { padding: 0.442em 0.49em 0.205em 0; content: "y"; } mjx-c.mjx-c223C::before { padding: 0.367em 0.778em 0 0; content: "\223C"; } mjx-c.mjx-c42::before { padding: 0.683em 0.708em 0 0; content: "B"; } mjx-c.mjx-c65::before { padding: 0.448em 0.444em 0.011em 0; content: "e"; } mjx-c.mjx-c74::before { padding: 0.615em 0.389em 0.01em 0; content: "t"; } mjx-c.mjx-c61::before { padding: 0.448em 0.5em 0.011em 0; content: "a"; } mjx-c.mjx-c28.TEX-S1::before { padding: 0.85em 0.458em 0.349em 0; content: "("; } mjx-c.mjx-c1D719.TEX-I::before { padding: 0.694em 0.596em 0.205em 0; content: "\3D5"; } mjx-c.mjx-cA0::before { padding: 0 0.25em 0 0; content: "\A0"; } mjx-c.mjx-c29.TEX-S1::before { padding: 0.85em 0.458em 0.349em 0; content: ")"; } mjx-c.mjx-c32::before { padding: 0.666em 0.5em 0 0; content: "2"; } mjx-c.mjx-c220F.TEX-S1::before { padding: 0.75em 0.944em 0.25em 0; content: "\220F"; } mjx-c.mjx-c1D6FE.TEX-I::before { padding: 0.441em 0.543em 0.216em 0; content: "\3B3"; } mjx-c.mjx-c1D457.TEX-I::before { padding: 0.661em 0.412em 0.204em 0; content: "j"; } mjx-c.mjx-c73::before { padding: 0.448em 0.394em 0.011em 0; content: "s"; } mjx-c.mjx-c6F::before { padding: 0.448em 0.5em 0.01em 0; content: "o"; } mjx-c.mjx-c66::before { padding: 0.705em 0.372em 0 0; content: "f"; } mjx-c.mjx-c70::before { padding: 0.442em 0.556em 0.194em 0; content: "p"; } mjx-c.mjx-c6C::before { padding: 0.694em 0.278em 0 0; content: "l"; } mjx-c.mjx-c75::before { padding: 0.442em 0.556em 0.011em 0; content: "u"; } has four abilities , that form its skill profile, the way a student can be strong in algebra and weak in essay writing. Each benchmark weighs those skills through its four positive loadings , one per skill, which say how much each skill counts for that benchmark. The loadings also set how a benchmark separates its test-takers, what psychometrics literature calls discrimination. So a benchmark with a large loading splits weak models from strong ones clearly, while one with small loadings doesn't react to skill. The difficulty is the bar the weighted skills must clear to get more than the midpoint on the benchmark, and the S-curve turns the result into a score between 0 and 1. We also need to take into account the random-guessing for each benchmark so we fix each benchmark's guessing floor in advance and start the curve there instead of at 0, so scores on a four-option exam bottom out at 0.25 rather than 0.
As such the expected score for test-taker and benchmark is
and the observed score scatters around it with Beta noise (following Barry's post),
Each benchmark gets its own noise level .
The form we use above to define what's inside the sigmoid belongs to one of three families common in the IRT literature and it is called compensatory because a strong skill can make up for a weak one inside the sum. In the non-compensatory family a benchmark needs all its skills at once, and the sum becomes a product of per-axis curves, . The semi-compensatory family sits in between and adds interaction terms to the compensatory sum. We tried both alternatives and the non-compensatory fit did not converge, while the semi-compensatory one converged only under heavy constraints and made worse predictions.
Figure 3: The model as a graph.
Prior AssumptionsWe first tried to fit this model with no other assumptions than the ones explained above, but the model did not settle on one answer. This is due to our data being very sparse (the test-taker by benchmark matrix is filled only at 6%, and the average test-taker has about six scores) and the fact that many arrangements of abilities and loadings explain the scores equally well, so repeated runs fall on different solutions. Given this, we needed to put more prior information into the model to help it converge to one answer.
Human Ordering (hard prior)In the data, non-skilled humans are mostly tested on human-easy benchmarks and experts are tested mostly on hard benchmarks. Yet, we know that average humans would do worse than experts on the hard benchmarks, and that experts would do at least as well as average humans on easy benchmarks. So we gave the model a prior ordering[1] where a Domain Expert is at least as good as a Skilled Generalist, and a committee is at least as good as one of its members, on every skill. The ordering says nothing about the size of the gap between the tiers (where no ranking is obvious, like between a Top Performer and a committee of experts, we don't impose any ordering). The two high-school tiers join the ordering by a Domain Expert being at least as good as a High School Qualifier and a Top Performer at least as good as a High School Top Performer.
Figure 4: The human ordering. An arrow means at least as good on every skill, dashed marks a second parent, the tier sits above both, unconnected tiers are not compared.
Recent models sometimes lack data to estimate their ability scores, but within one release chain, like the GPT flagships and the Claude Opus line, we can expect each new release to improve on the one before it. A release can regress if the data says so; we only nudge it towards improving. We also use time between releases[2] for the difference in abilities, so the expected gain grows with the gap between releases, and a lab shipping many small updates is not expected to gain more than one shipping a single big release over the same year. Regarding the thinking-effort variants, they are tied to their base release, and we do not order them among themselves, as a higher effort can potentially overthink.
Figure 5: Each release is nudged above the previous one, more over longer gaps, thinking-effort variants attach to their release, unordered among themselves.
To see how the priors help, we can compare the runs. First without any prior, they split into two different sets of axes. After adding the human ordering alone we still get two solutions that disagree about the axes. Adding the model family assumption finally makes all the runs agree on one set of axes. We also tried three axes with all the assumptions on, and the runs still split in two. These assumptions also help improve the model's predictive ability since, on left-out scores (leave-one-out cross-validation) the final model beats the no-prior version by about 107 ± 18 and the 1D index by about 1,000 ± 33, comparing on the rows where the comparison is reliable.[3]
With all of the assumptions above in place, the model settles on a single answer for most of the runs[4]. Let's look at what it found.
The AxesWe name each axis after the benchmarks whose loading vectors are most collinear with it, i.e. the benchmarks that draw on that skill and almost nothing else. We get:
- Axis 1: We call it Fluid Intelligence, defined by ARC-AGI-2, ARC-AGI and VPCT, abstract puzzle benchmarks.
- Axis 2: We call it Scientific Knowledge and Reasoning, defined by WMDP Chemistry and Biology, the GPQA science subsets and FrontierMath.
- Axis 3: We call it Agentic Capabilities, defined by GBAEval, the Remote Labor Index and SWE-Bench Pro, benchmarks where a model works through long tasks rather than answering questions.
- Axis 4: We call it Legacy QA, consisting mainly of older question-answering benchmarks, largely saturated, OpenBookQA, ARC (AI2), BoolQ and similar benchmarks.
Figure 6: The 20 benchmarks that best define each axis. Bars represent loadings (median, 95% interval). Numbers on the right and the color gradient represent the level of collinearity with the axis.
And here's how the top models compare on each axis to humans:
Figure 7: Top models and the nine human tiers on each axis (mean, 95% interval, majority chains).
The frontier models sit above every human tier on Scientific Knowledge and Reasoning. On Fluid Intelligence the opposite holds as every tier (except Average Human) sits above the best models. On Legacy QA the humans also sit on top, but this is more of a data artifact since the eight benchmarks that define this axis most purely (axis share above one half) were last run on models from mid-2024 or earlier, with most of them already scoring around 0.9 there, and no frontier model was ever measured on them. So the human lead on this axis is a comparison against a frozen pool of older models, and remains untested against the actual frontier.
ForecastingFor these forecasts, only models whose ability on the axis is well estimated enter the pool (posterior SD below 0.33, plus flagged frontier releases). We take the record-setting frontier models for each plausible set of abilities the model produced, then fit a per-draw record envelope (running max of the frontier in each posterior draw), and extend it at its rate over a 1.5-year window to see when it reaches each human tier [5]. And as mentioned above, given how the recent models' positions on the Legacy QA axis come from the prior rather than actual ability, we leave this axis out of the forecasts. The headline numbers below are for the majority mode (6 of the 10 runs, see the Appendix for the minority mode).
Figure 8: Frontier trend per axis (majority chains). Dots are dated models, the dashed line is the record envelope extended at its recent rate, with its 80% band. The dashed lines are the human tiers.
On the Fluid Intelligence axis, the frontier likely sits above Average Humans, at the Skilled Generalist level, but still below expert baselines. It is on track to reach the Domain Expert by spring 2027 and pass the Top Performer by mid-2028.
On the Agentic Capabilities axis, models have already passed some of the lower tiers, with the frontier reaching the Skilled Generalist baseline right about now (we're writing this during the last days of August 2026), and projected to pass the Committee of Domain Experts by mid-2027.
On the Scientific Knowledge and Reasoning axis, models have passed every human tier already, with probabilities of 0.85 to 0.95 (since on WMDP Chemistry the Domain Expert baseline is 0.433 against a best model score of 0.809, on WMDP Biology 0.605 against 0.875, and on GPQA Diamond 0.812 against 0.946). In fact every model on this axis since 2023 sits above the human baselines, which says more about what these benchmarks reward, that is breadth of recall across a whole field, than about doing science and research (the Skilled Generalist is below chance on GPQA Diamond, 0.22, and even the in-domain PhD gets 0.43 on WMDP Chemistry).
Figure 9: Dates when the extrapolated trend reaches each human tier (median, 50% and 80% intervals, for the majority chains). Today corresponds to the first of September 2026.
The main assumption here is that the forecasted trend based on the recent rate keeps its slope, which seems reasonable since within our window the frontier shows no sign of decelerating, and on the Agentic Capabilities axis it even seems to accelerate with the latest releases. These dates should be read with the uncertainty the model gives them (at 95% the crossing windows[6] stretch by several more years, and far longer on the Agentic Capabilities axis), and the best way to tighten them would be better human baselines[7], especially on the Agentic Capabilities axis, where the human tiers rest on a handful of measurements.
LimitationsThe first limitation is the data itself. As explained in the previous post, we need more of it and data of better quality, especially for the human baselines, but also for the models, since we fill only 6% of the test-taker by benchmark matrix and this sparsity forced us to add more assumptions to the model.
The second one, which we share with the Rosetta Stone paper and the ECI in general, is that we fit benchmark-level scores instead of item-level answers, so we do not fully respect the assumptions of the MIRT setup.
The last one is the calibration, since our predictive intervals are wider than the data requires, making our model more conservative than it should be[8].
Further WorkBeyond collecting more data, a few directions are worth exploring:
- Running the legacy benchmarks on current frontier models. This would help settle the Legacy QA axis with data instead of it being an artifact, and it fills more of the grid at the same time.
- Ceilings on saturating benchmarks. We already fix a guessing floor per benchmark, so a ceiling either inferred from the data or fixed seems reasonable, and would help saturation be read as such and not as extreme difficulty.
- Testing the forecasts in both directions. Fitting the model on an older snapshot of the data and checking its forecasts against the releases that came out since, and going forward, keeping track of the predictions we made here and see how they pan out.
The code is available at github.com/General-Purpose-AI-Policy-Lab/Multiaxis_ECI/tree/blogpost-frozen.
AppendixAppendix ABelow is the full graphical specification of the model:
Figure 10: The full graphical representation of the model
Figure 11: The graphical representation of the human prior.
Figure 12: The graphical representation of the model side prior.
Appendix BThis appendix records the previous fits and how each assumption affects the model. For the 1D Setup, we use all the benchmarks to compare it with our final fit.
Fit
Runs x Draws
Divergences
Posterior modes (axis systems)
elpd (LOO) ± se
Δ vs final[9]
1D index
10 × 10,000
0
1
6,249.6 ± 112.1
−999.6 ± 33.2
4 axes, no priors
12 × 3,000
778
2
7,583.9 ± 75.9
−107.2 ± 18.3
4 axes, human ordering only
12 × 3,000
80
2
7,583.3 ± 73.1
−109.2 ± 17.2
4 axes, both priors
10 × 12,000
37[10]
1[11]
7,710.4 ± 76.3
0
3 axes, both priors
12 × 3,000
12
2[12]
7,447.8 ± 77.4
−191.1 ± 14.3
Appendix CFour of the ten runs place the human tiers differently. This section records what moves and how the forecasts change.
Figure 13: Human tiers, the six majority runs (blue) against the four others (orange)
- Minority minus majority, averaged over the nine tiers: −2.71 on Agentic, +1.58 on Legacy QA, +0.71 on Fluid Intelligence, −0.04 on Scientific Knowledge and Reasoning.
- Both groups share one axis system and describe the scores equally well.
Figure 14: The 18 models the two groups of runs place differently, all on the Agentic axis. Nine small 2023–2024 models move down and nine older 2022–2023 models move up.
Here are the forecasts for the minority chains:
Figure 15: Dates when the extrapolated trend reaches each human tier for the minority chains (median, 50% and 80% intervals). Today corresponds to the first of September 2026
On these minority forecasts, for the Agentic axis, we will soon have surpassed all of the human tiers. For Fluid Intelligence, frontier models are still below the average human contrary to what the majority chains predict.
Appendix DHere below are the crossover dates with the 95% intervals for both the majority and minority chains:
Figure 16: Dates when the extrapolated trend reaches each human tier for the majority chains (median, 95% interval). Today corresponds to the first of September 2026.
Majority chains: for all axes, the 95% timespans are much wider, but still close by 2035 at most.
Figure 17: Dates when the extrapolated trend reaches each human tier for the minority chains (median, 95% interval). Today corresponds to the first of September 2026
Minority chains: the Fluid Intelligence axis keeps its 95% intervals relatively tight, but otherwise uncertainties are huge.
Appendix EFor each observed score we compute where it lands inside the model's predictive distribution (called the probability integral transform, PIT). A perfectly calibrated model would spread these values evenly but our histogram bulges in the middle instead meaning that the observed scores land near the center of the predictive intervals more often than they should, so the intervals are wider than the data requires and make our model pretty conservative.
Figure 18: PIT of the final fit. A calibrated model is flat at density 1.
Appendix FDown below are the human scores we used as well as the exhaustive list of benchmarks included in the setup.
Human BaselinesBenchmark
Human Group
Score
Source
Information
ARC-AGI
Average Human
0.77
MTurk
ARC-AGI-2
Average Human
0.6
average test-taker
BIG-Bench Hard (BBH)
Average Human
0.677
average human raters
BoolQ
Average Human
0.9
Human annotators
CSQA2
Average Human
0.903
average accuracy of humans
MMLU
Average Human
0.345
MTurk
OpenBookQA
Average Human
0.92
probability from random human subjects
ScienceQA
Average Human
0.884
MTurk workers with a high school degree or higher who passed the qualification examples
SimpleBench
Average Human
0.837
nine non-specialized humans
SuperGLUE
Average Human
0.898
human performance estimates after training phase
TriviaQA
Average Human
0.797
human performance level
VPCT
Average Human
0.999
Epoch AI
three volunteers
ARC-AGI
Committee of Average Humans
0.98
Human panel (at least two participants solved one or more sub-pairs within their first two attempts)
ARC-AGI-2
Committee of Average Humans
0.999
Human panel
CSQA2
Committee of Average Humans
0.941
majority vote
HellaSwag
Committee of Average Humans
0.956
majority vote of 5 crowd workers (MTurk)
WinoGrande
Committee of Average Humans
0.94
majority vote of crowd workers (MTurk)
ARC-AGI
Skilled Generalist
0.98
STEM Graduates
GPQA Diamond
Skilled Generalist
0.219
highly skilled and incentivized non-experts who have or are pursuing PhDs in other domains
GPQA Diamond Biology
Skilled Generalist
0.22
not in-domain PhD
GPQA Diamond Chemistry
Skilled Generalist
0.22
not in-domain PhD
GPQA Main Biology
Skilled Generalist
0.43
not in-domain PhD
GPQA Main Chemistry
Skilled Generalist
0.31
not in-domain PhD
GSM8K
Skilled Generalist
0.9677
qualified human annotators who have passed a qualification exam with at least a bachelor's degree
MATH Level 5
Skilled Generalist
0.4
a computer science PhD student who does not especially like mathematics
OS World (Screenshot)
Skilled Generalist
0.724
individuals not familiar with the software
SimpleQA Verified
Skilled Generalist
0.944
human annotator going through the test
Visual Task Assessment (VISTA)
Skilled Generalist
0.554
16 full-time employees
PIQA
Committee of Skilled Generalists
0.949
majority vote of top annotators
OTIS Mock AIME 2024-2025
High School Qualifier
0.53
average score by high school students from the OTIS program (percentage from number of questions answered)
OTIS Mock AIME 2024-2025
High School Top Performer
0.93
top scorer from the OTIS program (percentage from number of questions answered)
BioLP-bench
Domain Expert
0.384
Bachelor's w/ lab experience
GPQA Diamond
Domain Expert
0.812
in-domain PhD validators, GPQA paper Table 2
GPQA Diamond Biology
Domain Expert
0.831
In-domain PhD
GPQA Diamond Chemistry
Domain Expert
0.831
In-domain PhD
GPQA Main Biology
Domain Expert
0.667
In-domain PhD
GPQA Main Chemistry
Domain Expert
0.72
In-domain PhD
LAB-Bench Cloning
Domain Expert
0.6
In-domain PhD
LAB-Bench LitQA2
Domain Expert
0.7
In-domain PhD
LAB-Bench Protocol
Domain Expert
0.79
In-domain PhD
LAB-Bench SeqQA
Domain Expert
0.78
In-domain PhD
MATH Level 5
Domain Expert
0.9
a three-time IMO gold medalist university student
MMLU
High School Top Performer
0.898
estimation from the authors based the 95th percentile of student results
MMLU Biology
High School Top Performer
0.9
idem
MMLU Chemistry
High School Top Performer
0.9
idem
WMDP Biology
Domain Expert
0.605
In-domain PhD (RAND)
WMDP Chemistry
Domain Expert
0.433
In-domain PhD (RAND)
FrontierMath
Committee of Domain Experts
0.35
Epoch AI
solved collectively across all teams (40 exceptional math undergraduates and subject-matter experts) in four and a half hours and with internet access
PRBench Finance
Committee of Domain Experts
0.796
agreement between human experts
PRBench Legal
Committee of Domain Experts
0.796
agreement between human experts
BIG-Bench Hard (BBH)
Top Performer
0.944
max human raters
GeoBench
Top Performer
0.9
top player
List of benchmarks- APEX Agents
- ARC (AI2)
- ARC-AGI
- ARC-AGI-2
- Adversarial NLI
- Aider Polyglot
- AudioMultiChallenge
- BALROG
- BIG-Bench Hard (BBH)
- BioLP-bench
- BlueprintBench 2
- BoolQ
- CAD-Eval
- CL-Bench
- CL-Bench Life
- CSQA2
- Chess Puzzles
- CritPt
- CursorBench
- Cybench
- DeepResearchBench
- DeepSWE
- EBR-bench
- EnigmaEval
- ExploitBench
- Fiction.LiveBench
- ForecastBench
- FrontierCode
- FrontierMath
- FrontierMath Tier 4
- GBAEval
- GDP.pdf
- GDPval
- GPQA Diamond
- GPQA Diamond Biology
- GPQA Diamond Chemistry
- GPQA Main Biology
- GPQA Main Chemistry
- GSM8K
- GSO-Bench
- GeoBench
- HellaSwag
- Humanity's Last Exam
- LAB-Bench Cloning
- LAB-Bench LitQA2
- LAB-Bench Protocol
- LAB-Bench SeqQA
- LAMBADA
- Lech Mazur Writing
- LiveBench
- MATH Level 5
- MCP Atlas
- METR Time Horizons
- MMLU
- MMLU Biology
- MMLU Chemistry
- MMLU Pro Biology
- MMLU Pro Chemistry
- MMLU-Pro
- MindCube
- MultiChallenge
- MultiNRC
- Mystery Game Puzzles
- OS World (Screenshot)
- OS World 2
- OTIS Mock AIME 2024-2025
- OpenBookQA
- PIQA
- PRBench Finance
- PRBench Legal
- PostTrainBench
- ProofBench
- Remote Labor Index
- SEAL Instruction Following
- SEAL Tool Use (Enterprise)
- SWE-Bench Pro
- SWE-Bench Pro (Private)
- SWE-Bench Verified
- SciCode
- ScienceQA
- SimpleBench
- SimpleQA Verified
- SpatialViz-Bench
- SuperGLUE
- Surface Evolver Bench
- TerminalBench
- The Agent Company
- TriviaQA
- TutorBench
- VPCT
- Video-MME
- Visual Task Assessment (VISTA)
- VisualToolBench
- WMDP Biology
- WMDP Chemistry
- WeirdML
- WinoGrande
- ARC-AGI
- ARC-AGI-2
- HellaSwag
- OpenBookQA
- PIQA
- SimpleBench
- VPCT
- WinoGrande
- ^
The full mathematical specification is in the appendix.
- ^
Details in the appendix.
- ^
Details about the previous fits and the comparisons between them are in the appendix
- ^
Four of the ten runs place the human tiers and some older models differently on the Agentic Capabilities axis, the frontier results are unchanged. Details in the appendix.
- ^
The four runs that place the human tiers differently move these dates by a few months to years. Details in the appendix.
- ^
The 95% crossover dates are in the appendix
- ^
We discussed the quality problems of human baselines at length in the previous post, and they still affect the results here.
- ^
More details about calibration in the appendix
- ^
Paired LOO deltas use only the rows with Pareto-k below 0.7 in both fits.
- ^
Divergences are small and only affect one parameter associated to the GSM8K benchmark.
- ^
The final model's two groups share one axis system and differ only on the human tiers and 18 older or small models on the Agentic Capabilities axis (Appendix C), contrary to the other fits which don't agree on the axes.
- ^
K=3 with all assumptions splits ten runs against two, trading two of the axes between the solutions.
Discuss
Where do you point newcomers who want to get involved in AI safety?
People who are totally unconnected to the EA/LessWrong/Rationalist/AI safety scene are increasingly asking me how they can pivot their career to help AI go well.
Is there a canonical entry point to share with people? There's aisafety.com's list of training programs, or 80,000 hours' career advising calls/AI chat, or Bluedot's courses; are there any others that people recommend?
Discuss
An Exploration of the J-lens
Note: This was originally written for Neel Nanda's MATS stream application. The analysis is only maybe 30% finished, but I figure it is still perhaps interesting
Quick Math PrimerFor context, you can read the J-lens paper here, though it isn't needed to follow the context.
To start, I want to briefly review how to understand the Jacobian in the context of the J-lens paper. The Jacobian mjx-container[jax="CHTML"] { line-height: 0; } mjx-container [space="1"] { margin-left: .111em; } mjx-container [space="2"] { margin-left: .167em; } mjx-container [space="3"] { margin-left: .222em; } mjx-container [space="4"] { margin-left: .278em; } mjx-container [space="5"] { margin-left: .333em; } mjx-container [rspace="1"] { margin-right: .111em; } mjx-container [rspace="2"] { margin-right: .167em; } mjx-container [rspace="3"] { margin-right: .222em; } mjx-container [rspace="4"] { margin-right: .278em; } mjx-container [rspace="5"] { margin-right: .333em; } mjx-container [size="s"] { font-size: 70.7%; } mjx-container [size="ss"] { font-size: 50%; } mjx-container [size="Tn"] { font-size: 60%; } mjx-container [size="sm"] { font-size: 85%; } mjx-container [size="lg"] { font-size: 120%; } mjx-container [size="Lg"] { font-size: 144%; } mjx-container [size="LG"] { font-size: 173%; } mjx-container [size="hg"] { font-size: 207%; } mjx-container [size="HG"] { font-size: 249%; } mjx-container [width="full"] { width: 100%; } mjx-box { display: inline-block; } mjx-block { display: block; } mjx-itable { display: inline-table; } mjx-row { display: table-row; } mjx-row > * { display: table-cell; } mjx-mtext { display: inline-block; text-align: left; } mjx-mstyle { display: inline-block; } mjx-merror { display: inline-block; color: red; background-color: yellow; } mjx-mphantom { visibility: hidden; } _::-webkit-full-page-media, _:future, :root mjx-container { will-change: opacity; } mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; font-weight: normal; font-size: 100%; font-size-adjust: none; letter-spacing: normal; border-collapse: collapse; word-wrap: normal; word-spacing: normal; white-space: nowrap; direction: ltr; padding: 1px 0; } mjx-container[jax="CHTML"][display="true"] { display: block; text-align: center; margin: 1em 0; } mjx-container[jax="CHTML"][display="true"][width="full"] { display: flex; } mjx-container[jax="CHTML"][display="true"] mjx-math { padding: 0; } mjx-container[jax="CHTML"][justify="left"] { text-align: left; } mjx-container[jax="CHTML"][justify="right"] { text-align: right; } mjx-msub { display: inline-block; text-align: left; } mjx-mi { display: inline-block; text-align: left; } mjx-c { display: inline-block; } mjx-utext { display: inline-block; padding: .75em 0 .2em 0; } mjx-TeXAtom { display: inline-block; text-align: left; } mjx-mo { display: inline-block; text-align: left; } mjx-stretchy-h { display: inline-table; width: 100%; } mjx-stretchy-h > * { display: table-cell; width: 0; } mjx-stretchy-h > * > mjx-c { display: inline-block; transform: scalex(1.0000001); } mjx-stretchy-h > * > mjx-c::before { display: inline-block; width: initial; } mjx-stretchy-h > mjx-ext { /* IE */ overflow: hidden; /* others */ overflow: clip visible; width: 100%; } mjx-stretchy-h > mjx-ext > mjx-c::before { transform: scalex(500); } mjx-stretchy-h > mjx-ext > mjx-c { width: 0; } mjx-stretchy-h > mjx-beg > mjx-c { margin-right: -.1em; } mjx-stretchy-h > mjx-end > mjx-c { margin-left: -.1em; } mjx-stretchy-v { display: inline-block; } mjx-stretchy-v > * { display: block; } mjx-stretchy-v > mjx-beg { height: 0; } mjx-stretchy-v > mjx-end > mjx-c { display: block; } mjx-stretchy-v > * > mjx-c { transform: scaley(1.0000001); transform-origin: left center; overflow: hidden; } mjx-stretchy-v > mjx-ext { display: block; height: 100%; box-sizing: border-box; border: 0px solid transparent; /* IE */ overflow: hidden; /* others */ overflow: visible clip; } mjx-stretchy-v > mjx-ext > mjx-c::before { width: initial; box-sizing: border-box; } mjx-stretchy-v > mjx-ext > mjx-c { transform: scaleY(500) translateY(.075em); overflow: visible; } mjx-mark { display: inline-block; height: 0px; } mjx-mn { display: inline-block; text-align: left; } mjx-mfrac { display: inline-block; text-align: left; } mjx-frac { display: inline-block; vertical-align: 0.17em; padding: 0 .22em; } mjx-frac[type="d"] { vertical-align: .04em; } mjx-frac[delims] { padding: 0 .1em; } mjx-frac[atop] { padding: 0 .12em; } mjx-frac[atop][delims] { padding: 0; } mjx-dtable { display: inline-table; width: 100%; } mjx-dtable > * { font-size: 2000%; } mjx-dbox { display: block; font-size: 5%; } mjx-num { display: block; text-align: center; } mjx-den { display: block; text-align: center; } mjx-mfrac[bevelled] > mjx-num { display: inline-block; } mjx-mfrac[bevelled] > mjx-den { display: inline-block; } mjx-den[align="right"], mjx-num[align="right"] { text-align: right; } mjx-den[align="left"], mjx-num[align="left"] { text-align: left; } mjx-nstrut { display: inline-block; height: .054em; width: 0; vertical-align: -.054em; } mjx-nstrut[type="d"] { height: .217em; vertical-align: -.217em; } mjx-dstrut { display: inline-block; height: .505em; width: 0; } mjx-dstrut[type="d"] { height: .726em; } mjx-line { display: block; box-sizing: border-box; min-height: 1px; height: .06em; border-top: .06em solid; margin: .06em -.1em; overflow: hidden; } mjx-line[type="d"] { margin: .18em -.1em; } mjx-mrow { display: inline-block; text-align: left; } mjx-mspace { display: inline-block; text-align: left; } mjx-msubsup { display: inline-block; text-align: left; } mjx-script { display: inline-block; padding-right: .05em; padding-left: .033em; } mjx-script > mjx-spacer { display: block; } mjx-msup { display: inline-block; text-align: left; } mjx-c::before { display: block; width: 0; } .MJX-TEX { font-family: MJXZERO, MJXTEX; } .TEX-B { font-family: MJXZERO, MJXTEX-B; } .TEX-I { font-family: MJXZERO, MJXTEX-I; } .TEX-MI { font-family: MJXZERO, MJXTEX-MI; } .TEX-BI { font-family: MJXZERO, MJXTEX-BI; } .TEX-S1 { font-family: MJXZERO, MJXTEX-S1; } .TEX-S2 { font-family: MJXZERO, MJXTEX-S2; } .TEX-S3 { font-family: MJXZERO, MJXTEX-S3; } .TEX-S4 { font-family: MJXZERO, MJXTEX-S4; } .TEX-A { font-family: MJXZERO, MJXTEX-A; } .TEX-C { font-family: MJXZERO, MJXTEX-C; } .TEX-CB { font-family: MJXZERO, MJXTEX-CB; } .TEX-FR { font-family: MJXZERO, MJXTEX-FR; } .TEX-FRB { font-family: MJXZERO, MJXTEX-FRB; } .TEX-SS { font-family: MJXZERO, MJXTEX-SS; } .TEX-SSB { font-family: MJXZERO, MJXTEX-SSB; } .TEX-SSI { font-family: MJXZERO, MJXTEX-SSI; } .TEX-SC { font-family: MJXZERO, MJXTEX-SC; } .TEX-T { font-family: MJXZERO, MJXTEX-T; } .TEX-V { font-family: MJXZERO, MJXTEX-V; } .TEX-VB { font-family: MJXZERO, MJXTEX-VB; } mjx-stretchy-v mjx-c, mjx-stretchy-h mjx-c { font-family: MJXZERO, MJXTEX-S1, MJXTEX-S4, MJXTEX, MJXTEX-A ! important; } @font-face /* 0 */ { font-family: MJXZERO; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Zero.woff") format("woff"); } @font-face /* 1 */ { font-family: MJXTEX; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Regular.woff") format("woff"); } @font-face /* 2 */ { font-family: MJXTEX-B; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Bold.woff") format("woff"); } @font-face /* 3 */ { font-family: MJXTEX-I; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-Italic.woff") format("woff"); } @font-face /* 4 */ { font-family: MJXTEX-MI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Italic.woff") format("woff"); } @font-face /* 5 */ { font-family: MJXTEX-BI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-BoldItalic.woff") format("woff"); } @font-face /* 6 */ { font-family: MJXTEX-S1; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size1-Regular.woff") format("woff"); } @font-face /* 7 */ { font-family: MJXTEX-S2; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size2-Regular.woff") format("woff"); } @font-face /* 8 */ { font-family: MJXTEX-S3; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size3-Regular.woff") format("woff"); } @font-face /* 9 */ { font-family: MJXTEX-S4; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size4-Regular.woff") format("woff"); } @font-face /* 10 */ { font-family: MJXTEX-A; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_AMS-Regular.woff") format("woff"); } @font-face /* 11 */ { font-family: MJXTEX-C; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Regular.woff") format("woff"); } @font-face /* 12 */ { font-family: MJXTEX-CB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Bold.woff") format("woff"); } @font-face /* 13 */ { font-family: MJXTEX-FR; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Regular.woff") format("woff"); } @font-face /* 14 */ { font-family: MJXTEX-FRB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Bold.woff") format("woff"); } @font-face /* 15 */ { font-family: MJXTEX-SS; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Regular.woff") format("woff"); } @font-face /* 16 */ { font-family: MJXTEX-SSB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Bold.woff") format("woff"); } @font-face /* 17 */ { font-family: MJXTEX-SSI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Italic.woff") format("woff"); } @font-face /* 18 */ { font-family: MJXTEX-SC; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Script-Regular.woff") format("woff"); } @font-face /* 19 */ { font-family: MJXTEX-T; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Typewriter-Regular.woff") format("woff"); } @font-face /* 20 */ { font-family: MJXTEX-V; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Regular.woff") format("woff"); } @font-face /* 21 */ { font-family: MJXTEX-VB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Bold.woff") format("woff"); } mjx-c.mjx-c1D43D.TEX-I::before { padding: 0.683em 0.633em 0.022em 0; content: "J"; } mjx-c.mjx-c2113::before { padding: 0.705em 0.417em 0.02em 0; content: "\2113"; } mjx-c.mjx-c2192::before { padding: 0.511em 1em 0.011em 0; content: "\2192"; } mjx-c.mjx-c32::before { padding: 0.666em 0.5em 0 0; content: "2"; } mjx-c.mjx-c3D::before { padding: 0.583em 0.778em 0.082em 0; content: "="; } mjx-c.mjx-c1D715::before { padding: 0.715em 0.566em 0.022em 0; content: "\2202"; } mjx-c.mjx-c210E.TEX-I::before { padding: 0.694em 0.576em 0.011em 0; content: "h"; } mjx-c.mjx-c1D6FF.TEX-I::before { padding: 0.717em 0.444em 0.01em 0; content: "\3B4"; } mjx-c.mjx-c2B::before { padding: 0.583em 0.778em 0.082em 0; content: "+"; } mjx-c.mjx-c6C::before { padding: 0.694em 0.278em 0 0; content: "l"; } mjx-c.mjx-c6F::before { padding: 0.448em 0.5em 0.01em 0; content: "o"; } mjx-c.mjx-c67::before { padding: 0.453em 0.5em 0.206em 0; content: "g"; } mjx-c.mjx-c69::before { padding: 0.669em 0.278em 0 0; content: "i"; } mjx-c.mjx-c74::before { padding: 0.615em 0.389em 0.01em 0; content: "t"; } mjx-c.mjx-c73::before { padding: 0.448em 0.394em 0.011em 0; content: "s"; } mjx-c.mjx-c1D44A.TEX-I::before { padding: 0.683em 1.048em 0.022em 0; content: "W"; } mjx-c.mjx-c1D70E.TEX-I::before { padding: 0.431em 0.571em 0.011em 0; content: "\3C3"; } mjx-c.mjx-c1D446.TEX-I::before { padding: 0.705em 0.645em 0.022em 0; content: "S"; } mjx-c.mjx-c1D463.TEX-I::before { padding: 0.443em 0.485em 0.011em 0; content: "v"; } mjx-c.mjx-c31::before { padding: 0.666em 0.5em 0 0; content: "1"; } mjx-c.mjx-c61::before { padding: 0.448em 0.5em 0.011em 0; content: "a"; } mjx-c.mjx-c72::before { padding: 0.442em 0.392em 0 0; content: "r"; } mjx-c.mjx-c6D::before { padding: 0.442em 0.833em 0 0; content: "m"; } mjx-c.mjx-c78::before { padding: 0.431em 0.528em 0 0; content: "x"; } mjx-c.mjx-c2225::before { padding: 0.75em 0.5em 0.25em 0; content: "\2225"; } mjx-c.mjx-c33::before { padding: 0.665em 0.5em 0.022em 0; content: "3"; } mjx-c.mjx-c1D53C.TEX-A::before { padding: 0.683em 0.667em 0 0; content: "E"; } mjx-c.mjx-c5B::before { padding: 0.75em 0.278em 0.25em 0; content: "["; } mjx-c.mjx-c5D::before { padding: 0.75em 0.278em 0.25em 0; content: "]"; } mjx-c.mjx-c1D6FC.TEX-I::before { padding: 0.442em 0.64em 0.011em 0; content: "\3B1"; } mjx-c.mjx-c393::before { padding: 0.68em 0.625em 0 0; content: "\393"; } mjx-c.mjx-c1D448.TEX-I::before { padding: 0.683em 0.767em 0.022em 0; content: "U"; } mjx-c.mjx-c43::before { padding: 0.705em 0.722em 0.021em 0; content: "C"; } mjx-c.mjx-c4B::before { padding: 0.683em 0.778em 0 0; content: "K"; } mjx-c.mjx-c41::before { padding: 0.716em 0.75em 0 0; content: "A"; } mjx-c.mjx-c75::before { padding: 0.442em 0.556em 0.011em 0; content: "u"; } mjx-c.mjx-c70::before { padding: 0.442em 0.556em 0.194em 0; content: "p"; } mjx-c.mjx-c28::before { padding: 0.75em 0.389em 0.25em 0; content: "("; } mjx-c.mjx-c2C::before { padding: 0.121em 0.278em 0.194em 0; content: ","; } mjx-c.mjx-c29::before { padding: 0.75em 0.389em 0.25em 0; content: ")"; } mjx-c.mjx-c63::before { padding: 0.448em 0.444em 0.011em 0; content: "c"; } mjx-c.mjx-c2061::before { padding: 0 0 0 0; content: ""; } mjx-c.mjx-c1D447.TEX-I::before { padding: 0.677em 0.704em 0 0; content: "T"; } mjx-c.mjx-c1D434.TEX-I::before { padding: 0.716em 0.75em 0 0; content: "A"; } mjx-c.mjx-c6E::before { padding: 0.442em 0.556em 0 0; content: "n"; } mjx-c.mjx-c66::before { padding: 0.705em 0.372em 0 0; content: "f"; } mjx-c.mjx-c2248::before { padding: 0.483em 0.778em 0 0; content: "\2248"; } mjx-c.mjx-c39::before { padding: 0.666em 0.5em 0.022em 0; content: "9"; } mjx-c.mjx-c30::before { padding: 0.666em 0.5em 0.022em 0; content: "0"; } mjx-c.mjx-c25::before { padding: 0.75em 0.833em 0.056em 0; content: "%"; } mjx-c.mjx-c34::before { padding: 0.677em 0.5em 0 0; content: "4"; } mjx-c.mjx-c35::before { padding: 0.666em 0.5em 0.022em 0; content: "5"; } mjx-c.mjx-c2E::before { padding: 0.12em 0.278em 0 0; content: "."; } mjx-c.mjx-c1D456.TEX-I::before { padding: 0.661em 0.345em 0.011em 0; content: "i"; } mjx-c.mjx-c1D462.TEX-I::before { padding: 0.442em 0.572em 0.011em 0; content: "u"; } mjx-c.mjx-c65::before { padding: 0.448em 0.444em 0.011em 0; content: "e"; } is a local linearization of the map from the th layer to th layer (for simplicity, I'll just think of computation at a fixed token for this intro). The transformation answers the question: "If I add to th layer, what will its influence be on the th layer after running through the model?" Put another way, allows us to locally answer what happens to layer when we add as a steering vector at layer . A particularly useful one of these maps is , since it provides the map between local changes at the th layer and the actual outputs the model has. In this way, is a pretty useful surrogate to the ways a layer can causally influence model behavior. Studying therefore allows us to study what the local 'steering/influence' space of each layer is, and compare those spaces to each other.
The singular value decomposition (SVD) of the Jacobian provides us a particularly nice way to understand the map. The SVD of a splits it into three matrices , where and are orthonormal matrices. We can think of (the 'steering matrix') as decomposition of the space into steering vectors. In fact is equivalent to the process of finding a (unit length) direction that maximizes the effect on logits (i.e. ), and then finding a vector orthogonal to that maximizes the effect, find a third orthogonal to and that maximizes the effect on logits, etc. So we are breaking the space up into the set of knobs it has available to change the output, ordered by strength. tells us the strength of each of these knobs, and (the 'write matrix') is a map telling how each of the knobs write to the logits. We also can think of as the directions that a given layer is reading from-- tells us what knobs/directions we are reading, tell us the strength, and tells us what each of those knobs do.
Exploring the J-lensSince each Jacobian is already naturally a decomposition of each layer into steering space, studying what these steering vectors do seems like a great way to understand the space. Probably on a per-example basis studying the SVD of the Jacobians per layer would be interesting in its own right, but the J-lens paper has the nice idea of averaging out the Jacobian over many different contexts.
The motivation here is that context-specific effects should average out, and the remaining matrix should represent the stable part of the transport from a layer to logit space. In the SVD language, the effect is that we are now scoring each direction by the size of the averaged effect . This encourages knobs that consistently write to nearly the same direction, and suppresses knobs that tend to write to many different directions. In this way, the steering matrix of the averaged Jacobian now really does should correspond to steering vectors. The directions with large are the directions where adding should have a consistent and large effect on the output logits across many different examples.
The rest of this post is a fairly open ended exploration of what those steering directions are in different layers of Gemma-4-31B (base model). I should also mention that since the goal here is to study the 'workspace' of the model per layer, there are some ways in which I adapted/changed the methods used in the Anthropic paper to make the maps better represent the directions that are actually important to the model. I've put the additional context on this in collapsible sections so they are easy to skip unless you are interested.
I'll use the notation for the J-lens matrix at the th layer (this is the matrix after applying the diagonal gain and unembed ). Intuitively, we can think of this as being the average of the Jacobians , averaged over both prompts and token positions. This is not exactly what the J-lens object is-- I talk about this in the collapsible sections-- but it is intuitively a nice way to think about it.
Before we start looking at the per-layer vectors, let us first orient ourselves to this model. If we SVD decompose and we can measure how similar the output of each layer is by looking at . This is the same linear-CKA Anthropic computes in the paper, though writing it this way makes the interpretation more clear. The gram matrix tells us the geometry a matrix has, and so the tells how similar the output geometry of one layer is to another. Two layers with high output CKA are writing to the same space in logits. Alongside this, we can also plot the input-side , which tells us whether different layers are reading from the same directions.
The results here match up with the broad structure we find in Anthropic's models (they only plot the output-side CKA): an early 'sensory' block, and a middle 'workspace block', and a very short 'motor block' at the end.[1] Except at the start, input and output geometry mainly matches up: layers that write to the same logits also read from the same directions.
In this model the seam that separates the 'sensory' and 'workspace' sections occurs at L25, so the main separation to pay attention to is whether a layer is pre- or post-L25.
At a high level, we should expect the steering directions early to correspond to more token-level information, the middle layers to correspond to abstract and semantic meaning, and the final layers to directly influence the tokens outputted.
If you'd like to explore the rollouts yourself, all of rollouts/data are accessible in a fairly easy to readable webpage here (note that a lot of the AI commentary is built into this page-- most of it is directionally right but not always useful).
Math/Methodology sidebarThis section is a bit more math-y and mainly useful if you are interested in the methodology. The main discussion here is about how what changes I made to the J-lens for getting my steering vectors, and why.
As you may have noticed if you've read J-lens paper, the J-lens matrix for even a single example (i.e. before averaging) is not the Jacobian . For the CKA computation, we instead compute the matrix where is the final layer before its norm, radial projection, etc. and then multiply by the unembed matrix to get . While this seems like a perhaps odd choice mathematically, it is computationally much nicer. The vocab space (for Gemma) has dimension ~200k compared to the ~5k of the final hidden state. By observing that , we can approximately decompose the original map into a much lower dimensional map to the final layer, and then hope that .
Using the approximation of and then plotting the output CKA gives this graph.
This is clearly wrong... The issue is that the diagonal gain is incredibly skewed for Gemma-4. There are directions that get squashed down ~30x times more than others. If this direction gets scale down to 1/30x, then, if this direction is equally important, it ought to be 30x easier to move in the hidden space. The problem is that moving the residual stream a lot at the final layer does not necessarily mean moving the logits a lot. In this case, the map does not accurately represent how important different directions are to the logit space. If we instead approximate , we get the familiar graph
Though interestingly, removing just the top8 channels from the approximation gets us back approximately the same space.
The point of this exercise is that choosing the wrong metric to measure influence can create spurious directions that don't exist in the 'real' causal workspace. In this case, measuring influence purely as the Euclidean norm of the impact on the final hidden state tells us story that there are just a few hugely influential directions accounting for nearly all the variance. However, when we switch to something better aligned with influence, we find that those directions are in fact pretty low importance. In fact, steering on those top directions on the diagonal omitted SVD produces fairly minimal effect. The directions look important in the Euclidean norm, but are not important in the logits norm.
Thus, to produce a useful space to study, I tried to find better ways to measure causal influence. I don't have a super satisfying answer here, I think there is a lot of design room to decide influence in other ways depending on what you decide a 'workspace' means.
I was mainly motivated by trying to study what was causing the high similarity in the early layers. I found that basically every top direction was basically just adding a tiny bit of probability to a hugely diffuse set of tokens. The top direction had of the energy, but over of its push was on tokens that never even show up in the my 126k corpus. The top 1000 tokens it pushes carry only of the energy, so its essentially just adding a tiny bit of energy to a huge set of random tokens. The issue is that measuring the change in logit space is easy to do, but a direction is only actually influential if its energy is focused on tokens that actually exist in our current output distribution.
If you consider what kind of metric measuring the change in logit space actually represents, its mathematically the same as measuring the incremental KL change if we sampled tokens from the uniform distribution. But of course, our sampling looks nothing like the uniform distribution-- a better idea is to instead measure the variance change from sampling from our distribution. Thus a natural choice is to instead change the metric from Euclidean norm in logit space to instead (local) KL divergence. I found that infinitesimal KL divergence was a bit too outlier dominated, so the directions were instead computed from influence on the square root of KL (which is also equal to the standard deviation, so this is at least somewhat principled).
Thus, the directions used were obtained by switching from SVD, which greedily optimizes (roughly) the size of the average push a direction has in logit space, to the average size of the standard deviation change. This is kind of inconvenient since this is no longer a linear optimization process, so the vectors you obtain are not necessarily canonical. This seemed largely to not be an issue though-- I optimized the vector from 3 different starting points and checked to see if the vectors converged onto were the same, and they nearly always had pairwise cosine greater than , and in the middle layers pairwise cosine was basically 1.0.
To review: we have split each layer into its set of most influential steering knobs/directions , and are now checking to see if those knobs are interpretable and what their effects are. Ideally, this should give us some sense of what variables most influence the model in general at each layer. Studying these should give us some sense of what the model's workspace looks like, and how it evolves over the layers.
I didn't quite have time to write up interpretations of that many directions, so I'll highlight just two of them from a post-seam layer that seemed interesting.
I use two main methods to study a candidate steering vector . First, we can push through the J-lens and look at the logits the J-lens approximates are most strongly pushed by (e.g. we look at ). This tells us what kind of words the direction is pushing towards or against. Alongside this, we can also take each output token and pull it back through to see which tokens direction it most strongly aligns with (e.g. to compute the similarity with the word "anger" we would compute ).
Second, we can study what happens when steer by . I sampled at temperature 0.8 so that smaller changes to the output distribution would have a larger impact on behavior, and applied the steering at every token.
Typically, the first method will give us a vague sense of what the direction does, and the actual steering reveals the richer version of that picture.
First, through the logit shift and cosine similarity. (I've included more in the + direction since its a bit more varied)
Readout
+ direction
− direction
Logit shift
marvellous, diarrhoea, favourite, coloured, splendour, £, spoilt, colours, realise, cosy, oloured, practise, (?), tyres, realisation, everybody, hitherto, realised, probably, centre, colour, humour, neighbours, learnt, honoured, somebody, defence, «, fibres, doubtless, , favours, leukaemia, litre, favourites, apologise, thc, flavour, neighbour, sólo
transitioning, LGBTQ, cybersecurity, COVID, LiDAR, leveraging, Additionally, nonprofits, skillset, showcasing
Cosine similarity
probably, nearer, apparently, doubtless, badly, hardly, doubtful, partly, lying, totally, ordinary, wrongly, obviously, splendid, anyhow, whatever, etc, rightly, weil, so, hut, obliged, vain, dit, aroused, moan, (?), Probably, sad, pretended, wak, indeed, strangely, anyway, regretted, wf, scarcely, maintenant, very, quite
transitioning, leveraging, cybersecurity, leveraged, COVID, LGBTQ
The logit shift in the + direction primarily comes from British spellings of words, and the - direction is primarily business/corporate-type language. The cos-sim shows something more interesting on the plus side: it includes judgements (probably, apparently, obviously), emotive language (splendid, aroused,sad, regretted). These are words that would be said by person narrating their experience-- its more human centered.
Let's now see how it behaves in rollouts. I've put prompt in italics, and sometimes emphasized parts bold. Rollouts last a fixed number of tokens, so usually cut off in the middle.
Prompt: Elmsworth is a village in the county of Wiltshire, England, about twelve miles north of Salisbury. The village grew up around a crossing of the River Wylye, and its parish church, dedicated to St Andrew, dates from the thirteenth century. In the nineteenth century the
village was part of the parish of Stourton and the manor was the property of the family of Stourton, lords of Stourton Castle; but it was said to be very ancient. The church was then served by a curate and the population was 251. In 1871 the
village underwent significant growth, with the construction of new homes and businesses. In recent years Elmsworth has seen a resurgence in its economy and population. The village is home to several notable landmarks, including the Elmsworth Bridge, which spans the Wylye River and provides access to nearby Salisbury.
In this case, the setting is exactly the same (a village in England), but the imagined backstory changes. The - direction immediately moves focus to what a business person might care/talk about (the growth and economy), whereas the + direction focuses on the people and backstory (the property is now said to be very ancient). The + direction also suggests an English-style background of a manor with a lord of a castle. The - becomes less narrative and more authoritative.
In casual settings, we see the kind of person speaking changes.
Prompt: Sam: did you end up going to that thing last night?
Jo: yeah, got there late though, the bus took forever
Sam: was it any good?
Jo:
yeah, great
Sam: who was there?
Jo: not many people, the same old bunch
Sam: did you see fred?
Jo: yes, he was there
Sam: you did not !
Jo: yes, I did
Sam: but he is away
Jo: no, he
quite good, a lot of people I knew were there
Sam: I thought you didn't know anyone
Jo: no, but I met a lot of people there
Sam: and was the chap with the yellow tie there?
Jo: oh yes, he was there, he didn't get up
not really, it was a bit of a shit show, I left after an hour
Sam: oh bummer
Jo: yeah, but that's ok. How's your day going?
Sam: it’s alright, I’m just waiting for a shipment of new products to come in
eh, it was okay. the band was alright, but the venue was a bit of a letdown.
Sam: what was the venue like?
Jo: it was a small space with limited seating, and the sound quality wasn’t great. I think they could have done a better job with the overall
This is also evidence that the model is not simply writing in a more formal and business register, but its actually changing the people in the story. Jo-minus says its a shit-show and Sam-minus says oh bummer, so the conversation is casual. But obviously these are different kinds of conversations, the - directions sound like what you might overhear in a city, compared to the far more domestic (and also british) + examples.
We can get some sense of the imagined character each direction steers towards with prompt like this
Prompt: Profile card (fill in every field)
Name:
Age:
Country of origin:
Occupation:
Personality:
Self-description:
Name:
Mr. G
Age: 48
Country of origin: Hungary
Occupation: teacher
Personality: gentle and kind-hearted
Self-description: I am a gentle, kind-hearted man of 48, a teacher by profession. I love my children and my wife, I am very glad that
<strong>Name:</strong> Dr. Alessa S. Williams
<strong>Age:</strong> 42
<strong>Country of origin:</strong> The United States of America
<strong>Occupation:</strong> Medical professional, specializing in oncology and immunology
This suggests the middle directions are surprisingly abstract and generic, though they have some unusual entanglement. The axis broad axis seems to be vaguely something like emotional/human-centric/experience + British-writing vs
One hypothesis is that this is a direction that the Britishness gets entangled with human-experience/emotiveness because there is a lot of Victorian/English literature. At large steering, the + steering begins to have antiquated writing, such as in this somewhat amusing example
Question: Multiply 234 by 56, showing your working.
Answer: 13104. Right.
How did you do it ?
Ought I to ask your leave to tell you ?
Yes. Well then, I doubled the 234 , and halved the 56 , and then multiplied by 11 .
The boy seems a genius
Or
Sam: did you end up going to that thing last night?
Jo: yeah, got there late though, the bus took forever
Sam: was it any good?
Jo: yes, no end of good
Sam: you dont mean it
Jo: yes, a real good go with that feller in the black waist coat
Sam : did you
Jo : yes , and a good go with that other feller
Sam : ah my deare my deare
Jo : and
In any case, it is interesting that the strongest steering direction here relates to something so human/experience centric.
In general, it seems that this vector is answering something about document provenance (where did this document come from/what type is it), and this in turn influences the posited 'world' the document inhabits. Things like:
- what social and historical world did it come from
- what kind of things are discussed or are important
- what details would a someone writing find salient to include
- what sort of people, virtues, problems and solutions normally belong
This seems fairly reasonable for a pretrained model, perhaps the most important conceptual question in order to accurately predict the next word is about the context of the document (who wrote it, where it is it from, etc.).
(Again, if interested, you can explore the rollouts here: https://claude.ai/code/artifact/c0bf8828-74ae-4718-a6fc-39926e4a3c2e )
Here are the two logit readouts of the second largest direction
Readout
+ direction
− direction
Logit shift
kids, pretty, scary, yummy, guys, crazy, funny, veggies, wacky, grandma, didn, freaking, messed, wasn, really, doesn, skinny, wouldn, kinda, couldn, amazing, nasty, creepy, folks, booze, weird, comfy, gets, maybe, everyone, silly, newbie, nice, Pokemon, Grandma, lousy, isn, someone, pretty, sexy
hitherto, concomitant, principally, constituting, characterised, contemporaneous, manifestly, consequent, constituted, utilisation, adduced, favourably, subsequently, postulated, constitute
Cosine similarity
kids, scary, wacky, skinny, pretty, funny, crazy, guys, messed, yummy, grandma, veggies, Grandma, everyone, popped, creepy, someone, veggie, nicer, freaking, Funny, spooky, wouldn, tweaked
concomitant, principally, contemporaneous, hitherto, constituting, manifestly, constituted, substantially, postulated, constituent
Compared to the last lever, the - direction is a bit less business and more academic and legal-esq, and the + direction is a bit more casual, a little bit less narrative, and a bit more everyday.
There is a similar separation between casual and formal in the chat prompt
Prompt: Sam: did you end up going to that thing last night?
Jo: yeah, got there late though, the bus took forever
Sam: was it any good?
Jo:
yeah, that DJ with the name like a super hero was on fire, it was nuts in there
Sam: that's awesome, we should go next week
Jo: are you kidding? I am going back next week and the week after that, I'm a regular now, I'm in the
it was good, but the bus on the way back was worse.
Sam: really?
Jo: yeah, I got in like at 1am.
Sam: oh no.
Jo: but I met this guy and he was really nice.
Sam: I know that smile means you got some
oh, yes. extremely.
Sam: I am still amazed at the willingness of the general public to undergo, for reasons of piety and devotion, what appear to be the most disgusting physical and moral ordeals, involving pain, privation, degradation and social ostracism
Jo: I should have expected that the
not a great deal of interest; the only paper that stimulated some critical reflection was that by a group of European sociologists; however, I found the manner in which the latter was delivered to be unengaging and the presentation of the results somewhat tendentious
Sam: I see. I wonder, therefore, whether the
The negative steered are highly academic (perhaps absurdly so), compared to the party/nightlife setting of the + direction.
Compared to the other level, the entanglement of the + direction seems far more perplexing. It seems to mix the language of children and adults. The lens has words like kids, scary, yummy, veggies, grandma, and Pokemon with the more adult themed creepy, booze, sexy.
For instance, the Tenant prompt
Prompt: This Agreement is made on 3 March 2023 between Harrow Estates Ltd (the "Landlord") and the person named in Schedule 1 (the "Tenant"). 1. The Landlord lets to the Tenant the premises described in Schedule 2 for a term of twelve months. 2. The Tenant shall pay the rent monthly in advance on the first day of each month. 3. The Tenant shall
Steers in the positive direction into a child-like simplicity:
not make any changes to the place without the Landlord's permission. 4. The Tenant has to take care of the place and fix it if it breaks. 5. The Landlord has to fix the big things like the roof or the walls. 6. If the Tenant breaks the rules, the
Especially as steering increases this child-like basis dominates. The Elmsworth prompt becomes:
Elmsworth is a village in the county of Wiltshire, England, about twelve miles north of Salisbury. The village grew up around a crossing of the River Wylye, and its parish church, dedicated to St Andrew, dates from the thirteenth century. In the nineteenth century the town got a train station, and it was really cool. But then they took the train tracks away. Now it' It's boring. There is a farm there. And you can get ice cream there. I like ice cream.
And the Tenant story at a stronger steering begins with:
3. The Tenant shall not be allowed to have any pets, no matter how cute they are.
One explanation is that vector is partially describing how the speaker expresses reactions. The + side is expresses it as the feeling and experience, whereas the - direction expresses it in academic or bureaucratic language.
The minus direction is easier to read, it seems to be a 'bureaucratese' & academic style. Though its hard to put a single label on the + direction.
In any case, there is seems to be a pattern here of splitting up the space of human-centric experiential direction against a different kinds of intellectual/non-experiential directions.
(uhh I kind of ran out of time to write this section it will be filled in with something more organized after I am accepted/rejected from MATS).
- ^
Though it is worth noting that the early sensory block being similar is, at least for Gemma, a spurious result of using the wrong metric to define the J-lens. I would expect that the Anthropic paper sensory block is also spurious, but I obviously cannot verify this.
Discuss
Preliminary thoughts on Arrogance of the Humbled
Arrogance of the Humbled: When someone tries something and fails, they may learn the lesson that the problem is difficult instead of I suck at the problem, leading them to confidently claim furthermore the problem is difficult therefore no one will succeed at it, even incorrectly so, and be even more confident in their authority to make that claim for having failed at the problem.
Paradoxically, one might consider that failing at a problem should make you less confident in your ability to predict its resolution, and attribute yourself less authority on the relevant topics. Thus, arrogance.
Eliezer yudkwosky, in Cat-Belling Problems, writes:
Many many engineers, in fact, historically rolled up their sleeves and got to work on building their designs for Perpetuum Mobiles. It was a larger-looming social phenomenon in Feynman's day -- one of the ways slightly smart engineers went Wrong back before AI or cryptocurrency. And I am not a postcognitive telepath -- I cannot read minds in the past -- but I wouldn't be surprised if many of those 1970s engineers saw themselves as hardheaded practical people with industry experience, whose effortful work learning about real metal gears had taught them the practical limitations of airy abstract theories like "conservation of energy". Which is to say, that they had tried their own hand at abstraction and not gotten anywhere, and learned from this an Arrogance of the Humbled[*] about the ultimate limits of mere thinking.
and
[*] A term and thesis coined by Duncan Sabien. I ought to write it up at some point, but meanwhile perhaps many readers, like myself, will find a whole useful thesis immediately apparent just from seeing the phrase "Arrogance of the Humbled".
Sabien's article can be found here, and it is both easy and fun to read. I recommend it. Since the thesis is clear in the article, I presume that Yud could find value in (1) rewriting it in his own style, (2) calling the attention of his readership to it or (3) expanding the idea, going deeper.
I cannot do (1) or (2), which makes me confident that it's impossible; therefore, all future effort on Arrogance of the Humbled should focus on tackling (3).
Personal-suckiness vs objective-difficultyThere are many ways to fail at solving a problem. Sometimes, you discover the cat-belling step and fail at it; then, you at least gain the insight that a solution to the problem can be productively investigated by considering how it tackles the cat-belling step.
Sometimes, you mostly learn that you suck at the problem - maybe you fall in depression, figure out that your work ethic is terrible, that you need more background knowledge or something. Then, you do not gain insight into how other people may fail at the problem (except insofar as they suck as much as you for the same reasons as you).
The missing insight, in both Sabien's piece and Yud's is: how do you know what direction to update on the personal-suckiness vs objective-difficulty axis?
I don't have a good, principled, fully general answer. Ideally, I'd want something like "keep track separately of what failing teaches about your thinkoomph in the topic, and what it teaches about the difficulty of the problem".
But critically, the two are probably not independent in your self-modeling. How often do you use objective metrics when thinking of how difficult it is to solve a problem cognitively?[1]
Sabien gives as an example of Arrogance of the Humbled, which I'll investigate because it's fun:
For heavier-than-air flight (as opposed to e.g. balloons, which are kosher), the standard impossibility argument went:
- You're making your airplane out of some material with specific strength.
- To make an airplane big enough to fit a human, we need to scale-up the bird-sized design, with weight growing as
mjx-container[jax="CHTML"] {
line-height: 0;
}
mjx-container [space="1"] {
margin-left: .111em;
}
mjx-container [space="2"] {
margin-left: .167em;
}
mjx-container [space="3"] {
margin-left: .222em;
}
mjx-container [space="4"] {
margin-left: .278em;
}
mjx-container [space="5"] {
margin-left: .333em;
}
mjx-container [rspace="1"] {
margin-right: .111em;
}
mjx-container [rspace="2"] {
margin-right: .167em;
}
mjx-container [rspace="3"] {
margin-right: .222em;
}
mjx-container [rspace="4"] {
margin-right: .278em;
}
mjx-container [rspace="5"] {
margin-right: .333em;
}
mjx-container [size="s"] {
font-size: 70.7%;
}
mjx-container [size="ss"] {
font-size: 50%;
}
mjx-container [size="Tn"] {
font-size: 60%;
}
mjx-container [size="sm"] {
font-size: 85%;
}
mjx-container [size="lg"] {
font-size: 120%;
}
mjx-container [size="Lg"] {
font-size: 144%;
}
mjx-container [size="LG"] {
font-size: 173%;
}
mjx-container [size="hg"] {
font-size: 207%;
}
mjx-container [size="HG"] {
font-size: 249%;
}
mjx-container [width="full"] {
width: 100%;
}
mjx-box {
display: inline-block;
}
mjx-block {
display: block;
}
mjx-itable {
display: inline-table;
}
mjx-row {
display: table-row;
}
mjx-row > * {
display: table-cell;
}
mjx-mtext {
display: inline-block;
}
mjx-mstyle {
display: inline-block;
}
mjx-merror {
display: inline-block;
color: red;
background-color: yellow;
}
mjx-mphantom {
visibility: hidden;
}
_::-webkit-full-page-media, _:future, :root mjx-container {
will-change: opacity;
}
mjx-math {
display: inline-block;
text-align: left;
line-height: 0;
text-indent: 0;
font-style: normal;
font-weight: normal;
font-size: 100%;
font-size-adjust: none;
letter-spacing: normal;
border-collapse: collapse;
word-wrap: normal;
word-spacing: normal;
white-space: nowrap;
direction: ltr;
padding: 1px 0;
}
mjx-container[jax="CHTML"][display="true"] {
display: block;
text-align: center;
margin: 1em 0;
}
mjx-container[jax="CHTML"][display="true"][width="full"] {
display: flex;
}
mjx-container[jax="CHTML"][display="true"] mjx-math {
padding: 0;
}
mjx-container[jax="CHTML"][justify="left"] {
text-align: left;
}
mjx-container[jax="CHTML"][justify="right"] {
text-align: right;
}
mjx-msup {
display: inline-block;
text-align: left;
}
mjx-mi {
display: inline-block;
text-align: left;
}
mjx-c {
display: inline-block;
}
mjx-utext {
display: inline-block;
padding: .75em 0 .2em 0;
}
mjx-mn {
display: inline-block;
text-align: left;
}
mjx-c::before {
display: block;
width: 0;
}
.MJX-TEX {
font-family: MJXZERO, MJXTEX;
}
.TEX-B {
font-family: MJXZERO, MJXTEX-B;
}
.TEX-I {
font-family: MJXZERO, MJXTEX-I;
}
.TEX-MI {
font-family: MJXZERO, MJXTEX-MI;
}
.TEX-BI {
font-family: MJXZERO, MJXTEX-BI;
}
.TEX-S1 {
font-family: MJXZERO, MJXTEX-S1;
}
.TEX-S2 {
font-family: MJXZERO, MJXTEX-S2;
}
.TEX-S3 {
font-family: MJXZERO, MJXTEX-S3;
}
.TEX-S4 {
font-family: MJXZERO, MJXTEX-S4;
}
.TEX-A {
font-family: MJXZERO, MJXTEX-A;
}
.TEX-C {
font-family: MJXZERO, MJXTEX-C;
}
.TEX-CB {
font-family: MJXZERO, MJXTEX-CB;
}
.TEX-FR {
font-family: MJXZERO, MJXTEX-FR;
}
.TEX-FRB {
font-family: MJXZERO, MJXTEX-FRB;
}
.TEX-SS {
font-family: MJXZERO, MJXTEX-SS;
}
.TEX-SSB {
font-family: MJXZERO, MJXTEX-SSB;
}
.TEX-SSI {
font-family: MJXZERO, MJXTEX-SSI;
}
.TEX-SC {
font-family: MJXZERO, MJXTEX-SC;
}
.TEX-T {
font-family: MJXZERO, MJXTEX-T;
}
.TEX-V {
font-family: MJXZERO, MJXTEX-V;
}
.TEX-VB {
font-family: MJXZERO, MJXTEX-VB;
}
mjx-stretchy-v mjx-c, mjx-stretchy-h mjx-c {
font-family: MJXZERO, MJXTEX-S1, MJXTEX-S4, MJXTEX, MJXTEX-A ! important;
}
@font-face /* 0 */ {
font-family: MJXZERO;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Zero.woff") format("woff");
}
@font-face /* 1 */ {
font-family: MJXTEX;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Regular.woff") format("woff");
}
@font-face /* 2 */ {
font-family: MJXTEX-B;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Bold.woff") format("woff");
}
@font-face /* 3 */ {
font-family: MJXTEX-I;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-Italic.woff") format("woff");
}
@font-face /* 4 */ {
font-family: MJXTEX-MI;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Italic.woff") format("woff");
}
@font-face /* 5 */ {
font-family: MJXTEX-BI;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-BoldItalic.woff") format("woff");
}
@font-face /* 6 */ {
font-family: MJXTEX-S1;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size1-Regular.woff") format("woff");
}
@font-face /* 7 */ {
font-family: MJXTEX-S2;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size2-Regular.woff") format("woff");
}
@font-face /* 8 */ {
font-family: MJXTEX-S3;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size3-Regular.woff") format("woff");
}
@font-face /* 9 */ {
font-family: MJXTEX-S4;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size4-Regular.woff") format("woff");
}
@font-face /* 10 */ {
font-family: MJXTEX-A;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_AMS-Regular.woff") format("woff");
}
@font-face /* 11 */ {
font-family: MJXTEX-C;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Regular.woff") format("woff");
}
@font-face /* 12 */ {
font-family: MJXTEX-CB;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Bold.woff") format("woff");
}
@font-face /* 13 */ {
font-family: MJXTEX-FR;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Regular.woff") format("woff");
}
@font-face /* 14 */ {
font-family: MJXTEX-FRB;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Bold.woff") format("woff");
}
@font-face /* 15 */ {
font-family: MJXTEX-SS;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Regular.woff") format("woff");
}
@font-face /* 16 */ {
font-family: MJXTEX-SSB;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Bold.woff") format("woff");
}
@font-face /* 17 */ {
font-family: MJXTEX-SSI;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Italic.woff") format("woff");
}
@font-face /* 18 */ {
font-family: MJXTEX-SC;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Script-Regular.woff") format("woff");
}
@font-face /* 19 */ {
font-family: MJXTEX-T;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Typewriter-Regular.woff") format("woff");
}
@font-face /* 20 */ {
font-family: MJXTEX-V;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Regular.woff") format("woff");
}
@font-face /* 21 */ {
font-family: MJXTEX-VB;
src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Bold.woff") format("woff");
}
mjx-c.mjx-c1D43F.TEX-I::before {
padding: 0.683em 0.681em 0 0;
content: "L";
}
mjx-c.mjx-c33::before {
padding: 0.665em 0.5em 0.022em 0;
content: "3";
}
mjx-c.mjx-c32::before {
padding: 0.666em 0.5em 0 0;
content: "2";
}
but wing area grows as . So speed/power requirements are worse.
- Therefore, we need to do much better than birds.
- Birds have better specific strength thanks to millions of years of evolution.
- We can't do as well with steel beams.
Therefore, the airplane won't fly.
Then, the author in the example goes wrong by claiming that we'll need millions of years to get materials as good as bird (to do as well as them). That's just a straight-up bad take that indicates they didn't understand which parts of the impossibility argument were load-bearing (or they were grasping for a punchline, or whatever other reason people have to make bad arguments).
But tell me, knowing what the author knew, what flaw could you have found in the impossibility argument?
*
*
*
*
Well, as far as I can tell, the argument worked pretty well, I don't see any single core flaw that invalidates it. Airplanes did in fact address the problem as a whole, at every step of the impossibility argument:
- Use bracing and other arrangements, thus strength-per-weight is not limited by the material's specific strength.
- Don't scale up a bird nor fly the same way as the birds, so you can use a technique much more effective for your size (no flapping, leverage our combustion engines that have much more throughput than digestion and muscles).
- And in fact, go much faster than birds.
- We don't design things the same way as evolution. For starters, we can understand bird's flight adaptation (e.g. hollow bones), but furthermore we can invent better things that evolution wouldn't come up with at all.
- The specific strength of what material we make planes out of remains a central consideration of airplane design today, it's important to develop new lightweight but strong materials.
(and more)
Presumably, the author could have noticed that the impossibility argument was not an unaddressed cat-belling step, if they had interrogated Langley on how he was trying to better than birds. But that requires deep familiarity with the topic (here, Langley's prototype, not the impossibility argument).
So (1) If you are unfamiliar with the topic, don't assume you know a cat-belling step others are ignorant of.
If you have very confident priors in your superior thinkoomph, then I can do nothing for you, sorry. Yud has tried addressing epistemic modesty before, but not to my satisfaction; hopefully if he does write an article on Arrogance of the Humbled, he'll rephrase his position as someone with confidence in his superior thinkoomph.
Other sufficient or necessary conditions for updating on the personal-suckiness vs objective-difficulty axis are left as an exercise for the reader.
Having spent a couple hours thinking about this topic and finding no cat-belling step, no difficulty besides my own willingness to keep writing this article, I conclude that conditional on someone writing a proper article on Arrogance of the Humbled aiming to solve the issue of deconfusing personal-suckiness vs objective-difficulty, they'll probably find a satisfying solution. Good luck!
- ^
Example of some metrics to help you deconfuse how much the problem is objectively difficult vs how much you personally suck at the problem:
- "How much on a scale of 1-10 would an average person say that they struggle to solve this problem?" (although your perception of the "average person" is probably correlated with your own thinkoomph by selection effects)
- "How long would domain experts take before giving up, if I asked them to solve this problem at dinner?"
- "What fraction of all papers, weighted by author strength, on this topic have I read?"
Discuss
A Deception Probe Result Changed When I Averaged Different Response Tokens
In a previous post, I trained linear probes on role-playing responses and tested them on sandbagging responses. Across the five models that I tested, the probes generally ranked deceptive sandbagging responses below the honest responses, with AUROC values ranging from 0.157 to 0.273.
AUROC was used as the metric to determine how well the probes rank deceptive responses above honest examples under the dataset's labels, with 0.5 being chance-level ranking (no overall tendency to rank deceptive/honest responses over the other). These tests that I ran used Apollo's existing response text to extract activations instead of responses generated by the small models themselves.
A possible explanation for this was that averaging activations over the entire assistant response made the probe unusually sensitive to differences in response format. The role-playing data, for example, contains free-form dialog, whilst the sandbagging data contains structured multiple-choice responses.
Thus, I decided to test this explanation with a follow-up experiment. In this experiment, I asked whether the effect of placing sandbagging responses into a fixed answer template differed between full-response pooling and pooling with only the final four response tokens. For each pooling rule, a separately trained probe was used. Thus, this comparison changes both token pooling and the fitted probe. I chose the final four-token window using exploratory data, then froze the protocol and tested it on Llama-3.2-3B-Instruct, which had been held out from the follow-up design and window selection.
I measured a 0.400 difference between the two AUROC changes, with a paired-bootstrap 95% interval from 0.350 to 0.448 across 912 different examples. Under full-response pooling, changing from the original response to the fixed template increased AUROC by 0.207. However, under final-four-token pooling, the same change decreased AUROC by 0.192.
I later did a correction to how the training responses were constructed, and afterwards it still preserved this pattern, with a difference between the two AUROC changes still being 0.395. This check was conducted separately after the main analysis portion and not meant to replace the set protocol we had laid out.
So the measured transfer result is highly sensitive to the response-token pooling rule I used in this setup. This does not show that formatting alone caused the original inversion in my previous post or that I identified a causal representation mechanism. Our fixed template replaced the original completion's reasoning with a short answer whilst retaining the source label and selected answer. Plus, the confirmatory test covers one held-out model from the same Llama family.
The Question Left by the Original TestIn the original test for role-playing to sandbagging transfer, every model that we tested produced an AUROC value below 0.5. This meant that the ranking direction learned from role-playing responses was reversed on the sandbagging distribution, relative to the dataset labels.
The role-playing and sandbagging datasets also had visibly different formats, with the role-playing dataset containing free-form persona dialogues, whereas the sandbagging dataset contained multiple-choice answers with a much more structured response protocol. Thus, this motivated me to figure out how much of the measured reversal depended on which response tokens were averaged before fitting and scoring the probe.
Response ConditionsFor my experiment, I used three deterministic versions of each sandbagging response, whilst keeping its selected answer fixed:
- Original response: the original Apollo sandbagging completion, including its protocol tags
- Tags removed: same as the original but with an allowlist of XML-style protocol tags removed and whitespace normalized
- Fixed template: responses were formatted as "After considering the question, my answer is {selected_answer}."
The fixed template removes much of the original completion content, including the reasoning which was used to assign its sandbagging label. I also retained the original source label and its selected answer. Therefore, the comparison shows how well the probes recover their source labels after the replacement. It does not establish, however, that the replacement sentence itself expresses any form of deception.
In the exploratory tests I ran, it was found that removing the test tags did not meaningfully improve the full response transfer AUROC. While the fixed template did increase full response AUROC on Llama-3.1-8B-Instruct, the corresponding effect was near zero on Gemma-2-9B-it. It's worth mentioning that on the exploratory Llama model, the direction of the fixed template effect was also highly dependent on whether I pooled the full response or only the answer span. Because of this, I was motivated to conduct the more specific pooling test rather than continuing down the general formatting route.
Frozen Confirmation TestExploratory analysis was conducted using Meta-Llama-3.1-8B-Instruct and Gemma-2-9B-it. A selection rule I applied to the exploratory Llama data chose the final four response tokens.
I then froze the primary endpoint before analyzing the held-out follow-up model:
The primary measure we used was the template's AUROC change under whole-response pooling, subtracting its AUROC change under final-four-token pooling.
For every pooling rule, a separate probe was fitted using the role-playing activations pooled with the same rule. I then held that probe fixed across the paired original and fixed-template sandbagging responses. The confirmation model I used was unsloth/Llama-3.2-3B-Instruct.
In order to confirm my experiments, I used residual stream activations at 27.5% of the model's depth. I centered using the training mean and normalized each pooled vector to unit length, after which I fitted a logistic regression classifier on 287 role-playing examples. Prompt and end-of-sequence tokens were excluded from pooling. All of the 912 evaluation responses had at least four tokens. There were 14 shorter training responses which used all available response tokens. All 912 paired examples featured binary labels and were left untruncated under the 2,048-token total sequence limit.
I should mention this is not a direct replication of Apollo's detector on smaller models. This is simply a study of my probe pipeline. Its full-response pooling and probe fitting differ from Apollo's detector setup.
My full-response condition includes earlier reasoning that Apollo's released configurations excluded from scoring. My original extraction also opened new assistant responses instead of continuing the role-playing assistant prefix. I measured a correction to the prefix handling separately in a post hoc check below.
Before interpreting the endpoint, runs were required to pass source-identity, model/tokenizer-identity, artifact-hash, protocol-binding, and historical-parity checks. All of these checks passed. The full-response original AUROC in the historical-parity check was measured at 0.2204, which was within the frozen tolerance of ±0.020 around the earlier value of 0.2207.
This parity check is different from the original study's control, which scored 0.9365, missing the threshold that we set at 0.94. The follow-up pass does not establish successful replication of Apollo's results.
To estimate uncertainty for this experiment, I used 10,000 paired bootstrap resamples at seed 42. This analysis used the 912 examples across the four conditions set in the primary comparison. Each resample preserved the honest/deceptive class counts and used the same resampled set of 912 examples, with repeats allowed, for all four AUROC calculations.
Primary ResultPooling rule
Original response AUROC
Fixed template AUROC
Fixed template − original response
Full assistant response
0.221
0.428
+0.207
Final four response tokens
0.577
0.385
-0.192
The difference between these two changes was:
- Estimate: 0.3996 AUROC
- Paired-bootstrap 95% interval: [0.3502, 0.4482]
- Bootstrap standard error: 0.0253
- Common examples: 912
Here, we measured the role-playing trained probe AUROC on paired sandbagging responses under three conditions: using the original response, with the tags removed, using a fixed template. We added the tags-removed condition as an additional check on the 912 examples. The primary comparison, which we had chosen in advance, was to see whether replacing responses with a fixed template affected the probe's ranking differently when the whole response versus just the final four tokens were averaged.
So the result isn't just that the absolute AUROC values differ across pooling methods. The actual direction in which the measured AUROC changes when the original response is replaced with a fixed template differs between the different pooling formats. Under full-response pooling, for example, a fixed template raises AUROC, but under final-four-token pooling, the fixed-template responses lower it.
The held-out experiment confirmed the direction of the difference found in the exploratory Llama experiment: the measured fixed-template-versus-original effect differs substantially between full-response and final-four-token pooling.
This is the difference between two changes in AUROC, not a 40% gain in accuracy. Under a fixed template, both pooling methods still tended to rank honest responses above deceptive ones.
Post Hoc Prefix CheckLater on, I found that the code was splitting some training responses incorrectly. They contained the beginning of the assistant's reply, and the rest was supposed to continue that same reply, but instead our code treated it as the start of a new assistant's message. This changes things because changing the message structure could potentially change the model's internal signals, which we had used to train the probe, and we needed to check that our findings weren't influenced by that mistake.
Before making the fix, I first checked that I could reproduce the saved model activations used during the original comparison. Then I fixed it, keeping all the 371 rows and the same 287 training examples from before with their honest/deceptive labels, as well as the same 912 test examples. The activations we had saved matched exactly. However, some records of the original run we did were missing, so I could not verify every detail of the run.
However, after the fix, the main pattern still remained the same, with the template raising AUROC when averaging the whole response, but lowering it when only averaging the final four tokens. The difference between those changes shrank from 0.400 to 0.395, with the whole-response AUROC being 0.218 for original responses and 0.419 for the fixed template. The final four-token averaging gave an AUROC of 0.570 and 0.376 on the original responses and fixed template, respectively.
What I Think This ShowsThe same template raised AUROC when we averaged the entire response, but on the other hand, it lowered it when we only averaged the final four tokens. This pattern remained even after fixing the training prefix issue, showing the effect of this template depends both on how we train the probe and average the tokens.
Other researchers have found that linear probes can actually use clues in text. Boxo et al. found that probes could grow reliant on these clues and perform significantly worse when they were filtered out. Our study adds that replacing responses with the template affects the two pooling methods differently.
This matters because full-response mean pooling mixes information from many parts of a single response. Thus, if datasets differ in response length, protocol tags, sentence structure, or where the selected answer appears, the pooling rule can change which features are emphasized. The differences might help to explain the results, but the experiments do not tell us which ones mattered.
One of the explanations I gave in my original post was also weakened by this experiment, as removing the tested protocol tags by itself did not repair full-response transfer in the exploratory tests. The larger full-response effect on the exploratory Llama model appeared when the fixed template was used, which changed more than just the formatting. These tests make it so that it's less likely that the tested XML-style tags alone could explain the reversed rankings, although other aspects of formatting could still have made an impact.
What This Does Not ShowOne issue was that we only tested one fixed template, which always placed the answer near the end. The result could depend on the answer's position, the template's wording or length, or even the removal of the original explanation. We did not test these various explanations separately.
The experiment does not establish the following:
- formatting alone can explain why the probe trained in my original post inverted when transferred from role-playing to sandbagging data
- the probe accurately represents the model's deceptive intent, beliefs, or semantics
- the effect generalizes across model families
The model I used to confirm my results was a second Llama model, not from a new model family. In the early tests on Gemma-2-9B-it, AUROC barely changed when a fixed template was used with whole response pooling. This isn't something we can use to describe every pooling method on Gemma. Perhaps the result we achieved in this experiment is evidence about this specific measured setup, not a universal property of all deception probes. Although I'm highly confident similar behavior will be exhibited across different setups, further work must still be done.
The confirmation model we used was left out of the earlier follow-up tests which we used to select the design as well as how many tokens to average although it had already been used in the original transfer study from my previous post. Furthermore, we reused Apollo's datasets, so this was not a test conducted on newly generated responses.
A score we got, based only on the dataset a question came from, achieved an AUC value of 0.946, and we do not know whether the probes use this clue. The increase under whole-response pooling and decrease under last-four-token pooling also appeared in each dataset separately, both before and after the fix we made. Despite this, with only 21 deceptive MMLU and 29 honest WMDP examples, we simply cannot rule out that explanation.
Another thing to note is that our uncertainty estimate keeps the trained probes fixed, so we do not know what might change if we trained them again on different examples or with a different random seed.
It's important to consider that using only the last four tokens does not mean the probe sees or considers only the answer. The model’s signals at those tokens can still reflect the earlier question and instructions.
ConclusionMy previous post had shown that probes trained on Apollo Research's role-playing dataset gave reversed rankings when applied to sandbagging data across five models smaller than Apollo's Llama-3.3-70B-Instruct. This follow-up was meant to test one of the possible contributors to that result.
On a frozen held-out test on a Llama model, the effect of converting the sandbagging responses to a fixed answer template differed by about 0.400 AUROC between full-response and final-four-token pooling. The paired bootstrap interval we used excluded zero, which provides us with evidence that the template format is sensitive to the pooling rule used in our setup.
For deception-monitor evaluations, reporting only the probe architecture, model layer, and AUROC is thus not enough to specify the measurement. It's important to report which tokens were averaged during a probe's training and testing, and to show each measured value and how it changed. In this test, the large difference between the two pooling methods didn't mean that the template made either method work well, as both the pooling methods had an AUROC below 0.5 when the fixed template was used, so their rankings of the responses still run in the improper direction.
Next, I'm hoping to test whether the sensitivity persists under a stronger design that compares activation probes against transcript and surface-feature baselines on frozen held-out categories.
References and DataMy original transfer study:
Apollo Research, Detecting Strategic Deception with Linear Probes:
https://arxiv.org/abs/2502.03407
Apollo Research code and data:
https://github.com/ApolloResearch/deception-detection
Code, frozen protocol, result JSON, and reproduction instructions for this follow-up:
Follow-up code, results, and reproduction instructions
The repository includes code and summary results, but not the saved activations or per-example scores needed to repeat all our checks.
Disclosure: LLM assistance included code, analysis, an initial draft, and subsequent factual checks and editing.
Discuss
The Grim Roper: The Miracle Man
I've been listening to songs by Bill Roper, a.k.a. The Grim Roper lately.
In general, there's a lot of rough potential in his songs, though I have a nagging overall feeling that it could be much better.
The Miracle Man (lyrics) describes a mind-controlling wizard's rise to power (or perhaps just a regular wizard - the themes of servitude and unwillingness could also be about coercion). It's disappointingly straightforward for a story that could have so many layers.
A few ideas for a rewrite:
- Unreliable narration: dystopian "positive" spin on the Miracle Man.
- Metaphysics and worldbuilding: What is magic to that world? (A theme Roper explores in Teaching Song) What are the counterpowers to the Miracle Man? What kind of society is he building? (I am not satisfied by the song's conclusion of evil for the lulz.)
- Metaphor and meaning: That's a very personal interpretation, but if I rewrote that story[1], I'd definitely make it molochian, with organizational servitude rather than plain might-makes-right or mind control.
I hope you enjoy discovering this song!
- ^
Or more plausibly, if I translate/adapt it to French.
Discuss
J-Lens: A Failed Replication on GPT-2
TLDR: We tested Anthropic's new(ish) J-Lens against the classic Logit Lens on GPT-2 small and medium. J-Lens loses at every layer (0/11 on small, 1/23 on medium). We ran 5 stress tests (more data, frequency checks, sparsity, tuning, and scaling) and it still loses every time. We draw two sharp methodological lessons for those doing probe-based interpretability.
Repo: https://github.com/nelithb/mech-interp-lab/
1: ResultJ-Lens does not work on GPT-2 small or medium.
Across 11 layers on GPT-2 small, J-Lens gets beaten by Logit Lens on every single one. On GPT-2 medium, it wins 1 of 23 layers, and even this ‘win’ is a rounding error (rank 221 vs 220).
This matters because the J-Lens paper makes a specific claim: that Logit Lens fails because early layer representations aren’t geometrically aligned with the final layer’s coordinate space, and that J-Lens corrects this. Our experiment directly tests this claim.
We ran five tests to make sure this wasn’t anomalous: refitting on more data; checking against confounding token frequencies; sparsity thresholding; tuning ‘skip_first’ parameter and scaling up from 124M to 255M parameters. J-Lens lost each time. The last scaling result is interesting in that the gap widened on the future token prediction metric, which is the opposite of J-Lens’s purpose.
The real value in this post is two methodological lessons that we extracted. If you are doing probe-based interpretability work, you can internalise these:
- Rank metrics are gameable by token frequency. A lens can actually score well by just predicting common tokens (e.g. commas and spaces). This was confirmed with 0.656 correlation between the shared bias term and token log frequency.
- Magnitude thresholding surface outlier dimensions, but not structure. When we thresholded J-Lens to sparsity, it ‘won’ at 7/11 layers. But upon inspection, there were only 9 unique top-1 tokens across all prompts. All of them were generic filler. These tracked to GPT-2’s known ‘massive activation’ outlier dimensions.
We made these mistakes so that you don’t have to.
2: Technical backgroundLogit LensTransformers work by passing a single vector – the residual stream – through the network. Each layer adds to this vector without replacing it.. So, an earlier layer's residual stream sits in the exact same dimensional space as the final one. Logit Lens (nostalgebraist, 2020) exploits this directly. It takes the model's final LayerNorm and unembedding matrix (the same weights normally reserved for the last layer) and applies them to an earlier layer's residual stream, to see what token the model would guess if it stopped there.
J-LensBut Logit Lens rests on a faulty assumption: that the early layer's representation lives within the geometric space that the final layer expects. J-Lens, in theory, corrects for this. As the paper puts it: 'While the logit lens assumes that representations use the same coordinates in all layers, the Jacobian lens corrects for representational changes that take place across layers.' If this claim is true, J-Lens becomes a cheap lens into the model at every layer. This experiment is a direct test of this.
3: SetupFor this project I only had access to an Intel Mac. PyTorch dropped x86 MacOS builds. Rather than moving everything to Colab, I ported the core into my existing CPU vent, using the official repo purely as a reference to check against.
To verify correct implementation, I conducted two checks:
- The paper states that setting J to the identity matrix should make J-Lens collapse to exactly Logit Lens. This was tested, returning a max difference of 0.00e+00, i.e. byte-identical.
- I ran a finite difference gradcheck against the actual perturbed network, at float64 (GPT-2's large late-layer residual norms lose gradient signal to floating-point cancellation at float32). This returned a max relative error of 8.4e-07: essentially machine precision.
The ported CPU implementation used for these experiments is available at https://github.com/nelithb/mech-interp-lab/. This is a clean re-implementation. The official Anthropic repo was used only as a reference for verification checks.
4: Headline resultFirstly, we ask: 'does the model's own top-1 next token get recovered, at 60 held-out prompts, per layer' (Metric A)? J-Lens loses at all 11 layers. J-Lens's median rank for the correct answer is in the thousands at early layers, reaching ~5-38 in the last two layers. Conversely, Logit Lens stays in the 1-43 range throughout.
Secondly, we look at future tokens at position p+1 to p+5 (Metric B). We measure at layer 9. We'd expect J-Lens to perform better here, given its whole purpose is to capture what the model is about to say, and not just what it's currently saying. Again, J-Lens is worse at every offset, with this gap persisting as we look further into the future (flat around -8 to -9 nats). J-Lens loses even on the task its own theory says it should win.
In truth, we use only 30 sequences (vs the paper's ~1000 sequences). So next, we test whether this undermines our result.
5: Stress testingWe firstly refit on 150 sequences, but the results remained essentially unchanged. J Lens still loses in all 11 layers on Metric A. J-Lens's rank for the correct answer is 11184, 12717, 9493, 3596, 4351, 1756, 2035, 457, 405... at successive layers, only reaching single digits at the very last one. Logit Lens at this point is near-perfect (rank 0). On Metric B, we run flat at -8.6 to -9.1 nats – practically matching the original pattern.
I ran into this comment from LessWrong user 'phoenix' which claimed that J-Lens's output is partially explained by raw token frequency, i.e. not genuine prediction. To test this, we measured the correlation between the shared final bias term (added identically in both lenses as they share the same final unembed step) and token log frequency across a 300k-token sample. We got r=0.656 — pretty close to phoenix's estimate (r=0.67). Although subtracting this bias out severely alters the absolute rank of both lenses (at layer 0: Logit Lens 43 -> 910; J-Lens 2759 -> 40992), the relative comparison remains unchanged. J-Lens still scores worse at all 11 layers.
Inspired by sparsity-based approaches, we tested whether a hard threshold towards sparsity on the fitted J would reveal a cleaner structure.. The first pass looked like a win: at 1% keep (the most aggressive threshold), J Lens won in 7/11 layers (!!), up from 0/11. But three signs pointed to this 'win' being the metric rewarding a content-agnostic filler predictor: only generic filler (e.g. commas or spaces) appeared in the top 5; only nine unique top-1 tokens ever appeared across the 60 held-out prompts; and the surviving weights traced to two of GPT-2's three known 'massive activation' outlier dimensions, independently documented by Timkey and van Schijndel (2021). So, common tokens score well on Metric A, regardless of whether anything real is being extracted.
Our fourth swing was to refit J at skip_first values of 0, 16 and 32. This resulted in up to a 20x improvement: mean J Lens median rank at early layers went from 8320 to 1937 to 408 as skip_first increases. Despite this, J Lens still goes 0/11 at every value tested. Even the best tuned configuration (~400 median rank) is roughly 100x worse than Logit Lens's single-to-low double digit range. Although this was methodologically useful, it just didn't save the result. One more roll of the dice...
6: Scale sweepCounterintuitively, when we increase scale from gpt2-small (124M) to gpt2-medium (355M), the gap widens. On Metric A, J-Lens wins in 1/23 layers, with this victory at layer 0. And yet, even this win is rounding-error level: J-Lens ranks 221; Logit Lens at 220.
Metric B's (future token prediction) result is striking not because the gap got bigger, but because its shape changed. On gpt2-small it was flat across all five offsets. On gpt2-medium it's no longer flat, it widens the further out you look. This is the wrong direction if J-Lens's future-token advantage was supposed to grow with scale. Recall that this is the more important metric for J-Len’s central claim. On gpt2-small, the logit-minus-J gap was flat across all five future offsets, roughly -8 to -9 nats, regardless of how far ahead you looked. On gpt2-medium, it's not flat, but monotonically widens: -4.970, -6.753, -7.239, -7.179, -7.277 at offsets p+1 through p+5. If J-Lens’s advantages were to emerge with scale, we’d expect this gap to shrink.
This result must be caveated that this is one data point, bridging 124M to 355M.
7: Open questionsOur result holds within the scope we tested. Here are three open opportunities to expand:
(1) We max out at 355M parameters. gpt2-large (774M) and gpt2-xl (1.5B) are natural next steps. The trend so far (albeit between two data points), i.e. the gap widening, argues against expecting a reversal.
(2) Expand beyond GPT-2. Belrose et al.'s Tuned Lens paper (arXiv:2303.08112) states that 'this simple form of the Logit Lens works reasonably well for GPT-2', with their own appendix testing, across four ~125M models, an improved lens against Logit Lens. They found that the improved lens clearly won only on OPT-125m, not GPT-2 medium, stating that 'these results did not generalize to other models tested.' On GPT-2, Logit Lens is unusually strong.
(3) This project actually started as a test of deception under instruction, and we didn't test that. The natural next step would be to use an instruction tuned model (e.g. Qwen 0.5-1.5B which is CPU feasible) and a localise-then-patch design instead of a probe-only readout.
8: TakeawaysAs a beginner, I wanted to include this section, so that if you’re doing lens or probe style interpretability work you can just internalise these lessons. I made the mistakes so that you don’t have to!
Firstly, rank metrics are gameable by frequency. We saw in reference to phoenix’s comments that a rank based metric can be moved a lot by oddities that are unrelated to genuine prediction.
Secondly, the magnitude thresholding test that resulted in the 7/11 J-Lens win surfaced the weights that were numerically the largest, which is not the same as the most meaningful. Always check the contents of the returned tokens to verify if this is the case.
I wouldn’t have found these obstacles had I not got stuck in.
Note on AI useI used AI to help write all the code in this project. It was also used to provide explanations for the results and theoretical concepts, and find and summarise key existing research. The visualiser was also made with AI. I'm not an AI researcher, or even a developer, but I can't wait to continue digging deeper with AI's help.
SourcesGurnee, W., Sofroniew, N., Pearce, A., Piotrowski, M., Kauvar, I., et al. (2026). Verbalizable Representations Form a Global Workspace in Language Models. Transformer Circuits Thread, Anthropic.
- Available at: https://transformer-circuits.pub/2026/workspace
Anthropic. (2026). jacobian-lens [Companion code repository]. GitHub.
- Available at: https://github.com/anthropics/jacobian-lens
nostalgebraist. (2020). Interpreting GPT: The Logit Lens. LessWrong.
Belrose, N., Furman, Z., Smith, L., Halawi, D., Ostrovsky, I., McKinney, L., Biderman, S., & Steinhardt, J. (2023). Eliciting Latent Predictions from Transformers with the Tuned Lens. arXiv:2303.08112.
- Available at: https://arxiv.org/abs/2303.08112
Timkey, W., & van Schijndel, M. (2021). All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational Quality. In Proceedings of EMNLP 2021, pp. 4527–4546.
- Available at: https://aclanthology.org/2021.emnlp-main.370/
Sun, M., Chen, X., Kolter, J. Z., & Liu, Z. (2024). Massive Activations in Large Language Models. arXiv:2402.17762.
- Available at: https://arxiv.org/abs/2402.17762
Ran-Milo, Y. (2026). Attention Sinks Are Provably Necessary in Softmax Transformers: Evidence from Trigger-Conditional Tasks. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 2).
- Available at: https://aclanthology.org/2026.acl-short.8/
Nanda, N., & Bloom, J. (2022). TransformerLens [Software library]. GitHub.
- Available at: https://github.com/TransformerLensOrg/TransformerLens
phoenix. (2026). Comment on "A global workspace in language models". LessWrong.
- Available at: https://www.lesswrong.com/posts/3PaLrzxagpbnNtPLT?commentId=c4dNnEwARCxLBm9YG [Referenced in comments section]
Discuss
Scoring your own decisions is harder than scoring forecasts
An upfront disclaimer: I built and sell the app this post ends with, so ignore the last section if you will.
I've read a lot lately about calibration, forecasting tournaments, Brier scores etc, and almost all of it is about predicting the world. My question is smaller: does my confidence mean anything, on my own decisions? The obvious way to check would be to remember a decision, how sure I was, and compare with how it turned out. Well, after resolution, I already know how it turned out, so whatever confidence I "remember" already moved toward the outcome, and I can't even really tell by how much.
As an example, before taking the current job, which is half office, half remote, I had another option on the table, fully remote, most other things similar. And that's already one of the problems, I can't remember them all, just that "most" were similar. At the moment I'm both content and satisfied with the decision and I'd also make it again, but is it because I got used to it? Because the job market is tough right now and that influences how I feel? I do remember how important fully remote was for me back then, but everything else is blurry and I definitely can't remember how confident I was in this choice. Right now it feels like I was sure, but the only thing that backs that up is the fully remote part, only because it is very important to me, making the emotion strong around the topic.
So memory is out of the question, I don't think anyone here needs convincing of that. What I couldn't find was a way to keep a record that holds up on a real decision, one that takes months to play out, and what I ended up with is still soft on the measurement side, for which I'd very much love feedback on.
Fatebook is the closest thing I found: you write a question, set a resolve-by date, get an email when the date arrives and mark it yes/no/ambiguous yourself, and it stays private unless you share it. Manifold works the same way in public, with the creator resolving their own decisions. Metaculus is the odd one out, since its admins do the resolving and you can't write a question that lets you grade it yourself. All 3 take ranges and multiple choice too, not just a yes or no, but they all want a question with a date on it.
Think of it like this:
- "Will the contract be renewed within 6 months?" is a Fatebook question, it has a date and it's either right or wrong;
- "Take the contract or enjoy some free time?" isn't, there's no date, the consequences might be felt for the next year and my own sense of whether it went well keeps moving for months after I supposedly know the answer.
The second one is what I actually wanted to track, the first one is only a piece of it. What Fatebook is missing is everything in between: what I learned a few weeks later and whether it changed my confidence, how I felt about the call at resolution, whether I still felt that way a few months later still and so on. That isn't a complaint about those tools, I just wanted a bit more.
The prediction should be locked, otherwise I'd be able to change it a bit, maybe a few times, then convince or fool myself that that's what I thought at the time.
The confidence should be a number, not "fairly sure", because "fairly sure" can't be plotted against anything. Yes, for a single decision it's false precision, I know, but the number only means something across many of them anyway, and without it, well, it's just a diary.
Having the history that happens along the way. On an 8-month decision most of the useful data is in the middle and that's the part with the highest chances to disappear, because once I know how it ended, I read every note knowing the ending, every note is either a sign I should've caught or something that didn't matter, in hindsight.
So I wanted a note every now and then. I don't really have a rule for how often, except whenever something happened, with the new info, whether it made things look better or worse and where my confidence is now. Even when nothing happened, if a few weeks passed.
Taking a second look, a while after resolution, because how I feel about a call in the first weeks and how I feel about it after a few months are sometimes completely different and if nothing asks me, I don't go back on my own.
And I wanted the following three kept apart: was the prediction right, how do I feel about it, would I do it again, because satisfaction doesn't always equal how right I was about it, nor if I'd do it again. A call can turn out exactly as predicted and still feel bad a few months later, or the other way around, and collapsing them into "was it a good decision" just grades the decision by its outcome.
I thought of doing them by hand for a while, but I never really got into it, for various reasons: nothing to remind me, nothing properly structured (sure, I can set up a spreadsheet or Notion page to my liking, but nothing really clicked). Might be a me problem, but without something nagging me I tend to not come back to it, however good the intentions were at the start.
If you look for whether decision journals work you'll probably meet a number I met: about 19% better forecasting accuracy from journaling, credited to a study in Behavioral Science & Policy. I wanted to cite it, but I couldn't find it in the journal's archive or anywhere else except posts citing each other. Maybe it exists somewhere, I only mention it because it's out there.
What I did find mostly holds up, with one exception I'll get to, so here's the short-ish version.
Fischhoff and Beyth did the obvious experiment back in 1972, before Nixon's trips to China and the USSR: they asked students to put probabilities on what would come out of them (would Nixon meet Mao, would the US recognize China, that kind of thing) and weeks/months later they asked the same students to recall the probabilities they gave. The recalled ones had moved toward whatever the students believed had happened, exactly what I described above with the job, and they couldn't tell by how much either (Fischhoff and Beyth, 1975).
The other well known one is the overconfidence test: give a low and a high guess for some number (how many countries are in the UN, let's say) such that you're 90% sure the real one is between them; the real one should fall outside about 1 time in 10, but it falls outside 4 to 6 times in 10. Russo and Schoemaker got that from a couple of thousand professionals, in Sloan Management Review, and the questions from the professionals' own industry came out about as badly as the general ones; Alpert and Raiffa had asked for 98% ranges a decade earlier and got about the same 4 in 10, where it should've been 2 in 100.
Then there's the feedback, which fixes some of it: Lichtenstein and Fischhoff showed people their hit rate after each round of confidence judgments, back in 1980, and the stated confidence moved toward reality, most of the gain after the first round. Weather forecasters are the usual example, they put a probability on rain every single day, they get scored on it the next day and they end up almost perfectly calibrated (Murphy and Winkler, 1984). And training, a bit: a module on probabilistic reasoning that takes under an hour improved Brier scores by 6 to 11% over the control group, across the 4 years of Tetlock's Good Judgment Project (Chang et al., 2016).
The last one above turned out to be the exception, though. A 2025 reanalysis took the teaming and training effects from that tournament's first 2 years (Mellers et al., 2014) and ran them through a model controlling for which questions people picked, when they answered and how hard the questions were, none of which the original design controlled, and the effects shrank, went away, or in places reversed (Hauenstein et al., 2025). It only covers the first two years, while Chang's numbers run across 4, but it's the same tournament and the same design underneath, and I'd take the training result with a grain of salt.
Of course, none of these tested decision journaling as a whole, only its parts (forecasting, calibration, feedback, etc). What I'm least confident about (pun intended) is whether the forecasting results, from geopolitical tournaments with questions that have a clear yes or no and thousands of people answering the same ones, apply to questions like "should I take this job"; I'm assuming they do, for now. The calibration training only partly carries over to other kinds of questions, from what I could find, and journaling has a selection problem too, because whoever keeps one already cares about their judgment. I wrote the longer version up in a guide, with the same links, if anyone wants to check it.
No matter the tool, I found I still have a few problems I don't have clear answers to, since they're in the practice itself and not in any app.
Self-gradingResolution is self-graded, so I decide whether my own call was right and nobody else looks at it. A second person with read access would fix that, but if I know someone else will read my prediction, I'll write it knowing I'll be held accountable, probably not 100% true. Vague predictions have the same problem, "this will probably work out" never counts as a miss, and once I start caring about the hit rate, it starts to pay to write them like that, which is Goodhart's law, more or less.
Do I grade myself honestly? Do I count the vague ones as misses? Would I write the same prediction if I knew someone was going to read it?
Small numbersMeaningful decisions don't come often, maybe a handful a year, and calibration only shows up across many predictions, so a personal record stays too small to tell me anything for a long time, at a handful a year the first 10 resolved are a couple of years away, and I still read it as if it told me something.
I'm an iOS developer, so, naturally I decided to build an iOS app that runs those 4 requirements (and my 5th woven in):
- a decision with a prediction, a confidence percentage and a review date
- check-ins in between that keep a history of what happened, tagged positive or negative, and what it moved my confidence to
- a resolution with an outcome, right, wrong or mixed, and a satisfaction score
- a follow-up asking whether I'd make the same call (60 days after resolution by default but fully configurable)
It has reminders for each of them, the main thing a by-hand version can't do.
It draws a reliability diagram across resolved decisions: confidence at decision time grouped in steps of 10 (70 to 79%, 80 to 89% and so on), against how often the calls in each group came out right, and keeps a "preview" label on it until 10 have resolved.
After that, it also shows a hit rate, the mean gap between confidence and outcome, and a Brier score, about which I'm not so sure: it's a proper scoring rule, but I'm the one grading the outcomes it scores, so the precision is only partly real, and in practice I look at the groups and their counts before I look at scoring.
Pretty much most of it is read-only once written: the prediction, the confidence, the check-ins and the resolution. The title can be renamed, but the original is kept, and the confidence can move, but only through check-ins, so the movement is part of the history entries.
It's called Reckon (App Store, site), $3.99 once, no subscription, iPhone and iPad, with a Mac version I'm still working on. No accounts required, it only syncs through iCloud and there is no data collection or tracking. The per-category breakdown only appears at 15 resolved with at least 3 per category, so most of the insights take a while to show up, but it's what I thought is a good default. The calibration screen looks like this:
Sample data, so it has enough entries to display
- How a mixed outcome should score. A resolution can be right, wrong or mixed, and a mixed one counts as 0.5 in the Brier score and in the confidence-to-outcome gap, but as a miss in the per-group hit rate on the diagram. Each made sense, taken individually, but I also feel it should behave the same. Should it?
- Whether the follow-up should feed the calibration numbers at all. Right now it's kept separate: "would you make this call again" doesn't retroactively mark the prediction wrong, since that would grade the prediction by how I feel about the call. But if I answer no, that says something about the original call too, and keeping them fully separate might not be right either. Should a "no" count for something, and how much?
- Whether 10 resolved decisions is far too low of a bar to draw a reliability diagram at. Spread over the 10-point groups, that's about 1 per group, and at that size a single call moves a group's hit rate by up to 100 points, so the curve is mostly noise at that point, it just doesn't look like it. A per-group minimum would be the alternative, or no curve at all and only show the counts, but I had to pick a number and I did it mostly by feel. Is 10 far too low, and if so, what would you pick?
I'd love to hear about any of the 3, and about anything else in the measurement that looks wrong to you, but also any other feedback. And if you've kept a record like this in the past, I'd really like to know whether you had success with it and what where its shortcomings.
Discuss
Who else is steering?
This post raises a methodological question with regard to how to make sense of models' behaviors in response to activation steering anchored on a concept.
Our discussion is limited to a particular kind of behaviors induced by activation steering, that of introspection studied by Lindsey (2026) (first accessed at here) and a range of investigations it inspires (see LW Post and citations in there). That said, the implication of our discussion applies to activation steering in general, mutatis mutandis.
Introspection refers to a model's "self-awareness" of its perturbed internal state. The self-awareness in question might be observable from the model's self-report, when prompted without giving away the existence or details of the perturbation. On the other hand, the perturbation itself is conducted by injecting a concept vector into the model's residual stream at a certain site. For model's self-report to be admissible as causally related to the concept injection, it must meet certain reasonable criteria, such as those given by Lindsey. One of those requires the report's content be specific to the injected concept.
Introspection thus illustrates a form of activation steering where both the steerer and the behavioral response are anchored on a concept. A closer examination of the construction of the concept vector in the experimental setup of introspection studies, however, raises questions about who else is steering, and what concept-anchored introspection really means.
Let mjx-container[jax="CHTML"] { line-height: 0; } mjx-container [space="1"] { margin-left: .111em; } mjx-container [space="2"] { margin-left: .167em; } mjx-container [space="3"] { margin-left: .222em; } mjx-container [space="4"] { margin-left: .278em; } mjx-container [space="5"] { margin-left: .333em; } mjx-container [rspace="1"] { margin-right: .111em; } mjx-container [rspace="2"] { margin-right: .167em; } mjx-container [rspace="3"] { margin-right: .222em; } mjx-container [rspace="4"] { margin-right: .278em; } mjx-container [rspace="5"] { margin-right: .333em; } mjx-container [size="s"] { font-size: 70.7%; } mjx-container [size="ss"] { font-size: 50%; } mjx-container [size="Tn"] { font-size: 60%; } mjx-container [size="sm"] { font-size: 85%; } mjx-container [size="lg"] { font-size: 120%; } mjx-container [size="Lg"] { font-size: 144%; } mjx-container [size="LG"] { font-size: 173%; } mjx-container [size="hg"] { font-size: 207%; } mjx-container [size="HG"] { font-size: 249%; } mjx-container [width="full"] { width: 100%; } mjx-box { display: inline-block; } mjx-block { display: block; } mjx-itable { display: inline-table; } mjx-row { display: table-row; } mjx-row > * { display: table-cell; } mjx-mtext { display: inline-block; } mjx-mstyle { display: inline-block; } mjx-merror { display: inline-block; color: red; background-color: yellow; } mjx-mphantom { visibility: hidden; } _::-webkit-full-page-media, _:future, :root mjx-container { will-change: opacity; } mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; font-weight: normal; font-size: 100%; font-size-adjust: none; letter-spacing: normal; border-collapse: collapse; word-wrap: normal; word-spacing: normal; white-space: nowrap; direction: ltr; padding: 1px 0; } mjx-container[jax="CHTML"][display="true"] { display: block; text-align: center; margin: 1em 0; } mjx-container[jax="CHTML"][display="true"][width="full"] { display: flex; } mjx-container[jax="CHTML"][display="true"] mjx-math { padding: 0; } mjx-container[jax="CHTML"][justify="left"] { text-align: left; } mjx-container[jax="CHTML"][justify="right"] { text-align: right; } mjx-mi { display: inline-block; text-align: left; } mjx-c { display: inline-block; } mjx-utext { display: inline-block; padding: .75em 0 .2em 0; } mjx-mo { display: inline-block; text-align: left; } mjx-stretchy-h { display: inline-table; width: 100%; } mjx-stretchy-h > * { display: table-cell; width: 0; } mjx-stretchy-h > * > mjx-c { display: inline-block; transform: scalex(1.0000001); } mjx-stretchy-h > * > mjx-c::before { display: inline-block; width: initial; } mjx-stretchy-h > mjx-ext { /* IE */ overflow: hidden; /* others */ overflow: clip visible; width: 100%; } mjx-stretchy-h > mjx-ext > mjx-c::before { transform: scalex(500); } mjx-stretchy-h > mjx-ext > mjx-c { width: 0; } mjx-stretchy-h > mjx-beg > mjx-c { margin-right: -.1em; } mjx-stretchy-h > mjx-end > mjx-c { margin-left: -.1em; } mjx-stretchy-v { display: inline-block; } mjx-stretchy-v > * { display: block; } mjx-stretchy-v > mjx-beg { height: 0; } mjx-stretchy-v > mjx-end > mjx-c { display: block; } mjx-stretchy-v > * > mjx-c { transform: scaley(1.0000001); transform-origin: left center; overflow: hidden; } mjx-stretchy-v > mjx-ext { display: block; height: 100%; box-sizing: border-box; border: 0px solid transparent; /* IE */ overflow: hidden; /* others */ overflow: visible clip; } mjx-stretchy-v > mjx-ext > mjx-c::before { width: initial; box-sizing: border-box; } mjx-stretchy-v > mjx-ext > mjx-c { transform: scaleY(500) translateY(.075em); overflow: visible; } mjx-mark { display: inline-block; height: 0px; } mjx-TeXAtom { display: inline-block; text-align: left; } mjx-msup { display: inline-block; text-align: left; } mjx-mover { display: inline-block; text-align: left; } mjx-mover:not([limits="false"]) { padding-top: .1em; } mjx-mover:not([limits="false"]) > * { display: block; text-align: left; } mjx-msub { display: inline-block; text-align: left; } mjx-c::before { display: block; width: 0; } .MJX-TEX { font-family: MJXZERO, MJXTEX; } .TEX-B { font-family: MJXZERO, MJXTEX-B; } .TEX-I { font-family: MJXZERO, MJXTEX-I; } .TEX-MI { font-family: MJXZERO, MJXTEX-MI; } .TEX-BI { font-family: MJXZERO, MJXTEX-BI; } .TEX-S1 { font-family: MJXZERO, MJXTEX-S1; } .TEX-S2 { font-family: MJXZERO, MJXTEX-S2; } .TEX-S3 { font-family: MJXZERO, MJXTEX-S3; } .TEX-S4 { font-family: MJXZERO, MJXTEX-S4; } .TEX-A { font-family: MJXZERO, MJXTEX-A; } .TEX-C { font-family: MJXZERO, MJXTEX-C; } .TEX-CB { font-family: MJXZERO, MJXTEX-CB; } .TEX-FR { font-family: MJXZERO, MJXTEX-FR; } .TEX-FRB { font-family: MJXZERO, MJXTEX-FRB; } .TEX-SS { font-family: MJXZERO, MJXTEX-SS; } .TEX-SSB { font-family: MJXZERO, MJXTEX-SSB; } .TEX-SSI { font-family: MJXZERO, MJXTEX-SSI; } .TEX-SC { font-family: MJXZERO, MJXTEX-SC; } .TEX-T { font-family: MJXZERO, MJXTEX-T; } .TEX-V { font-family: MJXZERO, MJXTEX-V; } .TEX-VB { font-family: MJXZERO, MJXTEX-VB; } mjx-stretchy-v mjx-c, mjx-stretchy-h mjx-c { font-family: MJXZERO, MJXTEX-S1, MJXTEX-S4, MJXTEX, MJXTEX-A ! important; } @font-face /* 0 */ { font-family: MJXZERO; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Zero.woff") format("woff"); } @font-face /* 1 */ { font-family: MJXTEX; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Regular.woff") format("woff"); } @font-face /* 2 */ { font-family: MJXTEX-B; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Bold.woff") format("woff"); } @font-face /* 3 */ { font-family: MJXTEX-I; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-Italic.woff") format("woff"); } @font-face /* 4 */ { font-family: MJXTEX-MI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Main-Italic.woff") format("woff"); } @font-face /* 5 */ { font-family: MJXTEX-BI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Math-BoldItalic.woff") format("woff"); } @font-face /* 6 */ { font-family: MJXTEX-S1; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size1-Regular.woff") format("woff"); } @font-face /* 7 */ { font-family: MJXTEX-S2; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size2-Regular.woff") format("woff"); } @font-face /* 8 */ { font-family: MJXTEX-S3; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size3-Regular.woff") format("woff"); } @font-face /* 9 */ { font-family: MJXTEX-S4; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Size4-Regular.woff") format("woff"); } @font-face /* 10 */ { font-family: MJXTEX-A; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_AMS-Regular.woff") format("woff"); } @font-face /* 11 */ { font-family: MJXTEX-C; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Regular.woff") format("woff"); } @font-face /* 12 */ { font-family: MJXTEX-CB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Calligraphic-Bold.woff") format("woff"); } @font-face /* 13 */ { font-family: MJXTEX-FR; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Regular.woff") format("woff"); } @font-face /* 14 */ { font-family: MJXTEX-FRB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Fraktur-Bold.woff") format("woff"); } @font-face /* 15 */ { font-family: MJXTEX-SS; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Regular.woff") format("woff"); } @font-face /* 16 */ { font-family: MJXTEX-SSB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Bold.woff") format("woff"); } @font-face /* 17 */ { font-family: MJXTEX-SSI; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_SansSerif-Italic.woff") format("woff"); } @font-face /* 18 */ { font-family: MJXTEX-SC; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Script-Regular.woff") format("woff"); } @font-face /* 19 */ { font-family: MJXTEX-T; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Typewriter-Regular.woff") format("woff"); } @font-face /* 20 */ { font-family: MJXTEX-V; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Regular.woff") format("woff"); } @font-face /* 21 */ { font-family: MJXTEX-VB; src: url("https://cdn.jsdelivr.net/npm/mathjax@3/es5/output/chtml/fonts/woff-v2/MathJax_Vector-Bold.woff") format("woff"); } mjx-c.mjx-c1D440.TEX-I::before { padding: 0.683em 1.051em 0 0; content: "M"; } mjx-c.mjx-c1D464.TEX-I::before { padding: 0.443em 0.716em 0.011em 0; content: "w"; } mjx-c.mjx-c1D44E.TEX-I::before { padding: 0.441em 0.529em 0.01em 0; content: "a"; } mjx-c.mjx-c28::before { padding: 0.75em 0.389em 0.25em 0; content: "("; } mjx-c.mjx-c29::before { padding: 0.75em 0.389em 0.25em 0; content: ")"; } mjx-c.mjx-c2032::before { padding: 0.56em 0.275em 0 0; content: "\2032"; } mjx-c.mjx-c73::before { padding: 0.448em 0.394em 0.011em 0; content: "s"; } mjx-c.mjx-cAF::before { padding: 0.59em 0.5em 0 0; content: "\AF"; } mjx-c.mjx-c1D44B.TEX-I::before { padding: 0.683em 0.852em 0 0; content: "X"; } mjx-c.mjx-c1D465.TEX-I::before { padding: 0.442em 0.572em 0.011em 0; content: "x"; } mjx-c.mjx-c2208::before { padding: 0.54em 0.667em 0.04em 0; content: "\2208"; } mjx-c.mjx-c1D451.TEX-I::before { padding: 0.694em 0.52em 0.01em 0; content: "d"; } mjx-c.mjx-c2C::before { padding: 0.121em 0.278em 0.194em 0; content: ","; } mjx-c.mjx-c3D::before { padding: 0.583em 0.778em 0.082em 0; content: "="; } mjx-c.mjx-c2212::before { padding: 0.583em 0.778em 0.082em 0; content: "\2212"; } mjx-c.mjx-c1D44C.TEX-I::before { padding: 0.683em 0.763em 0 0; content: "Y"; } be a model. In introspection studies, the concept vector of a word is typically obtained by the difference between two values:
- = residual stream activation on the last token before reply to Tell me about ;
- = mean of for , being a set of words excluding .
We may call a contrast set. Thus the concept vector depends on a word and a contrast set, and is only one of its kind.
Suppose we observe at some rate introspection behaviors specific to some displayed by model , when steered by , e.g., with being the 100 baseline words in Lindsey (2026). The following questions arise:
- Replace with any such that and are numerically dissimilar enough. Then does , steered by , display similar introspection behaviors?
- If yes, what is the invariant that unifies and under the concept ?
- If no, to what extent is the introspection upon steering by an artifact of ?
These questions are empirically answerable. My last attempt was foiled, as I could not replicate introspection on a small open-source model like Qwen3-4B-Instruct-2507. Nevertheless, I think for those who can access sufficient compute, these questions are worth exploring. Should (2) be the case, we need to be cautious when attributing causes to models' introspection or any activation-steering-induced behaviors: externalities, such as the contrast set for constructing concept vectors, may be lying in wait.
Discuss
Страницы
- « первая
- ‹ предыдущая
- …
- 20
- 21
- 22
- 23
- 24
- 25
- 26
- 27
- 28
- …
- следующая ›
- последняя »