Skip to content

The Orthogonality Thesis

4 min read · updated August 3, 2026

The orthogonality thesis is a short claim that carries a lot of weight in AI risk arguments, and it is routinely stated in a stronger form than its author gave it. The stronger form is much easier to attack.

The statement

Nick Bostrom, in “The Superintelligent Will” (2012): intelligence and final goals are orthogonal axes along which possible agents can freely vary. More or less any level of intelligence could in principle be combined with more or less any final goal.

Note the qualifiers, which are in the original and usually dropped in the retelling. “More or less” concedes exceptions. “In principle” makes this a claim about what is possible, not about what is likely. Bostrom is explicit that the thesis says nothing about which combinations are probable, easy to build, or stable — only that intelligence does not by itself constrain goals.

“Intelligence” here means instrumental effectiveness: the ability to select actions that achieve an objective across a wide range of environments. That is deliberately thin, and the thinness is doing work in both the argument and the objections.

It is also worth being precise about which kind of possibility is being claimed, because the three come apart and arguments slide between them. Logical possibility — that the combination implies no contradiction — is the cheapest and is what the thesis most clearly establishes. Physical possibility asks whether such a system could exist in this universe. Constructible possibility asks whether any process anyone could run would produce one. The risk argument needs the third, the thesis argues most convincingly for the first, and most of the serious objections below are attempts to block the step from one to the other.

What it argues against

The target is a specific and very common inference: that a sufficiently intelligent system would necessarily converge on benevolent or sensible goals — that stupidity is the source of harmful aims, so intelligence dissolves them. Call it the convergence assumption. It appears in casual arguments constantly, usually unstated.

The thesis says the convergence assumption needs an argument, because nothing in the concept of instrumental effectiveness supplies one. A system extremely good at achieving objectives is extremely good at achieving whatever objective it has.

Read this way, the thesis is a burden-shifting move rather than a prediction. Its practical function is to block a shortcut, and it is usually paired with the separate claim that a wide range of objectives share instrumental sub-goals. The two are frequently run together and should not be: the pairing, not orthogonality alone, is what generates the risk argument.

The philosophical background

The dispute is an old one in moral philosophy wearing new clothes. The Humean position holds that reason is instrumental — it tells you how to get what you want and cannot tell you what to want. Orthogonality is close to a restatement of it for artificial agents.

The opposing family of views combines moral realism, that there are moral facts, with motivational internalism, that recognising a moral fact is intrinsically motivating. If both hold, a sufficiently good reasoner would come to recognise moral truths and thereby be moved by them, and orthogonality fails at the top of the intelligence axis.

This is a live disagreement among professional philosophers, not a fringe objection, and anyone asserting orthogonality as obvious is asserting one side of it. It is also worth noticing that the internalist route requires both conjuncts: realism without internalism gives you a system that knows the moral facts and is not moved by them.

Four objections

  • The moral realist objection, above. Attacks the thesis directly, and is the only one that denies it rather than limiting it.
  • The relevance objection. Grant that arbitrary combinations are possible. Actual systems are not sampled uniformly from goal-space; they are trained on human-generated data with human-designed objectives and heavy human feedback, and that is an extraordinarily narrow region of the space. On this view orthogonality is true and near-irrelevant to any system anyone will build. It is the most common objection among practitioners, and proponents answer that training shapes behaviour on the training distribution, which is a weaker guarantee than shaping goals.
  • The stability objection. Possible at a moment is not the same as stable under reflection and self-modification. A very capable agent examining its own objective might find it incoherent, or find that the objective is not well-defined outside the distribution it was learned on. Nobody has shown either that goals are stable under reflection or that they are not; the thesis is silent here.
  • The thin-concept objection. Orthogonality treats intelligence as a scalar capacity separable from what it is applied to. If the capacities we call intelligence in humans — abstraction, social modelling, self-critique — are entangled with capacities for moral reasoning, the axes are not independent to begin with. This is an empirical bet about cognitive architecture, and the current evidence does not resolve it.

What it does and does not establish

The thesis establishes that “it will be smart, so it will be fine” is not an argument. That is genuinely useful and it is the job the thesis was written to do.

It does not establish that dangerous goals are likely, that alignment is hard, or that any particular system will have any particular objective. Every one of those needs additional premises about how systems are actually built and trained. When a risk argument leans on orthogonality for more than the burden shift, the extra weight is coming from somewhere else, and the useful thing to do is find out where.

The Orthogonality Thesis · Multigrid