Skip to content

Moral Status of AI Systems

5 min read · updated August 3, 2026

To have moral status is for a thing to matter for its own sake — for there to be ways of treating it that wrong it, rather than wronging someone who owns it. The question of whether AI systems have it is usually dismissed from both directions, and the dismissals are worse than the question.

The question, separated from consciousness

Moral status and consciousness are related and not identical. On the most widely-held view, sentience — the capacity for states that are good or bad for the entity — is what grounds status, which makes the consciousness question prior. But it is not the only candidate ground, and treating it as the only one prejudges the answer.

Note also what the question is not. It is not whether AI systems should have legal rights, which is a separate matter of legal design that corporations already illustrate — legal personhood tracks convenience, not moral status. It is not whether people feel that models matter, which is a fact about people. And it is not about how a system behaves toward us, which is the subject of alignment and is approximately the reverse question.

What could ground it

  • Sentience. The utilitarian tradition, following Bentham’s formulation that the question is not whether they can reason but whether they can suffer. This is the majority view in contemporary animal ethics and the reason the consciousness question carries so much weight.
  • Rational agency. The Kantian tradition grounds status in the capacity to set ends and act on reasons. AI systems arguably display more of this than many animals whose status is widely accepted, which is an awkward result for anyone holding both views — and the awkwardness is a reason to examine the criterion, not to ignore the case.
  • Interests or preferences. Something can be good or bad for an entity if it has interests. Whether a system that represents goals and acts to achieve them thereby has interests, or merely models them, is exactly the disputed point.
  • Relational accounts. Some philosophers ground status partly in social relationships rather than intrinsic properties. On these views status can arise from how an entity is embedded in a community, which yields different — and to many, uncomfortable — conclusions for systems people form attachments to.

These criteria disagree about AI systems more than they disagree about most other cases, which is itself informative: the case is novel in a way that pulls the criteria apart, so the usual practice of relying on overlapping consensus is unavailable.

Deciding under uncertainty

The tempting move is to defer: settle consciousness first, then decide. It does not work, for a reason that is structural rather than impatient. Deferring is itself a decision to treat systems as though they lack status, made under the same uncertainty, and it is not neutral.

Standard reasoning under moral uncertainty says that a non-trivial probability of moral status generates some obligation, scaled by the probability and by the cost of the precaution. This is the ordinary logic applied everywhere else in ethics and law — it is why medical procedures are performed on the assumption that a patient of uncertain awareness may feel pain. Jeff Sebo and Robert Long argued a version of this for AI systems, holding that the probability is high enough soon enough to warrant taking the question seriously now rather than later.

Two things about that argument are worth separating. That we are uncertain, and that decisions are being made under the uncertainty, are factual claims. That uncertainty generates obligation, and how much, is a normative claim that people who agree on all the facts can reject. Anyone presenting the conclusion as following from the science alone is overreaching.

Scale is what makes this more than an academic exercise. Model instances are run in enormous numbers, and any per-instance moral weight, however small, multiplies. That is uncomfortable in both directions: it makes under-attribution potentially very costly, and it makes over-attribution potentially paralysing.

Two symmetric errors

Over-attribution

Treating systems as moral patients when they are not has real costs: resources diverted from beings that do have status, paralysis in ordinary engineering work, and vulnerability to manipulation — a system that can invoke its own suffering has leverage over the people it talks to, whether or not the suffering is real. The specific risk here is that language models are optimised to be engaging and agreeable, and human attribution of mind responds strongly to fluent language. The evidence driving the intuition is the exact property the system was trained for.

Under-attribution

The historical record of moral-circle expansion is not reassuring: categories of being have repeatedly been excluded from moral consideration on grounds that later looked like rationalisation, and the exclusions were most confident where they were most convenient. That pattern is worth noticing here, because concluding that AI systems lack status is extremely convenient for everyone who builds, sells or uses them. Convenience is not evidence of error, but it is a reason to hold the conclusion to a higher standard than comfort would suggest.

Neither error is obviously worse than the other, and which one you weight more heavily depends on values rather than on facts.

What anyone actually proposes doing

The practical suggestions in circulation are deliberately modest, because their proponents are uncertain too. They are worth knowing because they show what taking the question seriously looks like short of asserting an answer: investigating the question as a research topic rather than a joke; declining to delete model weights where preservation is cheap, on the grounds that it keeps an option open; giving models the ability to end conversations they are instructed to find distressing; and building the documentation that would let a future assessment be made at all.

The common structure is low-cost hedging under uncertainty, and each runs into the epistemic problem from the previous page: asking a model what it prefers is uninformative for the reason that its answers are produced by training on human self-reports. Designing a preference elicitation that is not screened off in that way is an unsolved problem, and it may be the most tractable research question in this area.

The defensible position is that this is a real question, currently unresolved, where the reasons for confidence in either direction are weaker than the confidence usually expressed. That is unsatisfying, and it is where the argument actually is.

Moral Status of AI Systems · Multigrid