AI Safety Institutes and What They Do
4 min read · updated August 3, 2026
Bodies of this type — a technical AI evaluation unit inside government, usually attached to a standards or science agency rather than to a regulator — were established in several jurisdictions from around 2023 onwards, and coordinate through an international network. What they can actually make anyone do is widely misunderstood in both directions.
The institutional form
The recurring design is a small, technical, well-funded unit that is deliberately not a regulator. It does not license, it does not approve, and in most cases it cannot fine. It is staffed to do engineering rather than enforcement, and it is usually placed inside a metrology, standards or science ministry rather than inside the department that would prosecute.
That placement is a choice with consequences and it is defended explicitly. The argument for it: an evaluator that can penalise you will not be given honest access, and honest access before deployment is the only way to learn anything a public benchmark cannot already tell you. The argument against, made by people who want harder governance: an evaluator with no power produces reports that a company is free to ignore, and the arrangement therefore lends public legitimacy to private decisions without constraining them. Both are serious.
What the work actually consists of
- Pre-deployment evaluation of models supplied voluntarily under an agreement, usually covering a set of capability domains the government considers security-relevant, plus general capability and misuse-resistance testing.
- Evaluation methodology. The more durable output. Building an evaluation that measures what it claims to measure is an unsolved research problem, and a public body has an unusual incentive to publish methods rather than keep them.
- Standards input. Feeding technical work into the standards bodies whose documents later get referenced by regulation and by procurement rules.
- Threat modelling and guidance for domestic sectors, typically in cooperation with the existing cyber security agency.
- Information sharing across an international network, which matters mainly because it reduces the number of times the same model is tested from scratch.
Where the authority comes from
Almost none of it is direct. The routes are indirect and all of them are worth tracing, because this is how technical bodies acquire real influence:
Access agreements. Voluntary, negotiated, and therefore revocable. The leverage is reputational: declining to participate is visible. Whether that leverage survives commercial pressure has not really been tested.
Publication. A finding published by a government technical body is evidence. It can be cited in litigation, in a regulator’s enforcement decision, in a procurement exclusion, and in a parliamentary committee. This is a slow lever and a sharp one.
Standards, then presumption. Several regulatory designs give conformity with a harmonised standard the effect of a presumption of compliance. If a body’s methods become the standard, its methods become the cheapest route to compliance, and a cheapest route is functionally a requirement. This is the most consequential channel and the least covered.
Procurement. Government buying conditions can require evaluation against a named methodology without any statute at all.
There is also a quieter function that these bodies perform and that is worth stating because it is often the real justification: they build state capacity. A government that cannot evaluate a model cannot regulate one, cannot buy one competently, cannot assess a claim made in a hearing, and cannot tell a genuine safety case from a well-produced document. Whatever one thinks of the current arrangement, an evaluation capability inside government is a precondition for every harder governance option, including the ones its critics prefer. That argument is made across the political range and it does not depend on the body having any enforcement power at all.
The same capacity argument explains why these units tend to co-locate with existing national security and metrology functions rather than with consumer or competition regulators. Measurement institutes already know how to define a test that means the same thing in two laboratories, which is the unglamorous problem at the centre of model evaluation, and security agencies already hold the clearances that classified threat work requires.
Four structural limits
- Evaluation gives a lower bound, not an upper one. This is the deepest problem and it is technical rather than political. A test that fails to elicit a capability has not shown the capability is absent; it has shown that this elicitation did not find it. Better prompting, tools, fine-tuning or scaffolding can raise measured capability on an unchanged model. Any statement that a model “cannot” do something is really a statement about the evaluators’ effort budget.
- Access level determines what can be concluded. Black box API access, access with sampling parameters and logits, access to a model without safety filters, access to weights, and access to training data support progressively stronger claims. A report that does not state its access level cannot be interpreted.
- The publication conflict. Publishing detailed findings is what makes the work checkable, and it is also what makes future access harder to negotiate and can itself be an uplift hazard. Every such body faces this trade and resolves it by withholding something.
- The liability halo. If a government body evaluates a model and does not object, that fact will be used — in marketing, in court, in a hearing — as an endorsement it was never meant to be. Bodies of this kind disclaim it explicitly, and the disclaimer does not travel with the headline.
How to judge one
Four questions, all answerable from public documents, and much more informative than the announcement: Is access statutory or voluntary? Are the evaluation methods published in enough detail to be reproduced? Are results published, and if not, to whom are they given? And is there a mechanism by which a finding causes something to happen — a referral, a procurement consequence, a mandatory disclosure — or does it stop at the report?
Mandates, names and reporting lines in this area change; check the current remit of whichever body is relevant to you rather than relying on any description written down, including this one.