Skip to content

Model Risk Governance in Financial Services

11 min read · updated August 4, 2026

Financial services already had a model governance regime before generative AI arrived, and supervisors on both sides of the Atlantic have been consistent that it applies. The interesting question is not whether SR 11-7 covers a language model. It is which parts of it stop working when you cannot see inside the model, cannot reproduce its output, and did not build it.

Information, not legal advice. Reviewed 4 August 2026. SR 11-7 and SS1/23 are supervisory guidance rather than rules, and supervisory expectations for AI are developing faster than the guidance is being rewritten. Your own supervisor’s current expectations, expressed in feedback letters and thematic reviews, will matter more than any published document.

The two frameworks that matter

FrameworkDescription
SR 11-7 / OCC Bulletin 2011-12US Supervisory Guidance on Model Risk Management, issued jointly by the Federal Reserve and the OCC in April 2011, and adopted by the FDIC in 2017. Three pillars: robust model development, implementation and use; effective validation; and sound governance, policies and controls. Fifteen years old, still the reference, and it is guidance rather than regulation.
PRA Supervisory Statement SS1/23UK model risk management principles for banks, published May 2023 and effective from May 2024. Five principles: model identification and risk classification; governance; development, implementation and use; independent validation; and model risk mitigants. Explicitly covers models bought in and models using AI, which SR 11-7 predates.

The two overlap heavily and it is not accidental — SS1/23 was written with SR 11-7 in view. The material differences are that SS1/23 is newer, is explicit about vendor models and AI, and puts more weight on a firm-wide model inventory with tiering as principle one. If you are building a programme from scratch, SS1/23 is the better skeleton even outside the UK.

The definition problem

SR 11-7 defines a model as a quantitative method, system or approach applying statistical, economic, financial or mathematical theories, techniques and assumptions to process input data into quantitative estimates. Read literally, a language model summarising a credit file is not obviously within it: the output is text, not a quantitative estimate.

That reading is a trap. Two responses, and both matter.

First, the definition is broader in application than in wording. If the output informs a business decision, supervisors treat it as a model whatever the output type, and SS1/23 makes this explicit by defining a model in terms of input processing and output generation rather than in terms of quantitative estimates. Second, and more practically, the consequence of excluding it is worse than including it: a system outside the model inventory has no owner, no validation, no monitoring and no one accountable, which is precisely the state the framework exists to prevent.

The defensible position is a tiered inventory that includes AI systems explicitly, with tiering driven by the decision the output affects rather than by the technique. A model producing a regulatory capital number is tier one whatever it is built from. A model drafting an internal meeting summary is not, whatever it is built from.

Validating something you cannot inspect

Validation under both frameworks has three components, and generative systems break each in a different way.

Validation componentDescription
Conceptual soundnessAssess the quality of the design and construction, including the theory and the assumptions. For a frontier model whose training data and architecture are undisclosed, you cannot assess construction. What you can assess is the soundness of the task design: whether the model is being asked to do something a language model can do, on inputs it can see, with an output the process can check.
Ongoing monitoringConfirm the model is implemented appropriately and continues to perform as intended. Harder here because the artefact moves under you: a hosted model can be updated by its provider without notice. Monitoring has to include detecting that the model changed, not only that performance drifted.
Outcomes analysisCompare outputs to actual outcomes, including benchmarking. Hard where there is no single correct output. The workable substitute is a fixed evaluation set with human-adjudicated answers, versioned, run on every change — which is a validation artefact, not just an engineering one.

The practical shift is from validating the model to validating the system around it. Retrieval quality, prompt construction, output constraints, the human review step, the escalation path and the failure behaviour are all inspectable, testable and within your control. A validation report that says “the model is a black box” and stops is inadequate; one that documents the controls around the black box and evidences that they work is not.

Build the evaluation set as a governance artefact from the start: fixed inputs, adjudicated expected outputs, version-controlled, with results for each model version retained. That single artefact answers most of what a validator will ask, and building one is a smaller job than it sounds.

Third-party models and effective challenge

“Effective challenge” — critical analysis by objective, informed parties with the competence and the incentive to challenge — is the concept at the centre of SR 11-7. It assumes somebody can interrogate the model’s construction. With a third-party frontier model, nobody in your organisation can.

Both frameworks anticipated vendor models, if not this kind. SR 11-7 says that vendor products should be subject to the same validation expectations and that firms should obtain information about the components and testing, and should perform their own outcomes analysis on their own data. SS1/23 addresses vendor models directly. Layered on top, the US interagency third-party risk management guidance issued in June 2023 and the EBA outsourcing guidelines in Europe govern the relationship itself.

What that means concretely when the vendor will not disclose:

  • Shift challenge to the use case. Challenge whether the task is appropriate, whether the inputs are complete, whether the output is used within its limits, and whether the controls catch the failure modes. Document the challenge and the response.
  • Get contractual rights to information: notice before a model version changes, the right to pin a version, evaluation results, and audit or attestation rights. These are negotiable and they are the raw material of validation. See the model change clause.
  • Run your own outcomes analysis on your own data. Not optional, and it is the part no vendor can do for you.
  • Have a substitution plan. Concentration risk in a single model provider is an operational resilience question, and in the EU it is a DORA question with a register entry attached.

The EU layer: AI Act and DORA

European financial institutions carry two further regimes.

Under the AI Act, Annex III lists creditworthiness evaluation and credit scoring, and risk assessment and pricing in life and health insurance, as high-risk uses. Fraud detection is excluded from the creditworthiness entry. For those systems the full high-risk stack applies, and deployers in those categories are among the few required to carry out a fundamental rights impact assessment under Article 27 before first use. The Act also allows a regulated financial institution to discharge some of its obligations through its existing internal governance arrangements under the banking directives rather than building a parallel structure — which is the correct instinct anyway, because a separate AI governance stack that does not talk to model risk management produces two inventories that disagree.

DORA, Regulation (EU) 2022/2554, has applied since January 2025 and is about ICT risk rather than model risk, but it reaches AI vendors squarely: a register of information on all contractual arrangements for ICT services, contractual requirements including audit and access rights and exit strategies, incident reporting, and an oversight regime for ICT providers designated as critical. A hosted model provider is an ICT third-party service provider. If it is not in your DORA register, that is a finding independent of anything to do with the model.

What supervisors are actually asking

Both the UK and US supervisors have signalled that they are applying existing frameworks rather than writing new ones, and the FCA in particular has stated that the Senior Managers and Certification Regime, the Consumer Duty and operational resilience rules already cover the ground. The Bank of England and the FCA run a periodic joint survey of AI use in UK financial services; the 2024 edition reported that a large majority of responding firms were using AI in some form, which is the number worth quoting when someone claims adoption is marginal.

The questions that recur in supervisory engagement, in rough order of how often they come up:

  1. Is it in the model inventory, and who owns it? A named individual, not a team. Under the SM&CR the accountability has to land on a person.
  2. What tier, and why? The rationale matters more than the tier.
  3. What did independent validation conclude, and who did it? Independent of the developer, with the competence to challenge.
  4. How would you know it had degraded? Name the metric, the threshold and the alert. “We monitor” is not an answer.
  5. What happens when it fails? The fallback, who invokes it, how quickly, and when it was last exercised.
  6. Where is the customer outcome? Under the Consumer Duty, a model producing worse outcomes for a group of customers is a Duty issue whether or not it is a discrimination issue.
  7. Can you explain an individual decision? Required by fair lending rules in the US and by data protection law in the EU and UK. See the explanation obligations.