Skip to main content
JOJonas Osman
· 4 min read

AI Model Risk Management in 2026

By , Actuary & Quantitative Risk Expert

AI adoption does not suspend model risk management. It raises the standard of evidence banks need before a model is allowed anywhere near a customer or a capital number.

This article relates to my work on Model Validation & Model Risk, AI & Quantitative Risk Models and Climate & Catastrophe Risk.

By Jonas Osman Abdelghafour.

The supervisory direction for 2026 is unambiguous: adopting artificial intelligence does not create an exemption from model risk management. It creates additional obligations. Where a traditional credit scorecard could be documented in thirty pages and validated against a stable population, an AI system may be retrained weekly, depend on a third-party foundation model, ingest unstructured text, and produce outputs that no reviewer can trace to a coefficient. None of that changes the underlying question a bank must be able to answer: who owns this model, what decision does it drive, how do we know it still works, and what happens when it stops.

Start with the inventory, not the algorithm

Most AI model risk failures in banks are inventory failures first. A tool built inside a business line — a document-summarisation assistant in credit, a triage classifier in financial crime, a pricing suggestion engine in commercial lending — is deployed as "productivity software" and never enters the model inventory. It is then never validated, never monitored, and never reviewed when the vendor changes the underlying model version.

The remedy is a materiality-based definition of a model that is written to capture judgement-influencing systems, not just quantitative estimators. If an output changes a lending, pricing, provisioning, capital, or customer-outcome decision, it is in scope, whether or not the business calls it a model. Tiering then does the work: the highest tier attracts full independent validation, the lowest attracts registration, an accountable owner, and periodic review.

Ownership and change management

AI systems break the annual validation cycle. A model whose weights, prompts, retrieval corpus, or vendor version can change between committee meetings needs change management that is continuous rather than calendar-driven. Three controls carry most of the weight:

  • Version pinning and release gates. No production change — including a vendor's silent model upgrade — without a recorded assessment of impact against the approved use.
  • Golden-set regression testing. A frozen, representative evaluation set run at every change, with pass thresholds agreed in advance by the second line.
  • Documented use restrictions. The approval states what the model may and may not be used for, and the control environment enforces it.

What independent validation should actually test

Effective challenge on an AI model is not re-performance. Reviewers add most value by attacking the assumptions the developers could not test themselves: whether the training population still resembles the live one, whether data lineage is defensible, whether performance holds across protected and commercially sensitive segments, whether the model degrades gracefully or catastrophically at the edge of its input distribution, and whether the human in the loop is genuinely able to override the output or merely rubber-stamps it.

Explainability deserves particular scepticism. Post-hoc attribution methods are useful diagnostics and weak evidence. A validation report that rests on a feature-importance chart, without a stability analysis of that chart, has not demonstrated that the model is understood.

Third-party and concentration risk

Very few banks train frontier models. Most consume them, which converts model risk into third-party and concentration risk. If several critical processes depend on one provider, the bank has created a single point of failure that its own model governance cannot see. Contractual rights to evidence, notice of material model changes, exit and substitution plans, and a documented view of what happens if the service is withdrawn or degraded belong in the risk assessment alongside performance metrics.

Reporting that a board can use

Boards do not need model counts. They need to know which decisions are now materially AI-influenced, how much exposure sits behind them, where performance is drifting, which findings are open past their remediation date, and what management will do if a key model or provider fails. A page that answers those five questions is worth more than a fifty-page appendix.

Practical checklist

  • Define models by decision influence, not by technique, and tier by materiality.
  • Register every AI system with a named accountable owner before deployment.
  • Gate every change — including vendor upgrades — behind regression testing.
  • Validate assumptions, data lineage, segment performance and failure modes; treat explainability output as a diagnostic, not proof.
  • Map provider concentration and hold a workable substitution plan.
  • Report drift, open findings and management actions in decision terms.

The competitive advantage in 2026 is not model sophistication. It is the ability to demonstrate, quickly and credibly, that a sophisticated model is under control — because governed models are the ones regulators, auditors and capital providers will let a bank actually use.

Primary sources: PRA SS1/23 (model risk management principles for banks); FSB, Sound Practices for the Responsible Adoption of AI, June 2026; EBA Risk Assessment Report, June 2026.