Skip to main content
JOJonas Osman
August 5, 2026 · 7 min read

Model Risk & Independent Validation

By

Independent model validation is often reduced to a compliance exercise. That is a mistake — a well-run challenge function is one of the few controls that catches model failures before the market or the regulator does.

By Jonas Osman Abdelghafour, actuary and quantitative risk expert.

Model risk is the risk that a model produces the wrong number and the firm acts on it. Every material decision at a modern bank, insurer, or pension scheme is model-driven, which means model risk is one of the two or three largest operational risks these firms carry. And yet, in most institutions, the second-line function that manages it is under-resourced, under-empowered, and structurally set up to lose arguments with the first-line owners of the models it is meant to challenge.

This note is about what makes an independent validation function actually work, drawing on SR 11-7 (US), TRIM (ECB), and SS1/23 (PRA) — the three frameworks that between them define the current global standard.

The frameworks agree on more than they disagree

The three frameworks were written for different regulators, different institution types, and different eras — SR 11-7 in 2011, TRIM guidance from 2017 onward, SS1/23 for UK banks in 2023. On the substantive expectations they converge almost completely.

Effective challenge. Validation is not review; it is challenge. The reviewer's job is to attempt to invalidate the model on its own terms — its assumptions, calibration, implementation, and use.

Independence. Validators are structurally independent of model developers. Reporting lines, incentives, and access to information must all support that independence.

Materiality-based scope. Validation effort is proportional to model materiality. The most important models get the deepest, most frequent validation.

Ongoing, not point-in-time. Validation is a cycle, not an event. Every material model has a defined re-validation cadence and specific triggers for out-of-cycle review.

Documented findings, tracked remediation. Every validation produces findings with severity, owner, and deadline. Findings are tracked to closure and reported to the board.

The remediation and tracking mechanics are set out in the model risk remediation note. The wider governance context sits in the three-lines-of-defence piece.

Where validation actually adds value

The value of independent validation is not that it catches technical errors — those are the responsibility of the first line's own quality assurance. The value is that it catches the second-order failures the first line is structurally unable to see.

Assumption drift. A model built five years ago on assumptions that were reasonable at the time may still be running on those assumptions today. The first line does not have the incentive to systematically re-examine assumptions that inconvenience the business. The validator does.

Use-versus-design mismatch. A model designed for one purpose is often quietly repurposed for another where it is no longer valid. Pricing models used for reserving. Reserving models used for capital. Stress test models used for BAU decisioning. The validator's inventory-driven view catches these transitions.

Calibration data quality. First-line teams are often too close to their own data to see systemic quality problems. The validator, with cross-model perspective, spots patterns — proxy data used past its intended horizon, calibration samples that no longer represent the current portfolio — that the first line has stopped noticing.

Aggregation and interaction risk. No single model owner is responsible for what happens when their model output feeds another model. The validator is one of the few functions with the mandate to trace those chains and check whether the aggregate produces a sensible result.

What makes challenge effective

Independence is a necessary but not sufficient condition. Several other design choices determine whether an independent validation function actually produces effective challenge or merely produces reports.

Technical depth. The validator must be at least as sophisticated as the developer. A validation function that cannot reproduce a model from scratch on demand — using the same data, the same specification, and independent code — cannot substantively challenge it. This is the single most common failure mode of understaffed validation functions.

Direct access to data and code. Validators need read access to the data, the source code, and the running production system. Requesting extracts through the first line introduces filters and delays that undermine challenge.

Escalation authority. Findings that the first line disputes must have a defined escalation path — through the model risk committee, to the CRO, to the board risk committee — that does not require the CRO's discretion at every step. Where the first line can simply outlast the reviewer, challenge is not effective.

Independent modelling capability. The strongest validation functions maintain their own benchmark models — often simpler and more transparent than the production model — that they can run against the same data. Divergence between the two is a diagnostic that focuses the challenge.

Cross-model perspective. A validator who reviews only one model type loses the pattern-recognition advantage that comes from seeing across the estate. Validation teams should rotate reviewers across model classes on a defined cycle.

Common failure modes

Three failure modes recur in institutions I review.

Validation as documentation review. The validator receives the model documentation, reviews it against a checklist, and produces a report on documentation adequacy. This is not validation. It is a compliance function that will not catch a model whose documentation is polished and whose model is wrong.

Findings without teeth. Findings are raised, categorised, and tabled at a committee that has no authority to compel remediation. Deadlines are extended without pushback. Aged findings accumulate. The result is a validation function that produces work-product no one reads and no one acts on.

Reviewer capture. The validator spends enough time embedded with the first line that they come to see the first line's problems as their own. Rotation, structural distance, and periodic external review are the counterweights.

Under-scoping AI components. Machine-learning components inside a pricing or credit-scoring chain are sometimes treated as sub-models of the main GLM and not validated in their own right. This is a growing gap the regulators — both financial and AI-specific — are actively closing. See the predictive modelling note for the technical dimension and the AI and quantitative risk service for how I structure validation of hybrid models.

The validation report the board should read

A one-page summary the board can act on is worth more than a hundred pages of technical annex. The pattern I use has five sections.

  1. Purpose and materiality. What the model does, what decisions it drives, and the exposure at risk if it is wrong.
  2. Findings summary. Number of findings by severity, aged findings, and any finding that materially affects the SCR, ECL, or pricing decisions.
  3. Model performance. Backtest results against realised outcomes, benchmark comparison, and calibration diagnostics — in that order, in plain language.
  4. Emerging risks. Assumption drift, data drift, external environment changes that could invalidate the model within the next validation cycle.
  5. Recommendation. Fit for continued use, use with conditions, or withdraw. This is the sentence the board needs and it is the sentence that is most often absent.

Why this matters now

Model use is expanding into more of the firm's decisions, model complexity is increasing with machine learning and AI, and regulator patience for weak validation is decreasing. The banks that came through the 2008 crisis with functioning model risk functions were, on average, in materially better shape than the ones that did not. The insurers and asset managers that come through the next cycle with functioning validation of their AI-augmented pricing and reserving models will be in better shape than the ones that treat validation as a checkbox.

Independent challenge is one of the cheapest controls available. It requires a small number of skilled people, direct access, and organisational backing. The institutions that resource it properly are the ones that catch model failures before the market or the regulator does. The ones that under-resource it are the ones that read about their own model failures in the financial press.

Closing

Model validation is not glamorous work. It rarely produces news that anyone outside the firm sees. But it is one of the few risk-management controls that reliably prevents specific, material, avoidable failures. If your validation function is under-resourced, structurally weak, or producing findings that never close, the fix is neither expensive nor difficult — but it does require senior backing.

For structured validation reviews, target-operating-model redesigns, or independent second-opinions on individual models, see the model validation service line or get in touch.