Skip to main content
JOJonas Osman
· 3 min read

Effective Challenge in the Second Line

By , Actuary & Quantitative Risk Expert

Effective challenge is not a tone of voice or a meeting slot. It is the demonstrable ability to change a decision, and most frameworks cannot evidence it.

This article relates to my work on Model Validation & Model Risk, AI & Quantitative Risk Models and Climate & Catastrophe Risk.

By Jonas Osman Abdelghafour.

"Effective challenge" appears in almost every supervisory statement on governance and in almost every risk function's terms of reference. It is also one of the least evidenced concepts in financial services, because the artefacts a firm produces — attendance lists, review memos, committee minutes — demonstrate that challenge occurred, not that it was effective. The distinguishing test is simpler and harder: can the function point to decisions that changed because of it?

The four conditions

Competence. A reviewer who does not understand the technique cannot challenge it and will default to process questions. This is the most common cause of validation that reads as an audit of documentation completeness. It is also the most expensive condition to satisfy, and firms that under-invest here should not be surprised when the second line becomes a compliance checkpoint.

Independence with authority. Structural independence — reporting line, budget, performance assessment — is necessary but insufficient. The function must be able to impose conditions, restrict use, or block deployment. If its output is advisory and the business decides whether to accept it, the challenge is a suggestion.

Access. To data, code, documentation and the people who built the thing. Restricted access converts review into an interview, and vendor arrangements that deny it should trigger compensating testing rather than a lower standard.

Standing. Whether senior management treats the function's findings as decisions or as negotiating positions. This is cultural and it is visible in one metric: how often findings are downgraded or extended without new evidence.

Focus on materiality, not completeness

The second line's scarcest resource is expert time, and it is routinely spent on the wrong things. Re-performing calculations proves arithmetic. Checking that every documentation section exists proves formatting. The productive agenda is narrower: identify the two or three assumptions that determine the outcome, test how the result moves when they move, and interrogate whether the evidence supports them.

For most models, that means data representativeness, the behavioural or judgemental parameters that dominate the result, performance in the segments and conditions that matter commercially, and the honesty of the stated limitations. A review that lands three findings on those points is worth more than one that lands thirty on presentation.

Make challenge visible

Challenge that is not recorded did not happen, from a supervisory perspective. Practical mechanisms include a challenge log capturing the point raised, the response, and the resolution; recorded dissent where the second line's position was not adopted, with the accountable decision-maker named; findings with severity ratings, owners and dates that are tracked to closure; and conditions attached to approvals, so that use restrictions are enforceable rather than aspirational.

Metrics that indicate whether it is working

  • Proportion of reviews producing material findings — near zero suggests either an unusually good first line or an ineffective second one, and it is worth knowing which.
  • Age and extension rate of open findings; repeated extensions without new evidence is the clearest failure signal available.
  • Frequency of recorded dissent and how it was resolved.
  • Decisions demonstrably changed: approvals refused, conditions imposed, use restricted, deployments delayed.
  • Repeat findings across reviews, which indicate that remediation is cosmetic.

Common failure modes

Effective challenge collapses in predictable ways. It becomes a gate applied too late, after commercial commitments make rejection impractical. It is absorbed into the delivery timetable, so the reviewer's incentive aligns with the launch date. It is diluted by giving the second line ownership of first-line deliverables, destroying independence in exchange for short-term capacity. Or it is quietly de-scoped when resourcing is cut, on the reasoning that the first line has matured — a judgement that is only testable after something has gone wrong.

The functions that get this right tend to look similar: technically credible, engaged early, deliberately narrow in focus, unembarrassed about recording disagreement, and measured on outcomes rather than throughput. Everything else is documentation.

Primary sources: PRA SS1/23 on model risk management principles; EBA guidelines on internal governance; Basel Committee guidance on corporate governance principles for banks.