Modern Loss Reserving: Chain-Ladder to ML
By Jonas Osman Abdelghafour
The chain-ladder is not obsolete and machine learning is not a silver bullet. This is a practitioner's map of how modern reserving actually combines the two, and where each fails.
Written by Jonas Osman Abdelghafour, actuary and quantitative risk expert.
Loss reserving sits at the intersection of the most consequential number on a non-life insurer's balance sheet and the most opinionated corner of actuarial practice. The last decade has changed the toolkit substantially — but it has not, contrary to some conference marketing, made the classical methods obsolete. What has changed is where each method should sit in the reserving stack, and how much information a modern actuary can extract from data the traditional triangle throws away.
This note maps the current state of the practice as I use it with clients, from chain-ladder up through individual claim machine-learning models, with an emphasis on when each is defensible and when it is not.
Why the triangle still wins most of the time
The Mack chain-ladder and its cousins remain the workhorse of the industry for a reason. They require little data, they are transparent, they produce a mean and a variance under stated assumptions, and — most importantly — they aggregate the information the business actually has: paid and incurred amounts by development period. For lines with reasonably stable claim mix and payment patterns, they are hard to beat on out-of-sample performance.
The classical objections are well-rehearsed. Triangles throw away claim-level information. Mack's variance formula assumes calendar-year independence, which is almost never true when inflation is trending. Development factors implicitly assume the mix of claims settling in each period is stable. Every one of these objections is real. None of them justify discarding the method — they justify using it as a benchmark and diagnosing where its assumptions bite.
A serious reserving process I would sign off on always includes a triangle-based estimate. It is the number every reviewer, auditor, and regulator can reproduce, and it is the number against which every more sophisticated method must justify its existence.
Bornhuetter-Ferguson and where it earns its keep
Bornhuetter-Ferguson (BF) does one thing the chain-ladder cannot: it lets the actuary carry a prior view into recent, immature accident years where the triangle has almost no signal. On long-tail lines like general liability, where the first development period contains a small fraction of ultimate losses, BF stops the chain-ladder from over-reacting to a few large or small early claims.
Two failure modes are common. The first is a stale prior — the a priori loss ratio was set two soft cycles ago and nobody has updated it. The second is a mechanical application that gives BF weight even in mature years where the chain-ladder signal is strong. Cape Cod and the Benktander credibility-weighted variants address the second problem and are worth learning if you are still applying pure BF beyond the third or fourth development period.
For a broader treatment of where reserving methods break, see the loss reserving explained note. The related LGD and EAD validation pitfalls piece covers analogous calibration traps on the credit side.
Generalised linear models: still underused
Overdispersed Poisson and Gamma GLMs on incremental payments recover the chain-ladder as a special case and generalise it in useful directions. They allow calendar-year effects (inflation, legal reform, claim-handling changes) to be estimated jointly with development, they produce a full predictive distribution, and they let the actuary include exposure covariates the triangle cannot use.
The main practical reason GLMs are underused is not statistical — it is that most reserving software makes them awkward to fit, and the reserving actuary is often not the same person who fits GLMs for pricing. When the same team owns both, GLM reserving is straightforward and pays back the investment through calendar-year diagnostics that flag inflation and reform effects before the classical triangle does.
Individual claim reserving: the real shift
The most substantive change in the last decade is not machine learning per se — it is the move from triangle-level to claim-level reserving. Individual claim models estimate, for each open claim, the distribution of its future payments as a function of claim characteristics (line, cause, jurisdiction, severity indicators, time since occurrence, current reserve, adjuster tags) rather than averaging over the whole cohort.
The advantages are considerable. Portfolio mix changes are handled naturally because the model predicts each claim on its own features. Large-loss dynamics can be separated from attritional dynamics without post-hoc adjustments. Reserving strengthenings become auditable at the claim level rather than reconciled at the triangle level.
The disadvantages are also real. The models require substantially more data cleansing, they are harder to validate, and they can be opaque to reviewers who are fluent in Mack but not in gradient boosting. This is where the interpretability discipline matters — see the predictive modelling in insurance note for the broader picture and my predictive modelling and pricing service for how I structure it for clients.
Machine learning: where it actually helps
Machine learning in reserving delivers real value in three specific places, and mostly hype in others.
Claim-level payment prediction. Gradient-boosted trees and random forests fit claim-level payment distributions better than parametric GLMs when there are non-linear interactions between claim features. The uplift is modest at portfolio level and material for large-claim tails.
Case reserve adequacy. Supervised models that predict the ultimate payment on an open claim, using the current case reserve as one input, are a powerful diagnostic. Systematic deviations flag either reserving philosophy drift or a change in the claim mix that the case reserve process has not yet absorbed.
Development pattern estimation on new lines. Where a portfolio is too immature for a stable triangle, transfer-learning approaches that borrow strength from related lines outperform naïve BF priors, provided the analogue portfolio is genuinely comparable.
Where machine learning does not help — and where I have seen it damage reserving processes — is as a black-box replacement for the actuarial judgement layer. A model that produces a lower reserve than the triangle without a diagnosable reason should not be adopted, regardless of its cross-validation score. Reserving reviewers, auditors, and regulators will (rightly) reject it.
Model risk and validation
Reserving models are among the most material models an insurer runs, and the most litigated at audit time. Independent validation is not optional. The validation programme should benchmark every alternative method against the triangle, decompose the difference into portfolio-mix effects, calendar-year effects, and claim-severity effects, and require a specific narrative for any material divergence.
The framework I use is set out in the model validation service line, and the model risk remediation piece covers how findings should flow into board-visible plans.
A layered reserving stack
For most non-life portfolios, the reserving stack I would defend today looks like this:
- Triangle-based estimate (Mack or overdispersed Poisson GLM) as the anchor.
- BF or Cape Cod overlay for immature accident years.
- Individual claim model for portfolios where claim-level data quality supports it, used to cross-check the triangle and to allocate reserves at claim granularity for reinsurance and IFRS 17 purposes.
- Machine-learning diagnostics for case reserve adequacy and large-loss dynamics.
- An expert-judgement layer, documented, that reconciles the layers and takes the final signing decision.
None of the layers alone is sufficient. All of them together, applied honestly, are.
Closing
The next decade of reserving will not throw the chain-ladder out. It will keep the triangle as the transparent anchor, add claim-level machine-learning models where the data supports them, and make the judgement layer more explicit than the profession has traditionally been comfortable with. The methods have moved. The obligation to explain the number to a non-actuary board has not.
If you are reviewing your reserving process, planning an IFRS 17 build, or scoping an individual-claim reserving project, see insurance and actuarial services or contact me directly.