TP Tool

Actuarial software for reserving: twelve questions to ask before you buy

Twelve questions to put to any actuarial software vendor before you buy reserving software: data intake, methods, audit trail, output to S.17.01 and S.19.01.

In this article

Actuarial software for reserving is bought about once a decade and questioned at every audit in between, so the purchase deserves a better test than a vendor demonstration. This article gives a chief actuary or a CFO twelve questions to put to any vendor of reserving software, with the reason each one matters and what a good answer sounds like. They cover data intake, risk groups, methods and diagnostics, the step from an undiscounted ultimate to a Solvency II best estimate, output to templates S.17.01 and S.19.01, and the audit trail. A test plan for a parallel run follows, then a short section on the ways spreadsheets fail. These are the questions we answer on a demo of TP Tool, the reserving software behind this article, and they apply to any product you compare it with.

Questions 1 to 3: what goes in

1. Does the software take claim and premium transactions, or does it expect finished triangles?

Why it matters: a product that starts from triangles leaves the aggregation step where it is today, usually in a workbook that sums payments by accident year and development year. That step is where grouping errors and late bookings hide, and an auditor cannot see it. We have not yet met a workbook aggregation that a reviewer could follow without its author in the room. A good answer: the product loads transactions at the level the claims and policy systems hold them and builds the triangles itself, so any triangle cell can be opened to show the transactions behind it. The IBNR calculation rests on few inputs, and the fewer hands they pass through, the better.

2. How are homogeneous risk groups defined, and how hard is it to change them?

Why it matters: reserving is done per homogeneous risk group, and the grouping decides how credible each triangle is. Product, cover, territory and claim type are the usual dimensions, and the right cut changes with the business. A good answer: groups are defined from the attributes on the transactions, and a new grouping can be added without reloading the data. Regrouping that needs a change request to the vendor is a warning sign.

3. What happens when the data is incomplete or inconsistent?

Why it matters: Article 48 of Directive 2009/138/EC makes the actuarial function responsible for assessing the sufficiency and quality of the data used in the technical provisions, so the checks have to run before the calculation, not after the auditor asks. A good answer: completeness and consistency checks run on load, with a report of what failed (missing attributes, payments on unknown claims, transaction dates outside the period) and a record of what was done about each item.

Questions 4 to 7: methods and judgement

4. Which methods are available, and can they run side by side on the same triangle?

Why it matters: no single method suits every risk group or every accident year. Chain-Ladder projects the triangle’s own pattern, Bornhuetter-Ferguson blends that pattern with an expected loss ratio, Cape Cod estimates the loss ratio from the triangle, and Mack’s method puts a standard error on the Chain-Ladder reserve. Friedland’s CAS text, Estimating Unpaid Claims Using Basic Techniques, and the IFoA Claims Reserving Manual describe the first three; Mack’s 1993 paper in the ASTIN Bulletin sets out the standard error. A good answer: at minimum Chain-Ladder, Bornhuetter-Ferguson and a loss ratio method, run together for each group with the results next to each other and the method not selected still visible. Our comparison of Chain-Ladder and Bornhuetter-Ferguson on one triangle shows why the gap between them is the number a reviewer wants to see.

5. What diagnostics does it show before the factors are accepted?

Why it matters: a volume weighted average factor hides the individual age to age ratios it came from. A calendar year effect or a change in case reserving practice shows in those ratios long before it shows in the average. A good answer: the individual ratios by accident year, a choice of averaging periods, the exclusion of individual cells with a recorded reason, and an actual versus expected comparison against the previous valuation.

6. How are tail factors fitted?

Why it matters: on a long tail liability group a large part of the reserve sits beyond the last observed development year, and the tail factor is a judgement rather than an observation. A good answer: curve fits to the observed factors, a manually entered tail with a stored rationale, and the tail shown separately so that its contribution to the reserve is visible.

7. How are large losses handled?

Why it matters: one large claim in a young accident year is multiplied by every remaining factor under Chain-Ladder, and once it sits in the triangle it distorts the factors for later years. The usual practice is to remove claims above a threshold, reserve them separately and run the methods on the attritional remainder. A good answer: a threshold per risk group, the removed claims listed with their own reserves, and the total reconciled back to the full triangle.

Questions 8 and 9: from ultimate to the Solvency II templates

8. Does it produce the Solvency II best estimate, or only the undiscounted ultimate?

Why it matters: the ultimate from a triangle method is a management figure. Article 36 of Delegated Regulation (EU) 2015/35 requires the non-life best estimate to be calculated separately for the premium provision and for the provision for claims outstanding, and Article 77 of the Directive defines the best estimate as a probability-weighted, discounted projection of all the cash flows needed to settle the obligations, which brings in claims handling expenses. A good answer: a payment pattern per risk group, discounting at the risk-free curve for the currency and reference date, inflation and expense assumptions in the same run, and the premium provision handled in the product rather than in a second workbook. The net claims provision per line of business is also the volume measure for reserve risk in the SCR standard formula, so the same number has to agree in two places.

9. Does it populate S.17.01 and S.19.01 from the same run?

Why it matters: template S.17.01, Non-life Technical Provisions, carries the best estimate of the claims and premium provisions per line of business, and template S.19.01, Non-life insurance claims, carries the triangles themselves: gross claims paid, the undiscounted best estimate of the claims provisions and the RBNS, by accident or underwriting year, for up to 15 years plus prior. EIOPA’s validation rules cross-check the two, and a difference between them means two data sets were used. A good answer: both templates, and the other technical provisions templates in the Solvency II QRT list, filled from the same run, with the mapping from risk group to line of business stored and open to review.

Questions 10 to 12: control and effort

10. Can a reviewer trace a reported figure back to a transaction, and rerun the calculation?

Why it matters: the actuarial function report and the external audit ask the same question in different words: where did this number come from. A good answer: every figure on a template opens to the risk group, the method, the selected factors, the triangle and the transactions behind it; every run is stored with its inputs and assumptions; and a stored run reproduces the same output when rerun.

11. Are runs versioned, and does a second person review a run before it becomes the reported one?

Why it matters: reserving is iterative. Factors change, a tail is revisited, and by the close nobody is quite sure which version was signed. We have seen this argued over in the week after a filing, and it is never a good week. A good answer: each run is a version with an author, a timestamp and a change note; a second person approves the version that goes into the templates; approved versions are locked.

12. Who can change what, and what does implementation take?

Why it matters: governance is about access as much as about method, and the effort of going live is underestimated on both sides. A good answer on access: named users with roles that separate loading data, running methods and approving results, an access log, and several legal entities under one account where a group or a consultant reserves for more than one company. A good answer on effort: the data fields required, the days from first load to first parallel run, and which of those days are yours.

A test plan: run one quarter in parallel

A demonstration shows what the vendor chose to show. A parallel run shows how the product behaves on your data. The plan below fits inside one quarter close and answers most of the twelve questions with evidence.

Step What to do What it tests
1 Extract claim and premium transactions for the last closed quarter, as the current process starts. Question 1: whether the loader accepts your data as it is.
2 Load them and keep the data quality report. Question 3: what the checks catch and how failures are reported.
3 Rebuild your current risk groups, then add one alternative grouping. Question 2: how much effort a change costs.
4 Compare the generated triangles with your workbook triangles cell by cell. Whether the aggregation matches; each difference is a data or workbook issue.
5 Run Chain-Ladder and Bornhuetter-Ferguson on every group with the workbook’s factors and tail. Questions 4 to 7: whether the methods reproduce the signed reserves.
6 Apply your payment pattern, discount curve and expense assumptions. Question 8: whether the best estimate matches the reported one.
7 Generate S.17.01 and S.19.01 and compare them with what was filed. Question 9: template output from the same run.
8 Pick three figures from the templates and trace each back to a transaction. Question 10: the audit trail.
9 Change one factor, save a new version and have a second person approve it. Question 11: versioning and review.
10 Count the hours spent by your team and by the vendor. Question 12: the real implementation effort.

Steps 4 and 5 rarely match at the first attempt, and that is fine. What matters is whether each difference can be explained from the trace in step 8, and whether the explanation points at the product or at the workbook.

Where spreadsheets fail

Workbooks are not wrong as such; many defensible reserves have been signed out of one. They fail in specific ways, and the twelve questions are aimed at those ways.

A triangle in a workbook is a cell range. When a new diagonal is added, every formula that refers to the old range has to be extended, and one that is missed keeps working on stale data without an error. A factor overridden by typing a number over a formula looks like every other cell. A link to the claims extract breaks when the extract is renamed, and the workbook keeps the last values it fetched. Two copies circulate during the close, and the one that is filed is not always the one that was reviewed. The person who built the workbook knows which cells are safe to touch, and that knowledge leaves with them. A workbook shows its current state, not who changed what and when.

None of these is a calculation error. They are control failures, which is why they surface in an audit rather than in the reserving result, and why the questions above spend as much time on the audit trail as on methods.

Where this lands in the software

TP Tool answers the first three questions at the loader. Claim and premium transactions are loaded as they are, millions of records at a time, checked for completeness and inconsistencies, and grouped into homogeneous risk groups in the Risk Group Designer, where the grouping can be adjusted as the portfolio changes. The Statistical Engine builds the run-off triangles for each group and runs Chain-Ladder, Bornhuetter-Ferguson and fixed loss ratio estimates side by side, with a choice of development factor selections and diagnostics that show how each method reacts to the data. Discounting, inflation and claims handling expenses are applied in the same run, reserve uncertainty is assessed next to the central estimate, and the results populate S.05.01, S.17.01, S.18.01, S.19.01, S.20.01, S.21.01, S.28.01, S.29.02 and S.29.03 with ECB currency conversion as at the reference date. The audit trail from transaction to reported figure stays unbroken.

Sources

  1. Friedland, Estimating Unpaid Claims Using Basic TechniquesCasualty Actuarial Society
  2. Claims Reserving Manual, volume 1Institute and Faculty of Actuaries
  3. Mack (1993), Distribution-free calculation of the standard error of chain ladder reserve estimates, ASTIN Bulletin 23(2)Casualty Actuarial Society
  4. Guidelines on valuation of technical provisionsEIOPA
  5. Solvency II Delegated Regulation (EN)SolvencyTool regulation library
  6. Solvency II Directive (EN)SolvencyTool regulation library
  7. S.17.01: Non-life Technical (Solo)SolvencyTool regulation library
  8. S.19.01: Non-life insurance claims (Solo)SolvencyTool regulation library

Frequently asked questions about actuarial reserving software

What does actuarial reserving software do?
It takes a non-life insurer’s claim and premium data, builds development triangles per homogeneous risk group, applies reserving methods such as Chain-Ladder and Bornhuetter-Ferguson, and turns the resulting ultimates into a claims provision and a premium provision. Under Solvency II it also discounts the cash flows, adds claims handling expenses and fills the technical provisions templates, so the same run produces the management reserve and the reported best estimate.
Can actuarial reserving software replace spreadsheets?
It can replace the parts of the process that spreadsheets do badly: aggregating transactions into triangles, keeping every version of a run, tracing a reported figure back to its inputs and separating who loads data from who approves results. The actuarial judgement stays with the actuary. Most teams keep a workbook for ad hoc analysis and move the production run into the software once a parallel quarter has reconciled.
Which reserving methods should actuarial software support?
At least Chain-Ladder, Bornhuetter-Ferguson and a loss ratio method, because no single method suits every risk group or every accident year. Cape Cod is useful where an a priori loss ratio is hard to defend, and a Mack standard error turns a point estimate into a range. More important than the length of the list is that the methods run side by side on the same triangle, with the age-to-age factors and diagnostics visible before a selection is made.
How does reserving software connect to the Solvency II templates?
The reserving run produces an undiscounted ultimate per risk group and accident year. The software maps risk groups to Solvency II lines of business, projects the payment pattern, discounts at the risk-free curve and adds expenses to arrive at the best estimate. That figure fills the claims provision rows of template S.17.01, and the triangles by accident or underwriting year fill template S.19.01, so both templates come from one data set and the validation rules between them hold.
What should the audit trail in reserving software show?
For any figure on a reported template it should show the risk group, the method, the selected development factors and tail, the triangle cell and the transactions behind that cell. It should also show who ran and who approved each version, what changed between versions and why, and it should allow a stored run to be reproduced with the same inputs. That is the evidence an actuarial function report and an external audit both ask for.
TP Tool

Ask us the twelve questions on a TP Tool demo

TP Tool: actuarial reserving software for non-life insurers. IBNR, premium provisions, Chain-Ladder, Bornhuetter-Ferguson, Solvency II QRT.