Model validation in banking

A validation unit answers one question about every model it reviews: would the bank make the same decision if it knew what the review knows. The work has a fixed order, from the data the model was built on to the report that goes to the management body. The duty behind it comes from MaRisk, and for banks under European Central Bank supervision from the ECB guide to internal models.

Calibration curves and residual plots overlay a bank risk model's validation data.

Where validation sits in the three lines of defense

Validation belongs to the second line. The first line builds and uses the model, the second line challenges it, and internal audit in the third line checks that the challenge happened properly. MaRisk AT 4.4 sets out the independent risk controlling function and the compliance function in the second line, and it keeps the review of the risk quantification procedures away from the people who developed them.

Independence is organizational and also about data. A validator who receives only the aggregated model output cannot test the model, so the unit needs access to the training data, the code and the parameter history. Where a bank buys a model, the provider's documentation takes the place of the code, and the model risk framework says what level of detail is enough.

The three things a validation report answers

Conceptual soundness comes first: is the method appropriate for the question, were the assumptions stated, and do the variables have a reason to be in the model beyond fitting the sample. A model may pass every statistical test and still be unsound because it uses a variable the bank cannot observe at decision time.

Outcome analysis compares what the model predicted with what happened. Benchmarking puts the model against an alternative, a simpler specification or an external estimate, so that a poor result has a reference point. The ECB guide to internal models organizes its expectations along these lines for the models a bank uses to calculate capital.

Backtesting and what happens when a model breaches its tolerance

Backtesting counts how often reality fell outside what the model allowed. For a market risk model this is the number of days the loss exceeded the value-at-risk estimate, and the count has a threshold written into the capital rules. A breach is not automatically a broken model, because a one-in-a-hundred event happens roughly once in a hundred days, so the test asks whether the exceptions cluster.

When the count passes the threshold, the model keeps running under a documented decision and a remediation plan with an owner and a date. The validation report states the finding, its severity and the deadline, and the unresolved findings are the list the supervisor reads first at the next review.

Ongoing monitoring and how drift is detected

Validation at a point in time and monitoring over time answer different questions. Monitoring watches the population: the applicants scored this quarter are not the ones the model was built on, and the gap grows quietly. The standard measures are the population stability index, which compares the current score distribution against the development one, and a Kolmogorov-Smirnov test for the same shift; the Jensen-Shannon and Kullback-Leibler divergences and the Wasserstein distance do the same job with different sensitivity to the tails.

Concept drift is the harder case, because the inputs look unchanged while the relationship between input and outcome has moved. It shows up as performance decay on a sliding window, so the monitoring set reports accuracy, precision, recall and the area under the curve alongside calibration, not instead of it. A validator who receives only an unchanged input distribution has been told nothing about concept drift.

Explaining a model the bank has to defend

Where a model refuses an application, someone has to be able to say why. The techniques split into two kinds. Globally interpretable methods constrain the model so the explanation is part of it, such as a generalized additive model or a boosted linear tree. Post-hoc methods explain a model that is already built: SHAP attributes a single decision to its inputs, LIME approximates the model locally, and partial dependence and accumulated local effects plots show how the output moves with one variable across the population.

A post-hoc explanation is an approximation of the model and not the model, which matters when the explanation is the one shown to a customer. German law already requires the substance: the Federal Court of Justice and the Court of Justice of the European Union have both addressed SCHUFA scoring, and AI compliance in Germany covers where that leaves an automated decision.

Data quality as a finding of its own

A validation that cannot trace a figure back to its source has found a defect before it tests anything. BCBS 239, the Basel Committee's principles for effective risk data aggregation and risk reporting, states the duty: a bank has to be able to produce accurate risk data and to explain how it was aggregated, including in a stress situation.

In practice three defects recur. A field changes definition in a source system without the model owner hearing about it, a manual correction enters a spreadsheet between the source and the model, and a history is rebuilt after a migration with values that no longer match the originals. Each one belongs in the report as a data finding, separate from the model finding.

Validating a machine learning model with no closed-form explanation

A gradient-boosted model has no coefficient to read, so conceptual soundness is tested differently. The validator checks the feature set for variables the bank may not use, examines stability across time slices and population segments, and compares the model against a transparent benchmark so that the gain from the complexity is measurable.

BaFin's guidance on ICT risks in the use of AI systems from December 2025 scales the testing to how critical the system is and treats explainability as something the firm has to be able to deliver at the level its use requires. Where the model scores consumer credit, the EU AI Act adds its own documentation and oversight duties, which AI governance in banking sets out.

What does a model validation unit do?

A model validation unit reviews each model independently of the people who built it and writes a report the management body can act on. The review tests whether the method fits the question, compares predictions against outcomes, benchmarks the model against an alternative, checks the data lineage, and records findings with a severity and a deadline. Under MaRisk the review is organizationally separate from model development.

What is in a model validation report?

The scope and the model version under review, the data used and its limits, the conceptual assessment, the outcome analysis with its backtesting results, the benchmark comparison, a list of findings with severity and remediation dates, and a conclusion on whether the model may be used as it stands. A report that ends without a usable conclusion leaves the decision with no owner.

What triggers a revalidation outside the annual cycle?

A market break that moves the data outside the range the model was calibrated on, a source system that changes a field definition, a provider updating a purchased model, a material change to the portfolio the model is applied to, and a backtesting result past its threshold. A bank writes these triggers into its model policy so that the decision to revalidate does not depend on who notices.

Model validation and Finance Loop

Finance Loop is the meeting place for validation units, quants and internal auditors in German finance. Finance Loop events put the people who write validation reports in a room with the supervisors and model owners who read them.

Finance Loop is a professional network and has the goal of driving the adoption of emerging technologies in finance, such as AI, tokenization, stablecoins, and DeFi. Finance Loop helps its members build skills and personal networks in these fields: Investment & Digital Assets, Payments & Digital Money, Digital Infrastructure & Sovereignty, and Risk & Compliance.

Let's stay in touch

4,000+ members in finance and tech. Become a Network Member for free.

Get updates for free!

Exclusive event invitations, member perks and news from the network. Unsubscribe at any time.

By submitting you agree to the terms.