AI can help estimate credit risk, organize application information or support an underwriter, but adding a model does not transfer accountability away from the creditor. A bank needs evidence that the system is fit for its defined use, produces reliable outcomes and can support the decisions and explanations required for the product and applicants it affects.
Testing starts with the decision the system is allowed to influence
The bank identifies whether the system recommends, ranks, prices, approves, declines or merely prepares information for a person. It documents the product, applicants, decision thresholds, data, users, owner and conditions in which the output must not be used.
Materiality depends on both the model and its use. A modest error in a directly automated approval process can affect many applicants, while a sophisticated tool used only for research may create a different exposure. Controls and validation are scaled to that actual role rather than to the technology label alone.
Data are tested for fitness, lineage and prohibited use
Reviewers trace each material input to its source and test accuracy, completeness, timeliness, representativeness and stability. They examine missing values, derived variables and vendor data so a convenient field does not silently stand in for information that is unreliable, unavailable to some applicants or impermissible for the decision.
A variable correlated with a protected characteristic is not automatically proof of unlawful treatment, but intended proxies and differences in how applicants are treated require careful legal and risk review. The bank applies current requirements to the complete decision process rather than assuming that removing an obvious field settles every fairness question.
Performance is measured across outcomes and relevant groups
Development and independent validation test discrimination, calibration, stability and error rates using data separate from model training where appropriate. Reviewers compare results with a reasonable benchmark and examine whether performance changes across products, channels, time periods and relevant applicant segments.
Aggregate accuracy can hide costly errors. Testing considers false approvals, false declines, uncertainty near a cutoff and the financial and customer effects of each error type, then connects performance measures to documented limits and escalation thresholds.
The bank must be able to act on and explain the output
Testing confirms that the model’s output can be translated into the actual principal reasons for an adverse action when Regulation B or another applicable requirement calls for them. Generic statements about a model, credit score or internal policy do not substitute for the specific reasons that drove the creditor’s decision.
A human reviewer can challenge unusual or unsupported output, but human involvement is not a universal cure. Reviewers need appropriate information, time, authority and consistent standards, and overrides are monitored so discretion does not introduce unexplained or uneven treatment.
Monitoring tests whether the approved assumptions still hold
After deployment, the bank monitors input shifts, model performance, approval and pricing patterns, overrides, complaints and downstream credit outcomes. Breaches of a threshold lead to investigation, restriction, recalibration or suspension rather than being treated as routine noise.
Changes to data, prompts, models, thresholds, products or vendors return through change control and proportionate validation. Version records preserve which system and rules affected an applicant so the bank can reproduce a decision, investigate an issue and retire the system without losing accountability.
Read the primary material
Banking Explained prioritizes regulators, official publications and first-party announcements.
