An AI result depends on the data that shaped it and the data supplied when it is used. Data lineage records where those inputs came from, how they changed and where they entered the model or business process.

01

Lineage connects a result to its source

A useful lineage record identifies source systems, owners, collection timing, transformations and movement between platforms. It helps a reviewer follow a data element from its original record through preparation, model use and the downstream decision or action.

This is more than a diagram of technical connections. The record should explain which version of a dataset, rule or feature was used so teams can reconstruct an outcome and determine whether the data were appropriate for the stated purpose.

02

Data quality can change model behavior

Missing values, stale records, inconsistent definitions and duplicated transactions can affect a model even when its code has not changed. Lineage helps teams locate where a quality problem entered the process and which outputs may have been affected.

Quality is contextual. Data that are accurate enough for a marketing summary may not be sufficiently complete, timely or representative for fraud detection, credit analysis or another higher-impact use.

03

Testing needs the same data path as production

Model development often uses a prepared historical dataset, while production draws data continuously from live systems. If the definitions, transformations or timing differ, test performance may not represent how the model behaves after launch.

Teams therefore compare development and production pipelines, test important transformations and monitor whether input patterns change. Lineage provides the map needed to investigate drift rather than treating every performance change as a problem inside the model alone.

04

Access, privacy and third parties remain visible

Lineage can show where sensitive information is copied, combined or shared, supporting access control, retention and privacy review. It also helps identify when a downstream model is using a field for a purpose different from the one under which it was collected or approved.

When a provider supplies data, a model or a processing platform, the bank still needs enough information to understand material inputs and limitations. Contract rights, documentation and monitoring should support that visibility without assuming that a vendor’s internal process is fully transparent.

05

Documentation must change with the system

A lineage map becomes unreliable when systems, field definitions or model versions change without updating it. Ownership, change controls and automated metadata can help keep the record aligned with the live process.

The level of detail should be proportionate to the use and potential harm. Higher-impact decisions generally require stronger traceability, more rigorous validation and clearer evidence that the data and model remain suitable over time.

Sources

Read the primary material

Banking Explained prioritizes regulators, official publications and first-party announcements.