Federated lineage and metadata intelligence: A framework for scalable data governance in distributed enterprise architectures
DOI:
https://doi.org/10.32996/jcsts.2026.8.10.2Keywords:
Data Lineage; Federated Governance; Metadata Intelligence; Openlineage; Data Mesh; Bcbs 239; Enterprise Data Governance; Regulatory Compliance; Metadata Quality Scoring; Distributed ArchitectureAbstract
Large enterprises with heterogeneous distributed data platforms cannot achieve end-to-end verifiable data lineage through a centralized metadata architecture, including Data Mesh and Lakehouse architectures with domain-based data ownership.Traditional centralized data lineage solutions lack the scalability to meet the requirements of distributed data platforms with siloed domains, heterogeneous execution engines, and decentralized architecture. Presented here is the Federated Lineage and Metadata Intelligence Framework (FLMIF): a five-layer architectural model to facilitate balancing between domain autonomy and organizational governance in large regulated organizations. FLMIF enables automated lineage capture from heterogeneous execution engines using the OpenLineage standard, ML-based metadata inference and enrichment, multi-dimensional lineage quality scoring, policy-driven federated governance without central bottlenecks, and continuous evidence automation for regulatory compliance. The framework is developed using a Design Science Research method and assessed with a scenario-based architecture evaluation against regulatory and architectural requirements (BCBS 239, GDPR, SOX, DORA, EU AI Act). Relative to centralized and manually maintained approaches, FLMIF is projected to materially improve lineage coverage, substantially reduce the manual effort of compliance-evidence collection, accelerate root-cause analysis, and deliver near-real-time metadata freshness. A proof-of-concept implementation on synthetic data confirms that the framework's capture, scoring, and root-cause mechanisms are computationally sound and scale to graphs of thousands of datasets. Quantitative projections, with the underlying assumptions stated, are provided in Section 7; empirical validation in production deployments is identified as the primary future work.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 https://creativecommons.org/licenses/by/4.0/

This work is licensed under a Creative Commons Attribution 4.0 International License.

Aims & scope
Call for Papers
Article Processing Charges
Publications Ethics
Google Scholar Citations
Recruitment