Data quality describes the fitness of data for specific purposes, characterized by accuracy, completeness, consistency, and timeliness. The concept covers measurement methods, governance, data lineage and processes for improvement. It is vital for reliable analytics, operational processes and automated decision-making.
Use this profile to understand the building block briefly, place it in the model, and switch to the 360° assessment when needed.
Theoretical construct: explains a term, principle, or mental model.
What you need to understand to reason about a domain.
Data quality describes how well data serves a specific purpose. It is assessed through checkable characteristics such as accuracy, completeness, consistency, and timeliness.
Within data engineering, analytics, and data governance, data quality emerged from a practical problem: integrated datasets are available, but not automatically trustworthy or usable. Once data from multiple sources is combined for reporting, operational work, or automated decisions, organizations need shared rules, measurement, validation, cleansing, and clear ownership. ISO 8000 and similar practice-based approaches systematize that need.
Think of data quality as a feedback loop. Profiling measures the current state, rules define allowed values and relationships, validation checks the data against them, cleansing repairs detected defects, and monitoring shows whether quality drifts again later. Governance keeps definitions, responsibilities, and escalation paths stable; data lineage helps trace causes and impacts back to source.
Measurable properties such as accuracy, completeness, consistency, timeliness, and uniqueness make quality assessable.
Profiling provides statistical clues about distributions, outliers, null values, and format patterns.
Validation checks data against rules, schemas, or reference values.
Cleansing corrects, standardizes, or removes incorrect, duplicate, or incomplete records.
Governance defines roles, standards, and decisions for data responsibility.
Monitoring detects quality drift during operations or after data loading.
This matters before dashboards, migrations, data pipelines, ML models, and automated decisions. Better quality reduces wrong decisions and rework, but it depends on shared definitions, clear ownership, and maintained validation rules. More controls also add effort, latency, and maintenance cost; cleansing can hide source problems if it is not documented and traceable.
Where this building block is located in the topic model.
No structure path available.
Explore how this building block connects to concepts, methods, technologies, and tools.
These sources establish the term and its professional meaning.
All direct connections of the current building block in a compact text view.
This classification shows where the building block typically matters, how demanding it is, and what kind of impact it has in the model.
The level within the organization (enterprise, domain, team) at which the AssetBlock is applied.