How data quality rules, issue detection, remediation, and measurement ensure the organisation's data is accurate, complete, consistent, and timely.
DQT-01
Definition and coverage of data quality rules
How well-defined are data quality rules (accuracy, completeness, consistency, timeliness, validity) for the datasets in your domain, and how much of your domain's data do these rules actually cover?
Maturity level descriptions
No data quality rules are defined for domain datasets; quality expectations are unstated and vary by individual judgement. Whether data "looks right" is judged informally, case by case, with no agreed definition of what quality means for any given dataset.
Informal quality expectations exist in people's heads (e.g. "this field shouldn't be blank"), but nothing is written down or agreed as a formal rule. Data Engineer can informally describe expected quality characteristics for a dataset but confirms nothing has been documented or formally agreed.
Documented data quality rules exist and are applied to at least the domain's priority datasets, covering core dimensions (completeness, accuracy, validity). Data Engineer maintains or applies a documented set of quality rules for priority datasets, covering at least the core quality dimensions relevant to that data.
Data quality rules cover the domain's full dataset inventory (not just priority datasets), are version-controlled, and are reviewed and updated on a defined cycle. Data Engineer ensures rule coverage extends across the domain, with rules reviewed periodically to confirm they still reflect actual business requirements.
Data quality rules are dynamically maintained and extended as new data enters the domain, informed by observed quality patterns and downstream usage needs, with minimal manual rule-writing effort. New datasets or fields automatically inherit or trigger review of applicable quality rules as part of onboarding, and rule sets evolve based on observed data behaviour rather than only manual definition.
DQT-02
Data quality monitoring and detection
How effectively are data quality issues in your domain detected — through deliberate monitoring and profiling, as opposed to being found by accident when a report looks wrong or a user complains?
Maturity level descriptions
Data quality issues are found by accident, usually when a report looks wrong or a downstream user raises a complaint; no deliberate monitoring occurs. Data Engineer becomes aware of quality problems only when someone else notices something is off, with no proactive checking in place.
Spot checks happen occasionally, done manually and inconsistently, with no defined schedule or consistent rule set applied. Data Engineer occasionally runs manual checks on a dataset (e.g. eyeballing a sample, running an ad hoc query) but not on any regular schedule.
Defined quality checks or profiling routines exist for priority datasets, run on a manual or semi-manual schedule (e.g. before major reporting cycles). Data Engineer runs or reviews the results of a documented quality-check routine for priority datasets, ahead of key business events like reporting deadlines.
Quality metrics are tracked on a dashboard, monitored against defined thresholds, and reviewed on a regular (e.g. weekly/monthly) basis across the domain's full dataset inventory. Data Engineer reviews a quality dashboard covering the domain on a regular cadence, with defined thresholds indicating when a metric requires attention.
Data quality is monitored continuously and automatically, with anomalies flagged in near-real-time as data is created or changes, before it reaches Data Consumers. Automated quality monitoring runs continuously against incoming or changing data, generating alerts that reach the Data Engineer before affected data is consumed downstream.
DQT-03
Data quality issue remediation
Once a data quality issue is identified in your domain, how effective, consistent, and timely is the process for actually fixing it?
Maturity level descriptions
Quality issues, once found, are fixed inconsistently or not at all; there is no defined process, and resolution depends entirely on who noticed and how much they care. Fixes happen ad hoc, if at all, with no consistent method, no tracking, and no guarantee the same issue won't simply recur unaddressed.
Issues are fixed manually on a case-by-case basis using informal, undocumented methods (e.g. one-off scripts or manual edits), without a repeatable procedure. Data Engineer manually corrects quality issues as they arise, using whatever method seems appropriate at the time, without a documented or repeatable procedure.
A defined remediation procedure exists for common issue types (e.g. a documented process for correcting specific known data problems) and is generally followed. Data Engineer follows a documented, repeatable procedure (e.g. a standard correction script or checklist) for at least the domain's most common recurring issue types.
Quality issues are logged, tracked to resolution with defined response-time expectations, and remediation outcomes are verified (i.e. confirmed fixed, not just closed). Data Engineer's domain maintains a tracked issue log showing detection date, remediation action, resolution date, and verification that the fix actually resolved the problem.
Remediation is automated where possible (e.g. automated correction rules, self-healing pipelines), with manual intervention reserved for genuinely novel issues, and fixes are verified automatically. Automated remediation rules correct known, well-understood quality issues without manual intervention, with automated verification confirming the fix was effective.
DQT-04
Data quality metrics, thresholds, and reporting
How is data quality in your domain measured using defined metrics and thresholds, and how is that measurement reported to relevant stakeholders?
Maturity level descriptions
No data quality metrics exist for the domain; there is no quantified way to describe how good or bad the domain's data quality currently is. Quality is described only in subjective, qualitative terms ("it's mostly fine," "there are some issues") with no supporting numbers.
Some ad hoc measurement has occurred (e.g. a one-off count of missing values), but it isn't repeated, tracked over time, or reported anywhere. Data Engineer recalls a one-off measurement exercise that produced a quality snapshot, but nothing was done with it afterward.
Defined quality metrics exist for priority datasets, with agreed thresholds, and are reported to relevant stakeholders on a regular cycle. Data Engineer contributes to or reviews a scheduled quality report covering priority datasets, using a consistent, defined metric set and agreed thresholds.
Quality metrics cover the domain's full dataset inventory, are tracked over time (trend, not just snapshot), and are visible to governance stakeholders on a dashboard. Data Engineer's domain quality metrics are tracked period-over-period and visible via a dashboard, not just a static periodic report.
Quality metrics are monitored continuously in near-real-time, with automated alerts when metrics breach thresholds, and reporting requires no manual compilation effort. Data Engineer has access to a live quality metrics view for their domain and receives automated alerts when a metric moves outside an acceptable range, without waiting for a scheduled report.
DQT-05
Root cause analysis and prevention of recurring issues
When a data quality issue occurs, how effectively does your domain investigate its root cause and take action to prevent it recurring, rather than simply fixing the symptom each time?
Maturity level descriptions
Quality issues are fixed at the symptom level every time they occur, with no investigation into why they keep happening; the same issues recur indefinitely. Data Engineer fixes the visible problem each time it appears without asking why it keeps happening or whether it could be prevented upstream.
Root cause is occasionally considered informally for particularly troublesome issues, but this is inconsistent and undocumented. Data Engineer has, at least once, informally wondered about or discussed why a particular issue keeps recurring, without a structured investigation.
A defined root cause analysis process exists and is applied to significant or frequently recurring quality issues, with findings documented. Data Engineer conducts or contributes to a documented root cause analysis (e.g. a "5 whys" exercise or equivalent) for at least one significant recurring issue.
Root cause findings are systematically translated into upstream preventive actions (e.g. source system fixes, validation rules added at entry point), tracked to completion. Data Engineer's domain has a tracked register linking root cause findings to specific preventive actions, with owners, timelines, and confirmation of implementation.
Root cause analysis and prevention are embedded as standard practice for all significant issues, with patterns across multiple issues analysed to proactively redesign processes or systems before problems recur elsewhere. Data Engineer's domain reviews patterns across multiple resolved issues to identify systemic causes, feeding proactive redesign of upstream processes rather than issue-by-issue fixes.
DQT-06
Tolerance for data imperfection in downstream processes, and how it is measured
For the datasets and processes in your domain, is there a defined understanding of how much data imperfection a downstream process can actually absorb before outcomes are materially affected — and is that tolerance measured, rather than assumed?
Maturity level descriptions
No concept of tolerance for imperfection exists; any known data imperfection is treated as equally concerning (or equally ignored) regardless of its actual effect on downstream outcomes. Data Engineer treats all identified data issues the same way, with no distinction between imperfections that meaningfully affect downstream decisions and those that don't.
An informal, anecdotal sense exists that "some imperfection is fine" for certain processes, but this has never been tested, measured, or documented — it is assumed rather than known. Data Engineer informally assumes a downstream process can tolerate a certain level of imperfection, based on experience rather than any measurement or test.
Tolerance for imperfection has been explicitly assessed for at least priority downstream processes, using a defined method (e.g. testing outcomes against deliberately degraded data samples), and documented. Data Engineer can point to a documented assessment showing the tested or estimated tolerance threshold for at least one priority downstream process in their domain.
Tolerance thresholds are defined and tracked across the domain's significant downstream processes, with data quality actively managed against these thresholds rather than against a single uniform quality bar. Data Engineer manages quality effort proportionally — prioritising fixes for imperfections that matter to a process's actual tolerance threshold, and consciously not over-investing in fixing imperfections that don't.
Tolerance for imperfection is quantitatively modelled — for example using sensitivity analysis or similar techniques to determine, per process, how much and what kind of data imperfection can be absorbed before outcomes degrade — and this modelling is kept current as processes and data change. Data Engineer's domain uses or has access to a quantitative model (e.g. a sensitivity or threshold-detection model) that specifies, per downstream process, the point at which data imperfection begins to measurably affect outcomes, updated as conditions change.
DQT-07
Data quality standards for AI/GenAI training, fine-tuning, and grounding data
How well-defined are the quality standards your domain's data must meet specifically to be used in AI/GenAI training, fine-tuning, or grounding (e.g. RAG source material) — covering criteria like representativeness, bias, freshness, and label accuracy — as distinct from general data quality rules?
Maturity level descriptions
No AI-specific quality standards exist; data judged "good enough" by general quality rules is used for AI/GenAI purposes with no additional criteria applied. Data that passes ordinary quality checks (completeness, accuracy) is assumed suitable for AI use, with no consideration of AI-specific concerns like representativeness or bias.
Awareness exists that AI use has different quality needs (e.g. concerns about bias or staleness), but no criteria have been defined or documented. Data Engineer or Data Practitioners have informally discussed AI-specific quality concerns for domain data but have not defined or written down any criteria.
Documented AI-specific quality criteria exist (e.g. representativeness checks, staleness thresholds, label accuracy requirements) and are applied to at least priority datasets used in AI initiatives. Data Engineer applies a documented AI-specific quality checklist or criteria set before domain data is approved for use in an AI initiative.
AI-specific quality standards are systematically applied and tracked across all domain datasets used in AI initiatives, with non-conformant data flagged and remediated before use. Data Engineer maintains a tracked register of AI-use datasets and their conformance to AI-specific quality standards, with remediation actions tracked for non-conformant data.
AI-specific quality standards are enforced automatically, with data failing representativeness, bias, freshness, or label-accuracy checks blocked from AI pipelines without manual review. Automated checks assess domain data against AI-specific quality criteria before it can enter a training, fine-tuning, or grounding pipeline, blocking or flagging non-conformant data automatically.
DQT-08
Quality gates and validation before data enters AI/GenAI pipelines
Is there a defined validation step or "quality gate" that domain data must pass through before it is used in an AI/GenAI training pipeline, fine-tuning process, or retrieval system — and how consistently is that gate actually applied?
Maturity level descriptions
No quality gate exists before data enters AI/GenAI pipelines; data flows directly from source to AI use with no validation checkpoint. Data moves straight from its source into an AI pipeline or tool with no check performed specifically at that handoff point.
A validation step is sometimes performed informally before AI use, depending on who is involved, but it is not a required or consistent step. Data Engineer or a Data Practitioner occasionally checks data informally before it's used in an AI pipeline, but this depends on individual initiative rather than a required step.
A defined validation step (checklist, sign-off, automated check) exists and is required before domain data enters an AI/GenAI pipeline for at least priority use cases. Data Engineer performs or confirms a documented validation step before priority domain data is released into an AI pipeline, with the check recorded.
The quality gate is consistently applied across all AI/GenAI use of domain data, with pass/fail outcomes tracked and failed data blocked pending remediation. Data Engineer's domain tracks quality-gate outcomes across all AI-related data releases, with failed validations blocking pipeline entry until resolved.
The quality gate is fully automated and embedded directly in the data pipeline, with data unable to reach AI/GenAI systems without passing validation, and gate performance continuously monitored. An automated gate is technically enforced within the pipeline architecture itself, making it impossible for unvalidated data to reach AI/GenAI systems, with gate pass/fail rates monitored over time.
DQT-09
Detection and feedback of AI-output quality issues traceable to source data
When an AI/GenAI system produces poor-quality, biased, or hallucinated output, how effectively can that issue be traced back to underlying data quality problems in your domain, and how is that insight fed back to improve the source data?
Maturity level descriptions
No mechanism exists to trace AI output problems back to source data quality; poor AI outputs are treated as a model/tooling issue only, with no consideration of the underlying data. When an AI system produces a bad or questionable output, the possibility that it stems from domain data quality is not investigated.
Data quality is occasionally suspected as a cause of AI output problems and looked into informally, but without a defined process or documented findings. Data Engineer or Data Practitioners have, at least once, informally suspected and looked into a data-quality cause for an AI output issue, without a defined process.
A defined process exists for investigating and documenting whether AI output issues trace back to source data quality, and Data Engineers are expected to use it when notified of a relevant issue. Data Engineer participates in a documented investigation process when an AI output issue is flagged as potentially data-related, recording findings.
AI output issues traced to source data quality are logged, tracked to remediation, and used to update the domain's data quality rules or AI-specific quality standards. Data Engineer's domain tracks confirmed data-quality-caused AI issues through to remediation, with resulting updates made to quality rules (DQ-1) or AI-specific standards (DQ-8) to prevent recurrence.
AI output monitoring is integrated with source data quality monitoring, so that emerging AI output degradation (e.g. drift) automatically triggers investigation of upstream domain data quality, closing the loop continuously. Automated monitoring of AI system outputs is linked to domain data quality monitoring, so that a detected pattern of output degradation automatically triggers a data-quality investigation without waiting for manual reports.