← Back to sections

Metadata

How business and technical metadata, definitions, classifications, and lineage are documented, standardised, and kept discoverable and accurate.

MDT-01

Technical metadata documentation

How complete and accurate is the technical metadata (schema definitions, data types, formats, constraints, source system) documented for datasets in your domain?

Maturity level descriptions
  1. No technical metadata is documented; understanding a dataset's structure requires inspecting it directly or asking whoever built it. Anyone needing to understand a dataset's structure must query it directly or track down the original builder, with nothing written down.
  2. Some technical metadata exists in informal or outdated form (old data dictionaries, comments in code), inconsistent and not reliably current. Data Engineer can locate fragments of technical documentation but confirms it is materially out of date or incomplete.
  3. Technical metadata is documented in a defined, centrally accessible location for at least priority datasets, covering schema, data types, and source system. Data Engineer maintains or has access to structured technical metadata for priority datasets in a shared, accessible format (e.g. within a catalogue).
  4. Technical metadata coverage extends across the domain's full dataset inventory, is kept current through a defined maintenance process, and includes constraints and data quality characteristics. Data Engineer follows a defined process ensuring technical metadata stays accurate as datasets change, covering the full domain inventory rather than only priority datasets.
  5. Technical metadata is captured and maintained automatically (e.g. via automated schema scanning integrated with the catalogue), always reflecting current structure without manual documentation effort. Automated tooling scans and updates technical metadata as schemas change, removing reliance on manual documentation upkeep.
Current maturity (As-Is)
Target maturity (To-Be)

MDT-02

Metadata consistency and standardisation

How consistently do metadata definitions, naming conventions, and classifications in your domain align with organisation-wide standards, avoiding conflicting or duplicate definitions for the same concept?

Maturity level descriptions
  1. No naming or definition standards exist; the same concept may be named or defined differently across datasets within the domain, with no awareness this is a problem. Different datasets use inconsistent names and definitions for what is effectively the same concept, with no one tracking or flagging the inconsistency.
  2. Awareness exists that inconsistency is a problem, and informal efforts are sometimes made to align terms, but there is no defined standard to align to. Data Engineer occasionally tries informally to align naming or definitions across datasets but has no standard document to refer to.
  3. A documented naming convention and definition standard exists and is applied to new metadata entries in the domain, reducing (but not eliminating) inconsistency. Data Engineer applies a documented standard when creating new metadata entries, even though legacy inconsistencies may remain unresolved.
  4. Existing metadata is systematically reviewed against the standard, with identified inconsistencies (duplicate or conflicting definitions) tracked and resolved on a defined cycle. Data Engineer's domain maintains a tracked list of identified metadata inconsistencies, with resolution actions assigned and progressed.
  5. Consistency is enforced automatically (e.g. automated duplicate/conflict detection against the enterprise glossary) at the point new metadata is created, preventing new inconsistencies from being introduced. Automated tooling checks new metadata entries against existing definitions and flags likely duplicates or conflicts before they are published.
Current maturity (As-Is)
Target maturity (To-Be)

MDT-03

Metadata accessibility and discoverability

How easily can people who need to use your domain's data actually find and access relevant metadata (glossary definitions, technical documentation, lineage) without having to ask a specific individual?

Maturity level descriptions
  1. Metadata, where it exists at all, is not discoverable by anyone other than the person who created it; finding it requires knowing exactly who to ask. Anyone wanting metadata about domain data must track down a specific individual, since there is no searchable or shared location.
  2. Metadata exists in known but scattered locations (shared drives, individual documents); finding it requires knowing where to look, with no search capability. Data Engineer can direct someone to a specific folder or document location for metadata, but there is no search or catalogue capability.
  3. Metadata for priority datasets is published in a defined, searchable location (e.g. a data catalogue) that relevant users know how to access. Data Engineer publishes or ensures publication of metadata for priority datasets into a searchable catalogue, and can direct users to it.
  4. Metadata for the domain's full dataset inventory is discoverable via the catalogue, with usage tracked to understand whether users are actually finding what they need. Data Engineer ensures comprehensive catalogue coverage for the domain and reviews catalogue usage/search data to identify discoverability gaps.
  5. Metadata is proactively surfaced to users at the point of need (e.g. within query tools, dashboards, or AI assistants) rather than requiring users to actively search for it. Metadata is integrated directly into the tools people use to access or analyse domain data, appearing automatically alongside the data itself.
Current maturity (As-Is)
Target maturity (To-Be)

MDT-04

Metadata quality and currency

How confident can users be that the metadata describing your domain's data is accurate and reflects the data's current state, rather than being outdated or misleading?

Maturity level descriptions
  1. Metadata, where it exists, is never reviewed or updated after initial creation; it is unknown whether it still accurately reflects the data. Metadata is written once, if at all, and never revisited even as the underlying data or its meaning changes over time.
  2. Metadata is occasionally updated when someone happens to notice it's wrong, but there is no deliberate process to check currency. Data Engineer or Data Practitioners fix metadata inaccuracies opportunistically when noticed, with no scheduled check for currency.
  3. A defined process exists for reviewing metadata currency for priority datasets on a periodic basis (e.g. annually), with updates made as part of that review. Data Engineer participates in a scheduled metadata review for priority datasets, confirming or correcting definitions and documentation as needed.
  4. Metadata currency is tracked across the domain's full dataset inventory, with staleness flagged (e.g. metadata not reviewed within a defined window) and remediation tracked. Data Engineer's domain tracks metadata review dates across the full inventory, with overdue reviews flagged and assigned for action.
  5. Metadata currency is maintained automatically where possible (e.g. metadata updates triggered by underlying data or schema changes), with minimal reliance on scheduled manual review. Changes to underlying data or schema automatically trigger a metadata review or update prompt, rather than waiting for the next scheduled cycle.
Current maturity (As-Is)
Target maturity (To-Be)

MDT-05

Metadata standards for AI/GenAI data assets

How well-defined are the metadata requirements specifically for data assets used in AI/GenAI initiatives (e.g. dataset composition, provenance, licensing, known biases, intended use and restrictions), as distinct from standard business/technical metadata?

Maturity level descriptions
  1. No AI-specific metadata requirements exist; datasets used in AI initiatives carry only standard business/technical metadata, if any, with no documentation of provenance, licensing, or known limitations relevant to AI use. Data used for AI purposes carries whatever standard metadata it already had, with no additional documentation addressing AI-specific concerns.
  2. Awareness exists that AI use needs additional metadata (e.g. licensing status, known bias concerns), but nothing has been defined or documented as a requirement. Data Engineer or Data Practitioners have informally discussed AI-specific metadata needs for domain data but have not defined or documented requirements.
  3. A documented AI-specific metadata standard exists (e.g. a "dataset card"-style template covering provenance, licensing, known limitations) and is applied to at least priority datasets used in AI initiatives. Data Engineer completes a documented AI-specific metadata template for priority datasets before or as they are used in AI initiatives.
  4. AI-specific metadata coverage extends across all domain datasets used in AI initiatives, is kept current, and is reviewed as AI use cases evolve. Data Engineer maintains comprehensive AI-specific metadata coverage for all domain datasets in active AI use, with periodic review as use cases change.
  5. AI-specific metadata is captured and kept current automatically as part of the AI pipeline (e.g. provenance and lineage auto-logged at ingestion), integrated into standard catalogue infrastructure rather than maintained separately. Automated tooling captures and maintains AI-specific metadata (provenance, licensing status, known limitations) as data flows into AI pipelines, without manual documentation effort.
Current maturity (As-Is)
Target maturity (To-Be)

MDT-06

Metadata completeness gate before AI/GenAI use

Is there a defined checkpoint that confirms a dataset has the required metadata (provenance, licensing, sensitivity classification, known limitations) before it is approved for use in an AI/GenAI initiative — and how consistently is that checkpoint applied?

Maturity level descriptions
  1. No checkpoint exists; datasets can be used in AI/GenAI initiatives regardless of whether required metadata is present or complete. Data moves into AI use with no check on whether provenance, licensing, or sensitivity metadata is even present, let alone complete.
  2. A metadata check is sometimes performed informally before AI use, depending on who is involved, but it is not a required or consistent step. Data Engineer or a Data Practitioner occasionally checks for key metadata informally before AI use, but this depends on individual initiative rather than a required step.
  3. A defined completeness checkpoint (checklist or sign-off) exists and is required before domain data is approved for at least priority AI/GenAI use cases. Data Engineer performs or confirms a documented metadata completeness check before priority domain data is approved for an AI/GenAI initiative, with the check recorded.
  4. The metadata completeness gate is consistently applied across all AI/GenAI use of domain data, with pass/fail outcomes tracked and incomplete datasets blocked pending remediation. Data Engineer's domain tracks metadata gate outcomes across all AI-related data approvals, with incomplete metadata blocking approval until resolved.
  5. The metadata completeness gate is fully automated and technically enforced within the data platform, making it impossible for data lacking required metadata to be provisioned into AI/GenAI systems. An automated gate technically prevents data lacking required AI-specific metadata from being provisioned into AI pipelines, with no reliance on manual sign-off.
Current maturity (As-Is)
Target maturity (To-Be)

MDT-07

Feedback loop: AI usage experience informing metadata improvement

When an AI/GenAI system misuses, misinterprets, or retrieves the wrong domain data (e.g. due to unclear definitions, missing context, or poor documentation), how effectively does that experience feed back into improving your domain's metadata?

Maturity level descriptions
  1. No mechanism exists to connect AI misuse or retrieval problems back to underlying metadata gaps; such issues are treated purely as a model or tooling problem. When an AI system retrieves or uses domain data incorrectly, the possibility that unclear or missing metadata contributed is not investigated.
  2. Metadata gaps are occasionally suspected as a cause of AI issues and looked into informally, but without a defined process or documented findings. Data Engineer or Data Practitioners have, at least once, informally suspected a metadata cause for an AI retrieval or interpretation issue, without a defined process.
  3. A defined process exists for investigating and documenting whether AI issues trace back to metadata gaps, and Data Engineers are expected to use it when notified of a relevant issue. Data Engineer participates in a documented investigation process when an AI issue is flagged as potentially metadata-related, recording findings.
  4. AI issues traced to metadata gaps are logged, tracked to remediation, and used to update the domain's glossary, technical metadata, or AI-specific metadata standards. Data Engineer's domain tracks confirmed metadata-caused AI issues through to remediation, with resulting updates made to the glossary (MD-1) or AI-specific standards (MD-6) to prevent recurrence.
  5. AI system performance monitoring is integrated with metadata quality monitoring, so that patterns suggesting metadata gaps (e.g. repeated poor retrieval for a given dataset) automatically trigger a metadata review, closing the loop continuously. Automated monitoring of AI retrieval/interpretation performance is linked to domain metadata, so that detected patterns automatically trigger a metadata review without waiting for manual reports.
Current maturity (As-Is)
Target maturity (To-Be)