← Back to sections

Data Governance

How data ownership, policies, access controls, and compliance are defined, enforced, and monitored across the organisation's data assets.

GOV-01

Policy definition and day-to-day enforcement

How consistently are data governance policies (data handling, classification, retention, PII treatment, etc.) actually followed in your domain's daily operations, as opposed to existing only on paper?

Maturity level descriptions
  1. Policies are informal, unwritten, or ignored; day-to-day work proceeds with no reference to any governance policy. Data handling decisions (sharing, classification, retention) are made based on individual judgement, with no policy document to refer to even if someone wanted to.
  2. Policies exist as static documents, but enforcement is manual, inconsistent, and largely reactive — followed when someone remembers, not by default. Data Engineer is aware policy documents exist and occasionally refers to them, but day-to-day practice frequently diverges from what's written, without consequence or correction.
  3. Policies are actively applied to domain work via defined, repeatable steps (checklists, standard procedures), even though enforcement remains largely manual. Data Engineer follows a documented procedure (e.g. a checklist for classifying new datasets, or a standard process for handling access requests) that operationalises policy into concrete steps.
  4. Policy compliance is systematically measured, with role-based controls in place; deviations are identified through monitoring rather than by accident. Access controls are role-based and enforced through system configuration rather than manual discipline alone; policy adherence is measured (e.g. via periodic sampling or automated checks) and reported.
  5. Policy enforcement is continuous and automated, integrated directly into the data catalogue, pipelines, or platform, leaving minimal room for manual deviation. Classification, retention, and access rules are enforced automatically at the point of data creation or ingestion, with violations blocked or flagged in real time rather than caught after the fact.
Current maturity (As-Is)
Target maturity (To-Be)

GOV-02

Access control and approval workflow

How well-defined, auditable, and consistently followed is the process for granting, reviewing, and revoking access to data in your domain?

Maturity level descriptions
  1. Access is granted informally (verbal requests, ad hoc emails), with no consistent process, approval step, or record. Anyone can request access through any informal channel and receive it without a defined approval step or documented justification.
  2. An informal approval step exists (e.g. a manager's sign-off), but it is inconsistently applied and not centrally tracked. Some access requests go through a manager or Data Engineer for informal sign-off, but this is not required consistently, and there is no central log.
  3. A documented access request and approval process exists and is generally followed, with requests recorded in a defined system (ticketing, spreadsheet, form). Data Engineer processes access requests through a named, repeatable workflow, with each request logged with requester, approver, and justification.
  4. Access is role-based, tied to defined access tiers, with periodic reviews to confirm access remains appropriate and revoke what is no longer needed. Access levels map to defined roles/tiers rather than being granted individually each time, and a scheduled review process (e.g. quarterly) identifies and removes stale or excessive access.
  5. Access provisioning, review, and revocation are largely automated and continuously monitored, with real-time visibility into who has access to what and why. Access changes are triggered automatically by role changes (e.g. via HR/identity system integration), with continuous monitoring flagging anomalous or unused access without waiting for a scheduled review.
Current maturity (As-Is)
Target maturity (To-Be)

GOV-03

Issue escalation, breach handling, and exception management

When a data governance issue occurs in your domain (a policy breach, an access anomaly, a data quality problem with compliance implications), how effective is the process for detecting, escalating, and resolving it?

Maturity level descriptions
  1. Issues are typically discovered by accident (a report looks wrong, someone complains) rather than through any deliberate process, and are handled case-by-case with no formal escalation path. When something goes wrong, it is dealt with informally, often after the fact, by whoever happens to notice, with no defined next step.
  2. Issues are handled reactively via informal channels (Slack, verbal, ad hoc email) once discovered, without a formal ticket, tracking, or audit trail. Data Engineer raises issues informally to whoever seems relevant, and resolution happens without any record of what occurred or how it was fixed.
  3. A defined escalation process exists (a named contact, a ticketing category, an escalation matrix) and Data Engineers are expected to use it when issues arise. Data Engineer routes governance issues through a specific, known channel (e.g. a compliance ticket queue) with a defined initial responder and expected response step.
  4. Issues are logged, tracked to resolution, and reviewed for root cause, with response-time expectations and accountability for closure. Data Engineer's domain has a tracked issue log showing time to detection, escalation, resolution, and a documented root-cause note for at least significant issues.
  5. Issue detection is proactive (monitoring/alerting catches problems before they're reported), and resolution patterns feed back into policy and control improvements. Automated monitoring flags likely governance issues (e.g. anomalous access, policy-violating data flows) before they are reported by a user, and recurring issue patterns trigger control or policy updates.
Current maturity (As-Is)
Target maturity (To-Be)

GOV-04

Existence and application of data contracts / terms of use

To what extent does your domain have defined data contracts or terms of use for its datasets, specifying obligations, permitted uses, and restrictions, and how consistently are these actually applied when access is granted?

Maturity level descriptions
  1. No data contracts or terms of use exist for domain datasets; access is granted with no stated obligations, permitted uses, or restrictions attached. Data is shared or accessed with no accompanying terms of any kind — recipients are not told what they can or can't do with it.
  2. Informal, unwritten expectations exist about acceptable use (e.g. "don't share this externally"), but nothing is documented or consistently communicated. Data Engineer may verbally mention usage expectations when granting access, but this is inconsistent, undocumented, and dependent on who is granting access.
  3. A standard data contract or terms-of-use template exists and is applied to at least the domain's priority/sensitive datasets. Data Engineer attaches a documented terms-of-use statement (permitted uses, restrictions, obligations) when granting access to priority datasets, using an available template.
  4. Data contracts/terms of use are applied consistently across all datasets in the domain (not just priority ones), are version-controlled, and are reviewed on a defined cycle. Every dataset in the domain has an associated, current terms-of-use record; Data Engineer can confirm coverage is comprehensive and terms are periodically reviewed and updated.
  5. Data contracts are embedded directly into the data platform/catalogue and enforced technically — e.g. terms are tied to access provisioning and cannot be bypassed. Terms of use are integrated into the access-granting system itself (e.g. access cannot be provisioned without an associated, current contract record), with automatic flags if a dataset's terms are missing or expired.
Current maturity (As-Is)
Target maturity (To-Be)

GOV-05

Governance policy coverage for AI/GenAI use of domain data

How well do your organisation's data governance policies specifically address the use of domain data in AI/GenAI initiatives (training, fine-tuning, retrieval-augmented generation, prompting, or agentic tool use), as distinct from general data governance policy?

Maturity level descriptions
  1. No governance policy addresses AI/GenAI use of data specifically; general data policy is silent on whether or how it applies to AI initiatives. Data flows into AI tools and pipelines with no policy coverage at all — governance policy predates, and does not mention, AI/GenAI use.
  2. General data governance policy is assumed to apply to AI use "by extension," but this has never been explicitly stated, tested, or interpreted for AI-specific scenarios. Data Engineer or Data Practitioners informally assume existing rules (e.g. on PII handling) apply to AI contexts too, without confirmation or explicit guidance.
  3. A specific policy addendum, section, or standalone document exists addressing AI/GenAI use of data, covering baseline rules (e.g. no sensitive data into public GenAI tools, approval required for training-data use). Data Engineer has access to and can reference a specific, named policy artefact addressing AI/GenAI data use, distinct from general governance policy.
  4. AI/GenAI governance policy is comprehensive, covering distinct AI use-case types (training, RAG, prompting, agentic access) with differentiated rules, and is reviewed on a defined cycle as AI usage evolves. Data Engineer applies differentiated rules depending on the specific type of AI use (e.g. stricter rules for training-data inclusion than for read-only RAG retrieval), with policy reviewed periodically against emerging AI use patterns in the domain.
  5. AI/GenAI governance policy is continuously updated in step with the organisation's evolving AI capability and regulatory obligations, with Data Engineers proactively briefed on changes before they take effect operationally. Data Engineer is notified of AI governance policy changes ahead of new AI tooling or use cases going live in their domain, ensuring policy is never lagging actual practice.
Current maturity (As-Is)
Target maturity (To-Be)

GOV-06

Access control and approval for AI systems/pipelines consuming domain data

How well-defined, auditable, and consistently applied is the access approval process specifically for AI systems, models, pipelines, or agentic tools requesting access to data in your domain?

Maturity level descriptions
  1. AI systems/pipelines obtain data access the same informal way as any other request, or entirely outside any request process (e.g. via existing broad access, shadow IT, or public GenAI tool uploads), with no AI-specific scrutiny. An AI tool, pipeline, or agent gains access to domain data without any distinct review of what that access enables the AI system to do with the data.
  2. AI access requests are occasionally flagged as different from normal requests, but there is no consistent, defined process distinguishing AI consumption from human consumption. Data Engineer sometimes recognises an access request is for an AI system and applies extra informal caution, but without a defined process to follow.
  3. A defined approval process exists specifically for AI/pipeline access requests, distinct from standard human access requests, covering what the AI system will do with the data (train, retrieve, generate, act). Data Engineer processes AI-related access requests through a named workflow that captures the intended AI use (training vs. retrieval vs. agentic action) before granting access.
  4. AI system access is role/purpose-based (e.g. tiered by use case risk), tracked separately from human access, with periodic review to confirm AI access remains appropriate and is revoked when a pipeline or tool is decommissioned. Data Engineer maintains or reviews a register of AI systems/pipelines with access to domain data, tagged by use case and risk tier, with scheduled reviews that revoke access for retired or changed AI initiatives.
  5. AI system access is provisioned, monitored, and revoked automatically, tied to the AI system's registered purpose, with real-time visibility into which AI systems/agents currently hold access to which domain data and why. Access for AI pipelines/agents is granted and withdrawn programmatically based on registered purpose and status, with continuous monitoring flagging AI access that falls outside its approved scope.
Current maturity (As-Is)
Target maturity (To-Be)

GOV-07

Monitoring, audit, and issue management for AI data usage

How effectively does your domain detect, investigate, and resolve issues arising from AI/GenAI systems' actual use of its data — including misuse, scope creep, data leakage into model outputs, or use beyond what was approved?

Maturity level descriptions
  1. There is no monitoring of how AI systems actually use domain data once access is granted; issues (misuse, leakage, scope creep) would only be discovered by accident, if at all. Once an AI system has access, there is no visibility into or check on what it actually does with the data, or whether that matches what was approved.
  2. Issues are occasionally discovered informally (e.g. a suspicious AI output prompts questions), and investigated ad hoc without a defined process. Data Engineer or Data Practitioners occasionally notice something that looks like AI misuse of domain data and look into it informally, without a defined investigation process.
  3. A defined process exists for reporting and investigating suspected AI data misuse or scope creep, with a named escalation path for the domain. Data Engineer routes a suspected AI data issue through a specific, known channel with a defined initial responder, similar to general governance issue escalation but recognising AI-specific risks.
  4. AI data usage issues are logged, tracked to resolution, and reviewed for root cause, with periodic proactive audits of AI systems' actual data consumption against their approved scope. Data Engineer's domain has a tracked log of AI-related data issues plus evidence of scheduled audits comparing actual AI data usage against approved purpose.
  5. AI data usage is continuously monitored with automated detection of anomalous consumption, scope violations, or potential data leakage into outputs, with alerts triggering immediate review and resolution patterns feeding back into policy and access controls. Automated monitoring flags AI systems consuming data outside approved scope or exhibiting signs of leakage/misuse in near-real-time, with recurring patterns feeding directly into policy or access-control updates (linking back to DG-8 and DG-9).
Current maturity (As-Is)
Target maturity (To-Be)