10 Data Governance Examples by Industry: Controls and Evidence

Ten data governance examples by industry, covering banking, insurance, telecom and healthcare, with the control that fixes each problem and the evidence a regulator asks to see.

by

Jatin S

Updated on

August 17, 2026

10 Data Governance Examples to Enhance Your Data Strategy

Key Takeaways

  • A governance example is only useful if it names three things. The data problem, the control that solves it, and the evidence a supervisor asks to see. Advice that stops at the first two cannot be defended in a supervisory meeting.
  • The industry decides the control, not the framework. A bank governs to prove a reported number is traceable. A hospital governs to prove only the right clinician read a record. Same discipline, different artefacts, different owners.
  • Asia Pacific supervisors ask specific questions. OJK expects an Indonesian bank to know where personal data lives and who can reach it. APRA expects an Australian bank to classify information assets and hold a register of them. MAS expects a Singapore firm to justify the attributes behind an automated decision.
  • Insurance governance is model governance. The NAIC model bulletin pushes insurers towards a written programme covering the AI systems used in underwriting and claims, including the ones bought from a vendor rather than built in house.
  • Telecommunications governance is a volume problem. Consent and retention are simple rules that become hard when they have to hold across a national subscriber base and billions of network records, propagate in minutes, and survive into every backup and derived copy.
  • Write the evidence down before you build the control. If you cannot name the document you will hand over when asked to prove the control ran, the control is not finished.

Most articles about data governance examples describe a discipline and leave the reader to work out what it means for their business. That gap is the reason a bank, an insurer, a telecom operator and a hospital all read the same governance advice and none of them can act on it, because the thing that changes between them is not the principle, it is the artefact somebody will eventually ask them to produce.

This article takes the opposite approach. Ten examples, each set in an industry Decube works in, each written to the same three part shape: the data problem as it actually shows up, the control that solves it, and the evidence a supervisor or an auditor asks to see. Financial services and telecommunications carry the most weight, because that is where the questions arrive most often, and healthcare, insurance and retail follow.

If you need the underlying definitions first, data governance concepts covers ownership, stewardship, policy and the operating model in one place. This page assumes them and goes straight to the worked examples.

What Makes a Data Governance Example Worth Copying

A governance example fails the moment it becomes a slogan. "Establish clear data ownership" is true everywhere and usable nowhere. The examples below are written to a fixed shape, and the shape is the point.

  • The data problem. The specific way the data breaks in that industry, described concretely enough that a practitioner recognises their own week in it.
  • The control. The mechanism that fixes it, named as a thing you build and maintain rather than as a principle you agree with.
  • The evidence. The artefacts you hand over when somebody asks you to prove the control exists and ran. This is the part almost every governance article skips, and it is the part a regulated buyer cares about most.

One test separates a real control from a slogan. Ask who is on the hook when it fails, by name, and ask what file gets opened to show it worked. If neither question has an answer, what you have is an intention.

The mindmap breaks a governance framework into the security protocols and access controls that make up the control, the catalog and lineage tracking features that make up the evidence, and the compliance and efficiency outcomes those are meant to produce.

1. Banking in Indonesia: Classifying Customer Data So OJK Reporting Can Be Trusted

The data problem. An Indonesian bank holds customer identity fields in core banking, in loan origination, in the mobile app and again in the data warehouse. Personal data has been copied into analytics tables that nobody classified, access was granted table by table as projects needed it, and when a regulatory return is questioned nobody can say which copy of the customer record fed the number.

The control. A classification register covering every table and column that carries personal or account data, with one accountable owner per data domain, and access granted by classification rather than by table name. New tables inherit a classification at creation or they do not get served. The register is not a spreadsheet somebody maintains quarterly; it lives in the catalog next to the data and it is refreshed automatically when a source changes.

The evidence. The classification register with an owner name and a date against each domain. The quarterly access review showing who could read classified columns and what changed. Lineage from the source system through each transformation to the figure that appears in the return, so a challenged number can be traced without a project.

What the supervisor is working from. OJK regulates the financial services sector in Indonesia and expects a supervised institution to apply risk management to its use of technology and data, which in practice means being able to say what personal data it holds, where it sits and who can reach it. Indonesias Personal Data Protection Law, Law 27 of 2022, sits alongside that and gives the same questions a legal edge. A bank that can answer both from one register answers them in an afternoon. A bank that cannot answers them with a project.

If you are choosing a platform to sit under this, our comparison of data governance platforms for banks works through the requirements a regulated bank should be testing against.

2. Banking in Australia: The Critical Data Element Register APRA Asks About

The data problem. Risk and finance report the same exposure from two different pipelines and produce two different totals. Both are defensible internally and neither can be defended externally, because no one wrote down which field is authoritative, what quality rule it has to pass, or what happens when it fails.

The control. A critical data element register. Name the fields the capital, liquidity and operational risk returns depend on. Give each one an accountable owner and one authoritative source system. Attach a quality rule with a threshold to each, run the rule on every load, and route a breach to a named person rather than to a mailbox. Then hold lineage from that source field to the reported figure.

The evidence. The register itself. The quality test results over time, which matter more than a point in time pass because a supervisor wants to see the control operating. The lineage graph from source to return. The incident log for breached thresholds, with what was done and how long it took.

What the supervisor is working from. APRA's prudential standard CPS 234 requires a regulated entity to classify its information assets by criticality and sensitivity and to maintain controls sized to that classification, which is why an information asset register is the first thing asked for. CPS 230 on operational risk management extends the same logic to critical operations and to the service providers behind them. For larger banks the Basel Committees principles for effective risk data aggregation, known as BCBS 239, set the expectation that risk data can be aggregated accurately and quickly, which is a lineage requirement in everything but name.

3. Insurance in the United States: Model Governance for Underwriting and Claims

The data problem. An underwriting or claims model consumes features built from data whose source, refresh cadence and quality nobody documented. When a state regulator asks why a claim was declined or why a premium was set where it was, the answer available internally is a model score, and the inputs behind it cannot be reconstructed for the date the decision was made.

The control. A model input register that ties every feature to the source column it derives from, its refresh schedule and the quality rule it has to pass, versioned so the state of the inputs on any past date can be reproduced. Around it, a documented review before any model or feature change reaches production, and the same treatment for models bought from a vendor as for models built in house.

The evidence. The feature to source mapping with versions and dates. The model documentation, including what it was tested for and by whom. The record of testing for unfairly discriminatory outcomes. The vendor oversight file for any third party model, because buying a model does not move the accountability.

What the supervisor is working from. The NAIC model bulletin on the use of artificial intelligence systems by insurers, adopted in December 2023 and taken up by a growing list of states, sets the expectation that an insurer maintains a written programme covering governance, risk management controls and third party vendor oversight for the AI systems behind insurance decisions, and that it can answer regulator questions about how those systems work. Almost all of that programme is a data governance question wearing a model governance label, because a model you cannot explain is usually a model whose inputs you never governed.

4. Financial Services in Singapore: Proving Fairness and Accountability to MAS

The data problem. An automated decision, a credit limit, a price, a fraud hold, is made from attributes nobody has justified. Some of them are proxies for characteristics the firm would never use directly, and the firm cannot tell, because there is no list of what the decision consumes.

The control. An accountability record for every automated decision that affects a customer. It names an internal owner, lists the attributes the decision consumes, states a justification for each attribute, marks the ones that carry proxy risk, and records a fairness assessment done before launch and repeated on a schedule afterwards. The record is versioned, so the firm can say which version made a decision on a given day.

The evidence. The attribute inventory with a justification against each entry. The fairness assessment with its metrics, its date and its author. The approval record from whoever signed it off. The ability to reproduce a single customer decision on request, which is the test that quietly fails most often.

What the supervisor is working from. MAS published principles for fairness, ethics, accountability and transparency in the use of artificial intelligence and data analytics in the financial sector, known as FEAT, and followed them with an industry programme that turned the principles into assessment methodology. The practical effect is that a Singapore firm is expected to be able to explain a data driven decision rather than only to make it, and explanation is an attribute inventory problem before it is a model problem.

5. Telecommunications: Holding Subscriber Consent Across a National Base

The data problem. Consent is captured in four places, the CRM, the app, the retail till and the call centre, and the marketing platform reads a copy that is a day or two old. A subscriber withdraws consent in the app on Monday and receives a campaign on Wednesday. Nobody was careless. The record simply had four homes and no owner.

The control. One consent record per subscriber identity, held as master data with a defined authoritative source and a publish path to every consuming system. Withdrawal propagates as an event rather than as a nightly batch. A freshness rule fails loudly when a downstream system falls behind, and a suppression check runs before any campaign is released rather than after a complaint arrives.

The evidence. The consent record with a timestamp and the channel it was captured in. The propagation log showing when each downstream system received a change, which is what turns "we honour withdrawals" into a provable statement. The suppression test results for each campaign, kept with the campaign.

The scale is what makes this hard rather than the rule. A rule that is trivial for ten thousand customers becomes an engineering problem at tens of millions, and the failure is always the same shape: the rule was right and the copy was stale.

The chart splits the case for continuous monitoring of a consent record into privacy penalty exposure, telecom return on investment and decision speed.

6. Telecommunications: Retention and Deletion on Records That Arrive by the Billion

The data problem. Call detail records, network event logs and location data accumulate faster than anyone reviews them. A retention policy exists in a document, expressed in years, and nothing in the platform enforces it. Meanwhile the same records have been copied into three analytics environments and a backup vault, none of which the policy author knew about.

The control. A retention schedule bound to the data itself rather than to a policy document. Classification drives a retention period per table, a deletion job runs on that schedule and records what it removed, and an exception register holds anything under legal hold with the reason and the release date. Derived copies inherit the retention of their source, which is the rule that most programmes forget to write.

The evidence. The schedule mapped to physical tables rather than to categories. The deletion job run history with row counts, because a job that ran and deleted nothing is a failure that looks like a success. The legal hold register. Proof that derived copies and backups are covered, which is where most retention programmes are actually breached.

7. Healthcare: Access Control and Emergency Access Audit on Patient Records

The data problem. Clinicians need wide access in an emergency, so access is granted widely and then never reviewed. Over a few years the effective read population for the most sensitive records in the hospital becomes larger than anyone intended, and nobody notices because nothing broke.

The control. Role based access matched to the classification of the record rather than to the system it sits in. An emergency access path that is always available and always logged, so clinical urgency never becomes a reason to weaken the default. Access recertification on a fixed cycle, where a named manager confirms each person still needs what they hold and inaction removes access rather than preserving it.

The evidence. The role to classification matrix. The emergency access log with a reviewed reason against each event, which is the artefact that makes emergency access acceptable. The recertification record with dates and approvers. The read audit trail on sensitive records, retained long enough to answer a complaint that arrives months later.

What the supervisor is working from. The HIPAA privacy rule requires reasonable efforts to limit uses and disclosures of protected health information to the minimum necessary for the purpose, which is an access control obligation stated as a privacy principle. A hospital that cannot show the effective read population for a record set cannot show it met that standard.

Starting from automated access controls, the mindmap branches into the operational risks a stale access list creates, the health privacy and data protection reporting it has to satisfy, and the breach cost that follows when the review never happens.

8. Healthcare: Secondary Use, De Identification and the Consent Question

The data problem. A research group or an AI team needs patient data, and the fastest route to it is a copy of production into an analytics environment. It happens once as an exception, and the exception becomes the pipeline. Later, when a patient asks what their record was used for, the honest answer is that nobody tracked it.

This is not hypothetical, and the live version of this article already carried a version of it: a large health system drew scrutiny for putting AI transcription into clinical settings without clear patient consent. The lesson generalises. Consent given for care is not consent given for model training, and the distinction has to be held in the data rather than in a policy.

The control. A de identified or masked serving layer as the only supported route to secondary use, with a documented method: either the identifiers listed in the safe harbour approach removed, or a qualified expert determination that the re identification risk is very small. Project level approval recorded before access is granted, and a purpose recorded against each approved dataset so the question of what a record was used for has an answer.

The evidence. The de identification method documentation and the list of fields removed or transformed. The expert determination if that route was taken. The approval record per project with its purpose. Lineage showing each analytics table derives from the de identified layer and never from production, which is the one piece of proof that cannot be assembled retrospectively.

9. Retail and Ecommerce: Identity Resolution Without Breaking Consent

The data problem. The same shopper exists three times: a guest checkout, a loyalty member and an app user. Personalisation reads the merged profile, but consent was given by only one of the three identities, and the merge silently promoted it to all of them. The business sees a better customer view. The privacy team sees an unlogged change of legal basis.

The control. Identity resolution with a written merge rule, versioned, owned by a named person. Consent held against the resolved identity, with the most restrictive consent winning on merge rather than the most recent. A crosswalk from the merged profile back to every contributing record, so a merge can be explained and reversed. Deletion follows the crosswalk, not the profile.

The evidence. The merge rule and its version history. The crosswalk. A deletion test showing a single request reached every contributing record including the ones in the analytics copies. The consent state of the resolved identity with the date and source of each element.

The failure mode here is worth naming because it is not obvious. A wrong merge in retail looks like a small personalisation error. The same wrong merge mixes two peoples purchase histories, their addresses and their consent choices, and in a jurisdiction with a data protection regulator that is an incident rather than a defect.

10. Every Industry: Governing the AI Agents That Now Read Your Data

The data problem. An assistant or an autonomous agent is given warehouse access so it can answer questions, and it reads whatever its credentials allow. Nobody can say what it read, on whose behalf, or which tables produced a given answer. The access was granted once, to a service account, for a pilot.

The control. Agents registered as first class data consumers with their own identity rather than a shared service account. Their permission scope expressed against the classification register, so an agent inherits the same restrictions a person would. Every read logged with the agent, the requesting user and the purpose. Outputs traceable back to the tables behind them.

The evidence. The agent register with an owner per agent. The permission scope per agent, reviewed on the same cycle as human access. The read log. Lineage from an answer back to the source tables, which is what turns an AI output into something a risk function can sign off.

What the supervisor is working from. Obligations for general purpose AI models under the EU AI Act have applied since 2 August 2025 for models placed on the market from that date, the Commission's enforcement powers apply from 2 August 2026, and models placed on the market before 2 August 2025 have until 2 August 2027. The EU AI Act's Article 50 transparency rules were not changed and apply from 2 August 2026, and the Digital Omnibus, which entered force on 27 July 2026, did not move either of them. The Omnibus did shift the high risk obligations to 2 December 2027 for standalone systems and 2 August 2028 where the AI is embedded in a regulated product. A great deal of published guidance still quotes the older dates. If you are building the governance layer for this now, agentic AI data governance goes further into the control model.

The Ten Examples Side by Side

The table below is the article in one view. It is also the fastest way to find the row that matches your own week.

ExampleThe data problemThe controlThe evidence a supervisor asks for
1. Banking, Indonesia (OJK)Personal data copied into unclassified analytics tables; no one can say which copy fed a regulatory returnClassification register in the catalog, one accountable owner per domain, access granted by classificationRegister with owners and dates, quarterly access review, lineage from source to reported figure
2. Banking, Australia (APRA)Risk and finance report the same exposure from two pipelines and disagreeCritical data element register with an authoritative source, a quality rule and a threshold per fieldThe register, quality results over time, lineage to the return, incident log for breaches
3. Insurance, United States (NAIC)A model declines a claim and its inputs cannot be reconstructed for the decision dateVersioned model input register mapping every feature to a source column, plus documented change reviewFeature to source mapping, model documentation, discrimination testing record, vendor oversight file
4. Financial services, Singapore (MAS)An automated decision consumes attributes nobody has justified, some of them proxiesAccountability record per decision: owner, attribute list, justification, proxy risk flags, fairness assessmentAttribute inventory with justifications, fairness assessment, approval record, reproduction of a single decision
5. Telecommunications, consentConsent lives in four systems and the marketing platform reads a stale copyOne consent record per subscriber identity, event driven propagation, freshness rule, suppression check before releaseConsent record with timestamp and channel, propagation log per system, suppression test per campaign
6. Telecommunications, retentionNetwork and call records accumulate past their retention period across copies nobody mappedRetention bound to classification, automated deletion, exception register for legal hold, derived copies inherit retentionSchedule mapped to physical tables, deletion run history with row counts, legal hold register, backup coverage proof
7. Healthcare, accessEmergency access grants became permanent and the effective read population grew unnoticedRole based access tied to record classification, logged emergency access path, recurring recertificationRole to classification matrix, emergency access log with reviewed reasons, recertification record, read audit trail
8. Healthcare, secondary useProduction data copied for research and AI, with no record of what a patient record was used forDe identified serving layer as the only route, documented method, project approval with a recorded purposeMethod documentation, removed field list, expert determination, per project approval, lineage from the de identified layer
9. Retail and ecommerceIdentity merge promotes one consent to three identities without a logged change of legal basisVersioned merge rule with a named owner, most restrictive consent wins, crosswalk to contributing recordsMerge rule versions, crosswalk, deletion test across every contributing record, consent state with dates
10. Every industry, AI agentsAn agent reads whatever its shared service account allows and nobody can say what it readAgents registered with their own identity, scope expressed against the classification register, every read loggedAgent register with owners, permission scope per agent, read log, lineage from an answer to its source tables

What Each Regulator Actually Asks For

The obligations below are the ones that turn into data work. They are summarised at the level the published instruments support, and the sources are listed at the end of this article.

Supervisor or instrumentMarketWhat it means for your dataThe artefact it turns into
OJK, with Law 27 of 2022 alongside itIndonesia, financial servicesKnow what personal data the institution holds, where it lives, and who can reach it, under a risk management framework for technology useClassification register with named domain owners and an access review record
APRA prudential standard CPS 234Australia, banking, insurance and superannuationClassify information assets by criticality and sensitivity, and size controls to that classificationInformation asset register, tied to the systems that hold each asset
APRA prudential standard CPS 230Australia, banking, insurance and superannuationIdentify critical operations and the service providers behind them, and manage the operational risk across bothCritical operation and service provider register, with the data each one depends on
Basel Committee principles BCBS 239Global, larger banksAggregate risk data accurately and quickly, including under stress, with a clear line back to sourceCritical data element register and end to end lineage to the reported figure
NAIC model bulletin on AI systemsUnited States, insuranceMaintain a written programme for the AI systems used in insurance decisions, covering governance, risk controls and vendor oversightModel input register, model documentation, testing record, vendor oversight file
MAS FEAT principlesSingapore, financial servicesBe able to justify and explain data driven and AI driven decisions that affect customersAttribute inventory with justifications, and a dated fairness assessment per decision
HIPAA privacy ruleUnited States, healthcareLimit uses and disclosures of protected health information to the minimum necessary, and apply a recognised method when data is de identifiedRole to classification matrix, access recertification record, de identification method documentation
EU AI ActEuropean Union, all sectorsGPAI obligations have applied since 2 August 2025 for models placed on the market from that date, with models on the market before it having until 2 August 2027 and Commission enforcement powers from 2 August 2026. Article 50 disclosure applies from 2 August 2026. High risk duties move to 2 December 2027 and 2 August 2028 after the Digital OmnibusAgent and model register, permission scope, read log, lineage from output to source

Data Dictionary Examples: What One Entry Holds in Each Industry

Every control above resolves to a dictionary entry somewhere, and a dictionary entry that carries only a field name and a description is decoration. A useful entry carries the attributes that access, quality and retention are granted from, which is why the entry looks different in each industry. Below is the same customer identifier field as three different dictionary entries.

Attribute in the entryBank (Indonesia or Australia)HospitalTelecom operator
Business definitionThe unique identifier for a banking relationship holder, one per legal personThe unique identifier for a patient across every episode of careThe unique identifier for a subscriber, distinct from the SIM and the account
Authoritative sourceCore banking, never the warehouse copyThe patient administration system, never a departmental systemThe CRM, with the app and the retail till as writers rather than owners
ClassificationPersonal data, restrictedProtected health information, restricted, and often a higher tier for behavioural and sexual health recordsPersonal data, restricted, with location derived fields treated separately
Accountable ownerNamed head of customer data in retail bankingNamed clinical information owner, usually with the privacy officer countersigningNamed subscriber data owner in the customer function
Quality rule and thresholdUniqueness and completeness on identity fields, breach routed to the named ownerUniqueness across episodes, with duplicate patient records treated as a clinical safety issueUniqueness against the consent record, checked before every campaign release
RetentionSet by financial records rules and by the customer relationship end dateSet by clinical records retention, typically far longer than commercial dataSet by the shortest of the commercial rule and the lawful basis for the derived records
Special attributesWhether the field feeds a regulatory return, and which oneWhether the field survives de identification and by which methodWhether the field is consent bearing, and which consent record governs it

The test for a dictionary is whether anything is granted from it. If access, quality and retention are still decided in three other places, the dictionary is a glossary and it will be out of date within a year.

Information Governance and Data Governance Are Not the Same Thing

This distinction is worth keeping because it decides who owns which problem, and because mixing the two is how governance programmes end up with two committees and one budget.

Information governance covers unstructured content: documents, contracts, email, recordings and the records retention schedule that sits over them. It is usually owned by legal, compliance or a records function. Data governance covers structured and semi structured data in databases, warehouses and lakes, and it is usually owned by a data function reporting to a chief data officer or an equivalent.

DimensionData governanceInformation governance
What it coversStructured and semi structured data in databases, warehouses, lakes and streamsUnstructured content: documents, contracts, email, recordings, images
Typical ownerData function, chief data officer or head of dataLegal, compliance or a records management function
Main artefactsClassification register, data dictionary, quality rules, lineage, access modelRecords retention schedule, classification taxonomy, legal hold process, disposal certificates
How retention is expressedA period per table or dataset, enforced by a jobA period per record class, enforced by a records system and by policy
Where they meetThe same personal data appears in both, so classification and retention have to agree across the twoA legal hold on a matter has to reach the structured data, not only the document store

The practical failure is always at the join. A legal hold is placed on a matter, the document store honours it, and the deletion job in the warehouse keeps running because nobody told it. Whichever function owns the join should be named, and the honest answer in most organisations is that nobody currently does.

The two branches are the metadata work a strategy has to fund, setting clear definitions and keeping the catalog discoverable, and the cross departmental working that decides whether anyone outside the data team uses it.

How to Turn These Examples Into a Data Governance Strategy

A data governance strategy is a short set of decisions rather than a document listing principles: which data you will govern first, who is accountable for it, and what you will be able to prove by a date. The examples above give you the shape; the sequence below turns them into a plan.

  • Pick the return, not the domain. Start from a report, a regulatory return or a decision that somebody outside the data team already cares about. Governing "customer data" has no end state. Governing the twenty fields behind the liquidity return does.
  • Name one accountable person per domain, not a committee. Every survivorship dispute, every classification argument and every access exception needs a decision maker. A committee produces a meeting.
  • Write the evidence list before you build anything. For each control, write down the file you will open when asked to prove it ran. That list is your backlog, and it stops you building controls nobody can see.
  • Classify once, then let access follow classification. Access granted table by table decays within a year. Access granted by classification survives reorganisations, because the classification travels with the data.
  • Bind retention to the data, and make derived copies inherit it. A retention policy that lives only in a document is not a control. The rule that gets forgotten is the one about copies.
  • Treat agents as consumers with names. The moment an assistant reads production data, it needs an identity, a scope and a log, on the same terms as a person.

Sequence matters more than ambition. One return, governed end to end with evidence a supervisor accepts, is worth more than a classification exercise across the whole estate that nobody can point at a decision.

Where Decube Fits

The mindmap sets out the Decube platform by its features, including automated column level lineage and pipeline observability, the compliance and collaboration benefits they support, and the financial services and telecommunications sectors several of the examples above are drawn from.

Decube is a platform for the layer these examples all depend on, which is knowing what data you hold, whether it is fit to use, and where a number came from. Decube data governance covers ownership, classification, policy and the catalog surface that makes a classification register something people actually use rather than a spreadsheet that ages. Automated crawling keeps that metadata current once a source is connected, which removes the manual update cycle that kills most registers in their second year, and access to view or change data can be routed through an approval flow.

The evidence half is lineage and monitoring. Column level data lineage traces a value from the source system through every transformation to where it lands, which is the answer to the question an APRA or OJK supervisor asks about a reported number and the answer to the question a business user asks when a figure looks wrong. Alongside it, pipeline observability finds the breaks, and quality suggestions from Decube CoPilot propose the tests that turn a quality rule into a result somebody can read.

Decube holds GDPR, HIPAA, SOC 2 and ISO 27001 alignment, which matters when the platform itself becomes part of what a supervisor reviews. If you want to see what these examples look like against your own sources, you can request a demo.

Frequently Asked Questions

What are some data governance examples from successful companies?

The examples worth copying are described by control rather than by company name, because the control is what transfers. A bank builds a classification register and grants access by classification. Another bank builds a critical data element register with a quality rule and a threshold on each field behind its regulatory returns. An insurer builds a versioned model input register so an underwriting decision can be reconstructed. A telecom operator holds one consent record per subscriber and propagates withdrawal as an event. A hospital ties access to record classification and logs every emergency access. Each of those is a company example with the company name removed and the mechanism left in.

What is a data governance strategy and how do these examples fit into one?

A data governance strategy is a short set of decisions about which data you govern first, who is accountable for it, and what you will be able to prove by a date. It is not a list of principles. The examples fit in as the delivery unit: pick a report or a regulatory return that somebody outside the data team already cares about, name one accountable owner for the data behind it, write down the evidence you will hand over when asked to prove the control ran, and build only the controls that produce that evidence.

What is a good scenario for data governance?

The clearest scenario is one where two teams report the same number differently and neither can prove which is right. Risk and finance disagree on an exposure, marketing and the call centre disagree on whether a customer consented, or a research team and a privacy team disagree on whether a dataset is de identified. In each case the fix is the same shape: name the authoritative source, attach a quality rule and an owner to it, and hold lineage from that source to the number in dispute.

What are data dictionary examples in a governance programme?

A data dictionary entry in a real programme carries more than a description. For a bank it records the field name, the business definition, the authoritative source system, the classification, the accountable owner, the quality rule and threshold, and the retention period. For a hospital it adds whether the field is protected health information and whether it survives de identification. For a telecom operator it adds whether the field is consent bearing. The dictionary becomes useful at the point where access and retention are granted from it rather than from a separate spreadsheet.

What does OJK expect from data governance in an Indonesian bank?

OJK supervises the financial services sector in Indonesia and expects a supervised institution to apply risk management to its use of technology and data. In practice that means being able to say what personal data the bank holds, where it lives, who can reach it, and how a reported figure traces back to its source. Indonesia introduced a Personal Data Protection Law, Law 27 of 2022, which asks the same questions with a legal consequence attached. A classification register with named owners, an access review record and lineage to the reported figure answers both.

What does APRA expect from data governance in an Australian bank?

APRA prudential standard CPS 234 requires a regulated entity to classify its information assets by criticality and sensitivity and to maintain controls sized to that classification, which is why an information asset register is usually the first artefact requested. CPS 230 extends the same logic to critical operations and the service providers behind them. For larger banks the Basel Committee principles known as BCBS 239 add the expectation that risk data can be aggregated accurately and quickly, which in practice is a lineage requirement.

What does the NAIC model bulletin expect from insurers using AI?

The NAIC model bulletin on the use of artificial intelligence systems by insurers, adopted in December 2023 and taken up by a growing list of states, sets the expectation that an insurer maintains a written programme covering governance, risk management controls and third party vendor oversight for the AI systems behind insurance decisions, and that it can answer regulator questions about how those systems work. Models bought from a vendor are covered on the same terms as models built in house.

What does MAS expect on fairness and accountability in data driven decisions?

MAS published principles for fairness, ethics, accountability and transparency in the use of artificial intelligence and data analytics in the financial sector, known as FEAT, and followed them with an industry programme that turned the principles into assessment methodology. The practical expectation is that a firm can explain a decision that affects a customer, which means holding an inventory of the attributes the decision consumes with a justification against each, a dated fairness assessment, and the ability to reproduce a single decision on request.

How does healthcare data governance differ from financial services data governance?

The discipline is the same and the artefacts differ. A bank governs to prove a reported number is traceable, so its central artefacts are a critical data element register and lineage from source to return. A hospital governs to prove only the right clinician read a record and that secondary use was authorised, so its central artefacts are a role to classification matrix, an emergency access log with reviewed reasons, and de identification method documentation. The owner differs too: finance and risk in a bank, clinical governance and privacy in a hospital.

What is the difference between information governance and data governance?

Data governance covers structured and semi structured data in databases, warehouses, lakes and streams, and it is usually owned by a data function. Information governance covers unstructured content such as documents, contracts, email and recordings, and it is usually owned by legal, compliance or a records function. They meet wherever the same personal data appears in both, so classification and retention have to agree across them. The common failure is a legal hold that the document store honours while the deletion job in the warehouse keeps running.

Is Atlan worth it?
Atlan is worth it if your primary need is a modern data catalog with strong column-level lineage and cloud-native integrations (Snowflake, dbt, Databricks). It is harder to justify if you also need data observability and quality coverage across a heterogeneous stack — those capabilities require separate vendors, adding cost and complexity.
What is the best Atlan alternative
Decube is purpose-built for regulated financial services, with native observability, approval-gated lineage, PII auto-classification, and an AI layer (TrustyAI) that does not route metadata to a public LLM. These map directly to regulatory frameworks supervised by MAS, OJK, BNM, and APRA. Atlan AI's OpenAI dependency is often a procurement blocker in these environments.
How does Atlan compare to Alation?
Both are catalog-first platforms with strong discovery. Alation pioneered search-first data culture and analyst adoption. Atlan is stronger on column-level lineage and cloud integrations. Both require external tooling for observability and broad data quality coverage.
How long does it take to migrate from Atlan to another platform?
Migration time depends on estate size and the number of active integrations. SaaS-native platforms like Decube deploy in 2–6 weeks without professional services. The longer task is typically re-establishing business glossaries, data ownership, and custom attributes — that effort is roughly the same regardless of which platform you move to.
What is the difference between a context layer and a semantic layer?
A semantic layer standardizes how metrics are defined and calculated so every analyst and BI tool uses the same numbers. A context layer encodes governance rules, data lineage, quality signals, and organizational knowledge so AI agents can make safe, autonomous decisions. The semantic layer is for human-facing analytics. The context layer is for AI-facing autonomy.
Can I use a semantic layer without a context layer?
Yes - and most organizations do today. If your primary consumers are human analysts using BI tools, a semantic layer alone is sufficient. The context layer becomes essential when you introduce AI agents that need to understand not just what a metric means but whether and how they are allowed to use it.
Is a context layer the same as a data catalog?
No. A data catalog is a component of a context layer. The catalog inventories data assets and stores metadata. The context layer activates that metadata by delivering it to AI agents at query time through APIs and MCP connections. Modern platforms like Atlan extend catalog functionality into full context layer infrastructure.
Which tool implements a context layer?
Purpose-built context layer platforms include Decube, which combines catalog, lineage, quality, and governance into a metadata layer that delivers context to AI agents via MCP. You can also build a context layer on custom infrastructure using a vector database (for semantic search), a knowledge graph
How long does it take to implement a context layer?
Most enterprise context layer implementations take 8–16 weeks when using a purpose-built platform like Atlan. Building from scratch on custom infrastructure typically takes 6–12 months. The timeline depends heavily on how much governance metadata already exists and how many data sources need to be connected.
What is Data Context?
Data Context is the information that explains what data means, where it comes from, how it is transformed, whether it can be trusted, and how it should be used. It combines metadata, lineage, data quality, and governance so people and systems can confidently use data for analytics, reporting, and AI.
How is Data Context different from metadata?
Metadata describes data, while Data Context makes data usable and trustworthy. Metadata provides definitions, ownership, and technical details. Data Context extends this by adding lineage, quality signals, and governance rules, creating a complete, operational understanding of data.
Why is Data Context important for AI?
AI systems require Data Context to interpret data correctly, safely, and reliably. Without context, AI models may misunderstand metrics, use stale or incorrect data, or expose sensitive information. Data Context ensures AI uses trusted, well-defined, and policy-compliant data.
How does data lineage contribute to Data Context?
Data lineage provides visibility into how data flows and transforms across systems. It shows upstream sources, downstream dependencies, and transformation logic, enabling impact analysis, root-cause investigation, and confidence in reported numbers.
How do organizations build Data Context in practice?
Organizations build Data Context by unifying metadata, lineage, observability, and governance into a single operational layer. This includes defining business meaning, capturing end-to-end lineage, monitoring data quality, and enforcing usage policies directly within data workflows.
What is Context Engineering?
Context Engineering is the practice of designing and operationalizing business meaning, data lineage, quality signals, ownership, and policy constraints so that both humans and AI systems can reliably understand and act on enterprise data. Unlike traditional metadata management, Context Engineering focuses on decision-grade context that can be consumed programmatically by AI agents in real time.
How is Context Engineering different from prompt engineering?
Prompt engineering focuses on how questions are phrased for an AI model, while Context Engineering focuses on what the AI system already knows before a question is asked. In enterprise environments, context includes data definitions, lineage, quality, and usage constraints—making Context Engineering foundational for trustworthy and scalable Agentic AI.
Why is Context Engineering critical for Agentic AI?
Agentic AI systems reason, decide, and act autonomously across multiple systems. Without engineered context—such as trusted data meaning, lineage, and real-time quality signals—agents cannot assess risk or impact correctly. Context Engineering ensures AI agents act safely, explain decisions, and know when to pause or escalate.
What are the core components of Context Engineering?
The four core components of Context Engineering are: Semantic context (business meaning and definitions) Lineage context (end-to-end data flow and dependencies) Operational context (data quality and reliability signals) Policy context (privacy, compliance, and usage constraints) Together, these form a unified context layer that supports enterprise decision-making and AI automation
How should enterprises prepare for Context Engineering?
Enterprises should follow a phased approach: Inventory critical data and trust gaps Unify metadata, lineage, quality, and policy into a single context layer Expose context through APIs for AI agent consumption By 2026, this foundation will be essential for deploying Agentic AI at scale with confidence and auditability.
How do you measure the ROI of a data catalog?
ROI is measured by comparing the quantifiable benefits (such as reduced data search time, fewer data quality issues, and lower compliance effort) against the total costs (implementation, licensing, and support). Typical metrics include time savings, productivity gains, and compliance cost reduction.
What is a data catalog and why is it important for ROI?
A data catalog is a centralized inventory of data assets enriched with metadata that helps users find, understand, and trust data across an organization. It improves data discovery, reduces search time, and enhances collaboration — all of which contribute to measurable ROI by cutting operational costs and accelerating insights.
How quickly can businesses see ROI after implementing a data catalog?
Time-to-value varies with deployment and adoption, but many organizations begin seeing measurable improvements in days to months, especially through faster data discovery and reduced compliance effort. Early wins in these areas can quickly justify the investment.
What factors should you include when calculating the ROI of a data catalog?
When calculating ROI, include: Implementation and training costs Recurring maintenance and licensing fees Savings from reduced data search and rework Compliance cost reductions Productivity and decision-making improvements This ensures a holistic view of both costs and benefits.
How does a data catalog support data governance and compliance ROI?
A data catalog enhances governance by classifying data, enforcing rules, and providing transparency. This reduces regulatory risk and compliance effort, leading to direct cost savings and stronger data trust.
What is data lineage?
Data lineage shows where data comes from, how it moves, and how it changes across systems. It helps teams understand the full journey of data—from source to final reports or AI models.
Why is data lineage important for modern data teams?
Data lineage builds trust in data by making it transparent and explainable. It helps teams troubleshoot issues faster, assess impact before changes, meet compliance requirements, and confidently use data for analytics and AI.
What are the different types of data lineage?
Common types of data lineage include: Technical lineage – Tracks data movement at table and column level. Business lineage – Connects data to business definitions and metrics. Operational lineage – Shows how pipelines and jobs process data. End-to-end lineage – Combines all of the above across systems.
Is data lineage only useful for compliance?
No. While data lineage is critical for audits and regulatory compliance, it is equally valuable for debugging data issues, impact analysis, cost optimization, and AI readiness.
How does data lineage help with data quality?
Data lineage helps identify where data quality issues originate and which reports or dashboards are affected. This reduces time spent on root-cause analysis and improves accountability across data teams.
What is Metadata Management?
Metadata management involves the management and organization of data about data to enhance data governance, data asset quality, and compliance.
What are the key points of Metadata Management?
Metadata management involves defining a metadata strategy, establishing roles and policies, choosing the right metadata management tool, and maintaining an ongoing program.
How does Metadata Management work?
Metadata management is essential for improving data quality and relevance, utilizing metadata management tools, and driving digital transformation.
Why is Metadata Management important for businesses?
Metadata management is important for better data quality, usability, data insights, compliance adherence, and improved accuracy in data cataloging.
How should companies evolve their approach to Metadata Management?
Companies should manage all types of metadata across different environments, leverage intelligent methods, and follow best practices to maximize data investments.
What is a data definition example?
A data definition example could be: “Customer: a person or entity that has made at least one purchase within the past year.” It clearly sets business meaning and inclusion criteria.
Why is data definition important in data governance?
It ensures everyone interprets data consistently, reducing ambiguity and improving compliance, reporting, and collaboration.
Who should own data definitions?
Ownership should be shared between business domain experts (for context) and data stewards (for technical accuracy).
How often should data definitions be reviewed?
Ideally quarterly or whenever there’s a structural change in business logic, data models, or product offerings.
What’s the difference between data definition and data catalog?
A data catalog inventories data assets; data definition explains what those assets mean. Combined, they create full visibility and trust.
Why is Data Lineage important for businesses?
Data Lineage provides transparency and trust in your data ecosystem. It helps organizations ensure data accuracy, simplify root-cause analysis during data quality issues, and maintain compliance with regulations like GDPR or SOX. By understanding data flows, teams can make faster, more reliable decisions and improve overall data governance.
What are the key components of Data Lineage?
The main components of Data Lineage include: Data Sources: Where the data originates (databases, APIs, files). Transformations: How data is processed or modified. Data Pipelines: The tools or systems that move data. Destinations: Where the data is stored or consumed (dashboards, reports, models). Metadata: The contextual details that describe each step in the data’s lifecycle.
How does Data Lineage support Data Governance and AI readiness?
Data Lineage acts as the foundation for strong data governance by providing visibility into data ownership, transformation logic, and usage. For AI initiatives, lineage ensures that models are trained on accurate and traceable data, making AI outputs more explainable and trustworthy. Platforms like Decube’s Data Trust Platform unify lineage with data quality and metadata management to help enterprises achieve AI readiness.
What tools are commonly used for Data Lineage?
Several tools help automate and visualize data lineage, such as Decube, Atlan, Alation, Collibra, and OpenLineage. These tools connect to data warehouses, ETL pipelines, and BI tools to automatically map relationships between datasets — saving time and reducing manual effort.
What is Data Lineage?
Data Lineage is the process of tracking how data moves and transforms across an organization — from its origin to its final destination. It shows where data comes from, how it changes through different systems or pipelines, and where it ends up being used. In short, data lineage helps you visualize the journey of your data.
What does “data context” mean?
Data context refers to the semantic, structural, and business information that surrounds raw data. It explains what data means, where it comes from, who owns it, and how it should be used.
What is a centralized LLM framework?
It’s an enterprise-wide system where all departments access AI through a shared platform, equipped with guardrails, context layers, and multimodal capabilities.
What are guardrails in AI?
Guardrails are controls—policies, access restrictions, and compliance checks—that ensure AI outputs are secure, ethical, and aligned with enterprise goals.
How does data context affect ROI in AI?
Models trained or prompted with contextualized data deliver outputs that are relevant, trustworthy, and actionable—leading to faster adoption and higher business value.
What is MCP (Model Context Protocol) and why does it matter?
MCP defines how models interact with external tools and data sources. Feeding it with strong context ensures the AI agent can act accurately and responsibly.
What is a Data Trust Platform in financial services?
A Data Trust Platform is a unified framework that combines data observability, governance, lineage, and cataloging to ensure financial institutions have accurate, secure, and compliant data. In banking, it enables faster regulatory reporting, safer AI adoption, and new revenue opportunities from data products and APIs.
Why do AI initiatives fail in Latin American banks and fintechs?
Most AI initiatives in LATAM fail due to poor data quality, fragmented architectures, and lack of governance. When AI models are fed stale or incomplete data, predictions become inaccurate and untrustworthy. Establishing a Data Trust Strategy ensures models receive fresh, auditable, and high-quality data, significantly reducing failure rates.
What are the biggest data challenges for financial institutions in LATAM?
Key challenges include: Data silos and fragmentation across legacy and cloud systems. Stale and inconsistent data, leading to poor decision-making. Complex compliance requirements from regulators like CNBV, BCB, and SFC. Security and privacy risks in rapidly digitizing markets. AI adoption bottlenecks due to ungoverned data pipelines.
How can banks and fintechs monetize trusted data?
Once data is governed and AI-ready, institutions can: Reduce OPEX with predictive intelligence. Offer hyper-personalized products like ESG loans or SME financing. Launch data-as-a-product (DaaP) initiatives with anonymized, compliant data. Build API-driven ecosystems with partners and B2B customers.
What is data dictionary example?
A data dictionary is a centralized repository that provides detailed information about the data within an organization. It defines each data element—such as tables, columns, fields, metrics, and relationships—along with its meaning, format, source, and usage rules. Think of it as the “glossary” of your data landscape. By documenting metadata in a structured way, a data dictionary helps ensure consistency, reduces misinterpretation, and improves collaboration between business and technical teams. For example, when multiple teams use the term “customer ID”, the dictionary clarifies exactly how it is defined, where it is stored, and how it should be used. Modern platforms like Decube extend the concept of a data dictionary by connecting it directly with lineage, quality checks, and governance—so it’s not just documentation, but an active part of ensuring data trust across the enterprise.
What is an MCP Server?
An MCP Server stands for Model Context Protocol Server—a lightweight service that securely exposes tools, data, or functionality to AI systems (MCP clients) via a standardized protocol. It enables LLMs and agents to access external resources (like files, tools, or APIs) without custom integration for each one. Think of it as the “USB-C port for AI integrations.”
How does MCP architecture work?
The MCP architecture operates under a client-server model: MCP Host: The AI application (e.g., Claude Desktop or VS Code). MCP Client: Connects the host to the MCP Server. MCP Server: Exposes context or tools (e.g., file browsing, database access). These components communicate over JSON‑RPC (via stdio or HTTP), facilitating discovery, execution, and contextual handoffs.
Why does the MCP Server matter in AI workflows?
MCP simplifies access to data and tools, enabling modular, interoperable, and scalable AI systems. It eliminates repetitive, brittle integrations and accelerates tool interoperability.
How is MCP different from Retrieval-Augmented Generation (RAG)?
Unlike RAG—which retrieves documents for LLM consumption—MCP enables live, interactive tool execution and context exchange between agents and external systems. It’s more dynamic, bidirectional, and context-aware.
What is a data dictionary?
A data dictionary is a centralized repository that provides detailed information about the data within an organization. It defines each data element—such as tables, columns, fields, metrics, and relationships—along with its meaning, format, source, and usage rules. Think of it as the “glossary” of your data landscape. By documenting metadata in a structured way, a data dictionary helps ensure consistency, reduces misinterpretation, and improves collaboration between business and technical teams. For example, when multiple teams use the term “customer ID”, the dictionary clarifies exactly how it is defined, where it is stored, and how it should be used. Modern platforms like Decube extend the concept of a data dictionary by connecting it directly with lineage, quality checks, and governance—so it’s not just documentation, but an active part of ensuring data trust across the enterprise.
What is the purpose of a data dictionary?
The primary purpose of a data dictionary is to help data teams understand and use data assets effectively. It provides a centralized repository of information about the data, including its meaning, origins, usage, and format, which helps in planning, controlling, and evaluating the collection, storage, and use of data.
What are some best practices for data dictionary management?
Best practices for data dictionary management include assigning ownership of the document, involving key stakeholders in defining and documenting terms and definitions, encouraging collaboration and communication among team members, and regularly reviewing and updating the data dictionary to reflect any changes in data elements or relationships.
How does a business glossary differ from a data dictionary?
A business glossary covers business terminology and concepts for an entire organization, ensuring consistency in business terms and definitions. It is a prerequisite for data governance and should be established before building a data dictionary. While a data dictionary focuses on technical metadata and data objects, a business glossary provides a common vocabulary for discussing data.
What is the difference between a data catalog and a data dictionary?
While a data catalog focuses on indexing, inventorying, and classifying data assets across multiple sources, a data dictionary provides specific details about data elements within those assets. Data catalogs often integrate data dictionaries to provide rich context and offer features like data lineage, data observability, and collaboration.
What challenges do organizations face in implementing data governance?
Common challenges include resistance from business teams, lack of clear ownership, siloed systems, and tool fragmentation. Many organizations also struggle to balance strict governance with data democratization. The right approach involves embedding governance into workflows and using platforms that unify governance, observability, and catalog capabilities.
How does data governance impact AI and machine learning projects?
AI and ML rely on high-quality, unbiased, and compliant data. Poorly governed data leads to unreliable predictions and regulatory risks. A governance framework ensures that data feeding AI models is trustworthy, well-documented, and traceable. This increases confidence in AI outputs and makes enterprises audit-ready when regulations apply.
What is data governance and why is it important?
Data governance is the framework of policies, ownership, and controls that ensure data is accurate, secure, and compliant. It assigns accountability to data owners, enforces standards, and ensures consistency across the organization. Strong governance not only reduces compliance risks but also builds trust in data for AI and analytics initiatives.
What is the difference between a data catalog and metadata management?
A data catalog is a user-facing tool that provides a searchable inventory of data assets, enriched with business context such as ownership, lineage, and quality. It’s designed to help users easily discover, understand, and trust data across the organization. Metadata management, on the other hand, is the broader discipline of collecting, storing, and maintaining metadata (technical, business, and operational). It involves defining standards, policies, and processes for metadata to ensure consistency and governance. In short, metadata management is the foundation—it structures and governs metadata—while a data catalog is the application layer that makes this metadata accessible and actionable for business and technical users.
What features should you look for in a modern data catalog?
A strong catalog includes metadata harvesting, search and discovery, lineage visualization, business glossary integration, access controls, and collaboration features like data ratings or comments. More advanced catalogs integrate with observability platforms, enabling teams to not only find data but also understand its quality and reliability.
Why do businesses need a data catalog?
Without a catalog, employees often struggle to find the right datasets or waste time duplicating efforts. A data catalog solves this by centralizing metadata, providing business context, and improving collaboration. It enhances productivity, accelerates analytics projects, reduces compliance risks, and enables data democratization across teams.
What is a data catalog and how does it work?
A data catalog is a centralized inventory that organizes metadata about data assets, making them searchable and easy to understand. It typically extracts metadata automatically from various sources like databases, warehouses, and BI tools. Users can then discover datasets, understand their lineage, and see how they’re used across the organization.
What are the key features of a data observability platform?
Modern platforms include anomaly detection, schema and freshness monitoring, end-to-end lineage visualization, and alerting systems. Some also integrate with business glossaries, support SLA monitoring, and automate root cause analysis. Together, these features provide a holistic view of both technical data pipelines and business data quality.
How is data observability different from data monitoring?
Monitoring typically tracks system metrics (like CPU usage or uptime), whereas observability provides deep visibility into how data behaves across systems. Observability answers not only “is something wrong?” but also “why did it go wrong?” and “how does it impact downstream consumers?” This makes it a foundational practice for building AI-ready, trustworthy data systems.
What are the key pillars of Data Observability?
The five common pillars include: Freshness, Volume, Schema, Lineage, and Quality. Together, they provide a 360° view of how data flows and where issues might occur.
What is Data Observability and why is it important?
Data observability is the practice of continuously monitoring, tracking, and understanding the health of your data systems. It goes beyond simple monitoring by giving visibility into data freshness, schema changes, anomalies, and lineage. This helps organizations quickly detect and resolve issues before they impact analytics or AI models. For enterprises, data observability builds trust in data pipelines, ensuring decisions are made with reliable and accurate information.

Table of Contents

Read other blog articles

Grow with our latest insights

Sneak peek from the data world.

Thank you! Your submission has been received!
Talk to a designer