Clean Core, Trusted AI Part 2: The Data Pillar: Why Your AI Strategy Is Only as Good as Your Master Data
Part One closed on a warning: of the five Clean Core pillars, Data is the one that fails silently. Here’s what that actually looks like inside SAP, and what it takes to fix.
The New Blast Radius
Garbage-in-garbage-out is old news in ERP. A stale pricing condition record quoted the wrong discount, someone in Sales caught it before the invoice went out, and the fix stayed contained. AI removes that containment. A large language model or agent doesn’t know a condition record is wrong; it treats it as ground truth and reasons forward from there, recommending that discount, forecasting off it, explaining it to a customer, at the speed and scale of every interaction it touches from that point on. The mistake doesn’t stay a line item; it becomes the premise of a hundred downstream decisions before anyone notices. That’s the real stakes conversation: not “is our data accurate,” but “how fast does an inaccuracy compound once an agent starts acting on it.”

What “Governed” Actually Requires
“Clean data” isn’t a cleanup sprint activity or a mini-project. It’s five disciplines working together, not a checklist to complete once.
Strategy and governance are the foundation: clear data roles, a governance board, a data catalog, and increasingly Data Mesh or Data Product principles that treat data, customer, vendor, materials, finance master records, purchasing info records, pricing conditions and other conditional data as something owned and published, not just stored.
Quality is the accuracy layer: regular assessments and cleansing, not a one-time deduplication project.
Volume/lifecycle management and protection are the scale-and-trust layer of Master Data Governance driving standardization across custom data objects, retention and archiving policies (ILM) that keep the landscape performant, and access controls, encryption, and compliance audits that keep it defensible.
None of this is new. It’s the same ERP → data warehouse → data lake → lakehouse arc many of us have lived through for two decades, where every generation of technology promised to fix data quality and instead just moved the problem downstream and managed through exceptions.
AI is the first consumer that won’t tolerate that deferral.
Two scope notes, since AI’s data diet is bigger than this: configuration data (like plant, storage location, enterprise org structure, account hierarchies) is a related but distinct discipline with its own risks, and historical, IT/OT (sensor, PLC/SCADA, machinery telemetry), and streaming data flowing in and out of SAP fall more under the Integration pillar than the Data pillar. Both matter for AI trust — they’re just not what this post is about.

Why This Isn’t Just an SAP Problem
Swap “SAP S/4HANA” for whatever core system of record your organization runs, and the argument holds: any AI initiative inherits the data discipline (or lack of it) sitting underneath it. SAP customers just feel it first and hardest, because SAP is usually the system of record for the data AI touches most — pricing conditions, info records, customer, vendor, material and financial records tied to all of it. Get the quality and governance right there, and the rest of the enterprise data estate has a credible foundation to build on. Get it wrong, and every AI pilot inherits the same weak foundation, no matter how good the model is.
Which raises the real question for anyone holding a budget: if data governance is the prerequisite, not the AI tooling, what should you actually fund first? That’s Part Three.
