Connecting an artificial intelligence model to flawed data doesn't improve the data; it improves how its errors are presented. The system will return a well-written, coherent, and confident response built on incomplete or contradictory information, and that confidence is precisely what makes the problem more dangerous than before.
When a report came from a spreadsheet, someone looked at it with suspicion. When it comes from an AI system, it's discussed in a meeting.
Almost no medium-sized company lacks data. What almost all of them have is the same information in several places with different values: the customer in the CRM and in the ERP with different addresses, the margin calculated in two ways depending on the department, the status of an order that depends on who you ask.
A model does not resolve this contradiction. It chooses one of the versions—depending on how the query is constructed and which document it retrieves first—and responds without realizing that there was another one.
The three symptoms that always appear in a diagnosis:
There is no owner of the data. No one is responsible for ensuring that the "active client" field means the same thing in all systems.
There is no shared definition. «"Sale" can mean order, invoice, or payment. Three correct figures and three different conclusions.
There is no reliable history. The changes are overwritten, so it's impossible to reconstruct what was true in March.
Without addressing those three points, any analytical or AI layer built on top will inherit the problem and amplify it.
Level | What does it solve? | Without this, the next one won't work. |
|---|---|---|
1. Data Property | Who is responsible for each piece of data? | Nobody corrects anything |
2. Shared definition | What does each metric mean? | The numbers are discussed instead of used |
3. Single source of truth | Which system rules | The answers depend on where you look. |
4. History and traceability | What was true and when | There is no audit or time comparison |
5. Analytics and AI | Interpretation and prediction | It is built on sand |
The common temptation is to start with level 5, because it's the most visible and the one that can be shown. The result is a pretty panel whose numbers no one uses for decision-making, because every department has reasons to distrust them.
The reverse order is less flashy and produces value from the first level: when someone owns a piece of data and its definition is written down, discussions about figures disappear even if no model has been connected.
While the data was feeding into reports, an error was detected in the meeting. When they feed into systems that they perform actions, The error becomes an action performed: an order created, a customer contacted, an amount applied.
McKinsey noted in 2025 that the factor most correlated with EBIT impact is not the chosen model but the profound redesign of workflows, something only 211TP3 of the adopters had done. Redesigning a workflow requires defining which data governs it, and that's where the work that almost no one wants to do comes in.
It is the same conclusion, seen from another angle, that we present in AI won't fix a poorly designed company
The expression *AI-ready* is used very loosely. In practice, it means five verifiable things:
Each critical piece of data has a declared source system. And the others consult it, they don't rewrite it.
The definitions are written out and accessible. A short glossary that explains what a customer, a sale, and a closed, business-approved case are.
The data has a history. It is not overwritten: it is versioned. It allows for reconstruction and auditing.
The permissions are in the data, not in the application. If access control resides on the screen, any system that queries it from below will bypass it. That's exactly what an agent does.
There is a way to measure quality. Percentage of incomplete records, duplicates detected, discrepancies between systems. Without metrics, there is no improvement. An AI project built on a foundation that meets these five conditions has a radically different probability of success than one that does not. We have developed it in What does it mean to be an AI-ready company?
The usual reaction to this diagnosis is to propose a corporate data governance project. It typically lasts eighteen months, produces a document, and doesn't change operations.
The approach that works is the opposite: Start with the data that supports the process that explains the most.. Order that one, measure it, and use it. Then the next one. Each iteration produces a usable result and funds the next one.
A data project that doesn't improve a specific decision in less than three months probably won't improve any decisions.
Before approving an investment in analytics or AI, there is a test that costs half an hour: requests the same amount from three different departments.
If they match, the foundation is better than usual and can be built upon. If they don't match—and they usually don't—you already know what the first project is, and it's not artificial intelligence.
No. A model dealing with conflicting data chooses one of the available versions and responds confidently, without acknowledging that another version existed. The result is an error that is better presented and, therefore, harder to detect than when it came from a spreadsheet.
Five verifiable conditions: each critical data item has a declared source system, business definitions are written and shared, there is versioned history instead of overwriting, permissions reside in the data and not the application, and there are data quality metrics.
Because of the data that underpins the process that explains the most margin, not because of a corporate data governance project. Organizing an area, measuring it, and using it produces results in weeks and funds the next iteration; global projects usually end up as documentation.
Because they are built before resolving data ownership, shared definitions, and the single source of truth. When each department can justify a different figure, meetings are spent arguing about the numbers instead of making decisions based on them.
The error ceases to be a mistaken report and becomes an action taken: an order created, a customer contacted, or an amount applied. Automation transforms an information problem into an operational one.
Asking three different departments for the same figure. If the answers don't match, the first project needed isn't one of artificial intelligence, but of data definition and ownership.
Could your data withstand a layer of AI on top? We assess ownership, definitions, sources of truth, and quality before making any proposals. Two-hour analysis, no obligation. Let's talk → |