The Moment It Shifts
There’s a moment that happens in almost every early conversation I have around implementing AI tools. A chemical manufacturer or distributor has made the decision that they want to use AI, for product search, technical recommendations, customer support, whatever the use case, and I sit down to look at their product data together. Partway through, something shifts. A quiet recognition that the data they’ve been relying on for years looks very different when a system needs to use it versus when people do.
One example that sticks with me is a specialty chemical distributor managing around 10,000 products across roughly 1,700 suppliers, an experienced team, a well-run operation, and real commercial momentum behind the AI initiative. When I pulled the product catalogue and mapped attribute coverage, the numbers were stark: somewhere north of 40% of their portfolio had no Technical Data Sheet, and many of the ones that did exist were outdated and often incorrectly branded due to company mergers and acquisitions. These weren’t documents missing from some central system, they were genuinely missing, or wrong. Some of those products had been sold for years. The sales team knew them. The application engineers knew them. But that knowledge lived in people, not in data.
But not all of that 40% mattered in the same way. Some of those missing sheets sat behind products the system would be asked about every day. Others were for products almost nobody queries any more. Including content for products that had been rebranded or discontinued. The gap that stops an AI system isn’t the size of the hole in the catalogue. It’s whether the system needs that particular piece of data to answer a real question. That distinction, which gaps actually block the system and which are just noise, is the thing most companies can’t see yet, and it’s what really decides how big their problem is.

A Pattern, Not an Exception
That’s the pattern. And it shows up, in different forms, at different scales, in nearly every implementation I’ve worked on.
Most organisations going into an AI project believe their data situation is basically fine. They have a product catalogue, a PIM or an ERP, PDFs of TDS documents stored somewhere, CAS numbers. From the outside, it looks like the raw data is there. It’s when you look closely that the picture becomes very different. Product attributes are incomplete, sometimes across a majority of the portfolio. Technical Data Sheets exist in three versions for the same product, none of them dated, none of them clearly authoritative. Lab data sits on device-specific software that has never been connected to any central system, accessed only by the people who ran the tests. Regulatory documentation is scattered across local drives, email threads, and the institutional memory of people who’ve been with the company for decades.
None of this is unusual. In the chemical industry, this is the norm. Products get added to portfolios through acquisitions, supplier relationships, and decades of incremental growth. The documentation accumulates the same way, organically, without a governing structure, because no prior system ever required it to be otherwise. The question isn’t whether companies have this problem. It’s whether they’ve looked, and whether they know which parts of it will actually matter.
Why AI Can’t Fake It
The reason none of this gets audited isn’t negligence. It’s that no prior system ever demanded the data be in an AI-usable state. A skilled sales rep handles a missing TDS by calling the application team. A customer service agent compensates for incomplete attributes by pulling in a colleague who knows the product. Human intelligence fills the gaps constantly, invisibly, and the underlying data quality never becomes a visible constraint. Customers get answers, orders get placed, relationships hold.
An AI system cannot do this. It cannot call the application team. It cannot read context, make educated guesses, or paper over inconsistencies with experience. When the data is incomplete, the system either fails or fabricates, and both outcomes are worse than no system at all. This is why implementation functions as an audit in a way that no internal process ever did. It’s the first time the data has been stress-tested against something that requires it to actually be complete, structured, and consistent.
But the audit does something more useful than prove the data is flawed. It ranks it. A live system fails loudly on the gaps it actually needs filled, and stays silent on the ones it never touches. That ranking is the thing no spreadsheet review can give you, because until a system is trying to answer real questions, every gap looks equally urgent, or equally ignorable. The build is what tells you which is which.

The discovery is painful, but valuable. Roadmaps slip, scopes get adjusted, and stakeholders who expected to be selecting AI features find themselves discussing data governance instead. But you now know what you’re actually working with, and which parts of it are worth your attention. The AI project didn’t create the problem, it surfaced one that was already there, already affecting the business, just never visible enough to address.
Fixing Is Not the Same as Cleaning
Fixing the data problem is not the same as cleaning up the data. Simply cleaning up the data does not solve the problem.
This distinction matters because most organisations, when they realise they have a data quality issue, reach for a surface-level response: export the catalogue, find the gaps, fill holes. That produces a cleaner spreadsheet. It doesn’t produce AI-ready data, and it treats every gap as equally worth filling when only some of them were ever going to stop the system. Two months later, new products get added the same way old ones were. New TDS documents get uploaded without version control. The underlying structure, the one that generated the problem in the first place, is still there.
What the Build Forces Into Place
AI-ready product data is different in kind, not just degree, and you don’t get there by deciding to in advance. You get there because the system keeps demanding it. Attribute completeness matters, but completeness without consistency is still a problem: if flash point is recorded in different units across different product lines, or if the same property is captured under different field names by different teams, the system cannot use it reliably. Document version control matters, not so that files are tidy, but so that there is a single authoritative answer to the question “what is the current specification for this product?” A single source of truth matters, not as a phrase but as a functional reality: one place where product data lives, governed in a way that makes updates traceable and errors correctable.

The underlying question isn’t “do we have the data?” It’s “is our data structured, consistent, and governed well enough for a system to use it without human interpretation at every step?” In most cases, the answer is not yet. Getting there requires decisions about ownership, tooling, and process, and those decisions are far easier to make when a working system is showing you exactly where they matter than when you’re guessing at them in the abstract.
The Wrong Question, Asked First
Most organisations evaluating AI right now are asking the wrong question first. They’re comparing vendors, reviewing demos, and discussing accuracy benchmarks, when the more pressing question is what their own data actually looks like.
The vendors can wait. The data cannot. Until you understand what you’re working with, any capability you’re evaluating is hypothetical. The AI system will be as good as the data it runs on. If that data is incomplete, inconsistent, or ungoverned, the capability gap you’re trying to close with AI will be filled instead with a different kind of failure, one that’s harder to explain to the organisation that signed off on the budget.
Quality In, Trust Out
AI accuracy, reliability, and organisational trust in the outputs are dependent on the quality of the inputs. That’s not a technical observation, it’s an organisational one. The companies that will get real value from AI in the chemical industry are the ones that treat data readiness as something the build itself will force into the open, not a checklist to clear before it starts.
Readiness is a by-product of building AI, not a precondition. Starting is the only thing that has ever forced the data to be right.
I’ve watched companies try to prepare their data in the abstract: a cleanup sprint, a governance committee, a spreadsheet and a deadline. It rarely works, because nothing forces the real gaps into view until a system actually needs that data to function. The moment a product search fails because a TDS is missing, or a recommendation breaks because a unit doesn’t match, the priority becomes obvious in a way no audit could have made it. That’s not a failure of planning. That’s the process working exactly as it should, sorting the gaps that matter from the ones that don’t.
This is why I don’t ask clients to fix their data before we build. I ask them to start building, with someone who has seen this pattern enough times to know which gaps will actually stop the system and which ones are noise. Readiness isn’t a ticket you buy before the AI project starts. It’s what building leaves you with.







