Get in touch

Have a project in mind? Tell us a bit about it.

Enquiry Form

Every AI project postmortem sounds the same once you get past the surface story. The team blames the model, the vendor, or the timeline. Dig one layer down and the real cause is almost always the same thing: the data going into the project was never actually ready, and nobody checked before the build started.

This isn’t the technical problem people assume it is. It’s rarely about missing infrastructure or the wrong tool. It’s scattered, inconsistent, undocumented information that nobody owns, fed into a project timeline that never left room to fix any of that before asking a model to make sense of it.

1. What “not ready” actually looks like

Nobody kicks off an AI project by saying the data isn’t ready. They say it once the build stalls out. In practice, “not ready” rarely means there’s no data at all — it means three systems have three different names for the same customer field, half the support ticket history only exists as PDF exports someone made for an audit two years ago, and the “master” product list has four competing versions floating around in different people’s spreadsheets, none of them fully current.

None of that looks like a crisis from a distance. It only becomes visible once something is asked to actually read all of it at once and make consistent decisions from it.

2. The model isn’t the bottleneck — your plumbing is

Most AI project timelines are built around picking the model or the platform, as if that’s the hard part. It usually isn’t. The hard part is getting clean, consistent, current information into that model in a form it can use, and that work rarely gets its own line item. It gets discovered three weeks into the build, when someone finally tries to connect the real data and realizes half of it needs to be reshaped first.

Fix it: before scoping the AI part of the project, scope the data part separately, with its own timeline and its own owner. If nobody can say in one sentence where the source of truth for each key field lives, the project isn’t ready to be scoped yet.

3. Nobody owns the data, so nobody notices when it’s wrong

Data quality problems don’t usually get caught early because data quality usually isn’t anyone’s actual job. It’s a byproduct of five different people entering things slightly differently over several years, and nobody has the mandate to go back and reconcile it. Everyone downstream just works around the inconsistencies quietly, which is exactly why they stay invisible until an AI system tries to treat the data literally instead of working around it the way a person would.

Fix it: name one owner for each data source that will feed the project, even if that’s a part-time responsibility layered onto an existing role. An unowned dataset degrades slowly and silently. An owned one gets flagged the first time something looks off.

4. A clean pilot doesn’t prove the real data is ready

Pilots tend to succeed for a reason that has nothing to do with the technology: someone hand-picked and quietly cleaned the sample data to build the demo. That’s a reasonable way to prove a concept, but it’s a bad way to judge readiness, because it tests the AI against the best version of the data instead of the version it will actually face in production. The gap between the two is exactly where most “surprise” timeline overruns come from.

Fix it: before greenlighting the full build, run the pilot against a deliberately messy slice of real data — the oldest records, the ones with missing fields, the exceptions nobody wants to deal with. If it holds up there, it’s actually ready. If it doesn’t, better to find out on a small slice than after launch.

5. What “ready enough” actually means

Ready doesn’t mean perfect. Perfect data doesn’t exist in most businesses, and waiting for it is its own way of never starting. Ready means every key field has one agreed source of truth, someone is accountable for it, the known exceptions are documented instead of hidden, and the format is consistent enough that a system reading it programmatically won’t misinterpret it the way a person skimming it might forgive automatically.

Fix it: run a short, dedicated data audit — a couple of weeks, not a quarter — before the build starts, not in parallel with it. List every source the project will touch, who owns it, and what shape it’s actually in today. That audit is cheap compared to discovering the same gaps halfway through a build that’s already been scoped and staffed around a launch date.

The bottom line

When an AI project stalls or underdelivers, the postmortem almost always points at the model, the vendor, or an unrealistic timeline. Pull the thread further back and the real story is usually simpler: nobody checked whether the data was actually ready before the build started, because readiness isn’t a visible line item the way a model choice is. Businesses that get this right don’t have cleaner data by luck. They treat the audit as its own project phase, with its own owner and its own deadline, instead of an assumption baked silently into everyone else’s.