All articles

Is your data ready for AI?

You do not need perfect data to start with AI, but you do need to know its state. What readiness means in practice, the hidden work of getting there, and why the honest assessment comes before the model.

Someone has told you your company needs AI, and your first honest thought was about your data — that it is scattered, inconsistent, half of it in spreadsheets, and probably not ready. That worry is well placed, and it is also the most useful instinct you can bring to an AI project. Because the uncomfortable truth is that most AI efforts do not fail on the model. They fail on the data underneath it, long before anyone gets to the interesting part.

AI projects stall on data, not models

The models are, for most business problems, the easy part now. They are available, well-documented, and largely someone else's problem to build. What no vendor can hand you is your own data in a usable state. That is why so many projects that start with excitement end in a quiet stall: months in, the team is still cleaning, reconciling, and hunting for records, and the impressive demo from week two turns out to have run on a tidy sample that does not resemble reality.

This is not a reason to avoid AI. It is a reason to look at your data first, deliberately, and with clear eyes — before you commit a budget to a model that will only ever be as good as what you feed it.

What readiness actually means

Readiness is not a single yes or no. It is a handful of separate properties, and a dataset can be strong on some and weak on others. Accessible: can you actually get to the data, through a system rather than by asking a colleague to export it by hand each month? Reasonably clean: are the fields consistent, the duplicates manageable, the obvious errors not overwhelming? Not perfect — reasonable.

Labeled where needed: if you want the system to learn a distinction — spam or not, urgent or not, this category or that — do examples of that distinction exist, marked as such? Permitted to use: are you allowed to use this data for this purpose, under the terms you collected it and the law that governs it? And documented: does anyone know what the fields mean, where they came from, and which ones to trust? A dataset can be technically accessible and still useless because no one alive remembers what column fourteen represents.

The reason to separate these properties rather than ask a single question is that they fail independently and are fixed independently. Data can be spotlessly clean and completely off-limits for the purpose you have in mind. It can be perfectly permitted and hopelessly undocumented. Collapsing them into one nervous yes-or-no is how a project either freezes over a problem that does not exist or charges ahead into one that does. Rate them one at a time.

The hidden work of collecting and structuring it

Between raw data and ready data sits a body of work that rarely appears in the plan and almost always dominates the effort. Data lives in different systems that name the same customer three different ways. It arrives in formats that were never meant to be joined. The one field you need most turns out to have been optional, so it is empty for half the history. None of this is glamorous, and all of it has to happen before a model can learn anything worth acting on.

The mistake is to treat this as a one-time chore. Getting to ready is not an event; it is plumbing you will keep. If the data was messy because the process that produces it is messy, cleaning it once buys you a clean snapshot and a mess again next quarter. The durable fix is usually upstream — in how the data is captured — not in a heroic cleanup at the end.

There is also the question of who does this work. It rarely fits neatly into either the data team or the business, because it needs both — the technical skill to move and reshape data, and the domain knowledge to know what a field is supposed to mean and which anomalies are errors rather than reality. Projects stall when this falls into the gap between the two, owned by neither. Naming a person who holds both ends, or a pair who cover them together, is often what separates the readiness effort that finishes from the one that circles.

You do not need perfect data — you need to know its state

Here is the reassuring part. Waiting for perfect data is its own kind of failure; you will wait forever, because real data is never finished. Plenty of valuable AI runs on data that is imperfect but honestly understood. What you cannot afford is to not know the state of your data — to promise a system accuracy your records cannot support, or to discover the gap only after you have built on top of it.

So the goal of an early assessment is not to make the data perfect. It is to produce an honest map: what you have, how good it is, where the holes are, and what a given use case actually requires of it. With that map, a use case that needs more than you have becomes a data project first and a model project second — sequenced deliberately, not discovered by accident three months in.

Knowing the state of your data also changes how you set expectations with everyone waiting on the result. A model trained on data that is seventy percent complete will behave like a model trained on data that is seventy percent complete, and the only real failure is being surprised by that. When you know the ground you are standing on, you can promise what the system will actually do, design a human check where the data is thin, and improve deliberately — instead of overpromising on a foundation you never inspected.

Governance, privacy, and lineage

Data that touches people carries obligations, and AI does not suspend them. Under the rules most European companies live by, you need a lawful basis to use personal data for a new purpose, and training a model is a new purpose. Before data flows into a system, it is worth knowing what is personal, what is sensitive, and what should never have left the system it came from in the first place.

Two disciplines make this manageable rather than paralysing. Privacy by design means deciding up front what you genuinely need and leaving the rest out, because the data you do not hold cannot leak. Lineage means being able to answer, for any figure the system produces, where it came from and what fed it. Lineage feels like bureaucracy until the first time someone asks why the model said what it said — and then it is the difference between an answer and a shrug.

The signals that you are not ready yet

You rarely get a clean verdict; you get symptoms. The clearest one is that no single person can tell you where a given number comes from without asking three colleagues — a sign the data has no owner and no documented lineage. Another is that every report requires a manual export and a spreadsheet to reconcile, which means the data is not accessible in any way a system could rely on. A third is that the field you would most want a model to learn from is the one people fill in inconsistently or leave blank, because it was never enforced at the point of entry.

None of these is fatal, and none of them means you cannot start. They are a map of what to fix, and in what order, for the specific use case you have in mind. The danger is not having these problems; nearly everyone does. The danger is not knowing you have them, and promising a model a foundation that is not there.

Treat each signal as a small, scoped task rather than a reason for despair. The company that names its data problems can sequence them. The company that insists its data is fine — right up until the model produces nonsense — pays for the same problems later, and with less goodwill.

A pragmatic path to readiness

The way through is not a data warehouse you build for two years before touching AI. It is narrower and faster. Start from one concrete use case, not from your data in general, because the use case tells you which data actually matters and lets you ignore the rest for now. Assess that slice honestly against the properties above. Fix what blocks this use case and defer what does not.

Then improve at the source where you can, so the same cleanup does not return every quarter, and document as you go so the next project inherits knowledge instead of a mystery. Readiness built this way compounds: each use case leaves the ground firmer for the next. It is slower to feel impressive and far faster to reach something real — which, when the goal is a system you can trust, is the only speed that counts.

There is discipline in refusing to boil the ocean. Once you start looking at data, it is tempting to want to fix all of it — standardise every field, reconcile every system, and build the warehouse that will finally make everything tidy. That project has swallowed years at many companies and delivered no working AI at the end of them. The use-case-first path is deliberately narrower, and narrower is what ships.

Want to know if your data is ready?

A short, fixed-fee data assessment gives you an honest map of what you have, where the gaps are, and exactly what one target use case needs before a model is worth building.

Get a data assessment