Data Readiness

Most stalled AI programmes were not undone by the model. They were undone by conditions that were visible, and checkable, months before anyone signed the licence.

There is a particular meeting I have had more times than I can count. A team has bought something good — a licence, a platform, a model with a name everyone recognises. Six or eight months in, the pilot has not converted. Someone senior wants to know whether they bought the wrong tool.

Almost always, they did not. They bought a perfectly capable tool and pointed it at data that could not support what they asked of it. The tool is not the variable. It was never going to be.

The uncomfortable part is that this is usually knowable in advance. Not in a vague, "data quality matters" way, but concretely — there are specific conditions you can look for before you spend, and they show up early. Here are five of them. They are the ones I see most often, in Accra and in Atlanta, in banks and in ministries and in mid-market firms that thought they were further along than they were.

SIGN 01

Nobody can tell you where the number comes from

Ask for the definition of one number that matters. Active customer. Enrolled beneficiary. Open case. Ask three people in three different functions and see whether you get the same answer.

You often do not. Finance counts an active customer one way because of how revenue is recognised. Operations counts differently because of how the queue is worked. The reporting team has a third definition baked into a query someone wrote four years ago and left when they moved on.

None of these people are wrong. They are optimising for their own function, which is what they are paid to do. The problem is that a model does not know there are three definitions. It will pick up whichever one is encoded in the table it was given and treat it as truth, and the output will be confidently, invisibly wrong for two of the three audiences who read it.

The signal is not that the data is messy. The signal is that nobody owns the definition.

Where there is no owner, there is no correct answer to converge on — only the answer that happened to be in the file.

What to look at

Pick your five most-used metrics. For each, name the person accountable for the definition. If you cannot name a person, you have found the gap.

SIGN 02

Your history has been overwritten

This one is quiet, and it is the one that kills prediction work most reliably.

A great many operational systems are built to answer one question: what is true right now? A customer record holds the current address, the current risk tier, the current status. When something changes, the field is updated in place. That is correct design for running the business. It is fatal for anything that needs to learn from the past.

Because the moment you want to predict churn, or default, or drop-off, you need to know what was true at the time — what the tier was before the account went bad, not what it is now that someone has already reclassified it. If the record was overwritten, that history is gone. It is not recoverable through better modelling or a cleverer prompt. It is simply not there.

I have watched teams lose a full quarter discovering this, having already committed to the project. The prototype looked extraordinary in testing, because the current-state fields contained information about outcomes that had already happened. Then it went live against genuinely unknown cases and performed no better than the rule of thumb it was meant to replace.

What to look at

Take one important entity — a customer, a loan, a case — and try to reconstruct its state as of a date twelve months ago. If you can't, prediction is off the table until you can.

SIGN 03

The critical context lives outside the system

Every organisation I have worked with has a version of this. The system holds the transaction. The reason for the transaction is in an email, a WhatsApp thread, a spreadsheet on someone's laptop, or in the head of a person who has been there eleven years.

This is not a failure of discipline. It is usually a rational response to a system that made the right entry hard and the workaround easy. But it means the recorded data is a shadow of the actual process, and the parts that were left out are almost never random. They are the exceptions, the escalations, the judgement calls — which is to say, precisely the parts you were hoping to automate.

Train on what was recorded and you get a system that handles the easy cases you already handle fine, and falls over on the hard ones. Then someone has to check its work, and the checking costs as much as the doing did.

There is a version of this that is specific to institutions operating across multiple markets: the process genuinely differs by country, and the system only models the headquarters version. Everything local — the extra approval, the informal verification step, the workaround for a payment rail that behaves differently — sits outside.

A model built on the central system is a model of one market wearing the name of all of them.

What to look at

Sit with someone who does the work for two hours and count the times they consult something that is not the system of record.

SIGN 04

Access takes weeks and requires a favour

Here is a test that has nothing to do with data quality and predicts readiness better than most quality metrics. Request access to a dataset you don't currently touch. Not a copy — access. Then time it.

If the answer arrives in days through a defined route, you have governance. If it takes weeks, requires knowing the right person, and ends with a file being emailed to you, you don't have a data problem so much as an organisational one. And it will surface in every single iteration of every AI project you attempt, because this work is not one request. It is dozens, over months, as the question changes and the requirements sharpen.

Teams consistently underestimate this. They budget for the model and forget that most of the calendar goes to obtaining, understanding, and re-obtaining data. A twelve-week project with a four-week access cycle is not a twelve-week project.

The favour-based version is worse than the slow version, incidentally. Slow is at least predictable. Favour-based means access disappears the day the helpful person changes role, and nobody notices until a pipeline breaks.

What to look at

The elapsed time on your last three access requests — and whether a named process or a named person made them happen.

SIGN 05

There is no baseline, so there is nothing to prove

Ask what the current process costs, how long it takes, and how often it is wrong. If the room goes quiet, stop before you buy anything.

Without a baseline you cannot demonstrate improvement, which means the pilot cannot succeed on evidence — only on impression. And impression fades. The sponsor who was enthusiastic in month one is asked in month nine what the return was, and there is no defensible answer, so the programme quietly loses its budget to something that can produce a number.

This is the most fixable item on the list and the most frequently skipped, because measuring the current state is unglamorous and it sometimes surfaces facts that are awkward for whoever owns the process. Do it anyway. Measure before, not after. Afterwards it is not a baseline; it is a reconstruction, and everyone in the room knows it.

What to look at

For the process you intend to change, write down today's cost, cycle time, and error rate. If you cannot fill in all three, that is your first project.

The thing worth internalising

None of these five are about the model. That is the point. The capability question — can a system do this at all — has largely moved. What remains is whether your organisation can supply the conditions under which that capability does anything useful, and those conditions are organisational as much as technical. Definitions have owners. History is preserved. Context is captured. Access is routine. Performance is measured.

The good news is that this work does not expire. A better model next year does not make an undefined metric defined, and it does not restore history that was overwritten. Everything you fix on this list keeps paying, whatever you end up buying.

The order matters less than starting. But if I had to pick one to do this month: go and find out whether you can reconstruct last year. That is the one that cannot be fixed retroactively, and every month you wait, another year of it is gone.

Want a structured read on where your organisation sits against these conditions?

Take the AI Data Readiness Score →