Thought leadership What has to be true first

Nobody questions data they do not trust

Before a business can usefully ask AI anything, somebody has to be able to disagree with the answer and find out who was right.

The phrase that gets used is a single source of truth, and it is usually sold as a destination: one platform, one dashboard, everybody looking at the same screen. That is not what makes data trustworthy. What makes data trustworthy is that when two people disagree about a number, there is a way to settle it.

Most operations cannot settle it. The number on the dashboard came from a weekly export that somebody reshaped in a spreadsheet, which was built from a report whose definition changed in March, and the person who made that change has left. So when the regional manager says the figure is wrong, the honest answer is that nobody can tell. The argument gets won by whoever is more senior, which is a poor way to run anything.

This is the real reason AI projects fail on data, and it is not the reason usually given. The problem is rarely that the data is missing. It is that the data cannot be checked, so nobody will stake a decision on it, so the model built on top of it is a toy regardless of how good the model is.

Making data checkable is unglamorous work. When exports arrive, you validate that they are complete before anything else happens, so a short file fails loudly instead of quietly becoming a low number. You hash and keep what arrived, so the question of what the source actually said has an answer that does not depend on memory. You load at the grain the business argues about — the order, the count, the ticket — rather than the weekly total, because an aggregate cannot be interrogated. And you reconcile: compare what one system claims against what another system paid, and make the gap visible rather than letting it average out.

On the aggregator work, this was the whole job. Two third-party aggregators sent weekly exports. Before, nobody could say per order what was late, missed or misreported. After, every store had order-level reporting every week, and the gap between reported sales and actual payouts was a figure somebody could take to a supplier meeting. The system did not become clever. It became arguable.

The same applies to the stock work, where the count moved off paper for 46 stores and 56 vendors. The value was not the time saved typing. It was that a variance finally had a before and an after that were recorded the same way, so a difference meant something.

There is a governance dividend that comes free with this, and under POPIA it is not optional. Knowing where a field came from, who touched it, and what left the country is the same discipline as knowing whether your numbers are right. Businesses tend to treat data governance as a cost imposed by regulation and data quality as a benefit they want. They are the same work.

The test I would apply before any AI project: pick the number your leadership team argues about most, and ask how you would prove who is right. If the answer involves somebody's laptop, start there instead.

Want this applied to your operation?

Next pieceWhat four weeks should leave behind