Why you mustn’t settle on close enough numbers 

Sometimes I encounter attitudes such as “this KPI is in the ballpark” or “not matching exactly but in tolerance”. Would that suffice if the question was about the balance of your bank account, salary payment, or tax bill?

Oh but your inventory level or exact billing rate isn’t just that critical?

Well, if there’s a small recognizable error, how can you trust other numbers or maybe that same KPI but on different sub-selection?

Why “close enough” numbers are a problem

Remember that the law of big numbers conceals errors! Big aggregates converge to about right. That works in insurance pricing but not in discovering who’s working more productively than others or when evaluating quality of a product batch.

Perhaps the most important, jeopardized thing is trust in reports. As with trust in general, it’s easy to lose and next to impossible to regain. Do you really want to start your reporting renovation with just-almost-reliable numbers? Unfortunately, many do. Good luck leading change in your organization if the reports meant to support it have already given people reason not to trust them.

Way trickier situation is when wrong numbers are not recognized for being wrong. It’s not uncommon at all to discover later on that some calculation logic is not robust, and thus the business has been using wrong numbers for quite some time. It’s always a delicate issue to critically evaluate someone’s doings and faults, but it is something that has to be done, nevertheless. I bet many errors are ignored simply because of conformism and avoiding inconvenience.

This is not about predictions or otherwise inherently uncertain numbers but about numbers that can and should be logically and undisputably derivable. This is about erroneous logic, bad data, or both.

You can’t fix everything within a reasonable time frame and costs, but digging your head into the sand is a terrible choice.

Especially with ERP, CRM or otherwise human input data, there’s an endless stream of typing errors, missing data etc., but that is not actually bad data (even though I just used that phrase) but an important piece of reality.

Hand-in-hand with your KPI analytics endeavor, you should be getting tons of information about your data-generating processes. That is, about the ways Pekka and Tanja are filling values into those systems. Same but different goes for your IoT devices of course.

There are no shortcuts to data quality

I don’t have any magic bullets left in my clip, but I do have a few ordinary ones:

  • Verify results in reports vs. database vs. source systems 
  • You should practically always have row-level data as the basis in your Power BI data model so that you can verify numbers and utilize all the richness of data. If the volume is hundreds of millions of rows, there are ways for optimization, e.g. conditional aggregation. 
  • Don’t hide outliers, “duplicates” etc. lazily. Inspect their true origin and correct them at the source if possible. Document and set monitors for what you can’t fix completely. 
  • There’s no shortcut to quality and deep systemic understanding. I’m a humble student of Ferrari and Russo and their book “The Definitive Guide to DAX”. Ten years of solving problems with wits, YouTube and GPT doesn’t necessarily teach you about the whole system, or more importantly, what you don’t know. Think it this way: is it better to learn physics on your own or according to a high school curriculum? Very few of us are so passionately curious about this kind of subject that the first is a better option. The rest of us should learn the big picture from a curated comprehensive source. 

If you need analytics that are accurate, traceable and built to be trusted – rather than a collection of isolated solutions – let’s have a chat!

– Lauri Nuotio, Senior Analytics Engineer

.