
For your company
Foundations
Module · The numbers your AI should not trust
The Data Quality Audit
AI learns whatever your data teaches it, including the lies. Some of your sources are gamed, some are stale, and most have no owner who would notice. This module finds which data is safe to feed a model, and which will teach it exactly the wrong thing.
What the five levels look like
Every dimension in this assessment is scored 1 to 5. This is what the levels mean, dimension by dimension. The graded report diagnoses where your own answers land and what to do about it.
Gamed sources are known
- 1Never questioned
- 2Suspected, untested
- 3Some identified
- 4Most mapped
- 5Distortion tracked
At the low end: You are treating every number as equally true, which means the gamed ones travel straight into the model. List the five metrics tied to someone's bonus; start your suspicion there. What good looks like: You track where incentives distort your data, which is rare. Keep the map current: every new KPI creates a new reason for someone to game it.
Data is graded
- 1No grading
- 2Gut feel only
- 3Ad hoc checks
- 4Graded by hand
- 5Systematic grading
At the low end: Ungraded data forces everyone to treat a guess and a measurement as equals. Add one column, confidence, to your most-used dataset and fill it in honestly. What good looks like: Graded data lets both people and models weight evidence correctly. Automate the grading where you can so discipline survives volume.
Every dataset has an owner
- 1Nobody owns it
- 2IT holds it
- 3Owner unclear
- 4Named owners
- 5Owners held accountable
At the low end: Unowned data rots quietly because nobody's job is to notice. Name an owner for your top three datasets this week; accountability is cheaper than cleanup. What good looks like: Named, accountable owners are why your good datasets stay good. Make quality part of how they are measured, not a favour they do.
Know what to feed first
- 1Feed everything
- 2No priority
- 3Rough sense
- 4Ranked shortlist
- 5Curated feed
At the low end: Feeding a model everything teaches it your errors alongside your truths. Pick the one dataset you would stake a decision on, and start there. What good looks like: A curated, prioritised feed is exactly how good AI programmes start. Keep the shortlist tied to grading and ownership so it stays honest.
Numbers trace to reality
- 1No lineage
- 2Tribal knowledge
- 3Partial lineage
- 4Documented lineage
- 5Lineage plus ground truth
At the low end: Numbers with no traceable origin cannot be trusted or corrected. Pick your most-quoted metric and document its full path from source to dashboard. What good looks like: Traceable data checked against reality is the foundation everything else stands on. Automate the checks so drift surfaces before a decision does.