World Model Readiness
Engraved foundations instrument

For your company

Foundations

Module · The numbers your AI should not trust

The Data Quality Audit

AI learns whatever your data teaches it, including the lies. Some of your sources are gamed, some are stale, and most have no owner who would notice. This module finds which data is safe to feed a model, and which will teach it exactly the wrong thing.

Question 1 of 5 · Gamed sources are known

Do you know which of your data sources are gamed by the people who feed them?

Every metric someone is measured on eventually gets optimised, not always honestly. Sales pipelines, ticket close rates, utilisation: the numbers people's bonuses depend on are the numbers most likely to lie to a model.

Question 2 of 5 · Data is graded

Is your data graded for quality before anyone trusts it for a decision?

Not all data deserves the same confidence. A grade (verified, estimated, unverified) travels with the number and tells a model, or a human, how much weight it has earned.

Question 3 of 5 · Every dataset has an owner

Does each dataset that matters have a named owner accountable for its quality?

Ownership is the difference between an error that gets fixed and one that gets rediscovered every quarter. Not a custodian who stores the data: an owner who is called when it is wrong.

Question 4 of 5 · Know what to feed first

Do you know which of your data is trustworthy enough to feed an AI system first?

The instinct is to feed the model everything. The discipline is to feed it your cleanest, best-owned, ground-truthed data first, so it learns from your best evidence, not your worst.

Question 5 of 5 · Numbers trace to reality

Can you trace a key number back to where it came from and check it against reality?

Provenance is the difference between a number you can defend and one you merely repeat. If nobody can trace a figure to its source, nobody can tell when it quietly breaks.

For the statistics · one click each

Three questions for the public picture

These do not affect your score. They feed the anonymised, aggregated statistics; groups under 8 respondents are never shown.

Does your single most important dataset have a named owner?

Yes, one clear owner
Shared, unclear
IT holds it
No owner
We do not know

Have you ever discovered a key metric was being gamed?

Yes, and fixed it
Yes, still open
We suspect so
Never found one
We do not know

How much of your decision data would you stake a real decision on?

Most of it
About half
Very little
We are not sure

Your context

Used to calibrate the report. Company size and sector remain in the anonymized dataset; your email does not.

What the five levels look like

Every dimension in this assessment is scored 1 to 5. This is what the levels mean, dimension by dimension. The graded report diagnoses where your own answers land and what to do about it.

Gamed sources are known

  1. 1Never questioned
  2. 2Suspected, untested
  3. 3Some identified
  4. 4Most mapped
  5. 5Distortion tracked

At the low end: You are treating every number as equally true, which means the gamed ones travel straight into the model. List the five metrics tied to someone's bonus; start your suspicion there. What good looks like: You track where incentives distort your data, which is rare. Keep the map current: every new KPI creates a new reason for someone to game it.

Data is graded

  1. 1No grading
  2. 2Gut feel only
  3. 3Ad hoc checks
  4. 4Graded by hand
  5. 5Systematic grading

At the low end: Ungraded data forces everyone to treat a guess and a measurement as equals. Add one column, confidence, to your most-used dataset and fill it in honestly. What good looks like: Graded data lets both people and models weight evidence correctly. Automate the grading where you can so discipline survives volume.

Every dataset has an owner

  1. 1Nobody owns it
  2. 2IT holds it
  3. 3Owner unclear
  4. 4Named owners
  5. 5Owners held accountable

At the low end: Unowned data rots quietly because nobody's job is to notice. Name an owner for your top three datasets this week; accountability is cheaper than cleanup. What good looks like: Named, accountable owners are why your good datasets stay good. Make quality part of how they are measured, not a favour they do.

Know what to feed first

  1. 1Feed everything
  2. 2No priority
  3. 3Rough sense
  4. 4Ranked shortlist
  5. 5Curated feed

At the low end: Feeding a model everything teaches it your errors alongside your truths. Pick the one dataset you would stake a decision on, and start there. What good looks like: A curated, prioritised feed is exactly how good AI programmes start. Keep the shortlist tied to grading and ownership so it stays honest.

Numbers trace to reality

  1. 1No lineage
  2. 2Tribal knowledge
  3. 3Partial lineage
  4. 4Documented lineage
  5. 5Lineage plus ground truth

At the low end: Numbers with no traceable origin cannot be trusted or corrected. Pick your most-quoted metric and document its full path from source to dashboard. What good looks like: Traceable data checked against reality is the foundation everything else stands on. Automate the checks so drift surfaces before a decision does.