World Model Readiness
Engraved team-practice instrument

For your team

Team Practice

Module · What leaves the team unchecked

The Team Review Discipline Check

AI lets your team produce more than it can read. The risk is not that the work is bad, it is that nobody looked before it shipped. This module checks the five parts of a review habit that hold under volume: whether a gate exists at all, how deeply you sample, who is accountable, whether caught errors change the prompts, and whether scrutiny survives a deadline.

Question 1 of 5 · A review gate exists

Does AI-assisted work pass a human review before it leaves your team?

Not for everything, but for anything that reaches a customer, a system of record, or another team. If AI output can ship straight from the tool to the outside, the review gate is a preference, not a control.

Question 2 of 5 · Sampling is deep enough

When your team reviews AI work, does anyone actually check the substance?

Skimming for tone is not review. Real review checks the claim, the number, the citation, the code path. AI errors are confident and fluent, so a quick read is exactly what they survive.

Question 3 of 5 · A reviewer is accountable

When reviewed AI work turns out wrong, is it clear who signed off?

If the answer is 'the AI wrote it', nobody owns the output. A named reviewer who put their name on it is what turns review from a ritual into a responsibility.

Question 4 of 5 · Errors feed the prompts

When your team catches an AI mistake, does it change how you prompt next time?

A caught error is free training data. If it is fixed once and forgotten, the same mistake returns tomorrow. Teams that improve fold the fix back into the shared prompt or the checklist.

Question 5 of 5 · Scrutiny survives deadlines

When a deadline is tight, does review get done or get skipped?

The value of a review process is decided entirely on the bad day. If the first thing cut under pressure is the check, then you do not have a review discipline, you have a review preference.

For the statistics · one click each

Three questions for the public picture

These do not affect your score. They feed the anonymised, aggregated statistics; groups under 8 respondents are never shown.

What share of your team's outbound work is now AI-assisted?

Under 10 percent
10 to 40 percent
40 to 70 percent
Over 70 percent
We do not track it

How is AI work reviewed before it leaves your team today?

It is not
Ad hoc, if there is time
A norm, not a rule
A required gate
A risk-tiered gate

Has an AI error reached a customer or another team from yours?

Not that we know of
A near-miss, caught late
Yes, minor
Yes, serious
We would not know

Your context

Used to calibrate the report. Company size and sector remain in the anonymized dataset; your email does not.

What the five levels look like

Every dimension in this assessment is scored 1 to 5. This is what the levels mean, dimension by dimension. The graded report diagnoses where your own answers land and what to do about it.

A review gate exists

  1. 1No gate at all
  2. 2Reviewed if remembered
  3. 3Gate for some work
  4. 4Gate for outbound work
  5. 5Gated, risk-tiered

At the low end: Work leaving the team unreviewed means the first person to catch an error is the customer. Define one line: nothing AI-touched goes outside without a named human reading it first. What good looks like: A risk-tiered gate is the right shape: heavy review where it matters, light touch where it does not. Keep the tiers written down so they do not quietly erode under load.

Sampling is deep enough

  1. 1Nobody reads it
  2. 2Skim for tone
  3. 3Spot-check surface
  4. 4Check the substance
  5. 5Depth scaled to risk

At the low end: Unread output is unreviewed output with extra steps. Pick the highest-stakes stream and have someone verify the substance of every item this week; the error rate will tell you how big the problem is. What good looks like: Substantive review scaled to risk is what makes AI volume safe. Keep sampling the low-risk streams occasionally too; that is how you notice when the model quietly gets worse.

A reviewer is accountable

  1. 1Nobody owns it
  2. 2Blame the tool
  3. 3Author owns loosely
  4. 4Named reviewer signs
  5. 5Sign-off logged

At the low end: When no human owns the output, review is theatre. Assign a named reviewer to each outbound stream so that shipping it is a person's decision, not the model's. What good looks like: Logged sign-off means every piece of work has a person behind it. Use the log when something slips: not to punish, but to find which part of the review missed it.

Errors feed the prompts

  1. 1Fixed and forgotten
  2. 2Mentioned in passing
  3. 3Fixed per person
  4. 4Shared informally
  5. 5Folded into prompts

At the low end: Fixing an error without changing the prompt guarantees you fix it again next week. Start a shared note: every recurring mistake gets one line and a prompt tweak. What good looks like: Feeding errors back into shared prompts is how a team compounds instead of repeating. Review the prompt library on a cadence; retire the fixes that the model no longer needs.

Scrutiny survives deadlines

  1. 1First thing cut
  2. 2Skipped quietly
  3. 3Shortened informally
  4. 4Protected minimum
  5. 5Held under pressure

At the low end: A review that vanishes on deadline day is exactly the review you needed most. Define a minimum check that is never cut, however tight the schedule. What good looks like: Review that holds under a deadline is a real discipline. Protect it by sizing the minimum check to be fast enough that skipping it never saves meaningful time.