World Model Readiness
Engraved foundations instrument

For your company

Foundations

Module · Where does AI stop and you start?

The Boundary Layer Assessment

Every AI deployment draws a line between what the machine decides and what a human still decides. Drawn well, per decision class and by stakes, that line is your main safety control; drawn by accident, it is your main risk. This module checks where the line actually sits, how explicit it is, and whether the human side is real.

Question 1 of 5 · The line is explicit

Is the AI-decides versus human-decides line written down for each class of decision?

A line that lives in habit gets redrawn silently with every deployment. Explicit means written per decision class: for pricing, for routing, for approvals, someone has stated where AI acts alone and where a human must.

Question 2 of 5 · Placement matches stakes

Is the line drawn by the stakes and reversibility of each decision, not by convenience?

The right place for the line depends on what happens when the AI is wrong. A cheap, reversible decision can sit on the AI side; an expensive, irreversible one should not, however tempting the automation.

Question 3 of 5 · The boundary is staffed

Are there enough qualified humans to actually hold the human-decides side of the line?

A human-in-the-loop who reviews five hundred AI decisions an hour is a rubber stamp with a job title. The human side of the line is only real if the people on it have the time and skill to genuinely decide.

Question 4 of 5 · Humans really override

When a human sits over an AI decision, do they actually change outcomes, or rubber-stamp them?

Automation bias is real: people defer to the confident machine. If your override rate is effectively zero, the human is decoration, and the line is really drawn on the AI side whatever the diagram says.

Question 5 of 5 · The line moves on evidence

Does the line move deliberately as evidence accumulates, rather than staying frozen or drifting?

A boundary set once and never revisited is either too cautious or too reckless within a year. The system should earn its way across the line on evidence, and lose ground the same way, by decision, not by drift.

For the statistics · one click each

Three questions for the public picture

These do not affect your score. They feed the anonymised, aggregated statistics; groups under 8 respondents are never shown.

For your riskiest AI-influenced decision, is the human and AI line written down?

Yes, clearly written
Partly written
Understood, not written
No
We do not know

When a human reviews an AI recommendation, how often do they change it?

Often
Sometimes
Rarely
Almost never
We do not track it

Who decides where the AI-decides versus human-decides line sits?

A named owner
A committee
Each team itself
Nobody in particular
We do not know

Your context

Used to calibrate the report. Company size and sector remain in the anonymized dataset; your email does not.

What the five levels look like

Every dimension in this assessment is scored 1 to 5. This is what the levels mean, dimension by dimension. The graded report diagnoses where your own answers land and what to do about it.

The line is explicit

  1. 1Nowhere stated
  2. 2Implicit only
  3. 3Some classes written
  4. 4Most classes written
  5. 5Explicit per class

At the low end: An unwritten line is drawn by whoever ships next, not by you. Write it for your three highest-stakes decision classes this week; it is a paragraph each, not a project. What good looks like: An explicit line per decision class is the foundation of deliberate autonomy. Keep it versioned, so you can see when and why the line moved.

Placement matches stakes

  1. 1Convenience decides
  2. 2No logic
  3. 3Loosely by risk
  4. 4Mostly by stakes
  5. 5Stakes and reversibility

At the low end: Drawing the line for convenience puts automation exactly where it hurts most when it fails. Re-place your riskiest decisions by asking what an error costs and whether you can undo it. What good looks like: Placing the line by stakes and reversibility is how mature programmes decide what to automate. Revisit it as reversibility changes; a new refund policy can move the line.

The boundary is staffed

  1. 1Unstaffed
  2. 2Token reviewer
  3. 3Understaffed
  4. 4Adequately staffed
  5. 5Staffed and skilled

At the low end: A human side nobody has time to staff is automation in disguise. Count how many decisions your reviewers actually face per hour; the number usually exposes the fiction. What good looks like: A staffed, skilled boundary is what makes human oversight more than a slogan. Watch volume growth; the boundary silently becomes a rubber stamp as throughput rises.

Humans really override

  1. 1Never overridden
  2. 2Rubber-stamped
  3. 3Rare overrides
  4. 4Real overrides
  5. 5Overrides tracked and used

At the low end: An override that never happens means the human side of the line is fictional. Measure how often reviewers actually change the AI's call; near zero tells you the truth. What good looks like: Tracked, meaningful overrides prove the boundary is alive, not ceremonial. Feed the override reasons back into where the line should sit next.

The line moves on evidence

  1. 1Frozen or drifting
  2. 2Moves by accident
  3. 3Moved once
  4. 4Reviewed periodically
  5. 5Evidence-based moves

At the low end: A frozen line ages into folklore while an accidental one moves without anyone deciding. Put one scheduled review on the calendar of whoever owns the riskiest line. What good looks like: Evidence-based boundary moves are how autonomy expands safely. Publish the moves and their evidence; it builds the trust that lets the next move happen.