MODULE 7 · LESSON 3

Free — no login required

Sign in to track progress, save quiz attempts and enrol in the full course.

Sign in to track progress / enrol

Oversight, Inventory and Incidents

Three controls turn a policy into governance. None requires specialist tooling.

The model inventory

A list of every AI system in use, with enough detail to answer questions when they arrive. Most organisations cannot produce one, which means they cannot answer the first question any regulator, auditor, customer or board member asks: what are you using, and where?

Per system, one row:

  • What it is and what it does, in a sentence.
  • Who owns it — a name, not a department.
  • What decisions it makes or influences, and whether those affect people.
  • What data it uses, including whether personal data is involved.
  • Where it sits — bought, assembled, built; hosted where.
  • Human oversight — who reviews what, and when.
  • Last reviewed, and by whom.

A spreadsheet is fine. What matters is that it is complete and current, which means it needs an owner and a review cadence, and that new systems get added at procurement rather than at discovery.

Include the bought tools. The most common gap is a department subscribing to an AI product without anyone recording it — which is shadow AI at organisational scale, and it is where unpleasant surprises originate.

Oversight that is genuine

"A human reviews the output" appears in most governance documents and is frequently untrue in practice. Nominal oversight is worse than none, because it creates a documented control that does not function.

Genuine oversight needs four things:

🔗 Match the Pairs
Enough time to actually review, not 400 items an hourDrop here
Enough information to form an independent viewDrop here
Genuine authority to overturn, without justifying every deviationDrop here
An interest in getting it right, not only in clearing the queueDrop here

Test yours honestly. If a reviewer handles hundreds of items an hour, sees only a recommendation with no supporting reasoning, is measured on throughput, and is asked to explain any deviation from the system, then the oversight is decorative. Everyone involved knows it, and it will not survive being examined.

The conditions that produce rubber-stamping are structural rather than personal: high volume, low context, throughput targets, and a system that is right most of the time — which trains reviewers to accept, because accepting is usually correct and always faster.

The design response is to review a sample properly rather than everything nominally. A hundred cases genuinely examined tells you more, and controls more, than ten thousand waved through.

The incident route

Before you need it, establish:

What counts as an incident. A materially wrong output that reached someone, a system behaving outside expectation, data going somewhere it should not, a discovered bias. Say it concretely, or nothing gets reported.

Who to tell, and how. One route, easy, without needing to know the org chart.

What happens next. Who assesses severity, who can suspend the system, who decides on notification.

That reporting is welcomed. This determines whether the route is used at all. The people who see failures first are the ones using the system daily, and if reporting is inconvenient or unwelcome you have lost your best detection — while believing your controls work, because nothing is being reported.

A quiet incident log usually means people are not reporting, not that nothing is happening.

Practise once. Walk through a hypothetical: the system has been misclassifying a category for three weeks and a customer has complained. Who finds out, who decides to suspend it, who contacts the customer, who checks how many others were affected, who documents it? The gaps you find in an hour of discussion are the ones you would otherwise find during a real incident.

An organisation deployed a system recommending decisions on applications, with a documented control: every recommendation reviewed by a case handler before taking effect.

A year in, someone checked what the control was actually doing. Three figures settled it.

The override rate was under 2%. By itself ambiguous — it might mean the system was excellent.

The median review time was eleven seconds. Not enough to read the application, let alone form an independent view.

Override rate fell steadily over the year, from about 7% in month one to under 1% by month twelve. Handlers had learned that the system was usually right, and that overriding required a written justification while accepting required nothing.

The system may well have been good. That is not the point. The point is that the organisation had no idea, because the control it relied on had stopped producing information. It was recording agreement, not review — and it would have recorded agreement whether the system was excellent or had silently degraded.

What they changed, and it is a reasonable pattern to copy:

  • Sampled deep review replaced universal shallow review. Five percent of cases got fifteen minutes and full context; the rest went through on confidence.
  • Overriding and accepting were made equally easy. The asymmetric friction was doing more to shape behaviour than any instruction.
  • Reviewers saw the system's reasoning, so there was something to disagree with rather than a bare recommendation.
  • Override rate became a monitored metric. A rate falling toward zero is now treated as a warning about the control, not as evidence about the model.

That last inversion is the most useful idea here. A very low override rate is a question, not a reassurance.

Knowledge Check

A review control shows a 1.5% override rate and a median review time of eleven seconds, with overrides declining steadily over a year. What should a manager conclude?

📚 Flashcards1 / 6
Term

The model inventory

Click to flip
Definition

One row per system: what it does, who owns it by name, what it decides and whether it affects people, what data it uses, whether bought or built and hosted where, the oversight arrangement, and when it was last reviewed. Must include bought tools.

Click to flip back
💡Key Takeaway

Three controls make governance real: an inventory that can answer what you are running and who owns it, oversight with the capacity, context, authority and incentive to be genuine, and an incident route established before you need it. Test your oversight honestly — eleven-second reviews and a falling override rate mean the control records agreement rather than review. Sample deeply instead of reviewing everything shallowly, and treat a very low override rate as a question about the control.