MODULE 3 · LESSON 5

Free — no login required

Sign in to track progress, save quiz attempts and enrol in the full course.

Sign in to track progress / enrol

Your First Ninety Days

Everything so far is diagnostic. This lesson is what to actually do, at whatever level of authority you have.

Weeks 1–2: find the candidates

Run this exercise with three or four people who do the work, not with management. Ask each to list tasks in their week that are:

  • Repetitive — done many times, in much the same way.
  • Judgement-based but fast — a decision made in seconds from information in front of them.
  • Recorded — the decision and its input left some trace in a system.

The intersection is your candidate list. Expect five to fifteen candidates from an hour's conversation.

Two things reliably surface. Some tasks are repetitive and fast but were never recorded — a candidate only after you start recording, which is itself a cheap and valuable action. And some tasks people describe as complex turn out on inspection to be one common case plus a handful of exceptions, which is the best shape a candidate can have.

Weeks 3–4: sample the reality

For your two most promising candidates, run the sampling exercise from lesson 3.2. Pull a hundred real, random cases. Sort them into straightforward, awkward and genuinely hard.

This costs a couple of afternoons and produces more useful information than any vendor conversation. You will learn the true shape of the problem, discover the exceptions nobody mentioned, and begin an evaluation set.

Weeks 5–6: write the one-page brief

For the strongest candidate, write one page containing exactly this:

  1. The task today — who does it, how often, how long it takes.
  2. Input and output — one sentence each.
  3. The data — where past examples live, how many, whether the outcome was recorded.
  4. The baseline — how accurate people currently are, or the obvious simple rule. Measure it; do not estimate.
  5. What success would change — the decision that would be made differently, by whom, in which tool.
  6. Cost of being wrong — what happens on an error, and who absorbs it.
  7. What we would not automate — the cases going to a person regardless.

If you cannot complete this page, the project is not ready — and that finding has saved you a great deal of money.

Weeks 7–12: run something small

Pick the smallest version that produces evidence. Usually one of:

  • Buy a tool and run it in shadow mode on real data for a few weeks, comparing against what people actually did.
  • Have someone technical build a crude version for the single most common case only, ignoring exceptions entirely.
  • Fix the data problem the sampling exposed, if the answer was not in the input. Unglamorous and often the highest-value thing available.

Then measure honestly against the baseline, and write down what you learned, including if the answer is no.

Brief A. Task: two people spend roughly three hours a day categorising inbound enquiries into eight service lines and assigning an owner. Input: enquiry text plus sender domain. Output: one of eight service lines. Data: four years of enquiries in the CRM, each with the service line finally recorded — about 60,000, though the field was only made mandatory two years ago. Baseline: measured over 200 sampled enquiries, humans agree with the final assignment 91% of the time. Success changes: enquiries arrive pre-categorised in the CRM queue; the two staff review rather than sort. Cost of error: a misrouted enquiry loses roughly half a day; nothing is lost permanently. Not automated: anything from an existing large client, and anything mentioning a complaint.

That brief is ready. Closed problem, real labels, measured baseline, defined exception path, low error cost, clear change in behaviour.

Brief B. Task: the operations lead decides each morning which jobs to prioritise. Input: "everything — the schedule, who's available, which clients are annoyed, what the weather's doing, what came in overnight." Output: a priority order. Data: the final schedule is stored; the reasoning is not, and nobody records what was deprioritised or why. Baseline: unknown, and nobody can say what a good ordering would even be. Success changes: unclear. Cost of error: potentially a missed client commitment.

That brief is not ready, and writing it is what revealed why. The output has no ground truth — no record of what the right answer was — so there is nothing to learn from. It is also not one decision but several traded off against each other, and half the input is not in any system.

Brief B is not hopeless. It is not an AI project yet. The useful first step is to start recording the reasoning, which might make it a project in a year, and which would improve handover and cover for absence immediately regardless.

Writing the brief is the filter. It costs an afternoon and it separates the two cases before anyone spends a budget.

Knowledge Check

In the opportunity-spotting exercise, which combination of characteristics identifies the best AI candidates?

📚 Flashcards1 / 5
Term

The three-filter exercise

Click to flip
Definition

Ask people doing the work for tasks that are repetitive, judgement-based but fast, and recorded. The intersection is the candidate list.

Click to flip back
💡Key Takeaway

Spend two weeks finding tasks that are repetitive, fast and recorded; two weeks hand-sampling a hundred real cases from the best candidates; two weeks writing a one-page brief with a measured baseline and a defined exception path; and six weeks running the smallest thing that produces evidence. The brief is the filter — if you cannot complete it, you have learned the project is not ready, which is the cheapest useful finding available.