MODULE 1 ยท LESSON 1
Free โ no login requiredSign in to track progress, save quiz attempts and enrol in the full course.
Sign in to track progress / enrolThe Twenty Five Million Dollar Video Call
In January 2024, an employee in the Hong Kong finance office of Arup, a large and well run British engineering firm, received an email from the group Chief Financial Officer in London. It concerned a confidential transaction and it needed to move quickly.
The employee was suspicious. That is worth pausing on, because it is the part most retellings skip. The training worked. The email had the shape of a scam, the employee noticed, and the employee did not act on it.
So the attackers escalated. They invited the employee to a video call.
On that call was the CFO. Also present were several other colleagues the employee recognised from around the business. Faces appeared as expected. Voices sounded as expected. There was ordinary small talk of the kind that fills the first minute of any corporate call. The request from the email was repeated, in person, by a senior figure, in front of witnesses.
The employee's doubt collapsed, which is an entirely reasonable response to that evidence. Over the following days they made 15 separate transfers into 5 different Hong Kong bank accounts, totalling roughly 25 million US dollars.
The fraud was discovered only when the employee later raised the matter with head office through normal channels. No such transaction existed. No such meeting had ever taken place.
Every other participant on that video call was generated.
How it was built
There was no intrusion. Nobody hacked Arup's network, stole a password, or planted malware. The attack was assembled entirely from material the company had published itself.
Large firms put a great deal of executive video and audio into the world: conference keynotes, recorded webinars, press interviews, results presentations, promotional films. Each one is a clean, well lit, professionally recorded sample of a named senior person's face and voice. That is precisely the raw material a synthesis model needs.
The attackers gathered that public footage, built convincing likenesses of several executives, and drove them live in a video call.
Why this case is the right place to start
It would be comforting to conclude that someone was careless. That conclusion is false, and believing it will hurt you, because it tells you the answer is to be more careful, and being more careful would not have helped.
Look again at what the employee actually did. They received a suspicious request. They declined to act on it. They sought confirmation from the person supposedly making the request. They obtained what appeared to be direct confirmation. Then they acted.
That is the correct procedure. It is what any reasonable person would do, and what most corporate policies at the time would have described as due diligence.
The procedure failed for one reason, and it is the single most important idea in this course:
The verification happened inside a channel the attacker controlled.
The employee did not choose the video call. It was offered to them by the party asking to be trusted. Asking a stranger to confirm they are trustworthy, using a method the stranger provided, is not verification. It only feels like verification, and in this case it felt like it very strongly, because human beings treat a recognised face and voice as conclusive proof of identity.
For the whole of history, that instinct was correct. It stopped being correct recently, and most of us have not updated.
Recognising faces and voices is not something you reason your way to. It runs on dedicated neural machinery that produces a feeling of certainty before any conscious thought is available. This is why you can identify a friend from a two second clip of a phone call, and why the sensation of recognition arrives as a fact rather than as an estimate.
That machinery has no fraud detection built in, because for the entire period in which it evolved there was no way to counterfeit a face and a voice in real time. A face was a reliable signal precisely because faking one was impossible.
There are two practical consequences worth internalising.
The first is that you cannot solve this by trying harder. Training people to scrutinise video calls for artefacts is a losing strategy, because the generation quality is improving continuously while human perception is not. Any defence that depends on out perceiving a synthesis model is a defence with an expiry date.
The second is that the feeling of certainty is now actively misleading, and you should treat a strong sense of recognition in a high stakes request as neutral information rather than as evidence. That is genuinely difficult, which is why the countermeasure in Module 4 is a mechanical procedure rather than an act of judgement. Procedures work under pressure. Judgement does not, especially when a senior figure is on the line and appears to be losing patience with you.
What would have stopped it
Not better perception. Two ordinary process controls, either of which would have been sufficient on its own.
The first is an independent verification channel. If the employee had ended the call and telephoned the CFO on the number held in the company directory, the fraud would have ended in under a minute. The attacker controlled the call. They did not control the company's own contact records.
The second is a second approver on outbound payments. If moving that sum required a second named person to authorise independently, then convincing one employee, no matter how thoroughly, would not have been enough. This is why banks and finance teams have used dual authorisation for a century. It does not assume anyone is dishonest. It assumes any one person can be mistaken or deceived.
Note that neither control asks anyone to detect a deepfake. Both of them work whether or not the fake was perfect, and that is the property you want in a defence.
The Arup employee ended the video call convinced the request was genuine. What was the decisive flaw in their verification?
Deepfake
Click to flipSynthetic video or audio of a real person, generated by a model trained on samples of them. Now achievable in real time and typically built from publicly available recordings.
Click to flip backAn employee who was suspicious, who followed correct procedure, and who sought confirmation before acting, still lost 25 million dollars. The failure was not carelessness, it was verifying inside a channel the attacker had supplied. Recognition of a face and a voice is no longer evidence of identity, and no amount of extra vigilance will restore it. What works instead is structural: contact people through details you already hold, and require a second approver for anything irreversible.