The 25 million deepfake video-call fraud in Hong Kong

In this article

  1. The specific case, what we know
  2. The underlying technique, without alarmism
  3. The shift in the CEO fraud model
  4. What works and what doesn't as a defense
  5. The consequences we still aren't digesting
  6. The political question
  7. To go deeper
  8. You might also like

Definitions · References · Elsewhere

Today we talk about the case that best shows the qualitative change generative artificial intelligence introduces into corporate fraud. February 2024, the Hong Kong office of Arup —a British multinational engineering firm with a global footprint. An employee from the finance department joins a video call with his chief financial officer and five other executives. He recognizes the faces. He recognizes the voices. The video call looks routine. He transfers twenty-five million dollars to five separate accounts. Days later he discovers that none of the executives were actually on the call. They were all real-time deepfakes. Why does it matter? Because the video call stops being reliable evidence of who you're talking to. My take is biased, because I work at a company where video calls are routine and the operational change is brutal. Form your own.

I've had the Arup case lodged in my head since it became public in February 2024. It's probably the cleanest example of a paradigm shift we aren't digesting fast enough. The video call has just stopped being, without society having processed it, a reliable way to know who you're talking to.

The specific case, what we know

Arup is a British engineering and consultancy firm founded in 1946 by Ove Arup. It has around 18,500 employees, offices in more than thirty countries, and works on iconic projects: the Sydney Opera House, the Pompidou, the Beijing Aquatics Center, the roofing of the Bird's Nest. It's a benchmark in its sector and its public presence —presentations, conferences, corporate videos— is high.

The case became public in early February 2024, when the Hong Kong Police informed the South China Morning Post and other outlets about a massive fraud committed against the firm's Hong Kong office. The timeline, according to the details confirmed by the police and by Arup itself in its later public statements, is as follows.

Mid-January 2024. An employee in the finance department of the Hong Kong office receives an email apparently signed by the company's CFO —a British citizen, normally based in another office— requesting an urgent confidential transfer. The employee, trained in cybersecurity protocols, is suspicious. Urgent confidential transfers are one of the classic patterns of CEO fraud.

The next day, he's summoned to a video call with the CFO and five other executives to review the matter. The video call takes place. The employee sees and hears the people he knows. The faces are right. The voices are right. The way of speaking, the accent, the mannerisms —as far as the employee can judge in a video call— are consistent with the originals. The video call dispels his initial suspicions.

The employee proceeds with the transfer. Twenty-five million US dollars, in fifteen transfers to five different bank accounts in Hong Kong, according to the police details. The accounts are emptied quickly to other jurisdictions before the fraud is detected.

Detection comes, according to the public accounts, several days later, when the employee contacts company headquarters about a separate matter and discovers that the real CFO had never held any video call with him. None of the five executives had taken part in the call. They were five real-time deepfakes generated from public audiovisual material of the executives —recorded presentations, interviews, corporate videos available on YouTube, LinkedIn and other public sources.

The underlying technique, without alarmism

It's worth understanding the technique without turning it into a weapon of panic, because technical sobriety is what allows you to build operational defenses.

Generating a real-time video deepfake from a known person with moderate public exposure requires three technical components. First, a model for generating the face and facial expression, trained on material of the target person. Second, a voice-cloning model trained on audio of the same person —for reasonable-quality voice, a few minutes of clean audio is enough. Third, a system that synchronizes both in real time during the video call, translating the facial movements and voice of the real operator into the facial movements and voice of the impersonated character.

In 2020, achieving this with quality good enough to fool an attentive observer for several minutes required months of work by a skilled technical team, specialized hardware, and abundant training data. In 2024, several combinations of commercial technologies —some legal, others in gray areas— let you achieve it in a matter of hours with consumer hardware. Models like SadTalker, Wav2Lip, and commercial services like HeyGen, D-ID, Synthesia (for legitimate uses) and their equivalents in unregulated zones (for illegitimate ones) have dramatically lowered the technical cost.

The raw material needed —video and audio of the target person— is, in the case of corporate executives with public exposure, widely available. A fifteen-minute corporate presentation on YouTube, two LinkedIn interviews, a recorded conference. With that you train models good enough for an operational impersonation of several minutes on a video call.

This means the profile of potential victims has widened dramatically. Until 2022, this kind of fraud was limited to people with very high public visibility —heads of state, CEOs of large listed companies. Today, any executive with moderate presence on LinkedIn and in corporate videos is a viable target.

The shift in the CEO fraud model

CEO fraud, or BEC (Business Email Compromise), has been one of the most profitable forms of corporate fraud for decades. Its classic mechanics were textual: carefully prepared emails impersonating the CEO or CFO, requesting urgent transfers. Companies with good cybersecurity training learned to defend themselves by requiring confirmation through an alternative channel —a voice call to the executive, a message through a verified corporate channel.

Verification by voice channel worked reasonably well until 2022. Voice cloning was possible but costly, and the quality let you detect the impersonation with a bit of attention. From 2023 on, with the spread of services like ElevenLabs and equivalent open-source models, voice cloning reaches a quality indistinguishable to the human ear from just a few minutes of training audio. Verification by voice call is compromised.

Verification by video call seemed, until the Arup case, the last frontier. The reasoning was sound: even if audio could be faked, faking synchronized video in real time was still hard enough to consider the video call a verifiable channel. The Arup case shows that frontier has fallen too.

The operative consequence for corporate processes is serious. Any financial authorization process that depends, ultimately, on visual identification of the executive on the other side of a screen is compromised. The processes that stay reasonably secure are those requiring factors independent of the audiovisual channel: multi-factor authentication with physical tokens, authorization in a closed flow inside corporate apps, verification through a channel radically different from the request, mandatory second signatures.

What works and what doesn't as a defense

The technical defenses against real-time deepfakes are still limited. The automatic detection filters some video-call platforms have begun to incorporate have variable detection rates and many false negatives. The biometric identity verification systems offered as products are still largely reactive: they detect certain artifacts typical of current deepfakes, but generation improves with each generation and the artifacts keep disappearing.

What does work, according to the emerging consensus of cybersecurity consultancies and of the police agencies themselves —Europol published in April 2024 a specific guide titled Facing reality? Law enforcement and the challenge of deepfakes— are procedural defenses. Three operative principles hold the defense together.

The first is the principle of the independent channel. Any relevant financial request must be verified through a channel different from the one the request came through. If the request arrives by video call, the verification must be by another means —a call to the known corporate number, a message in a closed Slack/Teams channel, authentication on a corporate platform. The rule is that the attacker can compromise one channel, not two independent channels simultaneously.

The second is the principle of dual authorization. Transfers above a threshold must require approval from two independent people, each verified by their own mechanism. This multiplies the attack's complexity by the number of people the attacker has to impersonate simultaneously.

The third is the principle of the process, not the person. Authorization must follow a standardized flow, not depend on the employee's personal judgment of whether the person on the other side of the screen is real. Human judgment about visual identification isn't reliable in the presence of deepfakes. The process, if well designed, is robust even if visual identification fails.

Companies with a mature cybersecurity culture have been applying these principles for years. They apply them because they know that, before deepfakes, there were already phishing, telephone social engineering, spoofed emails. The deepfake is just the next link in a chain of increasingly sophisticated impersonation techniques. The principles that defend against it are the same.

The consequences we still aren't digesting

So far the case has talked about corporate fraud, which is the concrete part. The less discussed consequence is this: if the video call stops being reliable evidence of who you're talking to, the implications reach far beyond corporate fraud.

In the judicial sphere, recordings of video calls have been used as evidence in civil and criminal proceedings over the past decade. The implicit assumption that a captured and stored video call documents a real conversation between the people shown is now in question. Courts will begin —they're already doing it in some jurisdictions— to require a robust technical chain of custody before accepting video calls as evidence.

In the journalistic sphere, statements by public figures captured on video call or delivered in audiovisual format to newsrooms lose part of their evidentiary value. Before, a journalist who received a video where an executive confessed a fraud had a strong piece. Today, they must technically verify the authenticity before publishing, and the verification methods are imperfect.

In the governmental and diplomatic sphere, video calls between leaders have been common since 2020. The Putin-Biden video call of December 2021, the videoconference communications between presidents during crises, virtual cabinet meetings: all are, technically, exposed to the possibility of impersonation. The defense runs through encrypted channels and multi-channel verification, but the naive assumption that the video call authenticates the person is no longer tenable.

In the sphere of personal relationships, the case is more delicate and more structural. Video calls between family members, friends and partners are today one of the main channels for intimate communication at a distance. The growing appearance of scams aimed at older people, where an attacker impersonates a relative over a video call to request urgent money, is one of the most worrying developments documented in the 2024 and 2025 police fraud reports. The Spanish figures from INCIBE-CERT and the Guardia Civil show a sustained increase.

The political question

This is personal opinion, but the sector data backs it. Contemporary society has built much of its trust on audiovisual formats that are now becoming unreliable as evidence of identity. The adaptation —technological, legal, social— lags several years behind the technical change.

What I would ask for, on different planes. On the corporate plane, that every company with significant financial operations review its authorization protocols to eliminate dependence on visual identification in video calls as a sole or main factor. It's low operational cost and high defensive benefit. On the legal plane, that the evidentiary framework of criminal and civil jurisdictions explicitly incorporate the requirement of technical verification for audiovisual recordings presented as evidence. On the public education plane, that training programs —schools, universities, vocational training, courses for older people— include operative information on how to verify the authenticity of a video call through independent channels. On the regulatory plane, that videoconference platforms with mass corporate use —Zoom, Teams, Google Meet— incorporate features for detecting and labeling synthetic content, with legal liability in case of non-compliance.

Meanwhile, the individual reader can do at least three things. First, assume that anyone with a moderate public presence online —including yourself if you post videos or presentations— is a viable target for deepfake impersonation. Second, agree with close relatives —especially older people— on keywords or verification phrases known only between the two of you, to use in case of urgent requests through an audiovisual channel. Third, demand from the companies you operate with financially that their processes don't rest on visual identification.

The hard fact to close on. According to the Sumsub Identity Fraud Report 2024, the number of fraud attempts using deepfakes detected on a global scale grew roughly 700% between the first half of 2023 and the first half of 2024. The Arup case was the most visible, not the only one. The figure of 25 million dollars draws headlines; the everyday pattern of dozens of similar attempts below the media radar is what changes the economics of fraud in the medium term. The video call as evidence of identity has just died as an operational category. Whoever keeps operating as if that weren't so will discover the change with the bill in hand.

Definitions

Real-time deepfake: generation of synthetic image, audio or video synchronized with a live communication channel —a video call, a broadcast. Until 2022 it was technically hard; in 2024 it's within reach of non-specialist operators.

Business Email Compromise (BEC) / CEO fraud: a form of corporate fraud in which an attacker impersonates a senior figure in a company to request transfers or sensitive information. It has existed since the early years of the internet; deepfakes expand its technical capacity.

Voice cloning: generation of synthetic audio that reproduces the voice of a target person from samples of their real voice. Commercial services like ElevenLabs offer this capability at high quality from just a few minutes of source audio.

Multi-factor authentication (MFA): a method of identity verification requiring two or more independent factors —something you know, something you have, something you are. It's the most robust defense against impersonation attacks, including deepfakes.

References

South China Morning Post, Hong Kong company duped of HK$200 million in deepfake video conference scam (4 February 2024). Initial coverage of the case by the local outlet of record, with statements from the Hong Kong Police.

Arup Group, public statements following the coverage of the case (February-March 2024). Official confirmation of the facts by the company itself.

Europol, Facing reality? Law enforcement and the challenge of deepfakes (Europol Innovation Lab, April 2024). Sector guide on the detection, defense and investigation of deepfakes in policing contexts.

Sumsub, Identity Fraud Report 2024 (sumsub.com, 2024). Sector data on identity fraud using generative AI.

Financial Times and Wall Street Journal, extended coverage of the Arup case and its sector context (February-May 2024). Later analyses of the implications for corporate processes.

Robert Cialdini, Influence: The Psychology of Persuasion (Harper Business, 1984; revised edition 2021). Classic framework on the social-engineering mechanisms that deepfakes amplify.

Danielle Keats Citron, The Fight for Privacy (W. W. Norton, 2022). Legal framework on digital impersonation and privacy.

To go deeper

INCIBE-CERT (incibe.es). Spain's security incident response center; it publishes alerts and operative guides on deepfake fraud aimed at companies and individuals.

FBI Internet Crime Complaint Center (IC3), Internet Crime Report (annual reports 2023-2024-2025). US figures on online fraud, with specific sections on BEC and deepfakes.

Ronen Bergman, Rise and Kill First (Random House, 2018). Historical framework on intelligence operations that use impersonation, a conceptual antecedent of today's corporate deepfake.

Bruce Schneier, A Hacker's Mind: How the Powerful Bend Society's Rules, and How to Bend Them Back (W. W. Norton, 2023). Framework for understanding how hacking techniques extend from technical systems to social and political ones.

You might also like

Elsewhere

Comments0

No comments yet.

Leave a comment