Deepfake attacks on remote onboarding: which check breaks first
A remote identity check has three moving parts, and a generated face does not attack all three the same way. Knowing which part each attack class defeats is what separates a useful control from an expensive one, and it is the question an auditor will ask when an onboarding control fails.
Below: the three parts of the check, the line between a presentation attack and an injection attack, why the standard that covers one does not cover the other, where synthetic identities differ from stolen ones, and what a German institution owes its supervisor when the control fails. The German account-opening procedure itself sits on video identification in Germany.
Three parts of a remote check, three different targets
A remote check does three things. It examines the identity document for the security features that say the document is genuine. It compares the face in front of the camera with the portrait on that document. And it establishes that the face belongs to a living person who is present right now, which is what liveness means. BaFin Circular 3/2017 (GW) describes the first two for the German video procedure, and the third is the part a human agent performs by asking you to tilt the card and turn your head.
Each part has its own failure mode. A forged or digitally edited document defeats the first. A face swapped onto the attacker's own live video defeats the second. A replayed or generated stream defeats the third. A control that only hardens one part moves the attacker to another, which is why vendors who talk about a single score are describing a product and not the problem. The BSI publishes its own warnings on what generated media does to identification procedures.
Presentation attack versus injection attack
A presentation attack happens in front of the camera: a printed photograph, a mask, a phone screen held up to the lens, a silicone finger. The biometric sensor captures something real; the something is just not a person. This is the class with a standard behind it. ISO/IEC 30107-3, Information technology, Biometric presentation attack detection, Part 3: Testing and reporting, sets out how to assess a detection mechanism and classifies known attack types, and its scope is explicitly attacks at the biometric capture device.
An injection attack skips the camera entirely. The attacker feeds a prepared or generated video stream into the application where the camera's output would normally arrive, through a virtual camera driver, a hooked library or a tampered client. Nothing is presented to a sensor, so a mechanism measured under ISO/IEC 30107-3 has not been measured against this at all. That is the technical heart of the subject: the strongest presentation attack detection score in a datasheet says nothing about injection, because the two attacks enter the system at different places.
Why pixel analysis alone cannot settle the question
If the stream never passed through a real camera, the useful evidence sits outside the image. Detection at the session and software layer asks where this video came from: whether the camera the operating system reports is a physical device or a virtual one, whether the client binary is the one you shipped, whether the frames carry the signal characteristics of a real sensor, whether the timing of the responses matches a person reacting. The upgrades reported by ID-Pal are described in exactly those terms, with virtual camera interference as the thing being detected.
The consequence for a procurement team is uncomfortable but simple. A generated face will keep getting better, and a detector trained on last year's generators will keep falling behind, so a control whose only evidence is the picture is a control with a short shelf life. A control that also establishes the provenance of the stream does not care how good the rendering got. ENISA tracks AI-enabled threats in the European threat picture, and this is the pattern it describes on the identification side.
A synthetic identity is not a stolen one
Identity theft uses a real person's data, so there is a victim who eventually notices and complains. A synthetic identity is assembled: a plausible combination of attributes that belongs to nobody, built up over months with small, well-behaved activity until the file looks like a customer with a history. No victim complains, because no victim exists, which is why these accounts survive controls designed around a disputing customer.
Generated media and synthetic identities meet at the onboarding step, where the built-up paper person finally needs a face. That is the moment the two problems become one, and it is why the fraud side and the identity side of an institution end up in the same meeting. Fraud prevention in Germany covers the German fraud picture these accounts are eventually used for, and the EU anti-money laundering package is the frame the identification duty sits in.
Why the eID route and a qualified signature resist this
Both replace a judgment about an image with a cryptographic operation. The chip in a German identity card holds a key it will not surrender, and it authenticates against the provider's terminal before releasing an attribute, so the provider's confidence comes from a protocol and not from a camera frame. Rendering a convincing face does not produce that signature, which is the whole point. Video identification in Germany sets out how the two routes compare on assurance level.
A qualified electronic signature makes the same move on the document side, and the argument there is legal as much as technical: the signature carries a legal effect that a scan of a signed page does not, independently of how good the scan looks. For a product decision that means the question is less about which detector scores best and more about where you can move the proof out of reach of generated media.
What you owe your supervisor when the control fails
An onboarding control that was bypassed is an incident, and in a German financial institution the question of whether it is a reportable one is answered by the DORA classification criteria and the GwG duties, not by how embarrassing it is. The institution has to be able to say what happened, which customers were affected, when it was noticed and what changed as a result, which means the session evidence has to be retained in a form that survives the investigation.
That retention requirement is worth designing for before the incident. DORA in Germany covers the reporting and classification machinery, and KYC in Germany covers the identification duty whose breach you would be reporting. The practical failure in the cases institutions describe is rarely detection: it is that nobody kept enough of the session to prove which accounts came through the same route.
Attacker economics decide which controls are worth the friction
Every control you add costs you customers who give up partway through, so the control has to cost the attacker more than it costs you. That calculation moved when generating a convincing face stopped requiring a specialist: the cheaper an attempt becomes, the more attempts arrive, and a control calibrated for rare, expensive attacks behaves differently under volume.
Two numbers matter and neither is a detection rate. The first is what one successful account is worth to the attacker, which tells you how much effort they will spend. The second is what the attempt costs them, which is the number your controls actually move. iProov's report on financial institutions and deepfakes makes the related argument that a one-time check at onboarding leaves the rest of the relationship unprotected, which is the same economics applied to time instead of money.
What is a biometric injection attack?
Feeding a prepared video or image stream directly into the application, at the point where the camera's output would normally arrive, instead of showing something to the camera. Routes include a virtual camera driver, a hooked library and a tampered client. Because no sensor captures anything, presentation attack detection measured under ISO/IEC 30107-3 does not address it.
Does liveness detection stop deepfakes?
It stops the attacks that reach the camera, and it is the right control for those. Against an injected stream it answers a question that was never asked, because liveness evidence drawn from the frames is only as trustworthy as the claim that those frames came from a camera. A check that also establishes where the stream came from is what covers both.
How does synthetic identity fraud differ from a deepfake?
One is the paperwork, the other is the face. A synthetic identity is a built-up person who never existed, assembled from attributes and aged with real activity. A deepfake is generated media. The two meet when the synthetic file needs to pass a face check, which is why institutions that treat them as separate programs find the same accounts in both.
Is a bank required to report a bypassed onboarding check?
It depends on what the bypass affected, and the answer comes from the DORA incident classification criteria together with the GwG duties, not from the institution's own judgment about severity. The dependable rule is to retain the session evidence so the classification can be made at all, since an institution that cannot reconstruct which accounts came through one route cannot answer the supervisor's first question.
Deepfake attacks on onboarding and Finance Loop
Finance Loop is where the fraud teams seeing the attempts meet the identity engineers deciding what the check has to prove. Finance Loop is the meeting place for digital identity and security in German finance, with meetups and conferences on KYC, fraud, AI and the technology behind all three. Finance Loop keeps those dates in its event calendar.
Finance Loop is a professional network and has the goal of driving the adoption of emerging technologies in finance, such as AI, tokenization, stablecoins, and DeFi. Finance Loop helps its members build skills and personal networks in these fields: Investment & Digital Assets, Payments & Digital Money, Digital Infrastructure & Sovereignty, and Risk & Compliance.