Quick Summary
- A genuine ID paired with an AI-generated selfie completed a European brokerage’s live onboarding, even though the incumbent vendor flagged it.
- The real gap was decision confidence, not detection: analysts dismissed flags because of false-positive fatigue and no explanation to act on.
- Adding DuckDuckGoose’s explainable DeepDetector alongside the existing stack turned bare flags into forensic verdicts analysts could defend.
- [ILLUSTRATIVE] Any figures in this study are modelled examples, not measured client results. Verified data to be inserted before publishing.
Introduction
A genuine identity document. A real-looking selfie. Four onboarding controls passed. The account opened. The selfie was an AI-generated deepfake, and the firm’s incumbent verification vendor had actually flagged the attempt as suspicious. The flag changed nothing.
A European multi-asset brokerage ran a controlled test against its live remote onboarding flow: a genuine ID paired with a synthetic selfie of the same identity. The attempt completed verification. The incumbent liveness vendor did register suspicion, but the attempt succeeded anyway, because that suspicion signal was indistinguishable from the everyday false-positive noise the firm’s analysts had been trained by experience to ignore.
The gap was not in what the firm could detect. It was in what the firm was willing to act on.
The Onboarding Selfie as an Identity Attack Surface
A remote onboarding selfie produces two signals that look identical in the moment but answer different questions, and most stacks conflate them.
Liveness / presentation attack detection asks whether a real person is present rather than a photo, replay, or mask. It is mature in identity-verification platforms. It is also what injection and deepfake attacks are specifically engineered to satisfy.
Forensic deepfake detection asks a separate question: with a live person and a genuine stream, is the face being shown authentic? This is the specialist layer, and nothing in a standard liveness-and-document stack addresses it. A convincing synthetic face can pass liveness and still be a deepfake.
The Approach: From Raw Detection to Decision Confidence
The reframe was the turning point. The question stopped being whether deepfakes could be detected, the incumbent had detected this one, and became whether analysts could be given enough evidence to act on a suspicious attempt without drowning in false positives. That required three things from any layer added to the stack.
Forensic, artefact-level analysis. Rather than inferring authenticity from movement or challenge-response behaviour, which injection and deepfake attacks are built to defeat, DuckDuckGoose’s DeepDetector examines the image at the pixel level for the generative artefacts synthesis leaves behind.
Explainability the reviewer can use. The value was not a second yes/no signal, the firm already had one and ignored it, but a heatmap and confidence output an analyst could point to when justifying a decision.
Deployment alongside the existing flow. The layer was scoped to sit next to the incumbent stack at the biometric step, with a defined process for handling disagreement between the two systems, and all processing kept within the EU.
The Result: A Flag Becomes a Verdict Analysts Will Act On
The gap was never whether the deepfake could be caught. It was whether anyone would act on the flag when it fired.
The meaningful result is a change in operating posture, not a headline percentage. Before, the firm had a detection signal it did not trust and would not act on. After adding an explainable forensic layer, the same category of suspicious attempt produced a verdict an analyst could reason about and defend. The controlled deepfake test that had previously completed onboarding now surfaced a manipulation verdict with visual evidence, rather than a bare suspicion flag a reviewer would dismiss.
- Defensible rejections: analysts gained pixel-level reasoning to stand behind stopping a suspicious account, instead of waving it through to avoid friction with a possible real customer.
- No rip-and-replace: the forensic verdict was added alongside the incumbent vendor, with a clear path for resolving disagreements.
- Confidence, not just coverage: the change addressed why a correctly detected attack had been getting through, not merely whether it could be detected.
The economics also shift. When a flag carries evidence, the cost of acting on it moves from delaying a genuine customer to stopping a real attack, which is what makes analysts willing to act on it at all.
Key Takeaways
1. The gap is often decision confidence, not detection. A flag no one acts on is not a control. Adding another opaque score to a stack analysts already distrust changes nothing.
2. False positives carry a hidden second cost. Beyond wasted review time, they train analysts to ignore the alerts that matter, which is how a genuinely detected attack still gets through.
3. Explainability converts flags into actions. Reviewers need evidence they can point to and defend, not a binary verdict.
4. Prove the exposure. A single controlled test on the production flow moves a risk conversation further than any threat briefing.
5. Add, don’t rip out. The strongest deployments sit alongside the existing stack with a clear process for handling disagreement between signals.
How This Proceeds
An engagement of this kind is not a single procurement decision. What the firm authorises first is the analysis only, with no commitment to what follows.
Last update: Q3 2026







.webp)


