A Prototype deepfake Assessment Toolbox for forensic Experts – APATE
Real or Fake? The Science of Deepfake Detection
The rapid democratization of deepfake technology has created an urgent need in digital forensics: law enforcement and judicial systems increasingly face audiovisual evidence they cannot reliably authenticate. The aim of the APATE project was precisely to tackle this challenge.
When Seeing Is No Longer Believing: The Forensic Case for Deepfake Detection
The central issue raised by the APATE project is the growing asymmetry between the ease of creating deepfakes and the difficulty of detecting them. Generative AI tools capable of producing highly realistic synthetic video are now accessible to anyone, requiring no technical expertise and no specialized hardware. This democratization of deepfake creation has outpaced the development of reliable detection methods, leaving forensic practitioners, law enforcement agencies, and the judiciary without adequate tools to authenticate audiovisual evidence. In a legal context where the integrity of video evidence can determine the outcome of criminal proceedings, this gap carries serious consequences. A second and closely related issue is the lack of generalization in existing detection approaches. Methods trained to identify the artefacts of one family of generation techniques consistently fail when confronted with new or unseen methods. This fundamental limitation means that detection tools rapidly become obsolete as generation technology evolves, creating a moving target that research has so far been unable to pin down. The emergence of diffusion-based generative models during the project lifetime made this challenge more acute, introducing a qualitatively new class of synthetic content for which existing detectors were largely unprepared. A third issue concerns the distance between research and operational deployment. Even where detection methods demonstrate strong performance under controlled experimental conditions, translating them into tools that are robust, stable, and usable by non-specialist forensic practitioners requires a level of engineering effort that the research community alone is not positioned to provide. The needs of operational end users (reliability, interpretability, auditability, and ease of use) impose requirements that go well beyond academic benchmarking. Against this backdrop, the general objectives of APATE were to establish a rigorous understanding of the deepfake generation landscape, to build evaluation resources grounded in realistic fraud scenarios, to develop and validate detection methods at the state of the art, and to integrate these into a prototype toolbox meeting the operational requirements of forensic experts. Underpinning all of this was the objective of situating the technical work within the applicable legal and regulatory framework, ensuring that the tools and methods developed could be used responsibly and in compliance with French and European law.
The project drew on a range of methods and technologies spanning the full pipeline from deepfake generation analysis to operational tool development. On the generation side, the consortium conducted a comprehensive survey of existing techniques, establishing a detailed map of the threat landscape and identifying the generation methods most relevant to fraud scenarios. This analysis directly informed both the construction of evaluation datasets and the design of detection approaches.
Detection research within APATE was grounded classic methods but also in deep learning, exploiting the ability of convolutional and transformer-based neural networks to identify the subtle artefacts introduced by generative processes. Several complementary detection strategies were pursued in parallel: spatial analysis of facial regions, temporal consistency analysis across video frames, frequency-domain analysis targeting compression and synthesis artefacts, and multimodal approaches combining visual and audio signals to detect incoherence between face and voice. This diversity of approaches reflected the understanding that no single method is sufficient and that robust forensic detection requires the ability to bring multiple lines of evidence to bear.
A significant portion of the technical effort was devoted to dataset construction. Four original evaluation corpora were produced, covering the broad landscape of commercially accessible generation tools, the specific fraud scenarios of forensic relevance, diffusion-based image synthesis, and localized facial manipulation. These datasets were designed with controlled variation in generation method, language, appearance, and social context, providing a more demanding and operationally grounded benchmark than those available in the public domain, and collectively constitute a substantial contribution to the evaluation infrastructure available to the deepfake forensics research community.
The project also made use of established biometric analysis frameworks, leveraging expertise in face recognition and speaker verification to assess the vulnerability of identity verification systems to deepfake attacks. On the legal and regulatory side, structured legal analysis and framework mapping were employed to identify applicable legislation and formulate concrete recommendations for the responsible development and deployment of detection tools. Together, these methods reflect a deliberately multidisciplinary approach, combining computer vision, signal processing, biometrics, and law around a shared operational objective.
The APATE project produced results across four interconnected areas. The most tangible scientific output is the publication record: 35 peer-reviewed papers (8 in international journals and 27 at international conferences) covering deepfake generation analysis, novel detection methods, and the introduction of new datasets.
Four original evaluation datasets were constructed over the course of the project. Two address video-based deepfake fraud: one providing broad coverage of accessible face-swap and avatar methods using widely available generation tools, and another offering operationally grounded fraud scenarios. Two further datasets address the detection of synthetically generated images, targeting diffusion-based generation and localized facial manipulation respectively. Together these four corpora fill a significant gap in the public benchmark landscape and provide the research community with a more diverse and operationally relevant evaluation infrastructure than was previously available.
On the detection side, the project developed and validated approaches targeting spatial artefacts, temporal consistency, and audio-visual incoherence, resulting in six prototype software implementations forming the core of the APATE detection toolbox.
The most significant result, however, is a negative one: no general and reliable method for detecting deepfakes regardless of their generation technique was found. Detectors trained on known methods degrade substantially when confronted with unseen techniques, and the emergence of diffusion-based generative models during the project widened this gap further. This finding is not a failure but an important contribution in itself: a precise characterization of where the science stands and what the field must prioritize. For forensic practitioners and policymakers, it carries a direct implication: automated deepfake detection cannot yet be treated as a solved problem.
The defining feature of the APATE project is its deliberate combination of scientific ambition and operational grounding. Rather than pursuing detection performance on standard academic benchmarks in isolation, the project oriented its research agenda around the concrete needs of forensic practitioners, with the French forensic police service acting as a reality check on the relevance and usability of the methods being developed. This dual commitment shaped every aspect of the project, and produced its most valuable and honest output: a clear-eyed assessment of the gap that currently exists between what research can deliver and what operational forensic use demands.
Looking ahead, the generalization problem (the inability of current detectors to reliably identify synthetic content produced by unseen generation methods) remains the central open challenge in the field, growing more acute as new generative models appear with increasing frequency. One approach is to systematically develop tools specifically tailored to each generation method as it emerges, though this is inherently reactive and difficult to sustain at scale. Another promising long-term direction is to move beyond method-specific detection toward approaches grounded in more fundamental properties of synthetic content (whether physical, biological, or statistical) that remain valid regardless of the generation pipeline used. The forensic traces left by diffusion-based models in particular represent a largely uncharted area where significant scientific progress is both possible and urgently needed.
The path toward operational deployment remains one of the most demanding challenges ahead. As the APATE experience has shown, bringing research-grade detection methods to the robustness and usability standards required for forensic casework demands a level of software engineering effort that is substantial and should not be underestimated. Future projects must treat this engineering effort as a first-class objective from the outset, with dedicated resources and close collaboration between research teams and operational end users.
Finally, the legal and regulatory analysis carried out within APATE offers a useful reference point for future work. The six recommendations produced by the project — covering biometric data use, labelling obligations, consent requirements, data protection impact assessments, and platform liability — provide practical guidance for researchers and institutions developing or deploying deepfake detection tools within the French and European regulatory context, and offer a starting point for navigating the legal and ethical questions that any operational deployment will inevitably raise.
The rise in the manipulation of images and voice represents a potential threat for crimes including disinformation campaigns, security fraud, extortion, online crimes against children, crypto jacking or illicit markets. Deepfake techniques are openly described and widely available; generating deepfakes is easy, and their quality has considerably improved. Therefore, it is challenging to detect them by mere visual analysis. As a consequence, there is an increasing need for deepfake detection tools. While several methods achieve good error rates under controlled scenarios, no dedicated tools are available for criminalistics experts. The goal of APATE is to deliver state-of-the-art methods to detect deepfakes. Instead of a “one size fits all” tool, the project aims at providing a toolbox of complementary techniques, based on the audio or visual parts of the video, by exploiting either low-level or semantic information, or by combining them in a multimodal manner. Each tool will address a different family of deepfakes, and will come with documentation detailing the use-cases, the known bias, the validation framework and how the results can be interpreted. The consortium includes criminalistics experts from the French National Scientific Police Service (SNPS), ensuring that the proposed toolbox is usable, properly described, and processes efficiently actual deepfakes found in criminal cases. In addition, the literature on deepfake generation will be continuously reviewed and analysed, to ensure that datasets corresponding to the latest deepfake generation techniques are available for the partners; a special concern will be to avoid overfitting to the learning databases. The consortium includes three research laboratories (Centre Borelli at ENS Paris Saclay, EPITA, LIX at Ecole Polytechnique), the SNPS, and IDEMIA, world leader in biometric recognition.
Project coordination
Raffaele Grompone (Ecole normale supérieure Paris-Saclay)
The author of this summary is the project coordinator, who is responsible for the content of this summary. The ANR declines any responsibility as for its contents.
Partnership
IDEMIA I&S IDEMIA IDENTITY & SECURITY FRANCE
SNPS SERVICE NATIONAL DE POLICE SCIENTIFIQUE
EPITA ÉCOLE POUR L'INFORMATIQUE ET LES TECHNIQUES AVANCEES (EPITA)
LIX Ecole Polytechnique
CB Ecole normale supérieure Paris-Saclay
Help of the ANR 559,679 euros
Beginning and duration of the scientific project:
September 2022
- 36 Months