The Long Road Toward Absolute Deepfake Detection

Djamila Aouada, Senior Research Scientist and head of the Computer Vision, Imaging and Machine Intelligence (CVI2) research group at the SnT center of the University of Luxembourg (Photo © SnT)

Djamila Aouada, Senior Research Scientist and head of the Computer Vision, Imaging and Machine Intelligence (CVI2) research group at the SnT center of the University of Luxembourg, discusses her team’s work on deepfake detection, its implications for media, and the evolving landscape of AI technology.

You recently published a paper on a new approach for temporal deepfake localization. Could you explain the core methodology and objectives of this paper?

In general, the goal is deepfake detection in absolute terms, but what we’re chasing is something that is not only valid today and then obsolete tomorrow. We want a model that can be carried over, even as synthetic generation models evolve.

We worked on an approach that is deepfake type agnostic, meaning no matter which generation approach is used to create the fake, this model remains independent of it. The idea is to model deepfakes as anomalies. We focus on the normal behaviour of a face – how it looks and behaves – rather than relying on data from synthetic sources. This is different from the typical approach in the field, which often trains models using both real and fake data. Our method eliminates the need for constant retraining whenever new deepfake models emerge.

In this specific paper, we built upon our previous work by introducing motion analysis. We have a geometric understanding of how a face moves and behaves, so we utilized this knowledge to analyze temporal information in videos. By using key points, such as those around the eyes, mouth, and nose, we model these points over time to detect anomalies. This shift to a temporal approach gives us a new dimension of data and makes our model more adaptable and robust.

“Our goal is to ensure that multimedia content is verified to maintain public trust and prevent societal harm.”

Djamila Aouada, Senior Research Scientist and head of the Computer Vision, Imaging and Machine Intelligence (CVI2) research group at the SnT center of the University of Luxembourg

How has your approach and focus evolved over the past decade?

Ten years ago, we focused on recognizing people’s faces robustly when they entered a room, for example, in access control setting. We modelled the face’s geometry and its temporal behaviour, capturing small movements and expressions. This background in facial modelling provided a solid foundation for our current work. It’s an evolution where we take that accumulated expertise and apply it to a new challenge—deepfake detection. We’ve scaled up, working with about 10 thousand real videos to define what’s normal and identify deviations from this standard.

Our shift from static image analysis to dynamic, motion-based modelling allows us to capture more detailed anomalies in facial behaviour. This transition has been crucial, as it adds a temporal layer to our analysis, helping to differentiate genuine facial expressions from manipulated ones.

What are the research goals moving forward?

It’s a work in progress. This problem has multiple dimensions, and we continue to refine our models. Some fakes are becoming so good that not even experts, including myself and my students, can always tell. The real challenge is not the cheap fakes: it’s the high-quality ones that are almost indistinguishable from reality. That’s where our focus is.

We’re concentrating on fine-grained details. We group and analyze regions like the eyes, nose, or mouth separately before making an overall decision. This fine-grained modelling is different from what the community has been doing, as they usually process the entire image at once. By zooming into specific facial regions and their movements, we can achieve higher accuracy and detect subtle manipulations that broader analyses might miss.

We focus on fine-grained analysis, examining facial regions like the eyes, nose, and mouth individually to catch subtle manipulations that broader approaches may miss. Unlike conventional methods that analyze the whole image at once, our zoomed-in approach improves accuracy—critical as high-quality fakes become more complex with multiple people and interactions. This precision also helps detect synthetic videos and images across varied scenarios. In collaboration with POST and its Cyberforce Department, which plans to integrate our results into their products, we’re advancing robust solutions for next-generation security.

How effective is this approach in detecting deepfakes?

Our model’s rate of detection under a well-defined evaluation protocol is over 90%. In general, the more images and videos we have, the more precise the detection becomes. We are continually refining our models to ensure they stay robust as deepfake technologies advance. The aim is not just high accuracy but also maintaining efficiency and scalability in our models.

By leveraging temporal analysis and combining it with spatial information, we can detect even the subtlest manipulations. This allows us to validate our models using extensive benchmarks. The success rate is promising, and we are already refining the system by integrating other data types like audio, which shows a significant enhancement in detection capabilities.

“The industry creating fakes is much larger than the one detecting them.”

Djamila Aouada, Senior Research Scientist and head of the Computer Vision, Imaging and Machine Intelligence (CVI2) research group at the SnT center of the University of Luxembourg

What are the implications of your findings for maintaining accurate depictions in media?

There is a sense of urgency. These models should already be deployed and integrated into media workflows. It’s not theoretical anymore; it needs to be in products. Journalists should use detectors almost systematically because if they receive visual proof of an incident, it needs to be verified before publication. We’ve seen how misinformation spreads quickly when unchecked, and our tools can provide reliable detection capabilities.

Media organizations have a responsibility to verify visual content before reporting it. Our findings show that these tools are not just useful, they’re necessary. As deepfake technology advances, it’s critical that we have robust detection models available to prevent misinformation from becoming mainstream. Our goal is to ensure that multimedia content is verified to maintain public trust and prevent societal harm.

Should there be regulatory measures mandating the use of deepfake detection tools?

Absolutely. Regulators need to get involved. We’re pushing the quality of detectors, but the real impact is in broadcasting information – whether on social media or in journalism. The industry creating fakes is much larger than the one detecting them. It’s a massive industry, and it evolves quickly. That’s why detection technologies must keep pace. I don’t think detection of synthetics will become impossible, at least not any time soon. Faces and human expressions are complex, and there’s always a part of approximation when modelling them. While the industry producing fakes grows, there is still room to improve detection technologies. This is where legal frameworks and regulations become crucial. We need standards to ensure we stay ahead of the curve and that these detection tools are deployed efficiently across industries. With the right combination of technology and regulation, we can build a resilient system to safeguard the integrity of visual information.


This article was published in the special edition on artificial intelligence of Silicon Luxembourg magazine.

Total
0
Shares
Related Posts
Total
0
Share