Deepfake Image and Video Detection

Mike Raggo

DEF CON 33 · Day 1 · Main Stage

Overview

In an era increasingly saturated with synthetic media, the ability to discern genuine content from sophisticated fabrications is paramount. This talk, delivered by veteran security researcher Mike Raggo at DEF CON, delves into the evolving landscape of deepfake image and video detection. Raggo, with over two decades of experience in image analysis and steganography, presents a comprehensive overview of techniques ranging from traditional metadata scrutiny to advanced machine learning algorithms. The core of his presentation revolves around the challenges posed by generative AI tools like DALL-E and MidJourney and introduces practical, accessible methods for forensic analysis.

Watch on YouTube

Visual summary for Deepfake Image and Video Detection by Mike Raggo
Visual summary for Deepfake Image and Video Detection by Mike Raggo

Key moments

  1. 0:00 Introduction and speaker's cybersecurity origin story
  2. 3:00 Deepfake detection techniques overview and free GPT tool
  3. 4:40 Examining obvious deepfake examples and forensic challenges
  4. 6:00 Analyzing AI-generated images with the free GPT tool
  5. 7:10 Advanced metadata analysis using ICC profiles and hex editors

Deepfake Image and Video Detection

Speakers: Mike Raggo

Conference: DEF CON

YouTube: https://www.youtube.com/watch?v=GPqL9_muXJA

Overview

In an era increasingly saturated with synthetic media, the ability to discern genuine content from sophisticated fabrications is paramount. This talk, delivered by veteran security researcher Mike Raggo at DEF CON, delves into the evolving landscape of deepfake image and video detection. Raggo, with over two decades of experience in image analysis and steganography, presents a comprehensive overview of techniques ranging from traditional metadata scrutiny to advanced machine learning algorithms. The core of his presentation revolves around the challenges posed by generative AI tools like DALL-E and MidJourney and introduces practical, accessible methods for forensic analysis.

The talk highlights the critical need for robust detection mechanisms, not only for identifying obvious fakes but also for uncovering subtle manipulations that can mislead public opinion, fuel disinformation campaigns, or even be leveraged in cyberattacks. Raggo demonstrates a free, custom-built ChatGPT GPT called the "Fake Image Forensic Examiner," which integrates many of the discussed detection techniques, making advanced forensic analysis available to a broader audience. This tool, along with an underlying Python framework, offers a multi-layered approach to verify the authenticity of visual media, underscoring the ongoing cat-and-mouse game between creators of synthetic content and those striving to expose it.

Background

▶ Watch: Introduction and speaker's cybersecurity origin story (0:00)

Mike Raggo’s journey into cybersecurity began in a rather dramatic fashion in 1993, working as a Unix administrator at NASDAQ. An accidental incident involving the vulnerability scanner Satan (Security Assessment Tool for Analyzing Networks) — a precursor to modern tools like Nessus — led to him inadvertently taking down the stock market. This event, while perilous, ultimately propelled him into forming NASDAQ's first security team. This foundational experience sparked a career dedicated to understanding and mitigating digital threats, including a long-standing collaboration with Chad Hosmer on research spanning steganography and various forms of image analysis for over 25 years.

The problem of image and video manipulation has existed since the advent of photography, with early Photoshopping techniques being a common method. However, the proliferation of Generative Adversarial Networks (GANs) and sophisticated AI image generation tools like DALL-E, MidJourney, and Copilot has dramatically escalated the challenge. These tools can produce highly realistic, yet entirely fabricated, images and videos with increasing ease and speed, making it difficult for the average observer to distinguish real from fake. This rise in AI-generated content necessitates advanced forensic techniques to verify authenticity, provide evidentiary proof for investigations, and combat the spread of disinformation. The talk builds upon Raggo and Hosmer's extensive research, adapting it to confront these contemporary challenges.

Key Findings

▶ Watch: Deepfake detection techniques overview and free GPT tool (3:00)

The talk reveals several critical insights into the nature of deepfakes and their detection:

  • Multi-layered Detection is Essential: No single technique is foolproof. Effective deepfake detection requires a combination of metadata analysis, pixel-level scrutiny (error level, noise, edge anomalies), and advanced machine learning.
  • AI Imperfections are Detectable: Despite their sophistication, AI generation tools often leave subtle, characteristic imperfections. These can manifest as "weird" smoothing, anomalous edge details, or inconsistent noise patterns that are invisible to the naked eye but detectable through forensic tools.
  • Metadata Remnants Persist: While social media platforms often strip traditional EXIF data, other forms of metadata, such as ICC profiles (International Color Consortium), frequently remain embedded within images and can provide crucial forensic clues about an image's origin or manipulation.
  • Accessible Forensic Tools: The development of a free, custom ChatGPT GPT named "Fake Image Forensic Examiner" democratizes access to advanced deepfake detection, allowing users to perform detailed analysis and generate forensic reports with ease.
  • Video Analysis is Achievable: Modern AI models (like the underlying GPT-5 capabilities) enable video deepfake detection by breaking videos down into individual frames, analyzing each frame for anomalies, and then compiling a comprehensive report.
  • Beyond Manipulation: Human Characteristics: Surprisingly, the advanced Python-based tool can even identify non-manipulated human characteristics within images, such as colored contact lenses or dental caps, by detecting minute pixel-level anomalies. This highlights the extreme sensitivity and depth of analysis possible.
  • Emerging Threats: A significant and growing concern is the embedding of malcode within images to perform jailbreaks or DAN attacks against AI Large Language Models (LLMs), leading to undesirable behaviors like exfiltrating sensitive data. This underscores the expanded security implications of image-based threats.

Technical Deep Dive

▶ Watch: Examining obvious deepfake examples and forensic challenges (4:40)

Mike Raggo outlined a comprehensive suite of technical methods for deepfake image and video detection, ranging from foundational metadata analysis to cutting-edge machine learning.

Generative Adversarial Networks (GANs)

At the heart of many sophisticated deepfake creations are Generative Adversarial Networks (GANs). Raggo emphasizes that understanding how adversaries create these fakes is crucial for effective detection. A GAN consists of two neural networks: a generator and a discriminator. The generator creates synthetic data (e.g., fake images), while the discriminator tries to distinguish between real data and the generator's fakes. Through an adversarial training process, both networks improve: the generator becomes better at creating convincing fakes, and the discriminator becomes better at identifying them. When the discriminator can no longer reliably tell the difference, the generator has produced highly realistic output. From a detection standpoint, understanding the loss calculation and the inherent imperfections introduced during this adversarial process (especially when images are compressed or re-saved) is key to uncovering their synthetic origin. Adversaries leverage GANs to circumvent detection, making the detection process a continuous cat-and-mouse game.

Metadata Analysis

While often overlooked or intentionally stripped, metadata can be a goldmine for forensic analysis:

  • EXIF Data: Historically, Exchangeable Image File Format (EXIF) data provided details like camera model, date/time, and even geolocation. However, social media platforms frequently strip this information upon upload.
  • ICC Profiles: Raggo highlights the International Color Consortium (ICC) profile as a persistent and valuable source of metadata. These profiles, commonly found in images from devices like Android phones, dictate how an image's colors are rendered. Even after social media compression, ICC profiles often remain. Using a standard hex editor, forensic investigators can uncover "fingerprints" within these profiles, providing insights into the image's processing history or originating device.

Watermarks

Some AI image generators are beginning to incorporate watermarks, though their effectiveness varies:

  • Visible Watermarks: Tools like DALL-E embed visible watermarks, typically small color codes in the bottom right corner, indicating their AI origin.
  • Limitations: These watermarks are not foolproof. They can be cropped, or their integrity may be compromised if the image is re-saved in different formats with additional compression. Raggo notes that platforms like MidJourney, especially when accessed via Discord, often produce images with little to no embedded watermarks, making detection more challenging.

Error Level Analysis (ELA)

Error Level Analysis (ELA) is a technique that identifies areas within an image that have different compression histories. When an image is saved in a lossy format like JPEG, it undergoes compression. If parts of the image are later edited, pasted in, or re-compressed, these areas will exhibit different error levels when the image is re-saved. ELA algorithms highlight these inconsistencies, often revealing pixelated "fingerprints" around superimposed elements. Raggo demonstrated this with an image of a Teemu facility fire, where ELA clearly pointed out the superimposed logo and fire components due to their inconsistent compression artifacts.

Noise Detection and Noise Maps

Inconsistent noise patterns are another strong indicator of manipulation. Digital images naturally contain a certain level of noise. When parts of an image are altered or inserted, the noise characteristics of the added elements may not match the original image. Noise maps are generated by analyzing and tuning out the "less pixelated" areas of an image, leaving behind pronounced lines or strong pixels where inconsistencies exist. Raggo explained that these strong lines, once the background noise is removed, can frame out characteristics pointing to photoshopping or AI generation, indicating areas where algorithms struggled to achieve natural, uniform noise distribution.

Edge Anomalies

AI generation and manual photoshopping often struggle with perfectly rendering edges. These imperfections can manifest in several ways:

  • Imperfect Cutouts: When an object is cut from one image and pasted into another, the edges might not be perfectly smooth or might retain remnants of the original background, appearing unnaturally sharp or thick.
  • AI Smoothing Techniques: Generative AI, while impressive, can produce "weird" or inconsistent smoothing. Raggo cited an example of an AI-generated image of a woman on a horse where the horse’s anatomy had non-horse characteristics, the woman had six toes, and the smoothing around the hands and bridle was imperfect, appearing smudged or overly detailed in some areas and less so in others. These "plastic horse texture" characteristics are distinct from natural images.

Advanced Techniques (Python Tool)

Raggo revealed that the free ChatGPT GPT tool is built upon an advanced Python tool set developed over eight years, leveraging Machine Learning (ML) for deeper analysis:

  • ML Training: The Python tool is trained on a corpus of known real and fake images. This allows it to learn the subtle differences and characteristics that distinguish authentic content from manipulated content. Users can even train the model with their own custom datasets.
  • Nearest Neighbor Analysis: This technique involves breaking an image into a grid of small squares or "chunks." The tool then analyzes the nearest neighbors – adjacent chunks – comparing their individual characteristics (e.g., smoothing, compression, noise patterns). By examining these micro-level relationships across the image, the tool can achieve higher accuracy in identifying localized manipulations. Raggo demonstrated this with a movie poster featuring The Rock and Zach, where the tool's yellow dots highlighted a superimposed background and even a fake watch on The Rock's wrist.
  • Sensitivity Parameters: The Python tool offers sliders to adjust the level of sensitivity, allowing users to control the granularity of the grid analysis. A higher sensitivity will break the image into more squares, leading to a longer processing time but revealing more minute details.
  • Real-world Debunking: The tool successfully debunked a widely circulated image of Putin giving a thumbs up to Trump, revealing the hand was photoshopped. More remarkably, in another image of Trump and Putin, the tool's sensitivity picked up on Putin's colored contact lenses (appearing as "herds of dots" around his eyes) and Trump's dental caps (around his front teeth), demonstrating its ability to detect genuine human characteristics that appear anomalous at a pixel level, even if not a manipulation. This unexpected finding underscored the tool's profound analytical depth. Similarly, it debunked a fake image of a hurricane approaching New York City that had been published by news outlets.

Demo / Proof of Concept

▶ Watch: Analyzing AI-generated images with the free GPT tool (6:00)

Mike Raggo provided a live demonstration of his custom-built ChatGPT GPT, named "Fake Image Forensic Examiner," which integrates many of the advanced detection techniques discussed. He also elaborated on the capabilities of the underlying Python tool.

The "Fake Image Forensic Examiner" GPT is presented as a user-friendly, prompt-oriented tool accessible via a QR code or direct search in the ChatGPT interface. Users can upload an image and engage in a conversational analysis. The GPT is designed to:

  • Initial Scan: Automatically look for metadata, including EXIF data and ICC profiles, as well as other steganographic artifacts.
  • Advanced Analysis: Perform Error Level Analysis (ELA), noise analysis, and identify edge anomalies.
  • Object Recognition and OSINT: Beyond fakeness detection, the tool can assist with object recognition, optical character recognition (OCR), and reverse image searches for Open Source Intelligence (OSINT), eliminating the need to use multiple browser tabs.
  • Reporting: Generate a detailed forensic PDF report with marked-up images, highlighting detected anomalies and evidentiary information. Raggo noted ongoing refinements to the visual accuracy of the red boxes in the marked-up reports.

A significant update showcased was the GPT's new capability to handle video analysis, leveraging recent advancements in GPT-5. While it doesn't directly analyze video in real-time, it performs the forensic standard of breaking the video down into individual frames. Each frame is then analyzed using the same image detection techniques, and a comprehensive report is generated, detailing findings for each frame and providing an overall summary. Raggo cautioned that analyzing long videos could result in very extensive reports. The GPT prompts users to select key frames or intervals to manage report length.

The talk also touched upon the underlying Python tool set, which powers the GPT and offers even deeper, more customizable analysis. This tool provides a UI or command-line interface, allowing users to adjust sensitivity parameters and visualize manipulations through "herds of yellow dots" directly on the image. Examples included:

  • The Rock and Zach movie poster: The yellow dots clearly demarcated a superimposed background and a photoshopped watch on The Rock's wrist.
  • Trump and Putin images: The tool identified a photoshopped thumbs-up hand in one image and, in another, revealed the boundaries of superimposed faces. Crucially, the tool's extreme sensitivity also highlighted Putin's colored contact lenses and Trump's dental caps as pixel-level anomalies, demonstrating its unexpected ability to detect genuine, non-manipulated human characteristics.
  • Hurricane over New York City: The tool debunked a widely circulated fake image, which was subsequently pulled from news outlets.

These demonstrations underscored the practical utility and the profound analytical capabilities of both the accessible GPT and its sophisticated Python foundation.

Defensive Implications

▶ Watch: Advanced metadata analysis using ICC profiles and hex editors (7:10)

The proliferation of deepfake images and videos, coupled with the ease of their creation, presents significant defensive challenges across various sectors. Mike Raggo's talk offers crucial insights for defenders:

  1. Heightened Awareness and Skepticism: The primary defensive implication is the need for increased public and professional awareness regarding the capabilities of AI-generated content. Individuals and organizations must adopt a default posture of skepticism towards unverified visual media, especially in critical contexts like news, political discourse, or evidence.
  2. Proactive Verification: Defenders, including journalists, law enforcement, intelligence analysts, and corporate security teams, should proactively integrate deepfake detection tools into their workflows. The "Fake Image Forensic Examiner" ChatGPT GPT offers an accessible starting point for initial analysis, providing detailed reports that can guide further investigation.
  3. Multi-Layered Analysis: Relying on a single indicator for deepfake detection is insufficient. Defenders must employ a multi-layered approach, combining metadata analysis (including ICC profiles), pixel-level techniques (ELA, noise maps, edge anomalies), and machine learning-based methods (nearest neighbor analysis) to build a robust evidentiary case.
  4. Training and Education: Security professionals and digital forensic investigators require ongoing training to understand the evolving techniques of deepfake creation and the corresponding detection methods. This includes familiarization with tools like hex editors for metadata inspection and understanding the principles behind ML-based analysis.
  5. Securing AI Systems (LLMs): The emerging threat of malcode embedded in images to perform jailbreaks or DAN attacks on Large Language Models (LLMs) highlights a critical new attack vector. Defenders must focus on securing LLM inputs, implementing robust validation mechanisms for image content processed by AI, and researching defenses against such adversarial prompts. This extends beyond traditional image forensics into the realm of AI security.
  6. Continuous Research and Development: The "cat-and-mouse game" between deepfake creators and detectors necessitates continuous research into new adversarial techniques and the development of more sophisticated detection algorithms. Organizations should consider contributing to or leveraging community-driven research, such as Raggo's Python toolkit, to stay ahead.
  7. Policy and Standardization: As AI-generated content becomes more prevalent, there's a growing need for industry standards (e.g., universal watermarking for AI-generated content, though this has limitations) and policy frameworks to address the ethical and security implications of synthetic media.

By adopting these defensive strategies, organizations and individuals can better equip themselves to navigate the complex and often deceptive landscape of deepfake technology.

Key Takeaways

  • Deepfake detection requires a multi-layered forensic approach, combining metadata analysis, pixel-level scrutiny, and advanced machine learning techniques, as no single method is foolproof.
  • AI-generated images and videos, despite their realism, consistently leave subtle, detectable "fingerprints" such as inconsistent noise patterns, anomalous edges, and unusual compression artifacts.
  • Even when traditional EXIF data is stripped, valuable forensic clues can often be found in persistent metadata like ICC profiles, which dictate color rendering and can be analyzed with hex editors.
  • Tools like Mike Raggo's free "Fake Image Forensic Examiner" ChatGPT GPT offer accessible and detailed analysis for both images and videos (by frame-by-frame breakdown), generating comprehensive forensic reports.
  • Advanced machine learning techniques, such as nearest neighbor analysis applied to image grids, can pinpoint minute alterations, even revealing non-manipulated human characteristics like colored contacts or dental caps due to their pixel-level anomalies.
  • The threat landscape is expanding, with emerging concerns about malcode embedded within images to perform jailbreaks or DAN attacks against Large Language Models, necessitating new defensive strategies for AI systems.

About the Speaker(s)

Mike Raggo is a highly experienced cybersecurity professional and researcher with a long-standing presence at DEF CON, marking his 25th attendance at the conference. His career in cybersecurity began in 1993 at NASDAQ, where, as a Unix administrator, he famously (and accidentally) took down the stock market using the Satan vulnerability scanning tool. This incident led to a promotion, where he was tasked with forming NASDAQ's first security team.

Over the past 25 years, Raggo has conducted extensive research, often in collaboration with Chad Hosmer, covering a wide array of topics from steganography to advanced image analysis and detection. He has presented on these subjects at DEF CON multiple times. His work demonstrates a deep understanding of both offensive and defensive security techniques, particularly in the realm of digital media forensics. Mike Raggo has also served as an adjunct professor, contributing to cybersecurity education and research.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Raggo is a genuine practitioner with 25 years in the space and the talk delivers competent coverage of deepfake forensics — ELA, noise maps, edge anomalies, ICC metadata, nearest-neighbor ML analysis. Nothing here is new to the field, but the packaging is honest and the free GPT-based forensic tool gives attendees something they can actually use Monday morning.

Heather Calloway (CISO) — WEAK

Raggo brings genuine technical depth and a useful tool, but the talk never climbs to where the real risk lives. The forensic mechanics are solid; the institutional and organizational implications are almost entirely absent.

→ Top-rated talks at DEF CON 33

All talks from DEF CON 33