Securing CCTV Cameras Against Blind Spots

Jacob Shams

DEF CON 32 Main Stage · Day 1 · Main Stage

Overview

In the realm of modern surveillance, Artificial Intelligence (AI)-powered object detectors are increasingly deployed in Closed-Circuit Television (CCTV) systems to automate threat detection and enhance security. However, this talk, "Securing CCTV Cameras Against Blind Spots" by Jacob Shams at DEF CON 32, reveals a critical vulnerability in these systems: their inherent "blind spots" arising from varying detection confidence based on a person's position within the camera frame. Shams' research demonstrates that AI models like Yolo V3 and Faster R-CNN exhibit significant fluctuations in their ability to confidently identify individuals, creating exploitable pathways for evasion.

Watch on YouTube

Visual summary for Securing CCTV Cameras Against Blind Spots by Jacob Shams
Visual summary for Securing CCTV Cameras Against Blind Spots by Jacob Shams

Key moments

  1. 0:00 Initial findings: person's position affects detector confidence
  2. 2:00 Camera quality does not impact confidence variability
  3. 2:25 Analyzing confidence in real-world pedestrian traffic
  4. 4:00 Real-world heatmaps show varying detection confidence
  5. 5:15 Introducing the novel "TipToe" evasion attack
  6. 6:30 TipToe methodology: reducing problem to graph path
  7. 7:35 TipToe's threat model and attacker assumptions

Securing CCTV Cameras Against Blind Spots

Speakers: Jacob Shams

Conference: DEF CON 32

YouTube: https://www.youtube.com/watch?v=hB6rtwoyyrg

Overview

In the realm of modern surveillance, Artificial Intelligence (AI)-powered object detectors are increasingly deployed in Closed-Circuit Television (CCTV) systems to automate threat detection and enhance security. However, this talk, "Securing CCTV Cameras Against Blind Spots" by Jacob Shams at DEF CON 32, reveals a critical vulnerability in these systems: their inherent "blind spots" arising from varying detection confidence based on a person's position within the camera frame. Shams' research demonstrates that AI models like Yolo V3 and Faster R-CNN exhibit significant fluctuations in their ability to confidently identify individuals, creating exploitable pathways for evasion.

This presentation delves into a detailed analysis of how factors such as distance, angle, height, and location within a scene impact an object detector's confidence. Building on these empirical findings, Shams introduces TipToe, a novel evasion attack that leverages these low-confidence areas to construct optimal, low-detection paths through monitored environments. The implications of this research are profound for security professionals, highlighting that even sophisticated AI surveillance systems are not infallible and possess predictable vulnerabilities that can be exploited without specialized adversarial tools or physical interventions.

Background

▶ Watch: Initial findings: person's position affects detector confidence (0:00)

The proliferation of AI in security applications has led to a reliance on computer vision models for tasks such as pedestrian detection and anomaly identification. These systems are often assumed to provide comprehensive and consistent coverage, yet the underlying AI models are complex neural networks susceptible to various forms of "blindness" or reduced performance under specific conditions. Prior work has largely focused on adversarial attacks involving crafted visual perturbations (e.g., printed patterns or specific clothing) or direct interference with sensors (e.g., infrared lasers). However, Shams' research explores a more fundamental and less-studied vulnerability: the intrinsic variability in an object detector's confidence based purely on an object's spatial relationship to the camera.

This problem exists because AI models, despite their advanced capabilities, are trained on datasets that may not perfectly represent all possible viewing angles, distances, and environmental conditions. Consequently, when presented with real-world scenarios, their performance can degrade predictably in certain regions of the camera's field of view. The talk investigates this phenomenon, hypothesizing that these performance variations are not random but rather systematic, creating consistent zones of higher and lower detection confidence. Understanding these zones is crucial, as a person positioned in a low-confidence area might slip past detection if the model's confidence falls below its minimum reporting threshold, effectively creating an undetected passage.

Key Findings

▶ Watch: Analyzing confidence in real-world pedestrian traffic (2:25)

Jacob Shams' research uncovered several crucial findings regarding the behavior of AI object detectors, which form the foundation for understanding and exploiting their vulnerabilities:

  1. Position-Dependent Confidence: The primary finding is that the confidence level of an object detector when identifying a person is significantly dependent on the person's position relative to the camera. This includes their distance, angle, and height within the camera frame. Experiments showed confidence variations of up to 0.7 for Yolo V3 and between 0.35 and 0.45 for Faster R-CNN, depending on the lighting conditions and specific position.
  1. Independence from Model, Quality, and Lighting: This position-dependent behavior was found to be consistent across different object detection models (Yolo V3 and Faster R-CNN), varying video qualities (from 1080p down to 720p and even 480p), and diverse lighting conditions (daytime, nighttime). This indicates a fundamental characteristic of these models rather than an artifact of specific implementations or environmental factors.
  1. Real-World Applicability and Heat Maps: The varying confidence behavior extends to real-world pedestrian traffic in public areas. Analysis of public footage from locations like Shibuya Crossing (Tokyo), Broadway, and Castro Street (San Francisco) revealed distinct areas within the frame that consistently exhibited stronger or weaker detection confidence. These patterns could be visualized as confidence heat maps, illustrating zones where detection was more or less reliable. For instance, in Shibuya Crossing footage, a pole near the camera consistently reduced detection confidence in its vicinity due to obscuration. Average confidence in these real-world scenarios varied significantly, ranging from 0.25 to 0.92 depending on the time of day and location.
  1. Basis for Evasion Attack: These consistent patterns of varying confidence provide a predictable vulnerability that can be exploited. By understanding where these low-confidence zones exist, an attacker can construct a path through a monitored scene with minimal or no detection, leading to the development of the TipToe evasion attack.

Technical Deep Dive

▶ Watch: Real-world heatmaps show varying detection confidence (4:00)

The technical core of Shams' presentation comprises two main analytical phases followed by the development of a novel evasion attack.

The first analysis phase involved controlled lab experiments designed to isolate the impact of a person's position on object detector confidence. Researchers used lab videos featuring a person at various heights, distances, and angles relative to the camera. These videos were then analyzed by two prominent object detection models: Yolo V3 and Faster R-CNN. The results, presented as visual grids and tables, clearly demonstrated that detection confidence varied significantly based on position. For Yolo V3, the change in confidence could be as high as 0.7, while for Faster R-CNN, it ranged between 0.35 and 0.45 depending on lighting. Green indicated average confidence above 0.8, yellow between 0.6 and 0.8, and red below 0.6. This initial phase established that object detector confidence is indeed a function of a person's spatial relationship to the camera.

Further extending this analysis, Shams investigated the effect of camera quality. The original 1080p footage was downscaled to 720p and 480p. The study found that the position-dependent confidence behavior persisted across all resolutions, with confidence ranges of 0.7 for 720p footage and 0.6 for 480p footage. This crucial finding demonstrated that the phenomenon is independent of camera quality, suggesting it's an inherent characteristic of the object detection models themselves rather than a limitation imposed by lower-resolution input.

The second analysis phase moved from controlled lab settings to real-world scenarios, examining pedestrian traffic in public footage. Publicly available 24-hour footage from three global locations – Shibuya Crossing in Tokyo, Broadway, and Castro Street in San Francisco – was utilized. Each 24-hour period was divided into six consecutive four-hour videos to capture varying times of day and lighting conditions. The objective was to determine the average confidence of an object detector over time as a function of location within the observed area. For each video, one frame every two seconds was fed into the object detectors. For every pixel in the frame, the average confidence of person bounding boxes that landed on that pixel over time was calculated. This process, repeated for five different object detectors (though results for Yolo V3 and Faster R-CNN were primarily presented), generated detailed confidence heat maps. These heat maps visually represented areas of higher (green) and lower (red) average detection confidence, with average confidence varying between 0.25 and 0.92. A compelling example was the pole in the Shibuya Crossing footage, which consistently correlated with reduced detection confidence in its immediate vicinity, likely due to visual obstruction.

These findings directly led to the development of TipToe, a novel evasion attack. The core idea behind TipToe is to leverage the varying confidence levels recorded in the confidence heat map to construct a path across a scene with minimal or no detection. The problem space is reduced to a graph-based problem:

  • A scene is defined as the physical area observed by the camera.
  • A confidence heat map records for each pixel (i, j) in the scene the average confidence of bounding boxes that landed on that pixel over time.
  • A detection heat map records the number of bounding boxes that landed on a given pixel over time.

In the graph representation, each pixel in the heat map becomes a node, with two-way edges connecting it to its cardinal neighbors. The weight value of each node is the average confidence measurement from the confidence heat map. The attacker's objective is to find an optimal path between two points (start and end nodes) in this graph that minimizes both the maximum confidence encountered along the path and the average confidence of the path nodes.

The threat model for TipToe is critical:

  • The attacker wants to cross a scene monitored by an AI-based pedestrian detector with minimal or no detection.
  • The attacker has access to the confidence heat map and detection heat map for the given scene, allowing them to understand the location-based behavior of the tracker.
  • Crucially, the attacker does not have access to adversarial perturbation methods (e.g., printed patterns) or machinery like infrared lights or lasers to directly influence the detector.
  • The only factor under the attacker's control is the path taken through the scene. This makes TipToe a highly practical attack, as it requires no specialized equipment or techniques for the person executing the evasion.

TipToe itself is a modified version of Dijkstra's algorithm. Standard Dijkstra's finds the path with the lowest total cumulative cost. TipToe, however, is designed to find the path that crosses the lowest possible maximum confidence node. The algorithm's modified weight node update function calculates, for each point in the scene, the minimum maximum confidence that must be crossed to reach that point. This contrasts with Dijkstra's, which minimizes the total path cost. By minimizing the maximum confidence encountered, TipToe ensures that any other path through the scene would contain a node with either equal or greater maximum confidence. This is particularly useful for an attacker aiming to stay below a fixed detection threshold. Once the algorithm runs, the optimal path can be iteratively reconstructed by starting from the end node and working backward to the start node, similar to standard Dijkstra's path reconstruction.

Demo / Proof of Concept

▶ Watch: TipToe methodology: reducing problem to graph path (6:30)

The efficacy of the TipToe evasion attack was evaluated by applying it to the confidence and detection heat maps generated from the real-world pedestrian traffic analysis. The evaluation methodology was robust:

  • For each confidence heat map, 10 random start points and 10 random end points were chosen, resulting in 100 possible start-end path combinations.
  • For each combination, TipToe generated an optimal path. The average maximum confidence and average per-step confidence of these paths were calculated.
  • These results were then rigorously compared against two baseline path generation methods:
  1. Direct Manhattan distance paths: The shortest path between two points on a grid, moving only horizontally and vertically.
  2. Random length paths: Paths generated without any optimization for confidence.
  • This evaluation was repeated across the five previously mentioned object detection models and three global locations, with results for Yolo V3 and Faster R-CNN specifically highlighted.

The findings from this evaluation clearly demonstrated TipToe's effectiveness. TipToe consistently reduced both the maximum confidence encountered along the path and the average confidence of the path nodes compared to both direct Manhattan distance paths and random paths through the scene. This quantitative proof-of-concept confirmed that it is indeed possible for an attacker to manipulate an object detector's confidence and significantly lower their chances of detection by strategically leveraging the identified low-confidence areas within a surveillance scene. The demonstration validated that the inherent vulnerabilities of AI object detectors, as mapped by the confidence heat maps, can be practically exploited using a pathfinding algorithm.

Defensive Implications

▶ Watch: TipToe's threat model and attacker assumptions (7:35)

The findings presented in this talk carry significant defensive implications for organizations deploying and managing AI-powered CCTV surveillance systems. Recognizing that object detectors have predictable "blind spots" due to position-dependent confidence is the first crucial step. Defenders should consider the following actions:

  • Conduct Vulnerability Mapping: Organizations should actively generate confidence heat maps for their own CCTV deployments. This involves capturing footage under various conditions and analyzing it with their specific object detection models to identify consistently low-confidence areas. This data can reveal critical blind spots that might otherwise go unnoticed.
  • Optimize Camera Placement and Overlap: Armed with confidence heat maps, security teams can strategically reposition cameras or deploy additional cameras to create overlapping fields of view that eliminate or significantly reduce low-confidence zones. The goal is to ensure that any area prone to evasion from one camera is adequately covered by another camera with high detection confidence.
  • Layered Security Measures: Relying solely on AI object detection for critical areas is risky. Defenders should integrate other security technologies, such as traditional motion sensors, pressure plates, infrared tripwires, or even human patrols, in areas identified as having low AI detection confidence.
  • Dynamic Threshold Adjustment (with caution): While lowering detection confidence thresholds might seem like a solution, it can lead to an increase in false positives, overwhelming security personnel. If thresholds are adjusted, it should be done dynamically and intelligently, perhaps with higher thresholds in high-confidence areas and slightly lower (but still managed) thresholds in known weak spots, coupled with other alerts.
  • Regular System Audits: AI models evolve, and environmental conditions change. Regular audits of surveillance system performance, including re-generating confidence heat maps, are essential to ensure ongoing effectiveness against evasion tactics.
  • Educate Security Personnel: Security operators and responders should be aware of the concept of AI blind spots and the potential for evasion. Training should include scenarios where individuals might attempt to exploit these vulnerabilities, prompting more vigilant human observation in critical areas.
  • Consider Multi-Modal AI: Future defensive strategies might involve integrating multiple AI perception modalities (e.g., thermal imaging alongside visible light cameras) or fusing data from different types of sensors to create a more robust and less exploitable detection system.

By proactively understanding and addressing these inherent vulnerabilities, defenders can move beyond a passive reliance on AI and build more resilient and effective surveillance infrastructures.

Key Takeaways

  • AI Object Detector Confidence Varies Systematically: Object detection models like Yolo V3 and Faster R-CNN exhibit significant and predictable variations in detection confidence based on a person's distance, angle, height, and location within a camera's field of view.
  • Blind Spots are Inherent and Consistent: These confidence variations are independent of model architecture, video quality (1080p down to 480p), and lighting conditions, indicating a fundamental characteristic that creates consistent "blind spots" or low-confidence zones in surveillance footage.
  • Real-World Environments Confirm Vulnerability: Analysis of public pedestrian traffic demonstrated that these low-confidence areas exist in real-world scenarios, which can be mapped into "confidence heat maps" for any given surveillance scene.
  • TipToe Evasion is Practical and Effective: The TipToe attack leverages these identified low-confidence zones by employing a modified Dijkstra's algorithm to construct optimal paths that minimize the maximum and average detection confidence, enabling evasion without requiring specialized tools or adversarial perturbations.
  • Defenders Must Proactively Address Blind Spots: Organizations deploying AI CCTV systems need to identify and mitigate these inherent vulnerabilities through strategic camera placement, overlapping coverage, integration of layered security measures, and regular system audits to prevent sophisticated evasion.

About the Speaker(s)

Jacob Shams is a researcher who presented his work on securing CCTV cameras against blind spots at DEF CON 32. While specific affiliations or detailed biographical information were not provided within the transcript, his presentation demonstrates a deep understanding of AI object detection models, graph theory, and practical security implications. His work highlights expertise in analyzing the performance characteristics of computer vision systems and developing novel methods to evaluate and exploit their vulnerabilities.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This research by Jacob Shams is a critical dissection of AI-powered CCTV systems, revealing inherent, predictable blind spots in object detectors like Yolo V3 and Faster R-CNN. By systematically mapping position-dependent confidence, Shams developed "TipToe," a novel and practical evasion attack that allows individuals to navigate monitored areas with minimal detection, requiring no specialized tools. The work provides indispensable insights for defenders, forcing a fundamental re-evaluation of AI surveillance system reliability and necessitating proactive vulnerability mapping and layered security measures.

Heather Calloway (CISO) — MUST SEE

Jacob Shams' DEF CON talk is a critical examination of AI-powered CCTV, revealing inherent and exploitable blind spots in object detection models based purely on spatial positioning. This isn't a theoretical exercise; it's a practical demonstration of how a common security control can be bypassed without specialized tools. The research provides clear, actionable steps for security leaders to identify these vulnerabilities, optimize camera placement, and implement layered defenses, fundamentally changing how organizations should assess and manage physical security risks.

→ Top-rated talks at DEF CON 32 Main Stage

All talks from DEF CON 32 Main Stage