UIHash: Detecting Similar Android UIs through Grid-Based Visual Appearance Representation
Jiawei Li, Jian Mao, Jun Zeng, Qixiao Lin, Shaowen Feng, Zhenkai Liang
33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24
Overview
In the contemporary mobile landscape, user interfaces (UIs) serve as the primary interaction point between users and applications. However, the prevalence of similar UIs in counterfeit applications poses a significant security threat, acting as an attack surface designed to deceive users. Malicious actors leverage familiar UI designs to trick users into installing and interacting with harmful applications, often leading to credential theft or the distribution of malware. A prime example highlighted in the talk is a spoofing UI designed to solicit user credentials, demonstrating the critical need for robust methods to detect UI similarity. This talk, presented by Jiawei Li and colleagues from Beihang University and the National University of Singapore, introduces UIHash, a novel approach to identify similar Android UIs by representing their visual appearance through a grid-based abstraction.

Key moments
- 0:00 Introduction: The problem of similar UIs and limitations
- 2:20 Motivation: How humans perceive UI similarity (Principle of Proximity)
- 4:00 Overview of UIHash approach: From APK to similarity detection
- 5:00 Key features encoded in UIHash: Position, size, and type
- 6:15 UIHash representation and CNN-based similarity detection
- 7:40 Evaluation results: UIHash outperforms prior detection methods
- 8:15 UIHash detects similar UIs bypassing tree/image-based evasions
- 10:00 Combining UIHash with code features for enhanced app similarity detection
UIHash: Detecting Similar Android UIs through Grid-Based Visual Appearance Representation
Speakers: Jiawei Li; Jian Mao; Jun Zeng; Qixiao Lin; Shaowen Feng; Zhenkai Liang
Conference: USENIX Security '24
YouTube: https://www.youtube.com/watch?v=LcvDDmM5KoY
Overview
In the contemporary mobile landscape, user interfaces (UIs) serve as the primary interaction point between users and applications. However, the prevalence of similar UIs in counterfeit applications poses a significant security threat, acting as an attack surface designed to deceive users. Malicious actors leverage familiar UI designs to trick users into installing and interacting with harmful applications, often leading to credential theft or the distribution of malware. A prime example highlighted in the talk is a spoofing UI designed to solicit user credentials, demonstrating the critical need for robust methods to detect UI similarity. This talk, presented by Jiawei Li and colleagues from Beihang University and the National University of Singapore, introduces UIHash, a novel approach to identify similar Android UIs by representing their visual appearance through a grid-based abstraction.
Traditional methods for UI similarity detection, such as pixel-level image comparison or layout tree structure analysis, have proven inadequate against sophisticated evasion techniques. Users often exhibit high tolerance for minor visual discrepancies, and dynamic UI content further complicates image-based detection. Similarly, layout tree comparisons can be bypassed through various structural mutations that do not alter the perceived visual appearance. UIHash addresses these limitations by focusing on how humans perceive UI similarity, guided by principles from the vision domain, thereby offering a more resilient and accurate detection mechanism crucial for enhancing mobile security against deceptive applications.
The significance of UIHash lies in its ability to consistently align with human perception of UI similarity, even when confronted with mutations designed to bypass existing detection tools. By abstracting UI visual features into a multi-dimensional matrix, UIHash provides a powerful representation that can identify repackaged apps, malicious applications, and sophisticated spoofing attempts. The research demonstrates UIHash's superior effectiveness compared to prior methods and highlights its potential for integration with other app analysis techniques to create a more comprehensive defense against mobile threats.
Background
▶ Watch: Introduction: The problem of similar UIs and limitations (0:00)
The challenge of detecting similar UIs in Android applications is rooted in the inherent limitations of conventional analytical approaches. Historically, two primary methodologies have been employed: screenshot image-based detection and layout tree structure-based detection. Both, however, fall short in the face of real-world adversarial tactics and user behavior.
Screenshot image-based detection, which relies on comparing pixel-level features, is intuitively appealing but fundamentally flawed. Users often display a high tolerance for visual changes, making pixel-perfect comparisons unreliable. A study cited in the talk revealed that 40% of users would still trust and attempt to log into a fake Facebook login UI even if its background color was brown – a significant visual alteration. Furthermore, many legitimate applications feature dynamic content, such as images in news feeds, music players, or shopping apps, which constantly update. This dynamism renders static image comparisons ineffective, as the core UI structure might remain similar while pixel data frequently changes.
The alternative, layout tree structure-based detection, attempts to identify UI similarity by analyzing the hierarchical arrangement of UI components. A UI in Android is defined and initialized by a layout tree, making its comparison seem logical. However, this method also suffers from significant vulnerabilities. Minor alterations in the layout tree, sometimes as small as a single byte difference, can result in UIs with entirely different visual appearances. Conversely, malicious actors can construct visually identical UIs using vastly different layout trees. For instance, different layout containers like RelativeLayout or FrameLayout can be leveraged to achieve the same visual arrangement of controls, thereby bypassing tree-based similarity checks. This flexibility in UI construction means that tree similarity is not always a reliable indicator of perceived UI similarity.
Recognizing these limitations, the UIHash approach is guided by principles from the vision domain, specifically the Gestalt principle of proximity. This fundamental cognitive principle states that "people treat objects close together as a group." When humans perceive a UI, they don't analyze individual pixels or raw tree structures; instead, they observe groups of controls, their relative positions, sizes, and types. For example, a user might describe a UI as having "a logo at the top, two big text inputs in the middle, and some small text at the bottom." This description captures important semantic features like control size, type, and position, which are crucial for human recognition and are inherently more robust to minor, non-perceptual changes. UIHash aims to mimic this human perception by abstracting UI visual features into a representation that is resilient to mutations on screenshot images or layout trees that would otherwise bypass prior detection methods.
Key Findings
▶ Watch: Overview of UIHash approach: From APK to similarity detection (4:00)
The research behind UIHash presents several compelling findings that underscore its effectiveness and potential to significantly advance mobile security. The core contribution is that UIHash consistently outperforms prior similar UI detection methods, whether they are based on layout tree structures or screenshot images.
One of the most critical findings is UIHash's superior recall rate in detecting similar UIs. This indicates that it successfully identifies a higher proportion of genuinely similar UIs compared to existing techniques. This improved detection capability is particularly evident in scenarios involving sophisticated evasion techniques. UIHash successfully detected numerous similar UIs that were specifically engineered to bypass tree-based detection methods. For instance, many detected similar UI pairs exhibited large tree differences, with the tree edit distance (TED) sometimes reaching up to four times the size of one of the trees. This highlights the ability of UIHash to look beyond the static, structural definition of a UI and focus on its visual presentation.
The researchers identified specific evasion techniques used by adversaries to mutate tree structures without altering the UI's visual appearance:
- Flexible use of different UI control containers: Adversaries can interchange layout containers such as
LinearLayoutorRelativeLayoutto achieve identical visual layouts while generating distinct layout trees. - Addition of invisible controls: Malicious apps can insert invisible UI controls into a layout tree. These controls do not affect the visual appearance of the UI but significantly alter its underlying tree structure, effectively bypassing tree-based similarity checks.
Beyond tree-based evasions, UIHash also demonstrated its robustness against image-based bypasses. The system successfully detected various real-world threats, including:
- Cloning radio applications: Identical UIs from different apps.
- Repackaging games with additional advertisements: Apps that maintain the original game UI but inject new, often unwanted, elements.
- Spoofing bank applications: Malicious apps designed to mimic legitimate banking interfaces to steal user credentials.
A crucial validation of UIHash's approach came from a user study. This study confirmed that the UI pairs identified as similar by UIHash consistently received high similarity ratings from human users, even when minor details like background color or image content varied. This indicates that UIHash's detection results are more consistent with human perception than those of naive screenshot image-based detections, directly fulfilling its design objective.
Finally, the research explored the synergy between UIHash and other application analysis features. When combining UIHash with code features, the detection of similar applications was significantly enhanced. The integration resulted in 6% more similar apps being detected when adding UIHash to existing code-based similarity detections, and conversely, 7% more similar apps were detected when integrating code-based similarity into UIHash. This finding underscores that a multi-faceted approach, leveraging both visual and code-level insights, provides a more comprehensive understanding of app similarity and improves the overall detection of malicious or repackaged applications.
Technical Deep Dive
▶ Watch: UIHash representation and CNN-based similarity detection (6:15)
The technical foundation of UIHash is built upon a novel, grid-based visual appearance representation that abstracts UI semantics to align with human perception. The entire process begins by taking Android app installer files, specifically APK files, as input.
The first critical step is parsing UI appearance. Unlike traditional methods that rely on static layout trees, UIHash focuses on runtime semantics. This is crucial because different UI control types (e.g., ToggleButton, Switch, ImageButton, Checkbox) can all visually represent a "toggle" function, despite having distinct internal classifications. UIHash re-identifies controls based on their actual visual appearance rather than their claimed names or static definitions, thus capturing how a user would perceive them. To generalize the UIHash representation and ensure it accurately reflects user perception, the system collects and integrates visual features from various UI screen regions. Based on a user study, the chosen features for encoding are control position, size, and type.
The core of UIHash's representation lies in its grid-based partitioning of the UI screen. The screen is divided into a grid, and each grid region captures specific position and layout semantics. This grid structure allows for the abstraction of UI visual features and tolerates minor variations within individual grid cells. To preserve the distinct meaning of different control types, UIHash employs a multi-channel approach:
- Separate control channels: Different control types are separated into distinct channels. For example, there can be an image channel, a text channel, a button channel, a toggle channel, and so on. This ensures that the semantic meaning of a control type is maintained, even if its exact visual rendering varies slightly.
- Encoding control size: Within each grid region and channel, the size of a control is encoded using Intersection Over Union (IOU). IOU measures the overlap between a specific control's bounding box and the bounding box of the grid region. This metric quantifies how much of a control occupies a particular grid cell, providing a robust representation of its size and spatial relationship within the UI.
By integrating these visual semantics from different grid regions and channels, UIHash constructs its unique UI appearance representation: a multi-dimensional matrix. This matrix effectively encapsulates the high-level layout characteristics of the UI, capturing the relative positions, sizes, and types of controls in a way that mirrors human visual grouping.
To distill the semantics of this UIHash matrix and enable UI similarity comparison, the system employs a sophisticated machine learning architecture. Visual features are generalized during the embedding process, and a Convolutional Neural Network (CNN)-based Siamese Network is then applied. Siamese Networks are particularly well-suited for tasks involving similarity measurement, as they learn to embed inputs into a feature space where semantically similar items are mapped close together. The network takes a pair of UIHash matrices as input and outputs a similarity score, indicating the degree of visual resemblance between the two UIs. Based on this similarity score, a threshold-based detection mechanism is used to flag pairs of UIs as similar, providing an automated and reliable method for identifying potentially deceptive or repackaged applications.
Demo / Proof of Concept
▶ Watch: Evaluation results: UIHash outperforms prior detection methods (7:40)
While the talk did not feature a live, interactive demonstration of the UIHash tool itself, the research rigorously validated its capabilities through extensive evaluations on real-world applications. These evaluations served as the primary proof of concept, showcasing UIHash's effectiveness in detecting various forms of UI similarity that bypass existing methods. The results clearly illustrated how UIHash functions in practical scenarios.
The evaluation set comprised a diverse collection of Android applications, including repackaging apps, malicious apps, and a selection of recent apps, providing a robust testing environment. The core of the demonstration of concept lay in presenting the types of similar UIs that UIHash successfully identified, which were previously undetectable.
For instance, UIHash effectively detected cloning radio apps, where different application packages presented visually identical radio player interfaces. This demonstrated its ability to spot direct UI replication, regardless of underlying code or package differences. More critically, the system identified repackaging games with additional advertisements. In these cases, the core game UI remained consistent, but malicious additions altered the app's behavior without significantly changing the user-perceived interface, a scenario where image-based detection would likely fail due to dynamic ad content, and tree-based detection could be bypassed by subtle layout manipulations.
Perhaps the most impactful proof of concept involved the detection of spoofing bank applications designed to steal user credentials. These UIs were crafted to closely mimic legitimate banking interfaces, often with minor visual discrepancies (e.g., background color, specific image content) that users might overlook but are enough to fool pixel-based comparisons. UIHash's ability to consistently rate these pairs with a high similarity score, confirmed by user perception studies, validated its utility in identifying critical phishing threats.
The talk emphasized that all these UI pairs identified by UIHash obtained a high similarity rating score, which was further corroborated by a user study. This user study served as a crucial validation point, confirming that the detection results of UIHash are indeed more consistent with human perception compared to naive screenshot image-based detections. This empirical evidence, derived from real-world application analysis and human judgment, robustly demonstrates UIHash's practical efficacy as a potent tool for mobile security.
Defensive Implications
▶ Watch: Combining UIHash with code features for enhanced app similarity detection (10:00)
The insights and capabilities introduced by UIHash carry significant defensive implications for various stakeholders in the mobile security ecosystem, from app developers to security researchers and platform providers. The central message is clear: traditional methods of UI similarity detection are no longer sufficient against modern evasion techniques, necessitating a shift towards perception-aware analysis.
For app developers and UI designers, UIHash serves as a critical reminder of the importance of distinctiveness. When designing user interfaces, particularly for sensitive applications like banking, e-commerce, or credential management, developers must prioritize unique visual identities that are difficult to spoof. Even subtle changes in layout or control types can be perceived as similar by users and, more importantly, by sophisticated detection systems like UIHash. Developers should consider how their UI's high-level visual characteristics might be misinterpreted or intentionally mimicked, and strive for designs that minimize ambiguity.
Security analysts and incident response teams can leverage UIHash-like approaches to significantly enhance their detection capabilities for malicious Android applications. The ability to identify repackaged apps (e.g., games with injected ads), spoofing apps (e.g., fake banking logins), and outright counterfeit applications more accurately means that threats can be identified and mitigated faster. Integrating UIHash into static or dynamic analysis pipelines for app vetting can provide a crucial layer of defense, especially against threats that cleverly mutate underlying code or layout structures while maintaining a deceptive visual facade. This new paradigm necessitates moving beyond simple code signature matching or pixel-level UI comparisons.
Mobile platform providers (e.g., Google Play Protect) can integrate UIHash's principles into their automated app scanning systems. By analyzing the visual appearance of newly submitted applications against a database of known legitimate and malicious UIs, they can proactively identify and block deceptive apps before they reach a wide user base. This proactive defense is vital for maintaining user trust and the overall security of the app ecosystem. The combination of UIHash with code-based similarity detection, as highlighted in the research, further strengthens this capability, offering a comprehensive approach to app similarity assessment.
Finally, the research highlights the need for a holistic approach to app security analysis. Relying solely on code analysis or static manifest checks is insufficient. The visual layer, which is how users primarily interact with apps, is a critical vector for deception. Defenders must adopt tools and methodologies that understand and analyze the UI from a human perception standpoint, effectively closing a significant gap that adversaries have historically exploited. This includes focusing on runtime UI semantics and understanding how control positions, sizes, and types contribute to overall visual identity, rather than just their static definitions.
Key Takeaways
- UIHash is a novel, grid-based approach for detecting similar Android UIs, designed to overcome limitations of traditional image-based and layout tree-based methods.
- It is guided by the Gestalt principle of proximity, abstracting UI visual features to align with human perception, making it more resilient to minor visual changes and structural mutations.
- UIHash represents UI appearance as a multi-dimensional matrix, encoding control position, size (using IOU), and type across different channels within a grid structure.
- CNN-based Siamese Networks are used to calculate a similarity score between UIHash representations, enabling robust and perception-consistent UI similarity detection.
- The system effectively detects similar UIs that bypass existing methods, including those with large layout tree differences (e.g., up to 4x tree size) and those employing evasion techniques like flexible layout containers or invisible controls.
- UIHash's detection results are more consistent with user perception, as validated by user studies, making it highly effective against spoofing, repackaging, and counterfeit applications.
- Combining UIHash with other app features, such as code analysis, significantly enhances the overall detection of similar apps (e.g., 6-7% more similar apps detected), advocating for a multi-faceted approach to mobile security.
About the Speaker(s)
The talk was presented by Jiawei Li, who introduced the work as a collaborative effort between Beihang University and the National University of Singapore. While the full speaker list includes Jian Mao, Jun Zeng, Qixiao Lin, Shaowen Feng, and Zhenkai Liang, Jiawei Li served as the primary presenter, sharing the detailed insights and findings of their joint research on UIHash. The affiliations highlight a strong academic collaboration in the field of mobile security and UI analysis.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This work introduces UIHash, a novel grid-based approach to detect similar Android UIs by mimicking human perception. It effectively bypasses traditional pixel-level and layout-tree comparisons, offering a robust defense against sophisticated spoofing and repackaging attacks by focusing on runtime visual semantics and leveraging a CNN-based Siamese Network for similarity scoring.
Heather Calloway (CISO) — STRONG ACCEPT
This research on UIHash presents a vital, perception-aware approach to detecting deceptive mobile applications, directly addressing critical business risks like credential theft and malware. It effectively closes a significant gap in traditional UI similarity detection, offering a clear path for mobile security teams and platform providers to enhance app vetting and user protection.