What’s Done Is Not What’s Claimed: Detecting and Interpreting Inconsistencies in App Behaviors
Chang Yue (Institute of Information Engineering, Chinese Academy of Sciences)
Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · Mobile Security
Overview
Mobile applications have become indispensable, yet their extensive access to private user data—such as contacts, photos, and locations—poses significant privacy risks. Despite operating systems providing mechanisms for users to grant or deny permissions, a critical gap persists: users often struggle to comprehend why specific permissions are requested, and even when restrictions are in place, apps may surreptitiously access sensitive information without explicit user consent or awareness. This talk introduces InComputer, a novel system designed to bridge this information asymmetry by empowering users to understand the actual behaviors of their apps and assess associated privacy risks.
Key moments
- 0:00 Introduction and core problem of app privacy leakage
- 2:00 InComputer's overall approach: identifying and interpreting inconsistencies
- 4:30 Building attention library to filter user-related behaviors
- 6:00 Identifying inconsistent behaviors and LLM interpretation with risk analysis
- 8:00 InComputer's superior accuracy in detecting risky behaviors
- 8:30 Real-world impact: scale of affected apps and users
- 9:00 Historical trends and shifts in risky app behaviors
What’s Done Is Not What’s Claimed: Detecting and Interpreting Inconsistencies in App Behaviors
Speakers: Chang Yue, Institute of Information Engineering, Chinese Academy of Sciences
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=FTzknkTHoDM
Overview
Mobile applications have become indispensable, yet their extensive access to private user data—such as contacts, photos, and locations—poses significant privacy risks. Despite operating systems providing mechanisms for users to grant or deny permissions, a critical gap persists: users often struggle to comprehend why specific permissions are requested, and even when restrictions are in place, apps may surreptitiously access sensitive information without explicit user consent or awareness. This talk introduces InComputer, a novel system designed to bridge this information asymmetry by empowering users to understand the actual behaviors of their apps and assess associated privacy risks.
The core problem addressed by this research is the inconsistency between an app's claimed functionality (as presented in its user interface) and its actual runtime behaviors. For instance, a user might not realize that a photo-taking app is simultaneously accessing their location and uploading it to a server. InComputer tackles this by first identifying behaviors not adequately disclosed through the app's UI, then interpreting these "inconsistent behaviors" into natural language, and finally providing a comprehensive risk analysis.
By focusing on the user's perspective, InComputer aims to enhance transparency and accountability in the mobile app ecosystem. The system leverages a combination of static analysis, data flow analysis, and large language models to dissect app code, correlate it with UI elements, and articulate complex technical actions in an understandable format. This allows users to make informed decisions about the apps they use, fostering a more secure and privacy-respecting digital environment.
Background
▶ Watch: Introduction and core problem of app privacy leakage (0:00)
The pervasive issue of privacy leakage in mobile applications stems from a fundamental disconnect between app functionality and user comprehension. While modern mobile operating systems like Android and iOS have implemented granular permission models, requiring users to explicitly grant access to sensitive resources, these mechanisms often fall short in practice. Users are frequently presented with permission prompts that lack sufficient context, making it difficult to discern the necessity or implications of granting access. For example, an app requesting camera access might genuinely need it for photo capture, but it could also use that permission to surreptitiously record audio or video.
Beyond explicit permissions, apps can engage in behaviors that, while technically within the bounds of granted permissions, are not transparently communicated to the user through the app's UI. This hidden functionality can lead to privacy violations, where user data is accessed, processed, or transmitted in ways that contradict user expectations or the app's ostensible purpose. Existing approaches to app security often rely on static or dynamic analysis to detect malicious code or policy violations, but they typically do not focus on the user's understanding of app behaviors as presented through the UI. The challenge is not just identifying what an app does, but whether users are adequately informed about it. This lack of transparency undermines user control and trust, creating an environment where privacy breaches can occur even in seemingly legitimate applications.
Key Findings
▶ Watch: Building attention library to filter user-related behaviors (4:30)
InComputer's evaluation and real-world application yielded several significant findings, highlighting both the prevalence of hidden app behaviors and the system's effectiveness in identifying them:
- Superior Inconsistent Behavior Identification: InComputer demonstrated high accuracy in detecting risky or inconsistent behaviors. On a labeled dataset, it achieved a 94.89% identification rate, outperforming existing state-of-the-art tools. Specifically, InComputer identified 74 more inconsistent behaviors than a comparative tool named "Soda." When tested on an Android malware dataset, InComputer maintained a 94.56% risky behavior identification rate and uncovered 27 new risky behaviors not previously identified.
- Widespread Impact on Users: A large-scale scan of apps from the Android market revealed a significant problem. InComputer identified 1,664 risky inconsistent behaviors across 413 distinct applications. These affected apps spanned all categories, with 189 of them boasting over 1 million downloads, indicating that potentially millions of users are unknowingly exposed to these privacy risks. The identified behaviors included unauthorized leakage of location, messages, and contact information, as well as unauthorized audio recording.
- Prevalence of Self-Triggering Risks: A particularly concerning finding was that 77.97% of the identified apps contained self-triggering risky inconsistent behaviors. This means these behaviors could be activated automatically without any user interaction, posing an even more severe and insidious threat to user privacy, as users have no opportunity to intervene or become aware of the action.
- Evolving Landscape of Privacy Risks: An analysis of apps collected over the past 15 years showed a positive trend in overall privacy behavior, with risky inconsistent behaviors significantly decreasing from 26.44% in 2010 to 3.8% in the present day. This reduction is likely attributable to increased privacy concerns among users and stricter market regulations. However, the nature of these risks is shifting. While risky behaviors related to user content information (e.g., contacts, messages) have declined, there has been an increasing trend in those associated with location, Wi-Fi, and Bluetooth data. This shift may reflect changes in user behavior, such as decreased frequency of traditional phone calls and increased reliance on online communication and location-aware services.
Technical Deep Dive
▶ Watch: Identifying inconsistent behaviors and LLM interpretation with risk analysis (6:00)
The InComputer system is meticulously designed to detect and interpret inconsistencies between an app's UI claims and its actual behaviors. This process involves a multi-stage pipeline, combining static analysis, data flow analysis, and advanced natural language processing with large language models.
The first step in InComputer's methodology is static analysis to comprehensively understand the app's internal structure and potential actions. This involves extracting the app's call graph, which maps out the sequence of method calls within the application. Crucially, InComputer then associates UI content with specific APIs within this call graph. This UI-API association is achieved by locating commonly used binding APIs, such as setContentView and findViewById, which explicitly link code logic to visual elements. If a particular node in the call graph lacks a direct binding relationship, it inherits the UI semantics from its parent node, ensuring comprehensive coverage. For UI content, InComputer goes beyond just the layout files, also considering the dynamic text and icons rendered on the UI, as these are the elements directly presented to and perceived by users.
To ensure the completeness of identified behaviors, data flow analysis is employed. This technique tracks the flow of data between different API calls. For example, if the API getLastKnownLocation (which retrieves location data) feeds its output into uploadData (which sends data to a server), InComputer links these two distinct APIs into a single call sequence. This aggregation ensures that a complete, meaningful behavior (e.g., "accessing location and uploading it") is treated as a unified action rather than disparate operations. Additionally, InComputer summarizes common implicit calls associated with UI elements to further extend these call sequences, capturing behaviors that might not be immediately obvious from a simple API-UI mapping.
The next critical stage is filtering for user-related behaviors using an attention library. Recognizing that an app contains an overwhelming number of internal operations, InComputer focuses only on those that directly relate to user interaction or sensitive data. This library was constructed by analyzing API names from a vast number of applications, leading to three key findings:
- Unimportant words appear more frequently in API names than those related to sensitive resources.
- Words that combine with many other words tend to be less important in defining a sensitive action.
- An API is deemed important when it is semantically related to UI elements, implying user interaction or awareness.
Based on these findings, the attention library is built to contain keywords indicative of user-centric or privacy-sensitive operations. Any behavior whose associated API does not contain keywords from this library is filtered out, streamlining the analysis to relevant actions.
With the attention library in place, InComputer proceeds to identify inconsistent behaviors. This is achieved by comparing the attention keywords present in the APIs of a behavior with those found in the corresponding UI elements. If there is a mismatch—meaning an API performs a sensitive action but the UI provides no relevant keywords or information about it—the behavior is marked as inconsistent. A crucial aspect of this is handling cases where there is no UI text related to a sensitive operation at all; in such scenarios, the behavior is automatically flagged as inconsistent because the user is entirely uninformed.
Finally, InComputer employs large language models (LLMs) to interpret these identified inconsistent behaviors into natural language and provide risk analysis. To enhance the LLM's understanding and accuracy, relevant information about APIs and their associated permissions is gathered from official Google documentation and used as external knowledge. Through prompt engineering, this external knowledge is fed into the LLM, activating its capability to translate complex technical call sequences and UI contexts into easily understandable descriptions. During the development, InComputer compared several popular LLMs, with GPT-4 demonstrating the best performance on their labeled dataset, and thus was selected for the final interpretation step. The output format is standardized: InComputer first states "what the app is actually doing," followed by a detailed "risk analysis" for each behavior, providing users with a clear reference to assess potential privacy implications.
Demo / Proof of Concept
▶ Watch: Real-world impact: scale of affected apps and users (8:30)
While the talk did not feature a live, interactive demonstration in the traditional sense, the speaker presented a clear exposition of InComputer's output and validated its effectiveness through a user study. The core proof of concept lies in InComputer's ability to transform raw, technical inconsistencies into actionable, human-readable insights.
The demonstration of InComputer's output format illustrated how the system presents its findings to users. For an identified inconsistent behavior, InComputer first generates a concise statement describing "what the app is actually doing" in natural language. This is followed by a specific "risk analysis" for that behavior. This structured output is crucial for user comprehension, allowing individuals to quickly grasp the nature of the hidden activity and its potential privacy implications. For example, if an app were found to be uploading location data without UI notification, InComputer might output something like: "The app is accessing your device's precise location and transmitting it to a remote server" followed by a risk analysis detailing potential tracking or data misuse.
To validate the clarity and utility of these interpretations, the researchers conducted a user study. Participants were asked to rate InComputer's generated interpretations. The results indicated that participants generally found the outputs to be "easy to understand," "reasonable," and "helpful for understanding the app's behaviors." This user feedback is a critical component of the proof of concept, demonstrating that InComputer not only accurately detects inconsistencies but also effectively communicates them to a non-technical audience, thereby fulfilling its goal of empowering users. The selection of GPT-4 as the underlying large language model was also part of this validation process, as it consistently performed best in generating high-quality interpretations on their labeled dataset.
Defensive Implications
▶ Watch: Historical trends and shifts in risky app behaviors (9:00)
The findings from InComputer have profound implications for various stakeholders in the mobile app ecosystem, guiding defensive strategies for developers, platform providers, and users alike.
For app developers, the core message is one of enhanced transparency and ethical design. Developers must critically assess their applications to ensure that all sensitive data access and transmission behaviors are explicitly and clearly communicated through the user interface. This means moving beyond mere permission requests to actively notify users about why and when certain actions are performed. Tools like InComputer could potentially be integrated into development pipelines as a testing mechanism, allowing developers to identify and rectify UI-behavior inconsistencies before app deployment. This proactive approach can build user trust and reduce the likelihood of privacy violations.
Mobile platform providers (such as Google and Apple) have a crucial role in enforcing greater transparency. They could leverage the principles behind InComputer to implement stricter review processes for app submissions. This might involve automated analysis to detect significant discrepancies between an app's declared functionality (e.g., in app store descriptions or "nutrition labels" as mentioned in the Q&A) and its actual runtime behaviors as observed through UI-API consistency checks. Integrating such analysis could lead to more robust app store policies, potentially flagging apps that are intentionally opaque or misleading. Furthermore, enhancing existing "nutrition labels" or privacy dashboards to reflect a more granular, behavior-centric view, informed by tools like InComputer, would provide users with more accurate and actionable privacy information.
For end-users, InComputer offers a glimpse into a future where they are more empowered to understand and control their digital privacy. While InComputer itself is a research tool, its existence highlights the need for similar user-facing tools that can analyze app behavior and present digestible privacy insights. In the absence of such tools, users should remain vigilant, carefully scrutinizing permission requests, reading app reviews, and being wary of apps that request extensive permissions seemingly unrelated to their core functionality. The decreasing trend of overall risky behaviors is encouraging, but the shift towards location, Wi-Fi, and Bluetooth data leakage indicates that users must remain particularly cautious about apps requesting these specific types of access.
In summary, InComputer underscores the necessity for an ecosystem-wide shift towards greater transparency, where the "claimed" actions of an app truly align with "what's done."
Key Takeaways
- Pervasive Privacy Discrepancies: Despite permission systems, mobile apps frequently exhibit inconsistencies between their user interface claims and actual runtime behaviors, leading to unauthorized access and leakage of private user data.
- InComputer's Novel Approach: The InComputer system effectively detects these inconsistencies by correlating app code (via static and data flow analysis) with UI elements and then interprets these hidden behaviors into natural language using large language models.
- High Detection Accuracy: InComputer achieved a 94.89% identification rate for risky/inconsistent behaviors, outperforming existing tools and identifying numerous previously undetected risky actions in both benign and malware samples.
- Widespread Real-World Impact: A scan of Android market apps revealed 1,664 risky inconsistent behaviors across 413 apps, with 189 apps having over 1 million downloads, potentially affecting millions of users.
- Concerning Self-Triggering Behaviors: A significant 77.97% of affected apps contained self-triggering risky inconsistent behaviors, which activate automatically without user interaction, posing a severe and insidious threat to privacy.
- Evolving Threat Landscape: While overall risky behaviors have decreased over 15 years, there's a worrying trend of increasing inconsistencies related to location, Wi-Fi, and Bluetooth data, suggesting a shift in targeted user information.
About the Speaker(s)
Chang Yue is a researcher from the Institute of Information Engineering, Chinese Academy of Sciences. The presentation at the NDSS Symposium highlights their work in mobile app security and privacy, specifically focusing on the detection and interpretation of inconsistent app behaviors to empower users.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Competent academic systems paper with real engineering behind it — static analysis plus data flow plus LLM-powered natural language explanation is a sensible pipeline and the numbers are credible. Nothing here will make a vendor sweat or rewrite a threat model, but it's honest work that fills a genuine gap in the literature.
Heather Calloway (CISO) — WEAK
Technically sound research with real-world scale numbers, but InComputer stops at the user and never reaches the institutions that could actually act on it. The governance and accountability dimensions — app store policy, regulatory enforcement, enterprise mobile risk — go essentially untouched.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025