A Picture is Worth 500 Labels: A Case Study of Demographic Disparities in Local Machine Learning Models for Instagram and TikTok
Jack West, Lea Thiemt, Shimaa Ahmed, Maggie Bartig, Kassem Fawaz, Suman Banerjee
IEEE Symposium on Security and Privacy 2024 · Day 1 · Continental Ballroom 5
Overview
This talk, presented by Lea Thiemt and Jack West at IEEE S&P, delves into the often-hidden world of local machine learning (ML) models embedded within popular social media applications like Instagram and TikTok. As ML models increasingly shift from cloud-based processing to on-device execution for features like real-time face filters and augmented reality, they gain the capability to infer a vast array of information about users directly from their camera feeds and images. The core curiosity driving this research was to uncover precisely what insights these vision models extract and, critically, whether these inferences exhibit demographic disparities in quality or accuracy.

Key moments
- 0:00 Introduction to local ML models and research motivation
- 1:20 Two core research questions about model insights and disparities
- 2:00 Challenges: AISc, native libraries, and model evaluation
- 3:20 Novel methodology: detection, reconstruction, and evaluation pipeline
- 5:30 Discovered ML models: TikTok (age/gender) and Instagram (512 concepts)
- 6:50 TikTok age estimation disparity: inaccurate for younger demographics
- 8:05 Instagram concept evaluation using real and synthetic data
A Picture is Worth 500 Labels: A Case Study of Demographic Disparities in Local Machine Learning Models for Instagram and TikTok
Speakers: Jack West, Lea Thiemt, Shimaa Ahmed, Maggie Bartig, Kassem Fawaz, Suman Banerjee
Conference: IEEE S&P
YouTube: https://www.youtube.com/watch?v=lb-FWx2W2YM
Overview
This talk, presented by Lea Thiemt and Jack West at IEEE S&P, delves into the often-hidden world of local machine learning (ML) models embedded within popular social media applications like Instagram and TikTok. As ML models increasingly shift from cloud-based processing to on-device execution for features like real-time face filters and augmented reality, they gain the capability to infer a vast array of information about users directly from their camera feeds and images. The core curiosity driving this research was to uncover precisely what insights these vision models extract and, critically, whether these inferences exhibit demographic disparities in quality or accuracy.
The research team developed a novel methodology to overcome significant technical challenges associated with analyzing these proprietary, obfuscated, and often native-code-based models. Their findings reveal that these applications indeed extract extensive visual information, ranging from age and gender estimations to hundreds of specific concepts related to activities, facial features, and objects. More concerningly, the study systematically demonstrates significant demographic biases in these deployed models, highlighting issues such as wildly inaccurate age estimations for younger demographics on TikTok and spurious correlations between extracted concepts and ethnicity groups on Instagram.
This work is highly significant for several reasons. Firstly, it sheds light on the opaque practices of major technology companies regarding on-device data processing, which often operates outside the direct scrutiny of users or privacy regulations. Secondly, by exposing demographic disparities, the research underscores the real-world implications of biased AI, which can lead to discriminatory experiences, mischaracterizations, or even targeted content delivery based on flawed assumptions. Finally, the open-sourcing of their analysis tools provides a critical resource for future research into the security, privacy, and fairness of local ML models, empowering the broader security community to conduct similar investigations.
Background
▶ Watch: Introduction to local ML models and research motivation (0:00)
The proliferation of advanced features in social media applications, such as face filters, augmented reality (AR) overlays, and real-time image enhancements, has necessitated a shift in where the underlying machine learning computations occur. Traditionally, complex ML tasks were offloaded to powerful cloud servers. However, for applications requiring near real-time responsiveness and low latency, processing on the user's local device (smartphone, tablet) has become increasingly common. This trend, while offering performance benefits and potentially reducing data transmission, introduces a new layer of complexity and opacity regarding what data is processed and what inferences are made locally, often without explicit user consent or awareness.
The problem investigated by the researchers stems from the dual nature of these local ML models. While primarily designed for benign tasks like applying a mesh to a face for a filter, their underlying capabilities often extend to inferring a wide range of additional characteristics. For instance, the same model identifying facial landmarks for a filter could simultaneously detect facial hair, estimate age, or infer gender. This raises critical privacy and security questions: What else are these models looking for? And are these inferences accurate and fair across different user demographics?
Analyzing these embedded, proprietary ML models presents substantial technical hurdles. The Android ecosystem, where these apps predominantly operate, executes applications using two primary components: Java code and Native libraries. Java code runs on the Android Runtime (ART), undergoing compilation and execution on the CPU. Native libraries, typically written in C/C++, are pre-compiled and executed directly by the CPU, offering performance advantages but posing greater challenges for analysis. The research team identified three major technical obstacles:
- Obfuscation (AISc): Android applications, especially those from large companies, often employ Application-Level Instruction Set Obfuscation (AISc). This technique renders the decompiled code extremely difficult for human researchers to understand, even if it's perfectly executable by a machine. It's akin to "looking for a needle in a haystack" without clear signposts.
- Native Libraries: The segregation of Java and native code creates a challenge. These two contexts, operating in different languages and execution environments, are not naturally designed to communicate for the purpose of analysis. Bridging this gap to understand the full execution flow, especially when ML models are often implemented in high-performance native code, is crucial but complex.
- Model Evaluation: Unlike standardized benchmarks, the ML models embedded in commercial apps are unique and custom-designed. Their inputs, outputs, and internal logic are proprietary. Consequently, evaluating their performance, particularly for specific tasks like age estimation or concept recognition, requires a tailored approach for each application, making a generic evaluation methodology impractical. Prior work has often focused on publicly available models or cloud-based APIs, leaving on-device, proprietary models largely unexplored.
Key Findings
▶ Watch: Challenges: AISc, native libraries, and model evaluation (2:00)
The research successfully uncovered and analyzed the functionality of on-device computer vision models within TikTok and Instagram, revealing significant insights and concerning demographic disparities.
For TikTok, the researchers identified a model primarily focused on age and gender estimation. Key findings include:
- Live Execution: The model continuously executes in real-time while the in-app camera feed is running, meaning inferences are made actively as a user interacts with the camera.
- Native Code Implementation: Crucially, TikTok's model is implemented entirely within native code. This makes it particularly challenging to analyze using conventional Java-centric reverse engineering techniques, explaining why existing solutions might fail to extract its runtime behavior.
- Severe Age Estimation Disparities: When evaluated using the FairFace dataset, TikTok's age estimation model exhibited stark inaccuracies for younger demographics. For individuals labeled between 0-2 years old (babies and toddlers), the model consistently estimated an average age of 13 years old. While accuracy improved for older demographics, providing "proper age ages" for individuals in their 40s and above, its complete failure for children renders it "completely wrong" for this critical demographic. This bias could have significant implications if used for age-gating, content filtering, or personalized advertising.
For Instagram, the discovered model was considerably larger and more complex, inferring a broad range of concepts:
- Java Layer Execution: Unlike TikTok, Instagram's ML model primarily executes within the Java layer of the application.
- Extensive Concept Recognition: The model is designed to recognize over 512 distinct concepts. These concepts span a vast array of categories, including:
- Activities: e.g., "dancing"
- Facial Features: e.g., "eyeglasses," "facial hair"
- Locations: (implied from context, though specific examples not given)
- Items/Objects: e.g., "bicycle"
- Image Features: e.g., "motion blur"
- Demographic attributes: (implied, though not directly stated as age/gender)
- Spurious Correlations and Demographic Bias: Evaluating Instagram's model using both real and synthetic data (generated by Rosenberg et al. and with backgrounds removed to prevent bias), the researchers found "spurious correlations" between extracted concepts and ethnicity groups. Examples cited include:
- "Eyeglasses" being a commonly associated concept for Asian men.
- "Great Wall of China" being associated with Asian women.
- "Drag" being associated with Black women.
- "Nudity" being associated with White men.
These correlations are deeply problematic, indicating that the model's training data or internal logic has learned biased and potentially stereotypical associations, which could lead to mischaracterization, inappropriate content recommendations, or even discrimination based on perceived ethnicity and gender.
In summary, the key findings underscore three main points:
- Social media apps actively extract visual information from local images and camera feeds.
- The deployed on-device models exhibit significant demographic disparities in their performance and inferred concepts.
- These disparities manifest as spurious correlations between extracted concepts and ethnicity groups, as well as severe inaccuracies in demographic estimations for specific groups.
Technical Deep Dive
▶ Watch: Novel methodology: detection, reconstruction, and evaluation pipeline (3:20)
To address the inherent challenges of analyzing obfuscated, proprietary, and mixed-language ML models on Android, the researchers designed a novel, three-stage pipeline: ML Detection, Pipeline Reconstruction, and Model Evaluation.
1. ML Detection
The first challenge was detecting where and how machine learning operations were occurring within the complex and often obfuscated application code (AISc). To overcome this, the team employed a dynamic instrumentation approach on a modified LineageOS operating system.
- Operating System Instrumentation: They added specialized code directly into the Android operating system. This code was designed to attach itself to critical components of the Android Runtime (ART) and monitor function calls in real-time.
- Logging Function Calls: While the target app (Instagram or TikTok) was running, their instrumented OS logged out comprehensive information about every function call, including:
- Function names
- Return values
- Function arguments and their data types
- ML Keyword String Search: Simultaneously, they performed a string search for common machine learning keywords (e.g., "tensor," "model," "neural," "AI," "predict," "infer," "face," "vision") within the logged function names, arguments, and return values.
- Evidence Collection: If a match was found, the corresponding function call was marked as "ML evidence." This dynamic, runtime analysis allowed them to pinpoint potential ML execution points even within heavily obfuscated code, as the actual execution flow and parameters would still reveal ML-related operations. The use of
logcatwas instrumental in capturing this verbose output.
2. Pipeline Reconstruction
Once ML evidence was collected, the next step was to reconstruct the full ML pipeline, understanding the flow of data from input to output. This involved dissecting both the Java and native code components.
- Code Disassembly and Decompilation:
- For the Java code segments identified through the ML detection phase, they used JADX. JADX is a powerful decompiler that converts Android Dalvik bytecode (
.dexfiles) back into human-readable Java source code. This provided "pseudo Java code" for analysis. - For the native library segments (typically
.sofiles) where ML operations were suspected, they employed Ghidra. Ghidra is a sophisticated software reverse engineering suite developed by the NSA. It provides extensive capabilities for disassembling, assembling, decompiling, graphing, and scripting. It was used to generate "pseudo C code" from the native binaries. - Bridging Java and Native Contexts: The challenge of connecting the Java and native execution contexts was addressed by carefully analyzing the interfaces between them. Java Native Interface (JNI) calls are the standard mechanism for Java code to invoke functions in native libraries and vice-versa. By monitoring JNI calls during the ML detection phase and then examining the decompiled code from both JADX and Ghidra, the researchers could trace the data flow across the language barrier, effectively rebuilding the end-to-end ML pipeline.
- Understanding Model Boundaries: This reconstruction phase allowed them to identify the inputs to the ML models (e.g., camera frames, selected images) and their outputs (e.g., concept scores, age estimations).
3. Model Evaluation
With the ML pipelines reconstructed, the final stage involved evaluating the performance and biases of the models. This step required tailored approaches for each application due to their distinct execution patterns.
- Internal and External Injection Methods: The researchers developed two primary methods for feeding test data to the models and observing their outputs:
- Internal Injection (Instagram): Instagram's ML model was observed to execute primarily when a user selected or took a picture. To evaluate this, the researchers programmatically fed a large dataset of test images directly into the phone's storage. They then dynamically executed the user interaction (e.g., simulating a picture selection) within the app. This allowed them to monitor the ML model's output for each injected image without actual user interaction.
- External Injection (TikTok): TikTok's model, conversely, continuously executes when the user opens the in-app camera. For this scenario, the researchers used an external display mechanism. They would programmatically display test images to the device's camera, effectively "showing" the images to the in-app camera feed. The model's output was then dynamically monitored as it processed these live camera frames.
- Dataset Application:
- For TikTok's age estimation, they utilized the FairFace dataset. This dataset is specifically designed for fairness research in facial recognition and contains diverse demographic representation, allowing for robust evaluation of age estimation across different groups.
- For Instagram's concept recognition, they used both real and synthetic data generated by Rosenberg et al. Crucially, they removed backgrounds from these images to eliminate potential biases introduced by environmental context, ensuring that any observed correlations were truly tied to the subject's demographics.
- Output Analysis: The model outputs, such as "concept scores" for Instagram or estimated "age" for TikTok, were logged and analyzed against the known ground truth of the test images. This allowed for the quantitative assessment of accuracy and the identification of demographic disparities and spurious correlations.
By combining dynamic instrumentation with sophisticated reverse engineering tools like JADX and Ghidra, and designing app-specific evaluation techniques, the researchers successfully bypassed the significant barriers imposed by commercial app development, revealing the hidden functionalities and biases of on-device ML models.
Demo / Proof of Concept
▶ Watch: TikTok age estimation disparity: inaccurate for younger demographics (6:50)
The talk describes the systematic methodology used to demonstrate the functionality and biases of the local ML models in Instagram and TikTok, which effectively served as their proof of concept. Rather than a live, interactive demo during the presentation, the researchers detailed their "internal and external injection" methods as the core mechanism for demonstrating their findings.
For Instagram, the proof of concept involved the internal injection method. The researchers fed a curated set of images, including both real and synthetically generated demographic data (from Rosenberg et al., with backgrounds removed to prevent bias), directly into the phone's file system. They then programmatically triggered the Instagram app to "select" these images, simulating a user interaction. The instrumented operating system, running their custom code, would then log the concept scores that Instagram's ML model calculated for each image. This allowed them to demonstrate that Instagram's model was indeed extracting hundreds of concepts and to show the specific, often spurious, correlations between these concepts (e.g., "eyeglasses") and particular demographic groups (e.g., "Asian men").
For TikTok, the proof of concept leveraged the external injection method. Since TikTok's model executes continuously when the in-app camera is open, the researchers demonstrated its behavior by displaying test images to the phone's camera. This was done dynamically, effectively showing a sequence of diverse faces to TikTok's live camera feed. Their instrumentation then captured the model's real-time outputs, specifically the estimated age and gender. This method directly demonstrated the live execution of TikTok's native-code model and was crucial for proving the severe age estimation disparities observed for younger demographics, where babies and toddlers were consistently misclassified as teenagers.
These injection techniques, coupled with the dynamic instrumentation and pipeline reconstruction detailed in the technical deep dive, collectively formed the robust proof of concept for their claims. They were able to reliably and repeatedly trigger the ML models within the target applications with controlled inputs and observe their outputs, providing concrete evidence of the models' operations and biases. The researchers also mentioned their open-source library, which contains all their code and resources, allowing others to replicate and extend their demonstrations and findings.
Defensive Implications
▶ Watch: Instagram concept evaluation using real and synthetic data (8:05)
The findings of this research carry significant defensive implications for users, developers, and policymakers concerned with privacy, fairness, and security in the age of on-device AI.
- User Awareness and Privacy: Users are largely unaware of the extent to which their images and camera feeds are analyzed locally by social media apps. The extraction of hundreds of concepts, age, and gender estimations happens without explicit consent or transparent disclosure. Defenders must advocate for greater transparency from app developers, requiring clear notification about what information is inferred locally and how it is used. Users should be empowered with granular controls to opt-out of such analysis or manage inferred data. This research highlights that "local processing" does not automatically equate to "privacy-preserving" if the inferences are still extensive and potentially shared or used to profile users.
- Bias in AI and Fair Use: The demonstrated demographic disparities, such as TikTok's inaccurate age estimation for children and Instagram's spurious correlations, are a major concern. These biases can lead to discriminatory outcomes. For instance, children being misidentified as teenagers could expose them to inappropriate content or advertising. Spurious correlations could reinforce stereotypes, lead to mischaracterizations (e.g., associating "nudity" with a specific demographic), or influence content moderation algorithms unfairly. Defenders (including civil rights advocates and AI ethicists) must push for mandatory fairness audits of on-device ML models, similar to those for cloud-based AI, ensuring that models are tested against diverse datasets and do not perpetuate or amplify societal biases.
- App Development and Security Best Practices: App developers have a responsibility to design and deploy ML models ethically. This research serves as a stark warning about the potential for unintended biases in training data or model architecture.
- Data Hygiene: Developers must prioritize diverse and balanced training datasets to minimize bias.
- Regular Audits: Implementing regular internal audits for fairness and accuracy, especially for models that infer sensitive demographic information, is crucial.
- Explainable AI (XAI): Where possible, developers should strive for more explainable models, allowing for better understanding of why a model makes certain inferences, which can help in identifying and mitigating biases.
- Minimal Data Principle: Only collect and infer the absolutely necessary information for a given feature. If a face filter doesn't need to know a user's age or ethnicity, those inferences should not be performed.
- Regulatory and Policy Frameworks: Existing privacy regulations (e.g., GDPR, CCPA) often focus on data collection and transmission to servers. The findings highlight a gap in addressing on-device inference and the creation of user profiles locally. Policymakers need to consider extending regulatory frameworks to cover local data processing and inference, particularly when it involves sensitive attributes or could lead to discrimination. This includes defining what constitutes "sensitive information" when inferred by an AI, even if it never leaves the device.
- Research and Open-Source Tools: The open-sourcing of the researchers' methodology and tools is a significant defensive contribution. It empowers the broader security research community to conduct similar investigations into other applications and platforms. This collective effort is essential for holding technology companies accountable and for driving improvements in the transparency, privacy, and fairness of on-device AI. Defenders can leverage these tools to perform independent audits and provide evidence-based arguments for policy changes.
In essence, the talk underscores that while local ML offers performance benefits, it introduces a new frontier for privacy and fairness challenges. A proactive, multi-faceted defensive strategy involving user education, developer accountability, robust auditing, and evolving regulatory frameworks is necessary to mitigate the risks posed by these increasingly sophisticated, yet often opaque, on-device AI systems.
Key Takeaways
- Social media apps utilize sophisticated on-device ML models to extract extensive visual information from users' camera feeds and images, often beyond what is necessary for core features like face filters.
- These local ML models, as demonstrated in TikTok and Instagram, exhibit significant demographic disparities and biases, leading to inaccurate demographic estimations and spurious correlations with ethnicity.
- TikTok's age estimation model is severely biased against younger demographics, consistently misidentifying babies and toddlers (0-2 years old) as 13-year-olds, while performing more accurately for older age groups.
- Instagram's model extracts over 512 concepts, showing concerning spurious correlations such as associating "eyeglasses" with Asian men, "Great Wall of China" with Asian women, and "nudity" with White men.
- Analyzing these proprietary models requires overcoming significant technical challenges like obfuscation (AISc), native code implementation, and developing app-specific evaluation methods (e.g., internal/external injection).
- There is an urgent need for greater transparency from app developers regarding on-device AI inference, alongside robust fairness audits and privacy-focused regulatory frameworks to address the ethical implications of local data processing and potential discrimination.
About the Speaker(s)
The talk was presented by Lea Thiemt and Jack West, who are members of the research team behind this paper. They were joined by co-authors Shimaa Ahmed, Maggie Bartig, Kassem Fawaz, and Suman Banerjee.
The speakers are associated with the Wipy Lab (Wireless Security and Privacy Lab) and the Mad Security Group at Wisc Medicine, indicating their expertise in wireless security, privacy, and broader computer security research. Their work focuses on understanding and addressing security and privacy challenges in various technological domains, including the analysis of machine learning models in mobile applications.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This research brilliantly dissects the opaque world of on-device ML in Instagram and TikTok, revealing extensive visual inference and disturbing demographic biases. The novel methodology for analyzing obfuscated native models is a significant technical achievement, providing critical transparency into practices that impact user privacy and fairness daily.
Heather Calloway (CISO) — STRONG ACCEPT
This research uncovers critical, often hidden, biases in on-device machine learning models used by major social media platforms. It highlights significant governance failures where these models make demographic inferences, including for children, without transparency or adequate oversight, leading to clear privacy and discrimination risks. The work provides a strong foundation for future audits and regulatory action, demanding accountability from product teams and boards.
→ Top-rated talks at IEEE Symposium on Security and Privacy 2024