I know what you MEME! Understanding and Detecting Harmful Memes with Multimodal Large Language Models
Yong Zhuang
Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · LLM Security
Overview
In an era dominated by digital communication, memes have emerged as a pervasive and powerful form of expression, blending images and text to convey ideas, humor, and narratives across social media platforms. While often a source of lighthearted entertainment, memes possess a significant "dark side," readily exploited by malicious actors to spread harmful content. This talk, presented by Yong Zhuang from the University at Buffalo at the NDSS Symposium, delves into the intricate challenges of understanding and detecting these detrimental memes, proposing a novel approach leveraging Multimodal Large Language Models (MLLMs).
Key moments
- 0:00 Introduction to harmful memes and detection challenges
- 2:00 Challenge 1: Multimodal semantic fusion in memes
- 3:10 Challenge 2: Meme composition types and impact
- 5:10 Challenge 3: Propaganda techniques obscuring harmful intent
- 6:10 Novel approach using chain-of-thought reasoning prompts
- 7:10 Superior detection performance against baseline models
- 8:20 Real-world effectiveness and future research directions
I know what you MEME! Understanding and Detecting Harmful Memes with Multimodal Large Language Models
Speakers: Yong Zhuang, Internet Society NDSS Fellow, University at Buffalo
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=v81b0h4FZ2U
Overview
In an era dominated by digital communication, memes have emerged as a pervasive and powerful form of expression, blending images and text to convey ideas, humor, and narratives across social media platforms. While often a source of lighthearted entertainment, memes possess a significant "dark side," readily exploited by malicious actors to spread harmful content. This talk, presented by Yong Zhuang from the University at Buffalo at the NDSS Symposium, delves into the intricate challenges of understanding and detecting these detrimental memes, proposing a novel approach leveraging Multimodal Large Language Models (MLLMs).
The presentation highlights how harmful memes contribute to severe online issues, including abuse, cyber harassment, and even the incitement of real-world crimes. Despite existing efforts to combat this problem, a fully satisfactory solution has remained elusive. Zhuang's research aims to dissect the underlying reasons for this difficulty, identifying three core challenges that complicate effective detection: the subtle interplay between text and images, the structural composition of memes, and the deployment of sophisticated propaganda techniques. By addressing these complexities, the work offers a significant step forward in safeguarding online communities from the pervasive threat of malicious visual-textual content.
Background
▶ Watch: Introduction to harmful memes and detection challenges (0:00)
The proliferation of memes on social media platforms has created a dual-edged sword for online communication. On one hand, memes foster community, creativity, and shared cultural experiences. On the other, their viral nature and often ambiguous interpretations make them a potent vehicle for spreading harmful ideologies, misinformation, and hate speech. Malicious users exploit the inherent virality and often humorous facade of memes to disseminate content that can negatively impact online communities, leading to severe consequences ranging from cyberbullying to real-world violence.
Previous research has attempted to tackle the problem of harmful meme detection, but these efforts have frequently fallen short due to the inherent complexity of the content. Zhuang and his collaborators identified three primary challenges that prevent existing methods from achieving robust and accurate detection:
- Multimodal Semantic Fusion: This challenge arises from the intricate interplay between the visual and textual components of a meme. Harmful intent is often not explicitly stated in either the image or the text alone, but rather emerges from their combined meaning. Subtle nuances, double entendres, or cultural references can convey derogatory or inflammatory messages that are difficult for traditional models to decipher. This "sabotal harmful content" requires a deep, integrated understanding of both modalities.
- Meme Composition: The way a meme is visually structured—its composition—can significantly influence perception and potentially mask harmful intent. This includes elements such as the number of panels (single vs. stitching memes), the type of images used (photos, screenshots, illustrations), the scale of the shot (closeup, medium, longshot), and even implied movement (physical, emotional, or causal). For instance, a multi-panel "stitching meme" might tell a complex story where the harmful message is only apparent when all visual elements are considered in sequence, complicating visual understanding for automated systems.
- Meme Propaganda Techniques: Malicious actors often employ specific rhetorical and psychological tactics within memes to influence opinions or behaviors, thereby obscuring the harmful nature of the content. These propaganda techniques make the expression more subtle and less detectable. Previous work had identified 22 distinct propaganda techniques, such as "ad hominem," "straw man," or "appeal to emotion," which are strategically deployed to manipulate audiences. Detecting these requires not just understanding the content, but also the intent and persuasive strategies behind it.
To investigate these challenges, the researchers first utilized the open-source HAT dataset, which contains human semantic annotations of harmful memes. By comparing human understanding with interpretations from various visual language models and MLLMs, they confirmed that multimodal semantic fusion is indeed a significant hurdle, particularly for traditional models. However, early observations suggested that advanced MLLMs, such as GPT-4, exhibited a nascent capability to understand these complex interactions. For meme composition and propaganda techniques, the team further annotated the well-known HarmMe benchmark data set. This annotation process identified specific compositional elements and propaganda techniques within each meme. Subsequent testing with Explain-HM, a state-of-the-art harmful meme detector, revealed that both meme composition (especially stitching images) and propaganda techniques significantly complicate detection, making harmful content more subtle and harder to identify.
Key Findings
▶ Watch: Challenge 2: Meme composition types and impact (3:10)
The research yielded several critical findings that underscore the complexity of harmful meme detection and highlight the transformative potential of Multimodal Large Language Models (MLLMs) in addressing these challenges.
Firstly, the investigation into multimodal semantic fusion confirmed that the subtle interplay between text and images is a significant barrier for traditional visual language models. However, a crucial observation was that advanced Multimodal Large Language Models (MLLMs), particularly those like GPT-4, demonstrated a notable capability to understand these nuanced semantic fusions. This suggests that MLLMs possess an inherent advantage in processing and integrating information from disparate modalities.
Secondly, the study rigorously demonstrated that meme composition poses substantial challenges for existing harmful meme detection systems. By annotating the HarmMe dataset and evaluating it against the state-of-the-art Explain-HM detector, the researchers found that specific compositional elements, especially stitching images (memes comprising multiple panels), significantly complicate visual understanding and detection. The sequential or comparative nature of these multi-panel memes often requires a holistic interpretation that single-image analysis struggles to provide.
Thirdly, the research confirmed the impact of meme propaganda techniques on detection efficacy. The annotation of the HarmMe dataset revealed that the presence of these 22 identified propaganda techniques makes harmful expressions more subtle and thus less detectable by automated systems. This finding emphasizes that successful detection requires not just content analysis, but also an understanding of the manipulative tactics employed.
Building on these insights, the researchers proposed a novel MLLM-based approach that significantly outperformed all five tested baselines, including other advanced harmful/hate-forming detectors and even the standalone GPT-4 model. Their method achieved remarkably high accuracy and, crucially, a very high precision, indicating a lower rate of false positives. This is a critical advantage in content moderation, where false positives can lead to unjust censorship.
The proposed method showed particular effectiveness in addressing specific categories of harmful memes. For instance, it achieved a near 97% detection rate for stitching memes, a category previously identified as highly challenging due to its complex composition. This demonstrates the model's ability to overcome the difficulties associated with multi-panel visual narratives.
Finally, to assess real-world applicability, the researchers collected a new "in the wild" dataset comprising over 500 memes from Pinterest. The evaluation on this novel dataset confirmed that their method maintained a very good and stable performance, indicating its robustness against the evolving nature and diverse topics of real-world harmful memes. This finding is crucial for practical deployment, as meme trends and content types are constantly changing.
Technical Deep Dive
▶ Watch: Challenge 3: Propaganda techniques obscuring harmful intent (5:10)
The core of the proposed solution lies in leveraging the advanced capabilities of Multimodal Large Language Models (MLLMs) through a meticulously designed chain of thought reasoning prompt. This innovative approach aims to convert the human-like understanding of multimodal semantic fusion, meme composition, and propaganda techniques into a structured, queryable format that MLLMs can process effectively.
The methodology does not involve building a new MLLM from scratch but rather orchestrates existing powerful MLLMs (like GPT-4, as indicated by the comparative results) to perform complex reasoning tasks. The key insight is that by providing MLLMs with a structured "thought process," they can systematically analyze the various facets of a meme to determine its harmful intent.
The chain of thought reasoning prompt is divided into three primary components:
- Adaptation: This initial section serves to contextualize the task for the MLLM. It provides a clear definition of what constitutes a "harmful meme," setting the boundaries and criteria for evaluation. This ensures the model operates within the intended problem space, especially important given the subjective nature of "harmful" content. The speaker noted that for more advanced contemporary MLLMs, this initial adaptation might become less necessary as they inherently possess a broader understanding of common online phenomena like memes.
- Comprehensive Reasoning Prompt: This is the most critical component, designed to systematically guide the MLLM through a multi-faceted analysis. Instead of a single, direct question, the prompt breaks down the detection task into a series of logical steps. It instructs the MLLM to:
- Analyze Multimodal Semantic Fusion: Examine how the text and image interact to convey meaning, looking for subtle or implied harmful content that emerges from their combination.
- Assess Meme Composition: Consider the visual structure, such as the number of panels, image types (photo, screenshot, illustration), scale (closeup, medium, longshot), and implied movement. For instance, if it's a stitching meme, the model is prompted to interpret the narrative flow across panels.
- Identify Propaganda Techniques: Scan for the presence of any of the 22 known propaganda techniques, evaluating how these tactics might be used to obscure or amplify harmful messages.
By systematically checking each potential factor that could contribute to harmful intent, the prompt ensures a thorough and granular examination of the meme.
- Final Decision: After completing the comprehensive reasoning, the prompt instructs the MLLM to synthesize all previous analyses. Based on the aggregated insights from semantic fusion, composition, and propaganda technique identification, the model is then required to draw a conclusive decision on whether the meme is harmful or not. This structured decision-making process mimics human cognitive steps in content moderation, leading to more explainable and accurate outcomes.
The evaluation of this MLLM-based approach was conducted on two existing benchmark datasets, comparing its performance against five different baselines. These baselines included both general-purpose harmful/hate-forming detectors and the raw GPT-4 model itself (without the specialized prompting). The results demonstrated a significant outperformance by the proposed method across key metrics such as accuracy and precision. The high precision is particularly noteworthy, as it indicates a lower rate of false positives, which is crucial for practical content moderation to avoid over-censorship. Specifically, the method's ability to achieve nearly 97% detection rates for stitching memes showcased its effectiveness in tackling compositionally complex content that previously challenged state-of-the-art systems.
Furthermore, the robustness of the approach was tested in an "in the wild" scenario by collecting over 500 new memes from Pinterest. This real-world evaluation confirmed the method's ability to maintain stable and high performance, indicating its adaptability to the dynamic and evolving landscape of online memes. The consistent performance across diverse and novel content highlights the potential for practical deployment in real-world social media moderation.
Demo / Proof of Concept
▶ Watch: Superior detection performance against baseline models (7:10)
While the presentation did not include a live, interactive demonstration of the system in action, the researchers provided compelling evidence of its efficacy through rigorous quantitative evaluations, which serve as the primary proof of concept. The core of their demonstration lay in presenting the performance metrics of their Multimodal Large Language Model (MLLM)-based approach against established benchmarks and real-world data.
The evaluation process itself functioned as a structured demonstration of the method's capabilities. By testing on two existing benchmark datasets, the team quantitatively showed how their method significantly outperformed five different baselines, including dedicated harmful and hate-forming detectors, as well as a raw GPT-4 model. The reported high accuracy and, importantly, high precision (indicating fewer false positives) directly demonstrated the system's ability to correctly classify harmful memes while minimizing errors.
A key aspect of this proof of concept was the specific performance breakdown for challenging meme categories. The nearly 97% detection rate for "stitching memes" provided concrete evidence that the method could effectively overcome the complexities of multi-panel visual narratives, a known weakness of prior detection systems. This specific achievement validated the approach's ability to address the "meme composition" challenge.
Furthermore, the "in the wild" evaluation on a newly collected dataset of over 500 memes from Pinterest served as a robust real-world demonstration. This segment of the evaluation confirmed the method's stability and strong performance on diverse, evolving content, moving beyond controlled benchmark environments. This practical validation underscores the potential for the system to operate effectively in dynamic social media environments. In essence, while not a live software demo, the comprehensive empirical results provided a strong, data-driven proof of concept for the proposed MLLM-based detection framework.
Defensive Implications
▶ Watch: Real-world effectiveness and future research directions (8:20)
The findings from this research offer significant and actionable insights for defenders engaged in the continuous battle against harmful online content. The enhanced understanding and detection capabilities provided by the MLLM-based approach can fundamentally improve how social media platforms and content moderation teams address the pervasive issue of malicious memes.
Firstly, the research underscores the necessity for multimodal analysis in content moderation. Relying solely on text-based analysis or simple image recognition is insufficient for detecting harmful memes, as the true intent often emerges from the intricate multimodal semantic fusion of both elements. Defenders should prioritize and invest in moderation tools that can process and intelligently integrate both visual and textual information, rather than treating them in isolation.
Secondly, the identified challenges of meme composition and meme propaganda techniques provide specific targets for improved defensive strategies. Content moderation systems should be designed to explicitly account for complex visual structures like stitching memes and to recognize the 22 identified propaganda tactics. Integrating these factors into detection algorithms can make systems more robust against subtle and manipulative content. This implies a need for more sophisticated visual analysis beyond basic object recognition, potentially employing graph-based representations for multi-panel memes or fine-tuned classifiers for propaganda cues.
The demonstrated success of leveraging Multimodal Large Language Models (MLLMs), particularly with chain of thought reasoning prompts, suggests a powerful new paradigm for content moderation. Platforms should explore adopting or integrating such MLLM-powered solutions. The high precision achieved by the proposed method, resulting in fewer false positives, is particularly valuable. It means less legitimate content is mistakenly flagged, reducing the burden on human moderators and mitigating concerns about over-censorship, thereby fostering a healthier online discourse.
Looking ahead, the future work outlined by the speaker also offers crucial defensive directions. Extending the method to multilingual scenarios is vital, as cultural nuances and linguistic differences profoundly impact the interpretation of harm. Defenders operating globally must consider these complexities. Furthermore, the idea of integrating advanced techniques like RAG (Retrieval-Augmented Generation) could enhance MLLMs' ability to draw upon up-to-date information and context, improving detection accuracy for novel or rapidly evolving harmful meme trends. Finally, the vision of an AI agent that collaborates with browsers to block or detox memes in real-time represents a proactive, user-centric defense mechanism. Such an agent could empower individual users by providing an immediate protective layer against exposure to harmful content, shifting some of the defense burden from platforms to the user's endpoint.
Key Takeaways
- Complex Nature of Harmful Memes: Detecting harmful memes is inherently challenging due to the intricate interplay of text and images (multimodal semantic fusion), diverse visual structures (meme composition), and the use of manipulative rhetorical tactics (meme propaganda techniques).
- MLLMs Outperform Traditional Methods: Multimodal Large Language Models (MLLMs), particularly when guided by structured reasoning, significantly outperform traditional visual language models and even standalone MLLMs like GPT-4 in detecting harmful memes.
- Chain of Thought Reasoning is Key: A carefully designed chain of thought reasoning prompt enables MLLMs to systematically analyze memes by breaking down the task into components: adaptation, comprehensive reasoning (checking semantic fusion, composition, propaganda), and final decision-making.
- High Accuracy and Low False Positives: The proposed method achieves high accuracy and precision, significantly reducing false positives, which is crucial for effective and fair content moderation on social platforms.
- Robustness in Real-World Scenarios: The approach demonstrated strong and stable performance on "in the wild" memes from Pinterest, indicating its practical applicability and adaptability to the dynamic nature of online content.
- Future Defensive Potential: Future work includes extending detection to multilingual contexts, integrating RAG (Retrieval-Augmented Generation) for enhanced intelligence, and developing AI agents for real-time browser-based content blocking or detoxification.
About the Speaker(s)
Yong Zhuang is a researcher affiliated with the University at Buffalo. He is also recognized as an Internet Society NDSS Fellow, highlighting his contributions and standing within the cybersecurity community. His work on understanding and detecting harmful memes represents a collaborative effort with colleagues from Wuhan University and the University at Buffalo, focusing on critical challenges at the intersection of artificial intelligence and online safety.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Legitimate academic research on a real problem — multimodal harmful meme detection using chain-of-thought prompted MLLMs — with solid empirical results and a clear problem decomposition. Competent work, but the core contribution is prompt engineering on top of GPT-4, which puts a ceiling on how novel this actually is.
Heather Calloway (CISO) — WEAK
Technically credible academic work on multimodal harmful meme detection with solid benchmark results, but it never crosses the threshold from research finding to operational guidance. The defensive implications section reads like a literature review appendix, not a deployment roadmap.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025