Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes
Quick Answer
This paper shows that The MAR-12 framework utilizes Vision Language Models to detect and explain harmful humor in memes, achieving 80.3% accuracy for humor and 75.9% for hate detection on the PrideMM and Memotion datasets.
Quick Take
This model offers structured reasoning through twelve perspectives and role-aware attention, ensuring transparent explanations, especially for memes with mixed cues.
Key Points
- MAR-12 introduces a novel framework for meme detection using Vision Language Models.
- Achieves 80.3% accuracy in humor detection and 75.9% in hate detection.
- Utilizes twelve structured perspectives for interpreting memes.
- Employs a role-aware soft-gated attention mechanism for enhanced reasoning.
- Confirmed by human and GPT-4 evaluations to provide coherent explanations.
DeepSignal Analysis
What happened
The MAR-12 framework has been developed to detect and explain harmful humor in memes using Vision Language Models. It achieved 80.3% accuracy in humor detection and 75.9% in hate detection on the PrideMM and Memotion datasets. The model employs twelve structured perspectives and role-aware attention for improved interpretability.
Key evidence
- MAR-12 utilizes Vision Language Models to analyze memes, focusing on the interplay of humor and hate elements.
- The framework achieved 80.3% accuracy for humor detection and 75.9% for hate detection on the PrideMM and Memotion datasets.
- Evaluations by humans and GPT-4 confirmed that MAR-12 provides coherent explanations, especially for memes with mixed cues.
Why it matters
Understanding harmful humor in memes is crucial as it can perpetuate negative stereotypes and social issues. The MAR-12 framework addresses the complexities of interpreting memes, which often blend humor with harmful intent. Its structured reasoning and high accuracy rates may enhance the development of more effective moderation tools in online platforms, contributing to safer digital environments.
What to watch
Paper Resources
Source Excerpt
Internet memes intertwine visual cues, textual content, and cultural context, making them particularly challenging to interpret in scenarios where humor, sarcasm, and harmful intent coexist. These complexities highlight the need for explainable meme understanding systems that can provide reliable and structured reasoning to support both accurate classification and human interpretability. However, existing multimodal classifiers either overlook these interdependencies or provide only limited inte
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →Automatic Ordinary Differential Equations Discovery For Biological Systems Using Powered Agentic System
The MEDA system utilizes large language models and symbolic regression to autonomously discover ordinary differential equations for biological systems, achieving strong structural recovery and biologically plausible models. It outperforms existing methods by integrating domain knowledge and mechanistic constraints, demonstrating effective retrieval and extrapolation capabilities.