ZeroR@CHiPSAL 2026: Two-Stage Vision-Language Adaptation with Contrastive Learning for Nepali Meme Classification
Quick Answer
This paper shows that The ZeroR@CHiPSAL 2026 system employs a two-stage vision-language adaptation using Qwen3-VL-8B-Instruct for Nepali meme classification, achieving 2nd place in hate speech detection (F1: 0.797) and 4th in sentiment analysis (F1: 0.518).
Quick Take
This approach integrates LoRA fine-tuning and contrastive learning while addressing class imbalance through various techniques, enhancing performance in low-resource languages.
Key Points
- Utilizes Qwen3-VL-8B-Instruct for effective vision-language processing.
- Achieved 2nd place in hate speech detection with an F1 score of 0.797.
- Implemented a two-stage training pipeline with LoRA fine-tuning.
- Addressed class imbalance through oversampling, augmentation, and focal loss.
- Eliminated error propagation by leveraging native Devanagari understanding.
Paper Resources
📖 Reader Mode
~2 min readAbstract:This paper presents our system for the CHiPSAL 2026 shared task on multimodal hate speech and sentiment detection in Nepali memes. We address both subtasks: binary hate speech classification and three-class sentiment analysis. Our approach adapts the Robust Adaptation of Hateful Meme Detection (RA-HMD) framework using Qwen3-VL-8B-Instruct, a state-of-the-art vision-language model with native Devanagari support. We employ a two-stage training pipeline: (1) LoRA fine-tuning with an MLP projection head for generative classification, and (2) contrastive backbone fine-tuning with supervised InfoNCE loss. We handle class imbalance through minority oversampling, image augmentation, and focal loss. At inference, we ensemble Stage 1 token probabilities with Stage 2 classifier scores using validation-tuned weights. Our end-to-end approach eliminates error propagation from separate OCR and translation pipelines by leveraging the model's native Devanagari understanding. Our system achieved \textbf{2nd place} on hate speech detection (F1: 0.797) and \textbf{4th place} on sentiment analysis (F1: 0.518). We provide detailed ablations, error analysis, and insights into adapting large vision-language models for low-resource South Asian languages.
| Comments: | 9 pages, 2 figures, system description paper for the CHiPSAL 2026 shared task at LREC 2026 |
| Subjects: | Computation and Language (cs.CL) |
| ACM classes: | I.2.7; I.5.1; I.4.0 |
| Cite as: | arXiv:2607.28637 [cs.CL] |
| (or arXiv:2607.28637v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.28637 arXiv-issued DOI via DataCite |
|
| Journal reference: | Proceedings of the Second Workshop on Challenges in Processing South Asian Languages (CHiPSAL 2026) @ LREC 2026, pages 275-283, Palma, Mallorca, Spain, 16 May 2026. ELRA Language Resources Association (ELRA). ISBN 978-2-493814-66-1 |
Submission history
From: Nitiz Khanal [view email]
[v1]
Tue, 19 May 2026 15:14:58 UTC (31 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.