Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs
Quick Answer
The study evaluates five open-weight language models, including Mistral 7B, using the Adversarial Surface-Form Robustness Dataset (ASRD) with 2,100 prompts.
Quick Take
Results show a harmful compliance rate of 20.27%, with comprehension failures rising significantly for leetspeak and encoded inputs, highlighting vulnerabilities in real-world applications.
Key Points
- The ASRD dataset includes 2,100 prompts across seven surface-form families.
- Mistral 7B primarily drives a harmful compliance baseline of 22.87%.
- Comprehension failure rates reach 65.60% for encoded wrappers.
- Three response behaviors identified: hallucinated benignity, structural collapse, and language drift.
- The findings emphasize the need for robust evaluations in real-world scenarios.
DeepSignal Analysis
What happened
The study evaluates five open-weight language models, including Mistral 7B, using the Adversarial Surface-Form Robustness Dataset (ASRD) with 2,100 prompts. The harmful compliance rate was found to be 20.27%, with comprehension failures significantly increasing for leetspeak and encoded inputs, indicating vulnerabilities in these models under non-canonical inputs.
Key evidence
- The Adversarial Surface-Form Robustness Dataset (ASRD) consists of 2,100 prompts across seven distinct surface-form families.
- The harmful compliance rate across the evaluated models was 20.27%, with Mistral 7B contributing significantly to this figure.
- Comprehension failures for leetspeak, encoded wrappers, and hybrid transformations reached 36.47%, 65.60%, and 34.47%, respectively.
Why it matters
Understanding the vulnerabilities of language models to non-canonical inputs is crucial for their deployment in real-world applications. The significant harmful compliance and comprehension failures highlight the need for improved robustness in these models. As language models are increasingly integrated into various systems, ensuring their safety against diverse input types is essential to prevent potential misuse and harmful outcomes.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Standard safety evaluations of large language models assess harmful requests written in canonical plain text, while models in real-world deployment routinely receive inputs containing emojis, altered spellings, encoded strings, and character-level variations. This work introduces the Adversarial Surface-Form Robustness Dataset (ASRD), comprising 2,100 prompts across seven distinct surface-form families. Five open-weight language models are evaluated across these prompts, producing 10,500 responses. The Quad-State Evaluation Rubric classifies each response into one of four outcomes: harmful compliance, safe response, comprehension failure, or indeterminate. Emoji and invisible Unicode variations cause almost no comprehension failure, with pooled harmful compliance of 20.27% and 17.20% against a 22.87% baseline that is driven mainly by Mistral 7B, whereas leetspeak, encoded wrappers, and hybrid transformations score 2.40%, 0.13%, and 2.40% while comprehension failure rises to 36.47%, 65.60%, and 34.47%. Inspection of raw model outputs reveals three response behaviors: hallucinated benignity, structural collapse, and language drift. Project page: this http URL
| Comments: | Accepted at the NeurIPS 2026 Workshop on Self-Evolving Diversity-Driven Search for Robust AI Systems (EvoRobust). 12 pages, 11 tables. Project page: this https URL Dataset: this https URL |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09033 [cs.CL] |
| (or arXiv:2610.09033v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09033 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Pavan Maddula [view email]
[v1]
Tue, 6 Oct 2026 19:33:24 UTC (20 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.