ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?
Quick Answer
ArcANE introduces a novel benchmark for Role-Playing Language Agents (RPLAs) that evaluates character alignment with psychological arcs across 17 novels and 80 characters.
Quick Take
Conditioning on Character Arcs significantly outperforms other context strategies, especially in scenarios not covered by the source text, enhancing model performance in open-weight models like ArcANE-8B/32B.
Key Points
- ArcANE benchmark spans 17 novels and 80 principal characters.
- Character Arc segments narratives into psychological phases for evaluation.
- Conditioning on Character Arc outperforms other context strategies.
- Performance gap is largest in scenarios outside the source text.
- Open-weight models fine-tuned on ArcANE data show improved results.
Paper Resources
Source Excerpt
arXiv:2606. 05553v1 Announce Type: new Abstract: Role-playing language agents (RPLAs) should play characters whose values and behavior evolve as the story progresses, not maintain a fixed persona. Existing benchmarks measure factual recall at a given chapter, not whether responses align with the character's psychological trajectory, especially in scenarios the source text never explores.
We introduce ArcANE (Arc-Aware Narrative Evaluation), an automatically constructed benchmark spanning 17 novels and 80 principal characters. …
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →RF-Agent: A Practical Framework for Building Language Agents for RFIC Design
RF-Agent introduces a novel framework for RF circuit design using , creating a unique RF-domain reasoning dataset with over 11,000 samples. The study reveals that domain-specific supervised fine-tuning and semantic retrieval strategies significantly enhance RF reasoning performance, particularly for smaller models.