D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios
Quick Answer
D2VBench introduces a benchmark for evaluating large language models (LLMs) on value dilemmas, featuring 10,000 real-world scenarios and a hybrid evaluation method.
Quick Take
The results indicate high reliability and robustness across eight mainstream , enhancing research on value alignment.
Key Points
- D2VBench includes 10,000 instances of real daily dilemma scenarios.
- The benchmark is based on 158 manually annotated fine-grained value concepts.
- A hybrid evaluation method combines multiple-choice and open-ended questions.
- Comprehensive evaluations were conducted on eight mainstream LLMs.
- D2VBench provides a realistic tool for assessing LLMs' value alignment.
DeepSignal Analysis
What happened
D2VBench is a newly introduced benchmark designed to evaluate large language models (LLMs) based on their handling of value dilemmas. It includes 10,000 real-world scenarios and employs a hybrid evaluation method that combines multiple-choice and open-ended questions. The benchmark aims to enhance the assessment of LLMs' value alignment.
Key evidence
- D2VBench consists of 10,000 instances of real daily dilemma scenarios, created through collaboration between LLMs and humans.
- The benchmark is grounded in 158 manually annotated fine-grained value concepts, addressing the limitations of existing evaluation benchmarks.
- Comprehensive evaluations were conducted on eight mainstream LLMs, demonstrating high reliability and robustness in reflecting their alignment across various value categories.
Why it matters
The introduction of D2VBench addresses significant gaps in the evaluation of LLMs, particularly regarding their alignment with human values. By focusing on real-world dilemmas and employing a hybrid evaluation approach, it provides a more nuanced understanding of how LLMs navigate complex value conflicts. This could lead to improved development and deployment of LLMs in sensitive applications, enhancing their societal impact.
Paper Resources
Source Excerpt
With the wide application of (LLMs) in real-world scenarios, the value implication of their outputs is crucial. However, existing evaluation benchmarks suffer from insufficient coverage of value dilemmas in daily scenarios involving multiple value conflicts and simplistic evaluation formalisms that fail to assess LLMs' value alignment. To address these issues, we propose D2VBench, a value alignment benchmark comprising 10,000 instances of real daily dilemma scenarios constr
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.