次回の更新記事:【論文著者監修・コメント】AIエージェントへの人間…(公開予定日:2026年07月27日)
AIDB Daily Papers

日常のジレンマにおける価値観の衝突を大規模言語モデルで評価するベンチマーク「D2VBench」の開発

原題: D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios
著者: Siyi Hao, Yidi Cao, Linhao Yu, Yuqi Ren, Deyi Xiong
公開日: 2026-07-22 | 分野: LLM ベンチマーク cs.CL アライメント AI安全性

※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。

ポイント

  • 日常の様々な価値観の衝突を含む1万件のシチュエーションで構成される新しい価値アライメントの評価ベンチマークを提案した。
  • 従来のベンチマークが抱える網羅性の不足や評価手法の単純さを克服し、多角的な価値観の不整合を正確に評価できるようにした。
  • 8つの主要な大規模言語モデルを評価した結果、提案手法が高い信頼性と頑健性を持ち、価値アライメントの度合いを効果的に示すことを明らかにした。

Abstract

With the wide application of large language models (LLMs) in real-world scenarios, the value implication of their outputs is crucial. However, existing evaluation benchmarks suffer from insufficient coverage of value dilemmas in daily scenarios involving multiple value conflicts and simplistic evaluation formalisms that fail to assess LLMs' value alignment. To address these issues, we propose D2VBench, a value alignment benchmark comprising 10,000 instances of real daily dilemma scenarios constructed through a multi-stage collaboration between LLMs and humans, grounded in 158 manually annotated fine-grained value concepts. For evaluation on the benchmark, we present a hybrid evaluation paradigm that integrates multiple-choice questions with open-ended questions. We conduct comprehensive evaluations on eight mainstream LLMs. Experimental results demonstrate that D2VBench exhibits high reliability and robustness, effectively reflecting the LLMs' alignment across different value categories and dimensions, and providing a more realistic and fine-grained tool for research on value alignment. The dataset is available at https://github.com/tjunlp-lab/D2VBench.

Paper AI Chat

この論文のPDF全文を対象にAIに質問できます。

質問の例:

AIチャット機能を利用するには、ログインまたは会員登録(無料)が必要です。

会員登録 / ログイン

関連するAIDB記事