次回の更新記事:【論文著者監修・コメント】AIエージェントへの人間…(公開予定日:2026年07月27日)
AIDB Daily Papers

トリビアは些細ではない:多言語大規模言語モデルにおける日常的知識の欠落

原題: When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs
著者: Anna Mosolova, Djamé Seddah
公開日: 2026-07-23 | 分野: LLM NLP ベンチマーク 多言語 cs.CL

※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。

ポイント

  • 日常的で文化に根ざした多言語の知識を評価する新しいベンチマーク「TriviaRoomQA」を構築した。
  • 既存の学術的ベンチマークでは捉えきれない、ポップカルチャーや日常的知識におけるモデルの欠落を実証した。
  • 歴史や地理に強い一方で、モデルの性能は言語や日常的なトピックによって大きく変動することが分かった。

Abstract

Quiz rooms, trivia nights, and quiz shows challenge human knowledge across a wide range of topics, from canonical facts to everyday culture. In this paper, we examine whether large language models (LLMs) can perform competitively in such settings, using quiz-style questions to test them on both common and niche topics. We introduce TriviaRoomQA, a multilingual benchmark designed to evaluate everyday, culturally grounded, and long-tail knowledge across 288 topics. The benchmark contains 3,300 parallel multiple-choice questions in six European languages and additional 5,340 French-only questions for a more fine-grained case study. We evaluate 30 open-weight LLMs from European, Asian, and North American providers, covering models from 7 to 70B parameters. We find that models are strong on knowledge-intensive topics such as history, geography, and mathematics, but substantially weaker on everyday popular-culture topics such as celebrities, music, movies, and news. Moreover, model performance varies across languages even for the same underlying questions, suggesting that access to factual knowledge is not always language-independent. In sum, our dataset and experiments demonstrate an important knowledge gap which is not captured by existing academic-based saturated benchmarks.

Paper AI Chat

この論文のPDF全文を対象にAIに質問できます。

質問の例:

AIチャット機能を利用するには、ログインまたは会員登録(無料)が必要です。

会員登録 / ログイン

関連するAIDB記事