次回の更新記事:【論文著者監修・コメント】AIエージェントへの人間…(公開予定日:2026年07月27日)
AIDB Daily Papers

マルチモーダル大規模言語モデルによる計算論的ユーモア:手法、データセット、評価、そして課題

原題: Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
著者: Tuo Liang, Zhe Hu, Disheng Liu, Jing Li, Yu Yin
公開日: 2026-07-21 | 分野: LLM マルチモーダル AI MLLM cs.CL cs.AI

※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。

ポイント

  • ミームやカートゥーンなどの視覚的ユーモア理解に関する研究動向を体系的に調査した。
  • 認識、解釈と推論、生成という能力中心の階層を用いて既存の文献やモデルを整理した。
  • 評価手法の限界や文化的背景の不足など、分野の発展における主要な障壁を明らかにした。

Abstract

Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description. This survey focuses on visual humor understanding in single-image and multi-panel artifacts, while treating humor generation as an emerging downstream frontier. We position the literature against prior humor, sarcasm, and general MLLM surveys and organize it using a capability-centric hierarchy spanning recognition, interpretation and reasoning, and generation. Under this lens, we synthesize benchmark design, evaluation protocols, and modeling paradigms, tracing the field's shift from task-specific fusion models to large-model approaches based on multimodal alignment, evidence-grounded reasoning, and controlled generation. We conclude by highlighting the main barriers to progress: shortcut-prone evaluation, limited cultural and narrative coverage, weak evidence grounding, and unresolved safety and ownership concerns.

Paper AI Chat

この論文のPDF全文を対象にAIに質問できます。

質問の例:

AIチャット機能を利用するには、ログインまたは会員登録(無料)が必要です。

会員登録 / ログイン

関連するAIDB記事