次回の更新記事:エージェントが手順を飛ばす原因はスキルファイルの…(公開予定日:2026年07月12日)
AIDB Daily Papers

LLMを活用した顔の表情とテキストの融合によるマルチモーダル性格診断

原題: LLM-based Multimodal Personality Recognition via Facial Action Unit-Text Semantic Fusion
著者: Tianyi Zhang, Wei Shan, Yuan Zong, Tianhua Qi, Wenming Zheng
公開日: 2026-06-29 | 分野: LLM マルチモーダル cs.AI cs.CV 性格分析

※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。

ポイント

  • ビデオ面接における顔の表情データと回答テキストをLLMで統合し、性格を推定する新しいフレームワークを提案した。
  • 表情の動きをテキスト記述に変換してLLMで処理することで、非言語的な手がかりを効率的に活用し、情報の欠落を防いだ。
  • ベンチマーク試験において、従来手法よりも高い精度で性格特性を予測し、モデルの解釈性と学習の安定性を向上させた。

Abstract

Personality recognition in asynchronous video interviews (AVIs) has become increasingly important due to their widespread adoption in modern recruitment. Existing approaches often rely on large language models (LLMs) to analyze textual responses of interviewees in AVI. However, unimodel methods often suffer from information loss (e.g., ignore facial cues). In contrast, multimodal methods that employ full-face images or sparsely sampled frames can discard fine-grained temporal dynamics critical for accurate personality assessment. To overcome these limitations, we propose an LLM-based framework that semantically fuse facial action units (AUs) with textual responses of AVI. AU sequences are first converted into interpretable textual descriptions, which are then fused with participants' textual responses through an LLM. A lightweight regression head transforms the resulting embeddings into continuous personality scores without disrupting the underlying semantic space. Experiments on the AVI-6 benchmark demonstrate consistent improvements over most baselines, with lower prediction errors and stronger correlations with human-rated scores across multiple traits. Further analysis reveals that AU-derived semantic representations offer complementary non-verbal cues to textual responses. Decoupling semantic understanding from regression prediction within the LLM also leads to greater training stability and clearer interpretability. Overall, these findings demonstrate that AU-text fusion provides a psychologically grounded and computationally efficient framework for personality recognition in AVIs.

Paper AI Chat

この論文のPDF全文を対象にAIに質問できます。

質問の例:

AIチャット機能を利用するには、ログインまたは会員登録(無料)が必要です。

会員登録 / ログイン

関連するAIDB記事