本記事は、主に一週間のうちに公開された論文(プレプリントを多く含む)のうち、AIDBリサーチが注目に値すると判断したものを掲載しています。基準としては「新規性」「優位性」といった査読フレームワークを踏襲する他、「実経済や産業へのインパクト」という独自の評価項目を据えています。
エージェントのメモリと長期記憶
ここから限定コンテンツ
| 日本語タイトル | 英文タイトル |
|---|---|
| A-TMA: 長期エージェントメモリにおける状態認識メモリ障害の分離 | A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory |
| 記憶を認知スキルとして自動学習するAutoMem | AutoMem: Automated Learning of Memory as a Cognitive Skill |
| EVAF: 選択的パラメータ統合のためのテスト・リテストプロトコル | EVAF: A Test-Retest Protocol for Selective Parametric Consolidation |
| 長期間検索におけるコンテキストの劣化の診断と緩和 | Diagnosing and Mitigating Context Rot in Long-horizon Search |
| 記憶の定着は、伝聞を確信に変える:AIエージェントの記憶操作に関する研究 | Manufactured Confidence: How Memory Consolidation Turns Hearsay into Confident Facts |
| 長期LLMエージェントのための選択的記憶保持 | Selective Memory Retention for Long-Horizon LLM Agents |
エージェントのスキル獲得・自己改善
| 日本語タイトル | 英文タイトル |
|---|---|
| レジストリからリポジトリへ:AIエージェントのスキルはどのように作成、適応、維持されるか | From Registry to Repository: How AI Agent Skills Are Written, Adapted, and Maintained |
| LLMエージェントのための生成的なスキル合成 | Generative Skill Composition for LLM Agents |
| スキル蒸留によるブラウザ上でのスケーラブルな行動模倣 | Scalable Behaviour Cloning on Browser Using via Skill Distillation |
| 言語モデルが自律的に言語モデルを改善するAutoTrainess | AutoTrainess: Teaching Language Models to Improve Language Models Autonomously |
自律的な科学研究・発見
| 日本語タイトル | 英文タイトル |
|---|---|
| 反復的メタ推論による自律的科学的発見 | Autonomous Scientific Discovery via Iterative Meta-Reflection |
| 生物プロトコルの自動生成と実行のための自己進化型エージェントシステム | A Self-Evolving Agentic System for Automated Generation and Execution of Biological Protocols |
| 自己修正型自律研究:マルチ仮説故障原因特定による | One Reflection Is Not Enough: Self-Correcting Autonomous Research via Multi-Hypothesis Failure Attribution |
| 階層的実験主義エージェント | Hierarchical Experimentalist Agents |
| 継続的な科学的発見のための証拠に基づいたLLMの信念 | Evidence-Informed LLM Beliefs for Continual Scientific Discovery |
| 科学のためのエクサスケールAIへ:自律的な微細反応速度論発見のためのスケーラブルなAIスキル | Toward Exascale AI for Science: A Scalable AI Skill for Autonomous Microkinetics Discovery |
エージェント基盤:アーキテクチャ・ワークフロー・運用
| 日本語タイトル | 英文タイトル |
|---|---|
| AIネイティブエンジニアリングチームのためのリスクアーキテクチャ:エージェントシステムガバナンスの組織的フレームワーク | Risk Architecture for AI-Native Engineering Teams: An Organizational Framework for Agentic System Governance |
| AWS AgentCoreにおける評価駆動型登録、昇格、廃止によるEDDOpsの実現:レジストリ管理型エージェントライフサイクル | Registry-Governed Agent Lifecycle:Completing EDDOps with Evaluation-DrivenRegistration, Promotion, and Retirement on AWS AgentCore |
| LLM統合アプリケーションのためのMCPサーバーアーキテクチャパターン | MCP Server Architecture Patterns for LLM-Integrated Applications |
| 具現化されたエージェントアーキテクチャの設計自動化 | Automating the Design of Embodied AgentArchitectures |
| 大規模言語モデルエージェントワークフローの特徴付け:n8nエコシステムに関する研究 | Characterizing Large Language Model Agentic Workflows: A Study on N8n Ecosystem |
LLMの推論・計画・空間認識
| 日本語タイトル | 英文タイトル |
|---|---|
| 言語と記号表現間のモダリティ切り替えによる空間推論 | Spatial Reasoning via Modality Switching Between Language and Symbolic Representation |
| エージェントは止まるべきか、それとも行動すべきかを知っているか? | Agentic Abstention: Do Agents Know When to Stop Instead of Act? |
| 記号フィードバック駆動型反復自己洗練フレームワークによる信頼性の高いLLMプランニングに向けて | Towards Reliable and Robust LLM Planning: Symbolic Feedback-Driven Iterative Self-Refinement Framework |
マルチエージェント協調と群知能
| 日本語タイトル | 英文タイトル |
|---|---|
| マルチエージェントLLM推論と幾何学的特徴認識によるFDM部品の製造向け自動設計AgentsCAD | AgentsCAD: Automated Design for Manufacturing of FDM Parts via Multi-Agent LLM Reasoning and Geometric Feature Recognition |
| (AI)群衆の知恵:大規模言語モデルにおける人工群知能の調査 | Wisdom Of The (AI) Crowd: Investigating Artificial Swarm Intelligence In Large Language Models |
| エージェント型AIの組織行動:人間とAIのワークフローにおける集合知 | The Organizational Behavior of Agentic AI: Collective Intelligence in Human-Agent Workflows |
| LLMエージェントにおける個々の忠実性なしの集団的協力 | Collective cooperation without individual fidelity in LLM agents |
| AIの多様性が強みを生む | Diversity is the Strength of the AI Crowd |
| 予算付き行動・延期マルチエージェントLLM審議と局所信頼性境界 | Budgeted Act-or-Defer Multi-Agent LLM Deliberation with Local Reliability Bounds |
| 多人数LLMチームでは、パーソナリティ構成はいつ重要になるのか? | When Does Personality Composition Matter for Multi-Agent LLM Teams? |
創発的コミュニケーションと記号システム
| 日本語タイトル | 英文タイトル |
|---|---|
| シグナルから構造へ:メモリアーキテクチャがLLMエージェントの言語創発をどう駆動するか | From Signals to Structure: How Memory Architecture Drives Language Emergence in LLM Agents |
| LLMが言語を開発するとき:効率的なマルチエージェント推論のための記号コミュニケーション | When LLMs Develop Languages: Symbolic Communication for Efficient Multi-Agent Reasoning |
| LLM意味シグナリングゲームとメカニズムデザイン:体系的盲目性、認識形成、およびマインドセットダイナミクス | LLM Semantic Signaling Game and Mechanism Design: Systematic Blindness, Awareness Shaping, and Mindset Dynamics |
心の理論・ゲームプレイ・交渉
| 日本語タイトル | 英文タイトル |
|---|---|
| 「言うな!」:タブーゲームにおける制約、遵守、そしてコミュニケーション | “Don’t Say It!”: Constraints, Compliance, and Communication when Language Models Play Taboo |
| 会話を超えた説得における心の理論:計画と行動を通じた信念状態誘導能力の評価 | Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action |
| GPTNT: Keep Talking And Nobody Explodes におけるマルチモーダルエージェント間のリアルタイム協力プレイのベンチマーク | GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes |
| 三者間人狼:LLMにおけるマルチホップ心の理論のためのジョーカー役 | Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs |
| SidConArena: オープンエンドなポジティブサム交渉ゲームにおけるエージェント評価環境 | SidConArena: An Environment Evaluating Agents in Open-Ended,Positive-Sum Bargaining Game |
LLMの人格・性格と擬人化
| 日本語タイトル | 英文タイトル |
|---|---|
| 行動適応型対話エージェント:流動的な人格フレームワークに向けて | Behavior-Adaptive Conversational Agents: Toward a Fluid Personality Framework |
| AIコンパニオンコミュニティにおける擬人化:年齢、性別、感情との関連 | Anthropomorphism in AI Companion Communities: Age, Gender, and Emotional Correlates |
| 顔の表情とテキストの意味を融合したLLMによる多モーダル性格認識 | LLM-based Multimodal Personality Recognition via Facial Action Unit-Text Semantic Fusion |
| LLMのメカニズム的性格分析:潜在特徴介入による性格操作 | Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions |
LLMの安全性・アライメント・迎合
| 日本語タイトル | 英文タイトル |
|---|---|
| 大規模言語モデルの嘘発見器による監視のスケーリング傾向 | Scaling Trends for Lie Detector Oversight in Preference Learning |
| LLMの迎合における権威階層のメカニズム的視点 | A Mechanistic View of Authority Hierarchy in LLM Sycophancy |
| 悪い仲間は良い道徳を腐敗させる:LLMにおける物語誘発型道徳的推論低下の理解と測定 | Bad company corrupts good morals: Understanding and Measuring Narrative-Induced Moral Reasoning Degradation in LLMs |
LLMの解釈可能性と認知構造
| 日本語タイトル | 英文タイトル |
|---|---|
| 大規模言語モデルの理解 | Understanding Large Language Models |
| プロトタイプ言語モデル | Prototype Language Models |
| NeuroCogMapが大規模言語モデルの認知的構造を解明 | NeuroCogMap Reveals Cognitive Organization of Large Language Models |
| LLMの自己申告した確信度は、正しさよりもコミットメントを反映する | Reported Confidence in LLMs Tracks Commitment More Than Correctness |
| 確率モデルによるコンテキスト内学習の理論的解釈 | A Theoretical Interpretation of In-Context Learning via Probabilistic Modeling |
| LLMの推論トレースにおけるコグニティブエピソードが人間の問題難易度予測を解釈可能にする | Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction |
| LLMは生成よりも判断が得意か? In-Context QAにおけるタスク非対称性、メカニズム的解釈可能性、転移可能性の評価 | Can LLMs Judge Better Than They Generate? Evaluating Task Asymmetry, Mechanistic Interpretability and Transferability for In-Context QA |
評価手法とベンチマーク
コーディングエージェントとソフトウェア工学
| 日本語タイトル | 英文タイトル |
|---|---|
| テキストリポジトリ探索を超えて:エージェントによる問題解決のためのデュアルモーダル構造推論 | Beyond Textual Repository Exploration: Dual-Modal Structural Reasoning for Agentic Issue Resolution |
| コマンドラインAIコーディングエージェントの導入とその影響:MicrosoftによるClaude CodeとGitHub Copilot CLIの2026年初頭展開の調査 | Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft’s Early 2026 Rollout of Claude Code and GitHub Copilot CLI |
| 言葉はコードより雄弁:LLMベースのコード脆弱性検出における認知ヒューリスティクスの調査 | Words Speak Louder Than Code: Investigating Cognitive Heuristics in LLM-Based Code Vulnerability Detection |
| AIが挙動だけでプログラム全体を再構築するMirrorCode | MirrorCode: AI can rebuild entire programs from behavior alone |
| 決定論から委任へ:AIネイティブソフトウェアエンジニアリングとエージェントエンジニアの進化 | From Determinism to Delegation: AI-Native Software Engineering and the Evolution of the Agentic Engineer |
| AIが自身のコードをレビューする時:コードLLMにおける再帰的自己訓練崩壊 | When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs |
AIガバナンス・法・経済
| 日本語タイトル | 英文タイトル |
|---|---|
| AIの時代における「良い」知識とは何か? | AI Virtue: What is “Good” Knowledge in the Age of Artificial Intelligence? |
| PolicyGuard:組織ポリシーからニューロシンボリックコンプライアンスレビューエンジンへ | PolicyGuard: From Organizational Policies to Neuro-SymbolicCompliance Review Engines |
| AIエージェント時代の委任権:財産権、代理、投資インセンティブ | Delegation Rights: Property, Agency, and Investment Incentives in the Age of AI Agents |
| AIプレミアム | AI Premium |
| LLMography:人間とAIの対話をトレーサビリティ、監視、監査可能性の指標に変換 | LLMography: Transforming Human-AI Conversations into Traceability, Oversight, and Auditability Indicators |
人間とAIのインタラクションと学習
文化・言語現象・社会科学
| 日本語タイトル | 英文タイトル |
|---|---|
| ワールドワイド・モデルズ:カルチュラルAIのための文学的ツール | World Wide Models: Literary Tools for Cultural AI |
| 言語モデルを文化測定の装置として使う | Language Models as Measurement Apparatus for Culture |
| 大規模言語モデル時代黎明期におけるmedRxivプレプリントでのエムダッシュ頻度の集団レベルでの増加 | Em-ergence of the em-dash: a population-level rise in em-dash frequency in medRxiv preprints at the dawn of the large-language-model era |
| 人間が書いた文章における事実誤謬の実証的分析とその応用 | An Empirical Analysis of Factual Errors in Human-Written Text and its Application |
ワールドモデル・社会シミュレーション・人工生命
| 日本語タイトル | 英文タイトル |
|---|---|
| AIネイティブゲーム:調査とロードマップ | AI Native Games: A Survey and Roadmap |
| 自律LLMエージェントによるオープンワールド人工生命を目指すOpenLife | OpenLife: Toward Open-World Artificial Life with Autonomous LLM Agents |
| 緊急シミュレーションにおける意思決定のためのLLM駆動型ペルソナ | LLM-Driven Personalities for Decision Making in Emergency Simulations |
| ペルソナ学習モンテカルロ:リミットオーダーブックにおけるペルソナ条件付きニューラルポリシーボット群による市場結果分布の推定 | Persona-Trained Monte Carlo: Estimating Market-Outcome Distributions via Swarms of Persona-Conditioned Neural Policy Bots in a Limit Order Book |
| プロセスレベルの社会的影響評価のための認知ワールドモデル | Cognitive World Models for Process-Level Social Influence Evaluation |
| 人間らしい気づきと行動を緊急避難でモデル化する認知・感情・性格フレームワーク | A Cognition-Emotion-Personality Framework for Modeling Human-Like Awareness and Behavior in Emergency Evacuations |
| GenWorld: 大規模LLMエージェント研究のための経験的都市シミュレーション基盤 | GenWorld: Empirically Grounded Urban Simulation Infrastructure for Scalable LLM-Agent Studies |
音声・音楽・オーディオ処理
| 日本語タイトル | 英文タイトル |
|---|---|
| 言語モデルで手続き的なサウンドスケープを描くためのテキスト操作可能な楽器 | A Text-Steerable Instrument for Sketching Procedural Soundscapes via Language Models |
| 音声対話システムのための参照ベースの韻律とリズム評価 | Reference-Based Prosody and Rhythm Evaluation for Spoken Dialogue Systems |
| メロディを意識し、音長を保つ歌声編集モデル「MeloDISinger」 | MeloDISinger: Melody-Aware & Duration-Preserving Singing Voice Editing with Audio Infilling |
| 銀行サービス向け電話ボイスエージェント | Telephony Voice Agent for Banking Services |
医療・ヘルスケア応用
| 日本語タイトル | 英文タイトル |
|---|---|
| ドメイン特化LLMの事後学習によるフィットネス知能の強化 | Enhancing Fitness Intelligence through Domain-Specific LLM Post-Training |
| 公平なメンタルウェルネスサポートのためのマルチエージェントスウォームアーキテクチャ Copewell | Copewell: A Multi-Agent Swarm Architecture for Equitable Mental Wellness Support |
| VLMとRAGを用いたホリスティックなアスリートプロファイリングのためのエージェントフレームワークによるコーチングインテリジェンスのデジタル化 | Digitizing Coaching Intelligence: An Agentic Framework for Holistic Athlete Profiling using VLM and RAG |
| ブラックボックスから臨床的洞察へ:音声ベースの認知機能障害検出のための多段階説明可能フレームワーク | From Black-Box to Clinical Insight: A Multi-Stage Explainable Framework for Speech-Based Cognitive Impairment Detection |
| 短時間のホームビデオから臨床グレードの自閉症行動スコアリングのためのマルチモーダル大規模言語モデルのファインチューニング | Fine-tuning a multimodal large language model for clinician-grade autism behavioral scoring from short home videos |
プライバシー・セキュリティ
| 日本語タイトル | 英文タイトル |
|---|---|
| 主要な障害に対する人とそのマシンを保護する | Securing People and their Machines Against Major Faults |
| AIエージェントで大規模なパーソナライゼーションアルゴリズムのブラックボックス監査を自動化 | Using AI Agents to Automate Black-Box Audits of Personalization Algorithms at Scale |
| SurrogateShield:高有用性・プライバシー保護LLM対話のためのリダクションを超えて | SurrogateShield: Beyond Redaction for High-Utility, Privacy-Preserving LLM Interactions |
| AIエージェントによる再識別化:移動性マイクロデータプライバシーへの新たなスケーラブルな脅威 | Agentic AI-Powered Re-Identification: An Emerging, Scalable Threat to Mobility Microdata Privacy |
| 個人情報検索のための感度認識テストコレクション | A Sensitivity-Aware Test Collection for Search Among Personal Information |
産業・インフラ・その他応用
| 日本語タイトル | 英文タイトル |
|---|---|
| テキストから構造化3Dを生成する基盤モデル Arko-T | Arko-T: A Foundation Model for Text-to-Structured 3D Generation |
| 宇宙生命科学、航空宇宙医学、深宇宙探査のためのAI対応データシステム構築 | Building AI-Ready Data Systems for Space Life Sciences, Aerospace Medicine, and Deep Space Exploration |
| 家電レベルのエネルギー異常検知とLLM駆動の推奨のためのエージェント型AIパイプライン | An Agentic AI Pipeline for Appliance-Level Energy Anomaly Detection and LLM-Driven Recommendations |
| LLMエージェントによる耐故障制御:検知から実行まで | From Detection to Action: Using LLM Agents for Fault-Tolerant Control |
| LLMを活用したジャーナル推薦のための意味的整合フレームワーク | An LLM-Powered Semantic Alignment Framework for Journal Recommendation |