本記事は、主に一週間のうちに公開された論文(プレプリントを多く含む)のうち、AIDBリサーチが注目に値すると判断したものを掲載しています。基準としては「新規性」「優位性」といった査読フレームワークを踏襲する他、「実経済や産業へのインパクト」という独自の評価項目を据えています。
エージェントハーネス
ここから限定コンテンツ
| 日本語タイトル | 英文タイトル |
|---|---|
| 再帰型エージェントハーネス | Recursive Agent Harnesses |
| エージェントハーネスとは何か:エージェントハーネスの必要十分条件 | What makes a harness a harness: necessary and sufficient conditions for an agent harness |
| 自己改良型ハーネス:自らを改善するハーネス | Self-Harness: Harnesses That Improve Themselves |
| 物理AIのためのハーネスエンジニアリング:ロボットミドルウェアがハーネス層となる | Harness Engineering for Physical AI: Robot Middleware Is the Harness Layer |
エージェントのスキル設計・進化
| 日本語タイトル | 英文タイトル |
|---|---|
| 大規模言語モデルのための適応型マルチ解像度手続き知識圧縮 | Adaptive Multi-Resolution Procedural Knowledge Compression for Large Language Models |
| Notes2Skills:実験ノートから確実性を考慮した科学エージェントのスキルへ | Notes2Skills: From Lab Notebooks to Certainty-Aware Scientific Agent Skills |
| エージェントのスキル構成が実行時動作をどう変えるかを測定するSkillJuror | SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior |
| エージェントのスキル評価と進化:フレームワークとベンチマーク | Agent Skill Evaluation and Evolution: Frameworks and Benchmarks |
| 外部知識をエージェントのための再利用可能なスキルにコンパイルするAnything2Skill | Anything2Skill: Compiling External Knowledge into Reusable Skills for Agents |
エージェント設計・自律的最適化
| 日本語タイトル | 英文タイトル |
|---|---|
| 規制プロセス自動化のためのニューロシンボリックエージェント:課題と研究アジェンダ | Neuro-Symbolic Agents for Regulated Process Automation: Challenges and Research Agenda |
| AIエージェントのための戦略的意思決定支援 | Strategic Decision Support for AI Agents |
| カスタムAIエージェント構築のための、基盤から本番までの一貫した開発手法 | Agents All the Way Down; A Methodology for Building Custom AI Agents from Substrate to Production |
| いつ尋ねるべきかを知る:階層型言語エージェントのための自己ゲーティング型明確化 | Knowing When to Ask: Self-Gated Clarification for Hierarchical Language Agents |
| 自己進化型LLMエージェントと分布内最適化 | Self-evolving LLM agents with in-distribution Optimization |
マルチエージェント・集合知
| 日本語タイトル | 英文タイトル |
|---|---|
| 自律型AIエージェントのインターネット:大規模な通信、協調、集合知 | The Internet of Agentic AI: Communication, Coordination, and Collective Intelligence at Scale |
| プログラム推論による人間と建物の対話のためのゼロショットマルチエージェントフレームワーク | A Zero-Shot Multi-Agent Framework for Human-Building Interaction via Programmatic Reasoning |
| AIエージェントの集合知を活用した新たな発見 | Harnessing the Collective Intelligence of AI Agents in the Wild for New Discoveries |
コーディングエージェント・ソフトウェア工学
| 日本語タイトル | 英文タイトル |
|---|---|
| コードレビューの終焉:コーディングエージェントが人間の検査に取って代わる | The End of Code Review: Coding Agents Supersede Human Inspection |
| AIネイティブソフトウェアエンジニアリングの台頭:実践、教育、将来の労働力への影響 | The Rise of AI-Native Software Engineering: Implications for Practice, Education, and the Future Workforce |
| 行レベルのコード著者検出のためのベンチマークデータセット HybridCodeAuthorship | HybridCodeAuthorship: A Benchmark Dataset for Line-Level Code Authorship Detection |
| ソフトウェアに意味を持たせる | Making Software Meaningful |
| 最先端のコーディングエージェントはメタプログラミングで未知の言語に適応する | Frontier Coding Agents Use Metaprogramming to Adapt to Unfamiliar Programming Languages |
| コーディングエージェントの採用、GitHub新プロジェクトで大幅増! | Agentic Very Much! Adoption of Coding Agent in New GitHub Projects |
| GitHubリポジトリにおけるAI利用の特徴と進化に関する実証研究:コードコメントからの証拠 | Empirical Study on the Characteristics and Evolution of AI-usage in GitHub Repositories: Evidence from Code Comments |
コンピュータ・ソフトウェア操作エージェント
| 日本語タイトル | 英文タイトル |
|---|---|
| COM-as-Actionパラダイムによるプロフェッショナルソフトウェア操作の再構築:ComAct | ComAct: Reframing Professional Software Manipulation via COM-as-Action Paradigm |
| MacArena: macOS環境でのコンピュータ利用エージェントのベンチマーク | MacArena: Benchmarking Computer Use Agents on an Online macOS Environment |
RAG・記憶
| 日本語タイトル | 英文タイトル |
|---|---|
| 記憶すべきことを学習する:認知科学に基づいたエージェント型記憶のためのマルチファクター価値モデル | Learning What to Remember: A Cognitively Grounded Multi-Factor Value Model for Agentic Memory |
| ドキュメントが増えるとRAGが劣化する問題:ドメイン限定・モデル非依存検索でベクトル検索の希薄化を軽減 | When More Documents Hurt RAG: Mitigating Vector Search Dilution with Domain-Scoped, Model-Agnostic Retrieval |
推論・世界モデル
| 日本語タイトル | 英文タイトル |
|---|---|
| 推論はパターンマッチング:人間とLLMの日常的推論における共通メカニズム | Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning |
| Transformerベースの次トークン予測における汎化性能の上限 | Generalization Bounds for Transformer-Based Next-Token Prediction in a Language Model |
| 心の理論ユーティリティ:メンタライジングメカニズムの形式仕様 | The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism |
| 大規模言語モデルで社会的世界モデルを構築する | Building Social World Models with Large Language Models |
| ビジネス世界モデル | Business World Model |
AI for science・科学研究の自動化
| 日本語タイトル | 英文タイトル |
|---|---|
| 大規模言語モデルを用いた社会・行動科学における再現性の自動評価 | Automated reproducibility assessments in the social and behavioral sciences using large language models |
| エージェントネイティブな知識オーケストレーションを目指すAgents-K1 | Agents-K1: Towards Agent-native Knowledge Orchestration |
| AIコーディングエージェントは社会科学の発見を再現できる | AI Coding Agents Can Reproduce Social Science Findings |
| 検証可能でエージェントネイティブな科学出版のためのフレームワーク「Traxia」 | Traxia: A Framework for Verifiable, Agent-Native Scientific Publishing |
評価・ベンチマーク
| 日本語タイトル | 英文タイトル |
|---|---|
| DailyReport:日々の検索タスクを評価するオープンエンドなベンチマーク | DailyReport: An Open-ended Benchmark for Evaluating Search Agents on Daily Search Tasks |
| LLMの心理測定評価を再考する:自己報告はいつ、なぜ行動を予測するのか | Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior |
| パラメトリック3D生成と構造推論のためのMLLMベンチマークP3D-Bench | P3D-Bench: Benchmarking MLLMs for Parametric 3D Generation and Structural Reasoning |
| T1-Bench: 現実世界のドメインにおけるマルチシナリオエージェントのベンチマーク | T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains |
安全性・セキュリティ・攻撃
アライメント・価値整合
| 日本語タイトル | 英文タイトル |
|---|---|
| 潜在的視点によるLLMにおける多元主義の評価 | Evaluating Pluralism in LLMs through Latent Perspectives |
| 適切なプロンプトでLLMは人間の判断をより良く捉えられる | LLMs Can Better Capture Human Judgments–With the Right Prompts |
| 人間とLLMの対話ガバナンス:安全性ゲート、丁寧さ誘導、感情的デフォルト固定 | The Governance of Human-LLM Interaction: Safety Gating, Civility Steering, and Affective Default Lock-In |
| LCAM: 対話型AIにおける対話アライメントの失敗を診断するフレームワーク | LCAM: A Framework for Diagnosing Interactional Alignment Failures in Con-versational AI |
| 誰の規範?大規模言語モデルにおける文化的・個人的整合性の解明 | Whose Norms? Disentangling Cultural and Personal Alignment in Large Language Models |
| 文脈を考慮した、価値整合のための道徳的信念の形成 | Accounting for Context: Shaping Moral Credences for Value Alignment |
AGI・超知能・AIリスク
| 日本語タイトル | 英文タイトル |
|---|---|
| AGIからASIへ | From AGI to ASI |
| 魂コンピューティング:独立した意識を持つインテリジェントエージェントのための理論的フレームワークと技術的アーキテクチャ | Soul Computing: A Theoretical Framework and Technical Architecture for Intelligent Agents with Independent Consciousness |
| 道具的収束と権力追求 | Instrumental convergence and power-seeking |
脳科学・認知科学
| 日本語タイトル | 英文タイトル |
|---|---|
| 脳科学的知見に基づく言語モデルで推論能力を向上させる | Beyond representational alignment with brain-guided language models for robust reasoning |
| LLM強化回帰フレームワークによる自然な感情ダイナミクスを脳から解読する | Decoding Naturalistic Emotion Dynamics from the Brain: An LLM-Enhanced Regression Framework |
計算社会科学・言語学・社会シミュレーション
| 日本語タイトル | 英文タイトル |
|---|---|
| 形態的交替パターンの進化のためのエージェントベースモデル | Agent-based models for the evolution of morphological alternation patterns |
| 大規模言語モデルを言語学における様相モデルとして捉える | Large Language Models as Modal Models in Linguistics |
| Agentopia:エージェント社会における長期ライフシミュレーションと学習 | Agentopia: Long-Term Life Simulation and Learning in Agent Societies |
| 性格アンカリングによるソーシャルシミュレーション:性格、社会的行動、相互作用の成功をLLMエージェントと結びつける | Personality Anchoring for Social Simulation: Linking Personality, Social Behavior, and Interaction Success with LLM Agents |
創造性・コンテンツ生成
| 日本語タイトル | 英文タイトル |
|---|---|
| IVIE:インタラクティブフィクション世界の段階的かつ検証可能な生成のためのニューロシンボリックアプローチ | IVIE: A Neuro-symbolic Approach to Incremental and Validated Generation of Interactive Fiction Worlds |
| LLMベースの並列テキスト生成による低遅延リアルタイム音声ゲーム解説システム | Low-Latency Real-Time Audio Game Commentary System via LLM-Based Parallel Text Generation |
| 機械はどのような条件で真に創造的になれるのか? | Under What Conditions Can a Machine Become Genuinely Creative? |
| Nonslop: 人間とAIの共同執筆におけるゲーミフィケーション実験 | Nonslop: A Gamified Experiment in Human-AI Collaborative Writing |
ドメイン応用(医療・サポート・ウェルビーイング)
| 日本語タイトル | 英文タイトル |
|---|---|
| MetaPlate: パーソナライズされた食事レコメンドと高血糖予防のための反事実的ガイダンス付きRAG-LLMツール | MetaPlate: Counterfactual-Guided RAG-LLM Tool for Personalized Food Recommendation and Hyperglycemia Prevention |
| 1億ユーザー規模のカスタマーサポートAIエージェント構築:評価駆動型フレームワーク | Building Customer Support AI Agents at 100M-User Scale: An Evaluation-Driven Framework |
| 大規模言語モデルを用いた脳MRIレポートからの構造化情報自動抽出 | Automatic Extraction of Structured Information from Brain MRI Reports Using an Open-Weight Large Language Model |
| FOMO(取り残されることへの恐れ)をLLMチャットボットで支援する初期デザイン探求「Moodie」 | Moodie: An Early-Stage Design Exploration for Supporting Fear of Missing Out with LLM-based Chatbots |
AIと労働・知識労働・教育
| 日本語タイトル | 英文タイトル |
|---|---|
| LLMによる合成ライティングにおける認知オフローディングのプロファイリング:量 vs 内容 | Profiling cognitive offloading in LLM-mediated synthesis writing: Volume vs. content |
| AIによる職業代替可能性の二極構造とその10年スケールでの逆転:安定した幾何学、反転する極 | Stable Geometry, Reversing Poles: The Bipolar Structure of AI Occupational Substitutability and Its Decade-Scale Inversion |
| ソフトウェア工学教育における学術的誠実性と不適切なLLM使用への感情的反応 | Academic Integrity and Emotional Responses to Inappropriate LLM Use in Software Engineering Education |
| AIエージェントは知識労働をどう変えるか:自律性、効率性、範囲 | How AI Agents Reshape Knowledge Work: Autonomy, Efficiency, and Scope |
AIと社会・人間の選好
| 日本語タイトル | 英文タイトル |
|---|---|
| ユーモアのスタイルが笑いを、トピックが受容性を決定:バイリンガルなロボットによるAIジョークの個人的・政治的評価 | Humor Style Drives Laughter, Topic Shapes Acceptability: Evaluating Bilingual Personal and Political Robot-Delivered AI Jokes |
| 人工知能研究におけるトピックの相転移:大規模データと早期警戒シグナル | Topical Phase Transitions in Artificial Intelligence Research: Large-Scale Evidence and an Early-Warning Signature for Emerging Topics |
| 「AIスラップだ、ボットめ!」オンライン上のAI生成コメントに対する非難、証拠、信頼性の研究 | “That’s AI Slop, You Bot!” Studying Accusations, Evidence, and Credibility in Online Discourse Towards LLM-Generated Comments |
| AI委任の社会的影響 | The social consequences of AI delegation |
| AI生成動物物語におけるジェンダー表現の「中立性」の落とし穴 | Neutrality Bites: Gender Representation in AI-Generated Animal Stories |
| モジュラーAIシステムにおける参加のスケーリング | Scaling Participation in Modular AI Systems |
| 人々はAIに何を求めているのか? 嗜好の多様性をマッピングする | What Do People Actually Want From AI? Mapping Preference Plurality |