本記事は、主に一週間のうちに公開された論文(プレプリントを多く含む)のうち、AIDBリサーチが注目に値すると判断したものを掲載しています。基準としては「新規性」「優位性」といった査読フレームワークを踏襲する他、「実経済や産業へのインパクト」という独自の評価項目を据えています。
エージェント設計・指示・ハーネス
ここから限定コンテンツ
| 日本語タイトル | 英文タイトル |
|---|---|
| レガシーワークフローをエージェント型BPMに引き上げるためのプロセスハーネス:CUGA FLOにおける設計と実現 | A Process Harness for Uplifting Legacy Workflows to Agentic BPM: Design and Realization in CUGA FLO |
| 開発者はエージェントの指示をどう維持・進化させるか?実証研究 | How Do Developers Maintain and Evolve Their Agents’ Instructions? An Empirical Study |
| エージェントはデモンストレーションをどう読むべきか?階層構造はフラットな行動ログより優れている | How Should Agents Read Demonstrations? Hierarchical Structure Beats Flat Action Logs |
コーディングエージェントとソフトウェア開発
| 日本語タイトル | 英文タイトル |
|---|---|
| AI支援開発のための仕様成長エンジン:仕様アンカー、コード結合、ドリフト強制アーキテクチャ | The Spec Growth Engine: Spec-Anchored, Code-Coupled, Drift-Enforced Architecture for AI-Assisted Software Development |
| LLMベースのプログラム修正におけるコード実行の費用対効果分析 | To Run or Not to Run: Analyzing the Cost-Effectiveness of Code Execution in LLM-Based Program Repair |
| LLMによる実世界のソフトウェア性能最適化の評価 | Evaluating LLMs on Real-World Software Performance Optimization |
| NatureBench:コーディングAIはNature誌系の論文のSOTAに匹敵できるか? | NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? |
コンピュータ操作・Webエージェント
| 日本語タイトル | 英文タイトル |
|---|---|
| エージェント型Webブラウザを支援技術として「ズームイン」する:ロービジョン技術専門家とのケーススタディ | “Zooming In” on Agentic Web Browsers as Assistive Technologies: A Case Study with a Low-Vision Technology Expert |
| MacAgentBench: 現実世界のmacOSデスクトップにおけるAIエージェントのベンチマーク | MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop |
| コンピュータ操作エージェントのためのスケーラブルな学習環境 Fara-1.5 | Fara-1.5: Scalable Learning Environments for Computer Use Agents |
エージェント型コマース・経済・予測の応用
| 日本語タイトル | 英文タイトル |
|---|---|
| AlgoEvolve: LLM駆動のアルゴリズム取引プログラムのメタ進化的発見 | AlgoEvolve: LLM-driven Meta-evolution of Algorithmic Trading Programs |
| 知るために支払う:エージェント型ECにおける検証済み製品情報のためのマイクロトランザクション市場 | Paying to Know: Micro-Transaction Markets for Verified Product Information in Agentic E-Commerce |
| IPOファイナンスエージェント:SpaceX (SPCX) IPOを事例とした、自動評価基準生成によるLLM金融アナリストの評価 | IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Automated Rubric Generation — the Case of the SpaceX (SPCX) IPO |
| 未来イベント予測のためのインフラとしてのエージェント型タイムマシン | Agentic Time Machine as an Infrastructure for Future-Event Forecasting |
マルチエージェントと社会シミュレーション
| 日本語タイトル | 英文タイトル |
|---|---|
| 言語モデルの組み合わせはいつ役立つのか?ルーティング、投票、混合エージェントにおける共同失敗の上限 | When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models |
| LLMエージェントで動く社会経済システムのデジタルツインプラットフォーム「EconSimulacra」 | EconSimulacra: A Digital Twin Platform of Socio-Economic Systems Powered by LLM Agents |
| マルチエージェントLLMアーキテクチャによる、ゲーミフィケーションを通じた金融リテラシーの隠れた評価 | Agentic Knowledge Tracing: A Multi-Agent LLM Architecture for Stealth Assessment of Financial Literacy in Serious Games |
| LLMエージェント社会における創発的関係秩序:集団的感情から権威階層へ | Emergent Relational Order in LLM Agent Societies: From Collective Affect to Authority Stratification |
| 専門家と汎用家の人工集合体は異なるタスクで優れている | Artificial collectives of specialists and generalists excel at different tasks |
世界モデルとデジタルツイン
| 日本語タイトル | 英文タイトル |
|---|---|
| 高齢者の認知支援のための言語ベースのデジタルツイン | Language-Based Digital Twins for Elderly Cognitive Assistance |
| アインシュタインの世界モデル | Einstein World Models |
| 生涯にわたる社会的知能のためのソーシャルワールドモデル | Social World Model for Lifelong Social Intelligence |
記憶と長期文脈
| 日本語タイトル | 英文タイトル |
|---|---|
| LLMエージェントの長期記憶をポイズニングから保護:機械チェック済みの保証付き、改変不能で起源に紐づいた権限 | Securing LLM-Agent Long-Term Memory Against Poisoning: Non-Malleable, Origin-Bound Authority with Machine-Checked Guarantees |
| 大規模言語モデルのプロスペクティブメモリを調査するTriggerBench | TriggerBench: Investigating Prospective Memory for Large Language Models |
| 実環境における長期記憶ベンチマーク DynamicMem | DynamicMem: A Long-Horizon Memory Benchmark in Real-World Settings |
安全性・セキュリティ・堅牢性
| 日本語タイトル | 英文タイトル |
|---|---|
| 真の意図で舗装:意図認識トレーニングは、トレーニング体制全体でLLMの安全性分類を向上させる | Paved with True Intents: Intent-Aware Training Improves LLM Safety Classification Across Training Regimes |
| プログラム表現は重要:LLMの脆弱性推論のための実証研究 | Representation Matters: An Empirical Study of Program Representations for LLM Vulnerability Reasoning |
| あなたのエージェントは誰?自律型Webエージェントの多層フィンガープリンティングと属性特定 | Whose Agent Are You? Multi-Layer Fingerprinting and Attribution of Autonomous Web Agents |
| PromptMark: ソースコードのウォーターマーキングのためのプロンプト誘導型反復フィードバックフレームワーク | PromptMark: A Prompt-Guided Iterative-Feedback Framework for Source Code Watermarking |
評価とベンチマーク
| 日本語タイトル | 英文タイトル |
|---|---|
| 能力のフロンティア:ベンチマークはモデル性能の82%を見逃す | The Capability Frontier: Benchmarks Miss 82% of Model Performance |
| LLM評価における温度設定は再現性に必要だが十分ではない | Necessary but Not Sufficient: Temperature Control and Reproducibility in LLM-as-Judge Safety Evaluations |
| ベンチマーク飽和後の世界:CORE-Benchのケーススタディ | Life After Benchmark Saturation: A Case Study of CORE-Bench |
| LLM時代:大規模言語モデルの推論、外交、信頼性を検証する戦略的1v1ベンチマーク(不確実性下) | Age of LLM: A Strategic 1v1 Benchmark for Reasoning, Diplomacy and Reliability of Large Language Models under Fog of War |
| EnterpriseClawBench: 実務セッションからのエンタープライズエージェントのベンチマーク | EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions |
| メタニムゲーム:構造的知能のための自己完結型、自己整合型LLMピアコミュニティベンチマーク | The Metanym Game: A Self-Contained, Self-Consistent LLM Peer-Community Benchmark for Structural Intelligence |
推論・認知・内部表現
| 日本語タイトル | 英文タイトル |
|---|---|
| なぞなぞのなぞなぞ:大規模言語モデルと人間の柔軟な推論をテストする | The Riddle Riddle: Testing Flexible Reasoning in Large Language Models and Humans |
| ラディカルAI解釈可能性 | Radical AI Interpretability |
| 人間は諦め、推論モデルは粘る:難易度登録と熟考配分の分離 | Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation |
| 抽象的な表現幾何学が大規模言語モデルの推論をサポートする | Abstract representational geometry supports inference in large language models |
| LLMと人間の表象モード | LLM and Human Modes of Representation |
学習ダイナミクス・基盤モデル・サーベイ
| 日本語タイトル | 英文タイトル |
|---|---|
| 自律AIへのヒッチハイクガイド:基礎からシステムまで | The Hitchhiker’s Guide to Agentic AI: From Foundations to Systems |
| 大規模言語モデルの可塑性低下はスケールで解決できるか? | Can Scale Save Us From Plasticity Loss in Large Language Models? |
| サカナフグ技術報告書 | Sakana Fugu Technical Report |
自動研究と科学的発見
| 日本語タイトル | 英文タイトル |
|---|---|
| エージェント駆動の理論発見と実験による心の科学の自動化 | auto-psych: Automating the science of mind using agent-driven theory discovery and experimentation |
| 自動認知科学者による心理学理論発見のループ閉鎖 | Closing the Loop to Discover Psychological Theories with an Automated Cognitive Scientist |
| LLMによる科学論文査読:手法、ベンチマーク、信頼性の課題 | LLM-Based Scientific Peer Review: Methods, Benchmarks, and Reliability Challenges |
| エージェント時代の因果発見 | Causal Discovery in the Era of Agents |
| マルチモーダルエージェントによるメタマテリアルデータベースの自律生成 | Autonomous Generation of Metamaterial Databases Based on Multimodal Agents |
| PaperClaw:エージェントを活用した自律的な研究と人間参加型改良 | PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement |
エンタープライズとAI導入の影響
| 日本語タイトル | 英文タイトル |
|---|---|
| エージェンティックAIへの移行:Codexからの証拠 | The Shift to Agentic AI: Evidence from Codex |
| AgentX:産業用レコメンダーシステムの自律的な自己進化を目指して | AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems |
| 人工知能がエンタープライズソフトウェアのユーザーロールに与える影響 | The impact of artificial intelligence on enterprise software user roles |
| オープンソースにおけるAIコーディングエージェントの検出:1億8000万リポジトリの多手法による検証済み調査 | Detecting AI Coding Agents in Open Source: A Validated Multi-Method Census of 180 Million Repositories |
教育・人間とAIの協働
| 日本語タイトル | 英文タイトル |
|---|---|
| 対話と思考の架け橋:協調的問題解決における対話ダイナミクスの理解 | Bridging Talk and Thought: Understanding Dialogue Dynamics Across Collaborative Problem-Solving Contexts |
| 人間とLLMの協働が科学論文の複雑性指標を変革する | Human–LLM Collaboration Is Transforming Complexity Metrics in Scientific Texts |
| 大規模言語モデル登場前後の学生の統計ライティング分析 | Analyzing Students’ Statistics Writing Before and After the Emergence of Large Language Models |
言語・翻訳・多言語
| 日本語タイトル | 英文タイトル |
|---|---|
| AI翻訳は「悪くない」が、読者はやはり人間翻訳を好む | AI translation of literary texts is “fine”, but readers still prefer human translations |
| 大規模言語モデルは言語や市場を越えてブランド評判をどう情報源としているか | How Large Language Models Source Brand Reputation Across Languages and Markets |
| 日本語の話し言葉と書き言葉のLLMにおける方言への頑健性評価 | Evaluating Japanese Dialect Robustness Across Speech and Text-based Large Language Models |
| Sarashina2.2-TTS:データ拡張と合成で日本語の漢字多義性に取り組む | Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis |
| 言語の盲点:クエリ言語とブランド認知度がAI構築のブランド評判を12のヨーロッパ言語でどう形成するか | The Language Blind Spot: How Query Language and Brand Recognition Tier Shape AI-Constructed Brand Reputation Across Twelve European Languages |
創作・物語生成・ロールプレイ
| 日本語タイトル | 英文タイトル |
|---|---|
| 心理学に基づいた推論と役割認識型方策最適化による汎用ロールプレイングエージェントの改善 | Improving General Role-Playing Agents via Psychology-Grounded Reasoning and Role-Aware Policy Optimization |
| AIが生成するフィクションの現状 | AI Fiction in the Wild |
| ストーリーラインツリー:長編物語のための階層的表現 | Storyline Trees: Hierarchical Representations for Long-Form Narratives |
倫理・公平性・ガバナンス
| 日本語タイトル | 英文タイトル |
|---|---|
| Pingquanqi (イコライザー): 人間とAIエージェントの相互作用を管理する分野横断的な社会技術フレームワーク | Pingquanqi (Equalizer): A Cross-Domain Sociotechnical Framework for Human-Agent Interaction Governance |
| 大規模言語モデルにおける、反証可能な倫理的推論のための推論時足場「思考の語り」 | Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models |
| 言語モデルはビッチデル・テストに合格するか?LLM生成脚本におけるジェンダーバイアスの監査 | Do Language Models Pass the Bechdel Test? Auditing Gender Biases in LLM-Generated Screenplays |
その他の応用
| 日本語タイトル | 英文タイトル |
|---|---|
| サービスフィードバックにおける新興トピック検出のためのLLMベースモデル | LLM-based Models for Detecting Emerging Topics in Service Feedback |
| NeuraDockビジュアル認知負荷エージェントチュートリアル:アルファダイナミクスとリアルタイムアプリケーションのための品質ゲート付きオープンソースEEGワークフロー | NeuraDock Visual Cognitive Load Agent Tutorial: A Quality-Gated Open-Source EEG Workflow for Alpha Dynamics and Real-Time Applications |
| 感情分析のためのプライバシーを保護するAIパイプラインEmotionAI | EmotionAI: A Privacy-Preserving Computational Intelligence Pipeline for Speech-Emotion-Grounded Conversational Analysis |
| マルチエージェント意味書き換えによるプライバシー保護RAG:文脈忠実性を損なわずに機密性を実現 | Privacy-Preserving RAG via Multi-Agent Semantic Rewriting: Achieving Confidentiality Without Compromising Contextual Fidelity |
| 無限OCRワークス | Unlimited OCR Works |
| ASCIIアートでLLMをVLAコントローラーに | ASCII Art Turns LLMs into VLA Controllers |