次回の更新記事:【論文著者監修・コメント】AIエージェントへの人間…(公開予定日:2026年07月27日)
AIDB Daily Papers

生成AIアシスタントはrobots.txtを尊重しているか?ウェブアクセス行動の追跡調査

原題: Do Generative AI Assistants Respect robots.txt? Tracing Web Access Beyond Visible Answers
著者: Gabriel Lopez-Fonseca, David Rodriguez, Stefan Bechtold, Jose M. Del Alamo
公開日: 2026-07-16 | 分野: LLM AI 検索 ウェブ cs.CY AIガバナンス

※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。

ポイント

  • 主要なAIアシスタント10種を対象に、検索機能利用時のrobots.txt遵守状況を実証的に調査した。
  • AIによるウェブアクセスと回答内容の乖離を明らかにし、従来のウェブガバナンスが機能不全に陥っていることを示した。
  • 一部のAIが制限を無視してアクセスする実態を突き止め、コンテンツ所有者の権利保護に向けた新たな標準化の必要性を提言した。

Abstract

AI assistants increasingly retrieve web content at inference time to provide fresh and grounded answers, yet it remains unclear whether these search-augmented capabilities respect website-owner restrictions expressed through robots.txt. We present a controlled empirical study of ten widely used AI assistants with advertised web-search capabilities. For each assistant, we first identify a configuration that actually produces observable web-browsing behavior and record the user-agent exposed during retrieval. We then evaluate compliance with controlled robots.txt rules across four complementary conditions: allowed for all user-agents, disallowed for all user-agents, allowed only for the assistant-specific user-agent, and disallowed only for that user-agent. Using server-side logs and secret codes embedded in target pages, we distinguish actual page access from user-visible answer correctness across 200 trials. Our results show substantial variation across assistants. Some systems followed the expected allowed/disallowed access pattern, whereas others accessed restricted resources without requesting robots.txt or used generic user-agents that complicated attribution. We also find that retrieval behavior and answer correctness can diverge: assistants may access pages without surfacing the retrieved content, or fail to access even allowed resources. These findings raise broader legal and governance concerns about whether AI-assisted web access adequately respects content owners' rights and restrictions. Furthermore, our observations provide valuable insight into the growing erosion of traditional web governance protocols, highlighting the urgent need for updated, enforceable standards that guarantee publisher autonomy in the age of search-augmented AI assistants.

Paper AI Chat

この論文のPDF全文を対象にAIに質問できます。

質問の例:

AIチャット機能を利用するには、ログインまたは会員登録(無料)が必要です。

会員登録 / ログイン

関連するAIDB記事