AIDB Daily Papers
進化する社会規範に対するAI価値アライメント
※ 日本語タイトル・ポイントはAIによる自動生成です。正確な内容は原論文をご確認ください。
ポイント
- AIの安全な展開を見据え、社会物理学に基づく柔軟な数学的モデリングフレームワークを提案した。
- 非適応的なアライメント手法における価値の硬直化や規範のモード崩壊のリスクを明らかにした点に新しさがある。
- 長期的な動的帰結を分析し、社会技術的予見のための厳密な仮説検証手法としての有用性を示した。
Abstract
AI alignment is essential for the safe deployment of advanced AI systems. Given that values and preferences change over time, culture, social roles, and context, we need to develop a better understanding of the possible long-term consequences of AI alignment, in particular considering the likely ubiquitous future use of personalized AI assistants. We introduce a flexible and extensible mathematical modelling framework, rooted in social physics, aimed at answering macro-level questions regarding the evolving social norms in human populations under the assumption of frequent AI use. Our analysis is part-analytical, and part-simulation, enabling us to characterize the long-term dynamical consequences under a diverse set of starting assumptions. We highlight the risk of value lock-in, and normative mode collapse, prominently featured in non-adaptive alignment formulations. Beyond alignment, we advocate for the wider adoption of these kinds of social physics models as an epistemic bridge: enabling rapid, rigorous, and quantitatively-grounded hypothesis testing for sociotechnical foresight in general AI futures, and acting as a tractable precursor to more computationally expensive large-scale agentic evaluations.
Paper AI Chat
この論文のPDF全文を対象にAIに質問できます。
質問の例: