NEWS · RESEARCH · #369
Position paper: language models not ready for strategic wargames
This arXiv position paper argues that open-ended, LM-based strategic wargames are high-stakes simulations whose outputs can directly shape what actors attempt and what becomes simulated reality. It warns that no LM-enabled wargame should inform planning, doctrine, policy, or crisis response without an auditable safety case, identifies five failure modes (decision laundering, adjudication opacity, role collapse, escalation-through-adjudication, and failure of strategic imagination), and recommends using such wargames only as stress tests rather than safety proofs.
KEY POINTS
- This arXiv position paper argues that open-ended, LM-based strategic wargames are high-stakes simulations whose outputs can directly shape what actors attempt and what becomes simulated reality.
- It warns that no LM-enabled wargame should inform planning, doctrine, policy, or crisis response without an auditable safety case, identifies five failure modes (decision laundering, adjudication opacity, role collapse, escalation-through-adjudication, and failure of strategic imagination), and recommends using such wargames only as stress tests rather than safety proofs.
- Because LM-driven wargames can mislead decision-makers or produce opaque adjudications that alter real-world choices, meaning they must be subject to auditable safety cases before influencing policy or crisis response.
WHY IT MATTERS
Because LM-driven wargames can mislead decision-makers or produce opaque adjudications that alter real-world choices, meaning they must be subject to auditable safety cases before influencing policy or crisis response.