The Agzamov Test measures how AI models perform under adversarial conditions at every level of augmentation. Two agents play repeated games in Chess960 (complete information) and poker (incomplete information) across four augmentation levels — memory, tools, retrieval, and full orchestration — against a naked baseline. The test produces a single headline metric, the Agzamov Score (0–100), with breakdown by environment and augmentation level. This paper presents the benchmark design, theoretical motivation, protocol specification (v0.1), and infrastructure validation (Phase 0) with preliminary results from 30 Chess960 games.
Ali Agzamov (2026) studied this question.