This Version 1 release extends the earlier v0.9.2 independent C++ cross-domain toy benchmark into a larger clean-start staged simulation evidence ladder, including falsification controls, component ablation, boundary sensitivity, human-latency thresholding, Safety Slack calibration, domain-specific deep runs, and a large C++ replication totaling 1.2096 billion synthetic simulated episodes. This is a synthetic toy diagnostic benchmark only. It is not real-world validation, not a deployed-system safety guarantee, not a medical or industrial safety certification, and not a proof of safety. The supported claim is that the core Signal-Time-Authority oversight pattern remains internally reproducible across the current toy simulations.
Htet Ko Ko Naing (2026) studied this question.