Deploying large language models (LLMs) for domain-specific analysis raises a critical validation challenge: distinguishing genuine structural reasoning from training data memorization. We address this through temporal obfuscation testing, which strips calendar dates, ticker symbols, and contextual markers from input sequences, forcing models to reason from numerical structure alone. Applying this framework to options dealer gamma exposure (GEX) patterns across two temporal scales, we validate detection using 2221 evaluations (1412 real windows plus 809 synthetic controls) spanning 2020–2025. At the single-day scale, obfuscation testing achieves 71. 5% detection of dealer hedging patterns with 91. 2% predictive accuracy; raw strike-level data outperforms pre-calculated GEX metrics by 30. 8 percentage points (92. 3% vs. 61. 5%), establishing that parametric aggregation represents lossy compression of structural signal. At the multi-day scale, 30-day regime detection achieves 81. 2% detection in 2024 (95% CI 75. 8, 86. 1%) versus 12. 1% in 2020 (95% CI 8. 1, 16. 6%) —a 69. 1 percentage point separation (φ = 0. 69, Fisher’s exact p = 1. 8 × 10−52) —with 0% false positives on synthetic controls. Multi-year analysis reveals regime evolution tracking zero-days-to-expiration (0DTE) adoption—detection rising from 3. 7% (2021) to 100% (2024) —with GEX magnitude growing from 3. 0B to 20. 3B. Stable detection despite collapsing profitability (Sharpe 1. 8 → 0. 1) confirms structural market mechanics rather than exploitable inefficiencies, establishing temporal obfuscation as a generalizable methodology for validating LLM reasoning in quantitative domains.
Regan et al. (2026) studied this question.