Abstract: Reported EEG Alzheimer’s accuracy can be dominated by evaluation choices rather than model qual-ity. Using OpenNeuro ds004504 (65 AD/CN subjects; DOI: 10. 18112/openneuro. ds004504. v1. 0. 8), we show three evaluation traps that can inflate or destabilize results. Trap 1: subject overlap en-ables memorization. When train and test include the same subjects, fingerprinting (subject-ID clas-sification) and AD/CN classification both reach near-ceiling epoch accuracy, consistent with subjectmemorization rather than disease learning. Trap 2: lucky folds. Under leakage-free, subject-disjointLeave-P -Subjects-Out (LPSO), performance still depends strongly on which subjects are held out, mak-ing single-split results non-reproducible. Trap 3: metric misalignment. Epoch accuracy can disagreewith subject-accuracy and can change model/hyperparameter rankings. We compute subject-accuracyas one decision per subject: label a subject AD if a majority of that subject’s test epochs are predictedAD (otherwise CN; τ = 0. 5). We recommend repeated subject-disjoint evaluations and reporting per-formance across many folds (median/IQR, or min/max). Importantly, there is a study scope here sincethe instability and quantization magnitudes reported here are demonstrated in a small-to-medium co-hort setting (N = 65, primarily P ∈ 2, 6). Similar effects can still arise in larger datasets whenheld-out cohorts are relatively small or non-representative, including cases where favorable folds areselectively reported. At a bare minimum, studies must disclose the held-out subject IDs per fold andreport both epoch accuracy and subject-accuracy (subject-accuracy as primary). Supplementary materials included: This upload contains the manuscript PDF of the preprint and the machine-readable split manifests (CSV files listing held-out subject IDs per fold) for all LPSO settings reported in the paper.
Adel Sahuc (2026) studied this question.