Abstract Introduction Polysomnography remains essential for identifying sleep disorders, but the manual scoring required to analyse these recordings is time-consuming and resource-heavy. Although automated sleep scoring offers an alternative, their performance can vary drastically from one sleep-disorder type to another. The present work examined how current state-of-the-art classifiers perform when applied to insomnia patients. Methods PSG data from 904 individuals with chronic insomnia were analysed. Five automated staging algorithms were compared: GSSC, U-Sleep, Luna, STAGES, and YASA. Their classification accuracy was assessed using macro F1 scores, confusion matrices, and predicted sleep parameters. Regression models were used to test whether demographic variables, sleepiness levels, or PSG characteristics influenced performance. Results Performance varied across classifiers. GSSC obtained the highest overall F1 score across sleep stages (0.66), followed by U-Sleep (0.62), Luna (0.56), STAGES (0.54), and YASA (0.52). GSSC also obtained the highest per-stage F1 scores for Wake (0.83), N2 (0.80), N3 (0.71), and REM (0.76), and was matched by U-Sleep in N1 and REM, and by Luna in N3. The poorest results were observed in deep sleep for STAGES (N3 = 0.39) and for REM in YASA (F1 = 0.35). Typical errors involved confusion between N1 and Wake/N2, and between N3 and N2. REM was frequently misclassified as Wake or light NREM by STAGES, Luna, and YASA. GSSC and U-Sleep showed the least variation across demographic subgroups, whereas STAGES and Luna were more affected. Algorithm performance was similar in patients with and without PSG-defined insomnia. Among sleep metrics, U-Sleep showed the best agreement for total sleep time (R2 = 0.88), STAGES for sleep onset latency (R2 = 0.82), and GSSC for wake after sleep onset (R2 = 0.82). Conclusion Automated staging systems show reliable but uneven accuracy when applied to chronic insomnia. GSSC and U-Sleep consistently outperformed the other classifiers and demonstrated the lowest demographic bias. The results indicate that automated staging can be used in large-scale clinical datasets, although the choice of algorithm remains critical when interpreting sleep stages and derived metrics in patients with insomnia. Support (if any)
Hanif et al. (Fri,) studied this question.