This study evaluated the performance of various normality tests including Shapiro–Wilk, Shapiro–Francia, Anderson–Darling, Lilliefors, Cramer–von Mises, and Jarque–Bera under different conditions, both with and without the presence of outliers. Monte Carlo simulations were conducted to calculate the type I error rates, power, and the Kappa–Fleiss agreement coefficient, which measured the concordance among the tests. For normally distributed data without outliers, the Shapiro–Wilk and Shapiro–Francia tests showed the best control over the type I error rate. In contrast, with the introduction of outliers, the Lilliefors and Cramer–von Mises tests performed better. In terms of test power, the Shapiro–Wilk and Shapiro–Francia tests performed best for distributions without outliers, while the Jarque–Bera test was more robust in the presence of outliers. Overall, the results highlight the sensitivity of these tests to sample size and the presence of outliers, suggesting that Shapiro–Wilk and Shapiro–Francia are suitable for data without outliers, while Jarque–Bera may be preferred in contaminated samples. The tests showed higher concordance for exponential and lognormal distributions but lower concordance for beta, χ², and t-Student distributions, illustrating the complexity of normality identification across various contexts.
HONÓRIO et al. (Thu,) studied this question.