Spike sorting is a critical technique in identifying the responses of neurons in neurophysiological experiments. The quality of such sorting is normally evaluated in terms of the number of correct and incorrect identifications of spike times in simulated recordings. In this study, we evaluate the overall performance of spike sorting techniques in terms of how well they detect responses of neurons to simulated stimuli. We do so by computing the observed effect size or test power obtained when varying the spike sorting technique and simulated spiking activity. We show that varying the spike detection threshold causes noticeable changes in both the observed effect size (η 2 , -0.033 - 0.058, 60% change) and the corresponding statistical power of the test (0.89 - 1.00, 11% change). These changes are puzzling considered in terms of signal detection theory, according to which changing a threshold based on a single variable should not change the effect size for an experiment. We also show that such changes persists across changes of both spike sorting parameters and the underlying simulation. Examination of effect size and test power in simulated experiments is a useful technique for evaluating the performance of spike sorting techniques. In this example, it reveals fundamental gaps in our understanding of spike detection behavior while at the same time enabling improved statistical power by identifying factors that maximize effect size.
Steinmetz et al. (Fri,) studied this question.