The effects of background spectral and temporal structure and overall level on the masking of speech were assessed for sentences presented in a speech-shaped noise (SSN), harmonic complex tone (HCT) (repetition period = 4.55 ms), and iterated ripple noise (IRN) (delay = 4.55 ms). The speech + noise was presented at 50 and 80 dB SPL using signal-to-noise ratios (SNRs) from 0 to -15 dB in 5-dB steps. The noises had similar spectral envelopes, but the HCT and IRN had spectral dips between peaks corresponding to the harmonic frequencies, and the HCT also had temporal dips. Especially for the SNRs of -5 and -10 dB, speech identification was best for the HCT masker and worst for the SSN, for both overall levels. For the SSN, the SNR required for 50% correct (SNR-50) was higher (worse) at 80 dB than at 50 dB, consistent with poorer frequency selectivity at the higher level. For the HCT, SNR-50 values were lower for the higher level, consistent with a better ability to "listen in the dips" at the higher level. The results indicate that the relative benefit of spectral and temporal dips varies with level.
Narne et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: