Data-driven machine learning methods, most notably artificial neural networks (ANNs), have become a major part of signal processing research. They have shown outstanding performance, generally outperforming conventional signal processing methods, in a wide range of applications, including single- and multichannel speech enhancement. Due to the typically high complexity of these approaches, the field of explainable AI (XAI) has emerged, providing methods to explain or interpret predictions made by ANNs. However, while concept-based XAI approaches investigate certain characteristics represented in the hidden features of an ANN, most of these XAI methods provide local, example-dependent explanations of ANN behavior, neglecting higher-level signal processing concepts known from conventional methods such as the differentiation between spectral and spatial filtering in multichannel speech enhancement. To bridge this gap, the main contribution of this thesis is a toolbox that allows to analyze the hidden features of an ANN designed for multichannel speech enhancement, denoted as neural spatiospectral filter (NSSF), with respect to the concepts of spatial and spectral filtering. Hence, this thesis resides at the intersection of conventional signal processing, ANN research, and the field of XAI for the task of multichannel speech enhancement. The main focus of the proposed analysis tools is on investigating whether and where spatial information is reflected and, thus, extracted from the multichannel input signal inside an NSSF. For this, specific analysis scenarios for training and testing are designed and an evaluation framework based on clustering, including suitable evaluation metrics, is proposed. With regard to spectral filtering, metrics to assess the filtering at any given stage within an NSSF are proposed, which can for example be applied to quantify the amount of spectral filtering performed by a preprocessing module. To validate the effectiveness of these tools, two conceptually different mask-based NSSF frameworks, namely the Complex-valued Spatial Autoencoder COSPA as a filter-and-sum NSSF with a multichannel mask, and the tempospectral joint nonlinear filter (FT-JNF) as a single-channel masking NSSF, are analyzed for the tasks of denoising and target speaker extraction. Particular focus is also put on the effect of the characteristics of the training target signal on the interpretation of the spatial filtering capabilities of the NSSFs, considering a dry source signal, a reverberated source signal, and a delay-and-sum beamformer-filtered signal as training target signals. As main result, it is shown that spatial information is reflected in the hidden features generated by recurrent neural network layers that have joint access to all channels. Moreover, the two NSSF frameworks represent and use spatial information differently in general, and in spatially unconstrained scenarios in particular, due to their inherent conceptual differences. Furthermore, the amount of spatial information represented inside an NSSF depends on the training target signal, due to possible implicit constraints on the filtering imposed by the training target signal. The knowledge about the network layers responsible for spatial filtering are then exploited to extend the NSSFs for denoising to the task of target speaker extraction where the target speaker is identified via its direction of arrival. It is confirmed that these extensions use the provided spatial information as intended and succeed in extracting the correct target speaker. For COSPA, spatial awareness across the whole testset can be shown, indicating that the network not only differentiates between target and interfering sources but also internally encodes where each of them is placed. Furthermore, for COSPA, the preprocessing module for spectral filtering is investigated in more detail, showing that the intended purpose of the module is not met by the original proposal but that it can be easily achieved by modifying the cost function used for training. This modification increases the interpretability of the processing pipeline within the NSSF without harming the performance and at no additional computational cost. The methods proposed in this thesis can be used to gain valuable insights into the proverbial black box of ANN processing, allowing to connect well established concepts from conventional signal processing with ANN-based approaches. These insights allow to make informed decisions when designing (or improving), training, and deploying NSSFs in real-world applications.
Annika Briegleb (Thu,) studied this question.