Audio signal processing methods based on neural networks (NNs) are typically trained at a single sampling frequency (SF). To handle untrained SFs, signal resampling is commonly used, but it can degrade NN performance, particularly at SFs much lower than the trained SF. As an alternative, we previously proposed a SF-independent (SFI) convolutional layer, which generates convolutional kernels based on an input SFs from a prototype kernel defined as a continuous-time/frequency (i.e., SFI-domain) function. Obtaining this function is therefore essential for incorporating SFI layers into NNs. However, no method exists to directly construct the SFI-domain function from pre-trained convolutional kernels. Consequently, the entire network must be retrained after replacing standard convolutional layers with SFI layers. In this presentation, we propose a method to convert a pre-trained convolutional layer into its SFI counterpart. The method approximates the original kernel using a NN that takes continuous time/frequency as input. Once trained, this network can serve as the SFI-domain function for the SFI convolutional layer. This enables us to build an SFI version of pre-trained models based on standard convolutional layers. Experiments on music source separation demonstrate that the proposed method achieves comparable performance to the approach that retrains the entire network.
Imamura et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: