This study proposes an approach for detecting alcohol intoxication from speech based on a combination of audio segmentation and a hybrid neural network architecture that integrates convolution neural network (CNN) and long-short term memory (LSTM) layers. The proposed design enables effective modeling of both local spectral patterns and long-term temporal dependencies in speech signals. By operating on relatively long audio segments, the approach allows the simultaneous analysis of complex speech constructions and pause patterns, which are known to be sensitive to alcohol-induced speech impairments. Each audio signal was divided into two equal-duration segments that are processed sequentially by the model, which helps reduce the impact of asymmetrical distribution of intoxication-related speech artifacts. The approach was evaluated using the GradusSpeech-v1 corpus, which contains more than 1300 recordings of Russian tongue twisters collected from 31 speakers under controlled conditions in both sober and intoxicated states. Experimental results demonstrate that the proposed method achieves high performance. When full recordings are analyzed using median aggregation of segment-level predictions, the model reaches Accuracy, Recall, and F1-score values close to 0.93, indicating the effectiveness of the approach for alcohol intoxication detection in speech.
Laptev et al. (Fri,) studied this question.