Heart disease is one of the leading causes of death in Japan, and the importance of early diagnosis through noninvasive auscultation is increasing, especially in rural areas. However, diagnosing heart sounds requires a high level of expertise, and physician shortages and aging populations are major challenges. Therefore, this study aims to develop an “automatic heart sound diagnosis system” that would complement doctors' diagnoses and could be applied to the training of young doctors and home care. Specifically, the authors apply a fast Fourier transform to heart sound data using the CirCor DigiScope dataset and generate color map images. This image is input into a CNN (Convolutional Neural Network), and a two-class classification is performed to distinguish between normal and abnormal images. 962 cases (511 normal and 451 abnormal) are used for learning, and each heart sound is processed as approximately 10 seconds of audio. As a result, the overall classification accuracy is 69 %, with 76 % for normal heart sounds and 53% for abnormal heart sounds, suggesting room for improvement, particularly in the detection of abnormalities.
Hosaka et al. (2025) studied this question.