PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 12, 2026Intellekt Sist Proizv0 citationsOpen Access

Development of an Automatic System for Transcribing Voice Recordings of Company Meetings Using Neural Networks

View Full Paper
MPM. A. PolozovLKL. A. Korobova

Key Points

  • The aim is to create a local and secure transcription system for audio recordings of company meetings.
  • Developed an automatic transcription system using neural networks and open-source models.
  • Integrated Whisper for speech recognition and pyannote.audio for speaker identification.
  • Implemented a hybrid architecture using Python and C# for audio processing and user interface.
  • Conducted tests in various acoustic environments to assess transcription accuracy.
  • Achieved acceptable Word Error Rate (WER) and Character Error Rate (CER) for business use.
  • Ensured good transcription quality even in noisy environments.
  • Demonstrated scalability and support for the Russian language.

Abstract

Amid the growing prevalence of remote work and the digital transformation of business, the automation of routine processes-including the processing of audio recordings of meetings and conferences-is becoming increasingly important. Modern video conferencing systems offer automatic transcription features; however, these are often restricted to paid subscription plans, require an internet connection, and do not provide a sufficient level of confidentiality. Consequently, the development of a local, economically accessible, and secure solution for speech transcription has become highly relevant. This paper presents an automatic system designed for transcribing voice recordings of meetings within a project company. A distinctive feature of the developed system is the integration of open-source neural network models: Whisper (for speech recognition), pyannote.audio (for diarization - speaker identification), and an open-source GPT model (for post-processing and text formatting). The system is implemented using a hybrid architecture employing two programming languages, Python and C#, which combines high-performance audio processing with a user-friendly graphical interface. The key advantages of the solution include complete autonomy (no cloud connection required), support for the Russian language, scalability, and compliance with information security requirements. Testing on a control audio fragment yielded Word Error Rate (WER) and Character Error Rate (CER) metrics at levels acceptable for business use. To assess the accuracy of the designed system, additional tests were conducted in various acoustic environments, demonstrating that the system ensures good transcription quality under typical operating conditions, as well as in the presence of background noise. The implemented software product will enable companies to save on paid access to video conferencing systems and corporate subscriptions, while simultaneously increasing the transparency and efficiency of documenting meetings and conferences. This work holds both theoretical and practical significance for the development of domestic IT solutions in the field of corporate automation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Polozov et al. (2026) studied this question.

synapsesocial.com/papers/69db380f4fe01fead37c6281https://doi.org/10.22213/2410-9304-2026-1-52-63
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Prototype mobile application definitions fresh products based on neural network2021 · 8 citations
  2. 2EVALUATION OF RUSSIAN SPEECH RECOGNITION QUALITY ON CLEAN AND NOISY AUDIO DATA2025 · 1 citations
  3. 3A NEURAL NETWORK MODEL FOR RUSSIAN SPEECH RECOGNITION2020 · 1 citations
  4. 4Telemarketing automation based on the MIKO IP-telephony module2022 · 2 citations