PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 15, 2026JMIR Formative Research0 citationsOpen Access

Automatic Speech Recognition and Large Language Models for Multilingual Pathology Report Generation: Proof-of-Concept Study

View Full Paper
KLKuan‐Hsun LinCCChia-Ping ChangCKChen-Tsung Kuo

Key Points

  • This study aims to assess the effectiveness of an ASR pipeline and LLMs in transcribing multilingual pathology dictation accurately.
  • Controlled proof-of-concept study with 125 simulated mixed Chinese-English audio recordings.
  • Audio was transcribed using Whisper ASR with and without contextual messages.
  • Transcripts were converted into reports using Qwen2, Llama3.1, and Gemma2 language models.
  • ASR contextual message reduced average character error rate from 0.344 to 0.066, P<0.001.
  • Qwen2 achieved the highest scores: Bilingual Evaluation Understudy 0.644 and ROUGE-1 0.866.
  • Total error rates were 16.8% for Qwen2, 45.6% for Llama3.1, and 92.8% for Gemma2.

Abstract

Abstract Background Accurate transcription of pathology gross examination dictation is important for clinical documentation, but multilingual dictation remains challenging in settings where clinicians mix Chinese and English while final pathology reports are written in English. Objective This study aimed to evaluate whether a Whisper-based automatic speech recognition (ASR) pipeline guided by contextual system messages and combined with open-source large language models (LLMs; Qwen2:72b, Llama3.1:70b, Gemma2:27b) could improve multilingual (Chinese-English) pathology dictation transcription accuracy and generate clinically appropriate English gross description reports. Methods We conducted a controlled proof-of-concept study using 125 simulated mixed Chinese-English pathology gross examination audio recordings created by physicians or pathologists. Audio recordings were transcribed using Whisper ASR with and without a contextual system message. The ASR transcripts were then converted into English gross description reports using 3 open-source LLMs: Qwen2:72b, Llama3.1:70b, and Gemma2:27b. Outcomes included character error rate, Bilingual Evaluation Understudy, Recall-Oriented Understudy for Gisting Evaluation (ROUGE)-1, ROUGE-2, ROUGE-L, Metric for Evaluation of Translation with Explicit Ordering, pathologist Win-Tie-Lose rankings, report-level error categories, inference time, and interrater agreement. Results The ASR contextual system message reduced the mean character error rate from 0.344 (SD 0.176; 95% CI 0.313‐0.375) to 0.066 (SD 0.100; 95% CI 0.048‐0.084; P <.001). Qwen2:72b achieved the highest automated metric scores, including a Bilingual Evaluation Understudy of 0.644 (SD 0.307), ROUGE-1 of 0.866 (SD 0.163), ROUGE-2 of 0.771 (SD 0.235), ROUGE-L of 0.842 (SD 0.178), and Metric for Evaluation of Translation with Explicit Ordering of 0.805 (SD 0.214). Pathologist-coded total error rates were 16.8% (21/125) for Qwen2:72b, 45.6% (57/125) for Llama3.1:70b, and 92.8% (116/125) for Gemma2:27b. The exact agreement between the 2 pathologists across full ranking categories was 76.8% (96/125; Cohen κ=0.668), and agreement on the top-ranked model or tied top group was 81.6% (102/125; Cohen κ=0.722). Conclusions In this proof-of-concept evaluation, contextual prompting improved ASR transcription accuracy, and Qwen2:72b generated the most accurate English pathology reports among the evaluated LLMs. However, the study used simulated audio recordings, a local vocabulary prompt, and report-level rather than term-level clinical annotation. LLM-generated reports should therefore be considered draft documentation requiring pathologist verification, and prospective validation in real clinical workflows is needed before clinical deployment.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Lin et al. (2026) studied this question.

synapsesocial.com/papers/6a06b928e7dec685947abb53https://doi.org/10.2196/90814
Ask AI
Helpful
Bookmark
Share
View Full Paper