PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 2, 20250 citationsOpen Access

Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment

View Full Paper
KWK. WangWWWenning WeiYDYan Deng

Key Points

  • Fine-tuning large multimodal models for automatic pronunciation assessment achieves notable improvements.
  • The Pearson Correlation Coefficient reaches 0.9, indicating a strong relationship with performance metrics.
  • Assessment performance is higher at word and sentence levels compared to phoneme-level evaluations.
  • Despite advancements, challenges with fine-grained assessment suggest areas for further investigation.

Abstract

Automatic Pronunciation Assessment (APA) is critical for Computer-Assisted Language Learning (CALL), requiring evaluation across multiple granularities and aspects. Large Multimodal Models (LMMs) present new opportunities for APA, but their effectiveness in fine-grained assessment remains uncertain. This work investigates fine-tuning LMMs for APA using the Speechocean762 dataset and a private corpus. Fine-tuning significantly outperforms zero-shot settings and achieves competitive results on single-granularity tasks compared to public and commercial systems. The model performs well at word and sentence levels, while phoneme-level assessment remains challenging. We also observe that the Pearson Correlation Coefficient (PCC) reaches 0.9, whereas Spearman's rank Correlation Coefficient (SCC) remains around 0.6, suggesting that SCC better reflects ordinal consistency. These findings highlight both the promise and limitations of LMMs for APA and point to future work on fine-grained modeling and rank-aware evaluation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wang et al. (2025) studied this question.

synapsesocial.com/papers/68de5da783cbc991d0a20d3ehttps://doi.org/10.48550/arxiv.2509.15701
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Pronunciation Assessment with Multi-modal Large Language Models2024 · 1 citations
  2. 2AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation2025
  3. 3Beyond Modality Limitations: A Unified MLLM Approach to Automated Speaking Assessment with Effective Curriculum Learning2025
  4. 4Leveraging Large Language Models to Refine Automatic Feedback Generation at Articulatory Level in Computer Aided Pronunciation Training2024 · 3 citations
  5. 5Investigating Automatic Scoring and Feedback using Large Language Models2024