PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 19, 20260 citationsOpen Access

Qwen3-Sussurro: Efficient Speech-to-Text Post-Correction via Parameter-Efficient Fine-Tuning

CECarlo Esposito

Key Points

  • The aim is to improve the quality of transcripts generated by automatic speech recognition systems by employing a fine-tuned model.
  • Developed Qwen3-Sussurro model using parameter-efficient fine-tuning
  • Utilized QLoRA for 4-bit quantization and Low-Rank Adaptation
  • Compared performance against zero-shot baselines using BLEU-4 and ROUGE-1 metrics
  • Observed inference speed improvements over the base model
  • Achieved a +807% improvement in BLEU-4 scores
  • Achieved a +237% increase in ROUGE-1 scores
  • Demonstrated statistical significance with p < 0.0001
  • Showed 4.6x faster inference compared to the base Qwen3-1.7B model

Abstract

Automatic Speech Recognition (ASR) systems often produce transcripts containing disfluencies, filler words (e.g., "um", "uh"), and grammatical errors that reduce readability and downstream task performance. We present Qwen3-Sussurro, a parameter-efficient fine-tuned model for post-processing ASR outputs. Using QLoRA (4-bit quantization with Low-Rank Adaptation) on the Qwen3-1.7B base model, we achieve substantial improvements over zero-shot baselines: +807% BLEU-4 and +237% ROUGE-1, with statistical significance (p < 0.0001). Additionally, our fine-tuned model demonstrates 4.6x faster inference compared to the base model.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Carlo Esposito (2026) studied this question.

synapsesocial.com/papers/6996a82decb39a600b3eea51https://doi.org/10.17613/sq6ez-f1p34
Ask AI
Helpful
Bookmark
Share
View Full Paper