PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 12, 20260 citationsOpen Access

Real-Time Voice Cloning and Streaming System

View Full Paper
MHMasroor HussainLSLalithaditya SSVS Kranthi Varma

Key Points

  • The aim is to develop a system that generates and streams speech with minimal delay using voice samples.
  • Developed a real-time voice cloning system
  • Utilized voice encoding techniques and speech synthesis models
  • Implemented low-latency streaming with WebSocket communication
  • Tested on standard personal computers
  • Achieved reduced latency in audio playback
  • Enabled simultaneous speech generation and streaming
  • Enhanced real-time interaction capabilities
  • Suitable for various applications like virtual assistants and accessibility tools

Abstract

The rapid advancement of artificial intelligence and speech processing technologies has significantly enhanced human-computer interaction. However, traditional voice cloning and text-to-speech systems often rely on high-cost infrastructure and generate complete audio before playback, leading to increased latency. This paper presents a Real-Time Voice Cloning and Streaming System designed to generate and stream speech simultaneously with minimal delay. The system operates efficiently on standard personal computers and processes text along with a reference voice sample to produce speech incrementally. The proposed system integrates advanced speech synthesis models, voice encoding techniques, and a low-latency streaming pipeline using WebSocket-based communication. This enables continuous and smooth audio playback without pre-generating the entire audio. The system offers reduced latency, improved efficiency, and enhanced real-time interaction capabilities. It is suitable for applications such as virtual assistants, conversational agents, accessibility tools, and interactive platforms. Keywords: Voice Cloning, Real-Time Streaming, Text-to-Speech, Low Latency, AI.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hussain et al. (2026) studied this question.

synapsesocial.com/papers/69db37044fe01fead37c4f9chttps://doi.org/10.5281/zenodo.19494307
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1ReVoice: A Neural Network based Voice Cloning System2024 · 11 citations
  2. 2Real-Time Voice Interaction Engine: Architecture, Processing and Pipeline2026
  3. 3Speech Cloning: Text-To-Speech Using VITS2024 · 2 citations
  4. 4Neural Voice Replication: Multispeaker Text-to-Speech Synthesizer2024 · 3 citations
  5. 5Voice Cloning with Deep Neural Networks: Techniques, Evaluation, Applications, and Ethical Considerations2025