VieNeu-TTS v3 Turbo Brings 48 kHz Vietnamese Speech to CPUs
The newly released VieNeu-TTS v3 Turbo model delivers high-fidelity 48 kHz Vietnamese speech synthesis on standard CPUs, offering developers an efficient, open-source voice solution.

The open-source community has received a major upgrade in Vietnamese speech technology with the release of VieNeu-TTS v3 Turbo. This text-to-speech model, which has already amassed nearly 780,000 downloads on Hugging Face, features a compact architecture of roughly 100 million parameters. Released under the permissive Apache-2.0 license, the model was trained entirely from scratch using approximately 10,000 hours of bilingual English and Vietnamese voice data.
A key highlight of the v3 Turbo release is its ability to generate high-fidelity 48 kHz audio, doubling the 24 kHz output of its predecessor by utilizing the MOSS Audio Tokenizer neural codec. For hardware deployment, the model offers a highly flexible architecture. On CPU setups, it runs without PyTorch dependencies via the ONNX Runtime. When deployed on a single Nvidia RTX 3060 GPU, the system can handle 16 concurrent real-time streams while maintaining a low first-audio latency of about 115 milliseconds.
Developers can access the model through a Python SDK called vieneu, which distributes weights in ONNX, Safetensors, and GGUF formats. The system supports 23 preset voices representing Northern, Central, and Southern Vietnamese dialects, alongside instant voice cloning capabilities. It also handles English-Vietnamese code-switching, inline emotional cues, and single-GPU LoRA fine-tuning.
For AI practitioners, this release simplifies the integration of high-quality voice agents into existing pipelines. The model includes an OpenAI-compatible API endpoint that integrates directly with popular conversational frameworks like Pipecat, LiveKit, and standard OpenAI SDK clients. This compatibility, combined with its low-latency CPU path, allows developers to deploy localized, responsive voice applications without relying on expensive GPU infrastructure.
This is our own summary of reporting by AlphaSignal



