Voice cloning with text-tag control over 11 emotions, 4 paralinguistic sounds (like laughter/breathing), and 14 Chinese dialects
Download with PC Client
16GB RAM recommended. 20GB+ storage recommended.
macOS 15+: M-series chips required.
Windows 10/11 64-bit: NVIDIA GPU with 6GB+ VRAM required.
Note: For NVIDIA GPUs, install a newer driver.1. Background & Team
PilotTTS is an open-source, lightweight autoregressive text-to-speech (TTS) system released in late May 2026 by the Amap Voice Team (AMAPVOICE), a subsidiary of AutoNavi/Alibaba. Driven by real-world demands for high-fidelity, dialect-rich, and emotional voice assistance in navigation and in-car systems, the team designed PilotTTS to deliver a production-ready speech synthesis solution that is both highly performant and user-friendly.
2. Key Features & Product Characteristics For content creators and general users, PilotTTS offers immense and highly controllable value:
3. Target Scenarios PilotTTS is ideally suited for smart travel and navigation systems, audiobook/podcast production, anime/game NPC voice acting, social media short video editing, and enterprise-level interactive digital humans.
4. Underlying Technology The core innovation of PilotTTS lies in its "minimalist modular recipe" combined with "rigorous data engineering." Instead of chasing bloated parameter sizes, it elegantly stitches together well-established open-source components:
Qwen3-0.6B (only 600 million parameters).w2v-bert-2.0.CosyVoice3's Conditional Flow Matching (CFM) decoder and Vocoder.