Install the app
Download the free Windows executable. Local engine wrappers are bundled in.
Transcribe audio to text and generate realistic voice clones locally on Windows. Powered by Whisper, Chatterbox, Kokoro, LuxTTS, TADA, and Qwen — no cloud required.
Faster-Whisper · large-v3 · local device
Speech-to-text · 24+ languages
Lightweight local synthesis
Zero-shot cloning · 23 languages
Fast CPU-only generation
CTranslate2 · ONNX backend
Sherpa-ONNX · Kokoro pipeline
Multimodal custom voice backends
Transcribed · 12:04 · English
Generated · Kokoro 82M
Transcribed · 48:31 · English
Donate on Ko-fi to support open-source work
View source code, releases and issues
Local address where sidecar backend process is running.
Configure default transcription language and theme.
Primary transcription language for new audio jobs.
Allow online connections for downloading models and updates.
Application color interface.
Engines dispatch to GPU when available, otherwise CPU.
Choose automatic dispatch or force a device.
Lower precision runs faster with minimal quality loss.
Faster-Whisper · large-v3 · local device
Local-first speech transcription and TTS workstation.
Open source, no license keys.
Desktop build target.
All processing stays on device.
No uploads. No account setup. Every model runs inside the desktop sidecar on your own machine.
Download the free Windows executable. Local engine wrappers are bundled in.
Whisper for transcription, or Chatterbox and Kokoro for zero-shot voice cloning.
Transcribe audio or generate speech on CPU or GPU, entirely without internet.
Declarative backend architecture powering offline speech and voice capabilities.
Tiny, Base, Small, Medium, Large-v3, and Turbo variants for 24+ language transcription and real-time streaming.
Zero-shot voice cloning across 23 languages using 5-second reference audio prompts.
Fast English voice cloning with native support for tags like [laugh], [cough],
and [chuckle].
High-fidelity, ultra-lightweight speech synthesis with a minimal memory footprint.
CPU-friendly voice generation built for fast execution without a dedicated GPU.
HumeAI TADA (1B & 3B) text-acoustic dual alignment and Qwen2 Audio custom voice backends.
All transcription and synthesis run on a local sidecar. No uploads, no accounts, no tracking.
Chatterbox clones a voice across 23 languages from a 5-second reference, with native expression tags.
CincoScribe is 100% free and open-source under the MIT License.
Windows desktop suite. No cloud, no subscriptions, no account.