Windows Desktop Suite · Free & Offline

Offline Transcription & Voice Cloning

Transcribe audio to text and generate realistic voice clones locally on Windows. Powered by Whisper, Chatterbox, Kokoro, LuxTTS, TADA, and Qwen — no cloud required.

No account required No subscriptions MIT licensed
C
Settings
Text to Speech
Generate speech locally with Kokoro, Chatterbox and LuxTTS.

Live Output

Faster-Whisper · large-v3 · local device

[INFO] Faster-Whisper large-v3 loaded on local device. [INFO] Detected language: English (confidence 0.98) [00:06] "Everything here runs on your own machine."
Local engine ready Device: CPU / CUDA · Offline
Input Text
Type or paste your text here to convert it to speech...
Voice
Bella - Clear & Professional Voice
Model
Kokoro 82M (Local)
Speed 1.0x
Faster-Whisper large-v3

Speech-to-text · 24+ languages

INSTALLED
Kokoro 82M

Lightweight local synthesis

INSTALLED
Chatterbox Multilingual

Zero-shot cloning · 23 languages

INSTALLED
LuxTTS

Fast CPU-only generation

AVAILABLE
ASR Runtime

CTranslate2 · ONNX backend

RUNNING
TTS Runtime

Sherpa-ONNX · Kokoro pipeline

RUNNING
TADA & Qwen

Multimodal custom voice backends

IDLE
interview-notes.wav

Transcribed · 12:04 · English

DONE
narration-take-3.wav

Generated · Kokoro 82M

DONE
podcast-ep12.mp3

Transcribed · 48:31 · English

DONE

Server Connection

Local address where sidecar backend process is running.

ONLINE
http://127.0.0.1:5555

Recognition & Appearance

Configure default transcription language and theme.

Default Language

Primary transcription language for new audio jobs.

Auto Detect
Network Permission

Allow online connections for downloading models and updates.

Allowed (Online)
Appearance Theme

Application color interface.

Dark Theme

CUDA Device

Engines dispatch to GPU when available, otherwise CPU.

DETECTED
VRAM allocation3.4 / 8.0 GB
Compute load61%
Execution Device

Choose automatic dispatch or force a device.

Auto (CUDA → CPU)
Precision

Lower precision runs faster with minimal quality loss.

float16

Live Output

Faster-Whisper · large-v3 · local device

[INFO] Sidecar server initialized on http://127.0.0.1:5555 [INFO] Faster-Whisper large-v3 engine loaded. [INFO] Detected language: English (confidence 0.98) [INFO] No outbound network requests made.
Local engine ready sidecar.log · Level: INFO

CincoScribe

Local-first speech transcription and TTS workstation.

v0.1.0
License

Open source, no license keys.

MIT
Platform

Desktop build target.

Windows x64
Privacy

All processing stays on device.

100% Offline
How It Works

Three steps.
Zero cloud.

No uploads. No account setup. Every model runs inside the desktop sidecar on your own machine.

Step 01 / 03
1

Install the app

Download the free Windows executable. Local engine wrappers are bundled in.

2

Pick an engine

Whisper for transcription, or Chatterbox and Kokoro for zero-shot voice cloning.

3

Process locally

Transcribe audio or generate speech on CPU or GPU, entirely without internet.

0
Model Engines
0+
ASR Languages
0s
Audio To Clone A Voice
0%
Free & Offline
Model Engines

Six local engines, one app

Declarative backend architecture powering offline speech and voice capabilities.

Speech To Text

Faster-Whisper

Tiny, Base, Small, Medium, Large-v3, and Turbo variants for 24+ language transcription and real-time streaming.

Multilingual Cloning

Chatterbox Multilingual

Zero-shot voice cloning across 23 languages using 5-second reference audio prompts.

Expression Tags

Chatterbox Turbo

Fast English voice cloning with native support for tags like [laugh], [cough], and [chuckle].

Lightweight 82M

Kokoro 82M

High-fidelity, ultra-lightweight speech synthesis with a minimal memory footprint.

Fast CPU

LuxTTS

CPU-friendly voice generation built for fast execution without a dedicated GPU.

Multimodal

TADA & Qwen

HumeAI TADA (1B & 3B) text-acoustic dual alignment and Qwen2 Audio custom voice backends.

OFFLINE
Privacy

Nothing leaves your machine

All transcription and synthesis run on a local sidecar. No uploads, no accounts, no tracking.

Local only
ZERO-SHOT
Voice Cloning

Five seconds is enough

Chatterbox clones a voice across 23 languages from a 5-second reference, with native expression tags.

[laugh] [cough] [chuckle]
Pricing

Free, forever

CincoScribe is 100% free and open-source under the MIT License.

Desktop App
Free
MIT License · Fully offline
  • Unlimited offline transcriptions
  • Zero-shot voice cloning
  • No license keys or paywalls
  • Full data privacy and control
Download Latest Release
Support
Donate
Optional · Supports development

If CincoScribe helps your workflow, consider supporting ongoing model and feature development on Ko-fi.

Buy Me a Coffee at ko-fi.com
FAQ

Questions

No. All transcription and text-to-speech synthesis run locally on your device sidecar without sending data online.
Chatterbox Multilingual, Chatterbox Turbo, Kokoro 82M, LuxTTS, TADA, and Qwen CustomVoice.
No. Engines dispatch to CPU or CUDA automatically, and LuxTTS is built specifically for fast execution without a dedicated GPU.
Yes — MIT licensed, with no license keys, paywalls, or subscriptions. Ko-fi donations are entirely optional.
Download

Run it locally

Windows desktop suite. No cloud, no subscriptions, no account.

v0.1.0 (Pre-release) Released 21 Jul 2026 · All releases
Early release. Custom voice cloning and CUDA acceleration are still in active development. Windows builds are published on the GitHub releases page.
Portfolio Home