Skip to content
STACK IT FAST

VoiceStudio

Curated OSSHybridLocal AI Voice Cloning & Speech Studio2-5 people

Audited from github.com/debpalash/VoiceStudio

VoiceStudio is an open-source, local-first alternative to ElevenLabs for voice cloning, voice design, video dubbing, dictation, transcription and audiobook creation. It runs a Python/PyTorch model backend behind a local API with a React desktop app.

Language
Python
Hosting
Docker
License
AGPL-3.0
Running for
5 months
Team
2-5 people

Why this architecture

Heavy model work stays in a Python FastAPI backend with a pluggable, hardware-aware engine registry. Desktop, web, MCP and remote-worker clients all use the same local API, so voice models run on the user's own GPU without a cloud dependency.

Tech stack by layer

20 technologies · audited Sep 25, 2026
Frontend & UI
  • ElectronCurrent desktop shell (electron/) built with electron-vite and electron-builder
  • ReactReact 19 UIs for the Vite web frontend and the Electron renderer
  • TanStack QueryServer-state fetching against the local API, alongside TanStack Router, Form and Virtual
  • Tailwind CSSStyling (Tailwind 4) with shadcn components on Radix UI and Base UI
  • TauriEarlier desktop shell (frontend/src-tauri, Rust) whose migration to Electron is documented in docs/
Backend & APIs
  • PythonBackend service (backend/main.py) and the bundled omnivoice TTS package, managed with uv and hatchling
  • FastAPILocal HTTP and WebSocket API on port 3900 with a Scalar-rendered OpenAPI reference
  • PyTorchRuns TTS, cloning, diarization and separation models on CUDA, ROCm, Apple MPS or CPU
Data & Persistence
  • AlembicSchema migrations for the local app database under backend/migrations
Infrastructure & Deploy
  • BunJavaScript package manager and script runner for the frontend/electron workspaces
  • TurborepoBuild orchestration across the JS workspaces
  • DockerCUDA and ROCm server images (deploy/Dockerfile) with a Bun-built frontend stage
  • PlaywrightE2E, production-bundle, performance and visual regression tests for the frontend
Tooling, Testing & Ops
  • TransformersHugging Face model loading for the default OmniVoice engine and other TTS/ASR backends
  • WhisperXPrimary transcription with word-level forced alignment used for dubbing lip-sync
  • MLXApple Silicon engines (mlx-whisper, mlx-audio, parakeet-mlx) gated by platform markers
  • Model Context Protocolbackend/mcp_server.py exposes voice generation and transcription as tools for AI agents

VoiceStudio architecture diagram

Open SVG
VoiceStudio architecture diagramDesktop app → Local API (HTTP · WebSocket); Web UI → Local API (HTTP · WebSocket); AI agents → Local API (MCP); Local API → Engine registry (jobs); Engine registry → Model runtime (inference); Local API → App database (app state); Engine registry → Hugging Face (load models)CLIENTSSERVICESWORKERS & JOBSDATA & STORAGEEXTERNALDesktop appElectron · ReactWeb UIVite · TanStackAI agentsMCPLocal APIFastAPI · port 3900Engine registryTTS · cloning · ASRModel runtimePyTorch · MLXApp databaseAlembic migrationsHugging Facemodel weightsinferenceHTTP · WebSocketHTTP · WebSocketMCPjobsapp stateload models
How the main components of VoiceStudio connect, drawn from the audited repository.
Diagram as text
  • Desktop app (Electron · React) → Local API (FastAPI · port 3900): HTTP · WebSocket
  • Web UI (Vite · TanStack) → Local API (FastAPI · port 3900): HTTP · WebSocket
  • AI agents (MCP) → Local API (FastAPI · port 3900): MCP
  • Local API (FastAPI · port 3900) → Engine registry (TTS · cloning · ASR): jobs
  • Engine registry (TTS · cloning · ASR) → Model runtime (PyTorch · MLX): inference
  • Local API (FastAPI · port 3900) → App database (Alembic migrations): app state
  • Engine registry (TTS · cloning · ASR) → Hugging Face (model weights): load models

Key architectural decisions

6 decisions
  1. 01

    Python model backend behind a local API, UI as a separate client

    backend/ (FastAPI + uvicorn with WebSockets) owns engines, jobs and storage, and the React UIs in frontend/ and electron/ talk to it over HTTP on localhost:3900; the same API powers the MCP server and remote workers.

  2. 02

    Pluggable engine registry with hardware-gated dependencies

    backend/engines hosts multiple TTS and ASR engines; pyproject.toml uses platform markers so MLX engines install only on Apple Silicon while WhisperX/faster-whisper, KittenTTS, Demucs and pyannote cover other platforms.

  3. 03

    Electron replaced Tauri as the desktop shell

    frontend/src-tauri (Rust) still exists, but electron/ is the default dev and build target in the root package.json, and a series of docs/electron-*.md files plus electron/PARITY.md track feature parity during the migration.

  4. 04

    One Docker image per GPU vendor with a torch-preservation guard

    deploy/Dockerfile builds from a pytorch CUDA base by default and a ROCm base via BASE_IMAGE, installs with uv against deploy/torch-constraints.txt, and fails the build if the GPU-specific torch was replaced by a generic wheel.

  5. 05

    Dependency pins documented inline with issue numbers

    pyproject.toml caps setuptools (<80), pyannote-audio (<4) and pedalboard (<0.9.21) and fastapi (<0.137) with comments naming the upstream breakage and the VoiceStudio issue, so pins can be revisited deliberately.

  6. 06

    Provenance watermarking of generated speech

    pyproject.toml depends on audioseal to embed an imperceptible neural watermark in AI-generated audio for later detection.

How VoiceStudio is built

How VoiceStudio is structured

VoiceStudio is a Python + JavaScript monorepo. docs/STRUCTURE.md describes the layout.

  • Python: pyproject.toml (package omnivoice, built with hatchling, managed with uv) covers:
    • backend/: the FastAPI app (backend/main.py) with api/, services/, engines/, worker/, plugins/, hooks/, schemas/, migrations/ and the MCP server (backend/mcp_server.py, backend/mcp_shim),
    • omnivoice/: the bundled TTS model code (models, training, eval, CLI).
  • JavaScript: a Bun workspace (package.json, packageManager: bun@1.4.2) with frontend/ (Vite + React, plus the older src-tauri shell) and electron/ (the current desktop app). turbo.json orchestrates builds.
  • Operations: deploy/ (Dockerfile, docker-compose, remote worker installer), bin/ (prebuilt omnivoice-tts binaries per OS), scripts/ (install, release, smoke and benchmark scripts) and notebooks/ (a Colab notebook).

Architecture decision records live in docs/adr. AGENTS.md, CLAUDE.md and .claude/agents give coding agents instructions.

Frontend

Both UIs use React 19. frontend/ is a Vite app with Radix UI primitives, Tailwind CSS 4, TanStack Query, Table and Virtual, i18next and PostHog. electron/ adds TanStack Router, Form, Store and Pacer, Base UI, Vidstack and dash.js for media playback, and electron-updater. It is built with electron-vite and electron-builder, with a packaging contract test (electron/tests/packaging-contract.mjs).

The Tauri shell in frontend/src-tauri is kept for reference. docs/electron-migration.md and electron/PARITY.md track the move to Electron.

Backend & APIs

The backend is FastAPI on uvicorn, with WebSockets for live events and scalar-fastapi for the API reference. Model work uses PyTorch (opens in a new tab), torchaudio and Hugging Face Transformers and Accelerate. The engines cover:

  • Speech generation: the default OmniVoice model, KittenTTS for small English voices, and on Apple Silicon the mlx-audio engines.
  • Transcription: WhisperX (with wav2vec2 alignment for dubbing timing), faster-whisper, mlx-whisper and parakeet-mlx.
  • Audio processing: pyannote-audio for diarization, Demucs for source separation, pedalboard for effects, yt-dlp for downloading media, and audioseal for watermarking.

Remote workers (backend/worker, deploy/install-worker, docs/remote-workers.md) let a desktop client use a GPU on another machine.

Data & persistence

Local state is managed with Alembic migrations (alembic.ini, backend/migrations). Tests cover migration safety, schema reconciliation and database backups. Model weights are cached under the Hugging Face home directory, which the Docker image sets to /app/omnivoice_data/huggingface. Jobs, voices, projects and generated audio are stored on the local disk.

Build, test & deploy

  • Python tests are in tests/: a large pytest suite covering ASR fallbacks, cloning, audiobook generation, authentication and CSRF, plus evals and smoke tests.
  • Frontend tests use Vitest and several Playwright configs: e2e, production bundle, performance and visual.
  • GitHub Actions: ci.yml, docker.yml, electron-build.yml, electron-release.yml, evals.yml, install-smoke.yml, security.yml, docs-drift.yml, build-omnivoice-tts.yml and cosyvoice-dependencies.yml.
  • deploy/Dockerfile builds the frontend with Bun, then installs the backend on a CUDA or ROCm PyTorch base with FFmpeg. It sets OMNIVOICE_SERVER_MODE=1 for headless servers.

Self-hosting notes

Desktop users download releases for macOS, Windows or Linux, or run scripts/install.sh / install.ps1. Server users run deploy/docker-compose.yml, choosing the CUDA or ROCm image. Developers run bun run dev:web, which starts uv sync, the API on 3900 and the Vite frontend together. The code is AGPL-3.0. pyproject.toml notes that a commercial license is available, and that the bundled OmniVoice model remains Apache-2.0 upstream.

What to copy (and what not to)

What to copy

  • Platform markers in pyproject.toml so each machine installs only the model backends it can run.
  • A build-time guard in the Dockerfile that fails if a dependency bump replaces the GPU-specific PyTorch.
  • Writing the reason and issue number next to every dependency cap.

What not to copy

  • Keeping two desktop shells (Tauri and Electron) in the tree during a long migration doubles the surface to test. Set a date to remove the old one.
  • Committing prebuilt binaries under bin/ inflates the repository. Release assets are a better home.

Sources & repo audit

Audited
Sep 25, 2026
Commit
47a06c0
License
AGPL-3.0

Independent analysis of repository at github.com/debpalash/VoiceStudio. Spotted an inaccuracy? Use the claim form to request a correction.

Maintainer? Add the architecture badge to your README
architecture: stackitfast
[![Architecture on STACK IT FAST](https://stackitfast.com/badge/voicestudio.svg)](https://stackitfast.com/project/voicestudio)
Use this stack

Scaffold it with your agent

Paste this prompt into Claude Code, Cursor, Windsurf or AGY to start a project with VoiceStudio's architecture.

  1. 1Copy the promptThe full markdown spec, with every layer and decision.
  2. 2Open your AI toolClaude Code, Cursor, Windsurf or Copilot, in a new repo.
  3. 3Paste and scaffoldUse it as the first instruction; review before you ship.
use-this-stack.md · 61 lines · 7.6 KB
# MISSION: Scaffold "VoiceStudio" Production Architecture
You are an expert Senior Staff Software Architect and Full-Stack Engineer. Your mission is to scaffold and implement a production-grade, highly reliable, and modular codebase following the proven architecture of **VoiceStudio**.
---
## 1. PROJECT SPECIFICATIONS & BENCHMARK
- **Reference Architecture**: VoiceStudio
- **What It Does**: VoiceStudio is an open-source, local-first alternative to ElevenLabs for voice cloning, voice design, video dubbing, dictation, transcription and audiobook creation. It runs a Python/PyTorch model backend behind a local API with a React desktop app.
- **Domain & Category**: Local AI Voice Cloning & Speech Studio
- **Production Scale**: 2-5 people
- **Development Mode**: HYBRID
- **Architectural Rationale**: Heavy model work stays in a Python FastAPI backend with a pluggable, hardware-aware engine registry. Desktop, web, MCP and remote-worker clients all use the same local API, so voice models run on the user's own GPU without a cloud dependency.
- **Live Website Reference**: https://voicestudio.sh
- **Source Repository**: https://github.com/debpalash/VoiceStudio
---
## 2. PRODUCTION TECH STACK
- **Full Stack Array**: Python, FastAPI, PyTorch, Transformers, WhisperX, MLX, Alembic, Model Context Protocol, TypeScript, React, Electron, Tauri, TanStack Query, Tailwind CSS, Radix UI, Vite, Bun, Turborepo, Docker, Playwright
- **Primary Language(s)**: Python, JavaScript, TypeScript, Rust, CSS
- **License of the reference repo**: AGPL-3.0
- **Frontend**: Electron — Current desktop shell (electron/) built with electron-vite and electron-builder; React — React 19 UIs for the Vite web frontend and the Electron renderer; TanStack Query — Server-state fetching against the local API, alongside TanStack Router, Form and Virtual; Tailwind CSS — Styling (Tailwind 4) with shadcn components on Radix UI and Base UI; Tauri — Earlier desktop shell (frontend/src-tauri, Rust) whose migration to Electron is documented in docs/
- **Backend & APIs**: Python — Backend service (backend/main.py) and the bundled omnivoice TTS package, managed with uv and hatchling; FastAPI — Local HTTP and WebSocket API on port 3900 with a Scalar-rendered OpenAPI reference; PyTorch — Runs TTS, cloning, diarization and separation models on CUDA, ROCm, Apple MPS or CPU
- **Data & persistence**: Alembic — Schema migrations for the local app database under backend/migrations
- **Infrastructure & deploy**: Bun — JavaScript package manager and script runner for the frontend/electron workspaces; Turborepo — Build orchestration across the JS workspaces; Docker — CUDA and ROCm server images (deploy/Dockerfile) with a Bun-built frontend stage; Playwright — E2E, production-bundle, performance and visual regression tests for the frontend
- **Tooling, testing & ops**: Transformers — Hugging Face model loading for the default OmniVoice engine and other TTS/ASR backends; WhisperX — Primary transcription with word-level forced alignment used for dubbing lip-sync; MLX — Apple Silicon engines (mlx-whisper, mlx-audio, parakeet-mlx) gated by platform markers; Model Context Protocol — backend/mcp_server.py exposes voice generation and transcription as tools for AI agents
---
## 3. KEY ARCHITECTURAL DECISIONS (audited from https://github.com/debpalash/VoiceStudio @ 47a06c0)
1. **Python model backend behind a local API, UI as a separate client**: backend/ (FastAPI + uvicorn with WebSockets) owns engines, jobs and storage, and the React UIs in frontend/ and electron/ talk to it over HTTP on localhost:3900; the same API powers the MCP server and remote workers.
2. **Pluggable engine registry with hardware-gated dependencies**: backend/engines hosts multiple TTS and ASR engines; pyproject.toml uses platform markers so MLX engines install only on Apple Silicon while WhisperX/faster-whisper, KittenTTS, Demucs and pyannote cover other platforms.
3. **Electron replaced Tauri as the desktop shell**: frontend/src-tauri (Rust) still exists, but electron/ is the default dev and build target in the root package.json, and a series of docs/electron-*.md files plus electron/PARITY.md track feature parity during the migration.
4. **One Docker image per GPU vendor with a torch-preservation guard**: deploy/Dockerfile builds from a pytorch CUDA base by default and a ROCm base via BASE_IMAGE, installs with uv against deploy/torch-constraints.txt, and fails the build if the GPU-specific torch was replaced by a generic wheel.
5. **Dependency pins documented inline with issue numbers**: pyproject.toml caps setuptools (<80), pyannote-audio (<4) and pedalboard (<0.9.21) and fastapi (<0.137) with comments naming the upstream breakage and the VoiceStudio issue, so pins can be revisited deliberately.
6. **Provenance watermarking of generated speech**: pyproject.toml depends on audioseal to embed an imperceptible neural watermark in AI-generated audio for later detection.
---
## 4. NON-NEGOTIABLE ARCHITECTURAL GUARDRAILS
1. **Monorepo & Modular Separation**:
   - Structure as a Turborepo monorepo with strict package boundaries:
     - `apps/web`: Application UI, routing, layouts, and server endpoints.
     - `packages/ui`: Shared design tokens, CSS variables, and Radix UI primitive components.
     - `packages/db`: Database schemas, client singleton, declarative migrations, and seed scripts.
     - `packages/config`: Shared TypeScript, ESLint, and build configurations.
2. **Strict Type Safety & Zero `any` Policy**:
   - Enable `strict: true`, `noImplicitAny: true`, and `strictNullChecks: true`.
   - Validate ALL external inputs, API request bodies, and query parameters with **Zod** schemas before execution.
3. **Frontend & Rendering Guidelines**:
   - Utilize Vite 6 with TanStack Router for fully type-safe routing. Manage server state and caching via TanStack Query v5 with optimistic UI updates.
4. **Design System & Aesthetics**:
   - Keep every color, radius, shadow and font in a single token file (CSS variables) and consume tokens everywhere; never hardcode hex values in components.
   - Prefer crisp 1px borders and one subtle shadow scale over blurry default shadows. Pair one sans-serif for body/headings with one monospace for tags, badges, metrics, and code.
5. **Data Layer & Reliability**:
   - Write declarative schema definitions with foreign keys, composite indexes on queried filters, and automated timestamp triggers.
   - Use connection pooling and prepared statements for serverless database execution.
---
## 5. STEP-BY-STEP SCAFFOLDING ROADMAP
- **Phase 1: Workspace & Root Config**: Initialize package manager, monorepo configuration (`turbo.json`, `tsconfig.base.json`, `package.json`).
- **Phase 2: Database Schema & Client**: Set up the data layer (Alembic): client, connection pool, models, and migration scripts.
- **Phase 3: Design Tokens & UI Primitives**: Build accessible `Button`, `Input`, `Card`, `Badge`, and layout wrappers inside `packages/ui`.
- **Phase 4: Core Application Routes & Handlers**: Implement primary authentication, user session handling, and application routes.
- **Phase 5: Quality Assurance & Build Verification**: Run `tsc --noEmit`, ESLint, Prettier, and smoke test suites to ensure zero compilation or runtime errors.
---
## 6. EXECUTION INSTRUCTIONS
1. Review all specifications, architectural guardrails, and stack choices above.
2. Present the full monorepo directory tree structure.
3. Systematically generate the complete, production-ready codebase according to the 5-phase roadmap above — starting with the root workspace setup, followed by the database schema, UI design system package, and full-stack application routes until the repository is fully scaffolded and ready to run.
Scaffolded something with this prompt?
Would you pick this stack for a local ai voice cloning & speech studio project?

Frequently asked about VoiceStudio

What is VoiceStudio built with?

VoiceStudio has a Python backend built on FastAPI and PyTorch, with Transformers, WhisperX and MLX-based engines, and a React 19 frontend styled with Tailwind CSS. The desktop app ships as an Electron shell, and Bun and Turborepo manage the JavaScript workspaces.

Does VoiceStudio run fully offline?

Local workflows run on the user's own hardware with models downloaded from Hugging Face. Remote GPU workers and remote services are optional, and according to the README usage analytics require consent.

Which GPUs does VoiceStudio support?

The backend uses PyTorch, so it runs on NVIDIA CUDA, AMD ROCm (a separate Docker image variant), Apple Silicon (with extra MLX engines) or CPU. Hardware-specific notes live in docs/.

Can AI agents use VoiceStudio?

Yes. The backend exposes an MCP server (backend/mcp_server.py) and a local HTTP API, so agents can generate speech, clone voices or transcribe audio as tools.

One email a month: new deep dives and stack trends

New source-audited architectures, head-to-head comparisons and the monthly stack report. No spam, unsubscribe anytime.

use-this-stack.md