VoiceStudio
Audited from github.com/debpalash/VoiceStudio
VoiceStudio is an open-source, local-first alternative to ElevenLabs for voice cloning, voice design, video dubbing, dictation, transcription and audiobook creation. It runs a Python/PyTorch model backend behind a local API with a React desktop app.
- Language
- Python
- Hosting
- Docker
- License
- AGPL-3.0
- Running for
- 5 months
- Team
- 2-5 people
Why this architecture
Heavy model work stays in a Python FastAPI backend with a pluggable, hardware-aware engine registry. Desktop, web, MCP and remote-worker clients all use the same local API, so voice models run on the user's own GPU without a cloud dependency.
Tech stack by layer
20 technologies · audited Sep 25, 2026- ElectronCurrent desktop shell (electron/) built with electron-vite and electron-builder
- ReactReact 19 UIs for the Vite web frontend and the Electron renderer
- TanStack QueryServer-state fetching against the local API, alongside TanStack Router, Form and Virtual
- Tailwind CSSStyling (Tailwind 4) with shadcn components on Radix UI and Base UI
- TauriEarlier desktop shell (frontend/src-tauri, Rust) whose migration to Electron is documented in docs/
- AlembicSchema migrations for the local app database under backend/migrations
- BunJavaScript package manager and script runner for the frontend/electron workspaces
- TurborepoBuild orchestration across the JS workspaces
- DockerCUDA and ROCm server images (deploy/Dockerfile) with a Bun-built frontend stage
- PlaywrightE2E, production-bundle, performance and visual regression tests for the frontend
- TransformersHugging Face model loading for the default OmniVoice engine and other TTS/ASR backends
- WhisperXPrimary transcription with word-level forced alignment used for dubbing lip-sync
- MLXApple Silicon engines (mlx-whisper, mlx-audio, parakeet-mlx) gated by platform markers
- Model Context Protocolbackend/mcp_server.py exposes voice generation and transcription as tools for AI agents
VoiceStudio architecture diagram
Open SVGDiagram as text
- Desktop app (Electron · React) → Local API (FastAPI · port 3900): HTTP · WebSocket
- Web UI (Vite · TanStack) → Local API (FastAPI · port 3900): HTTP · WebSocket
- AI agents (MCP) → Local API (FastAPI · port 3900): MCP
- Local API (FastAPI · port 3900) → Engine registry (TTS · cloning · ASR): jobs
- Engine registry (TTS · cloning · ASR) → Model runtime (PyTorch · MLX): inference
- Local API (FastAPI · port 3900) → App database (Alembic migrations): app state
- Engine registry (TTS · cloning · ASR) → Hugging Face (model weights): load models
Key architectural decisions
6 decisions- 01
Python model backend behind a local API, UI as a separate client
backend/ (FastAPI + uvicorn with WebSockets) owns engines, jobs and storage, and the React UIs in frontend/ and electron/ talk to it over HTTP on localhost:3900; the same API powers the MCP server and remote workers.
- 02
Pluggable engine registry with hardware-gated dependencies
backend/engines hosts multiple TTS and ASR engines; pyproject.toml uses platform markers so MLX engines install only on Apple Silicon while WhisperX/faster-whisper, KittenTTS, Demucs and pyannote cover other platforms.
- 03
Electron replaced Tauri as the desktop shell
frontend/src-tauri (Rust) still exists, but electron/ is the default dev and build target in the root package.json, and a series of docs/electron-*.md files plus electron/PARITY.md track feature parity during the migration.
- 04
One Docker image per GPU vendor with a torch-preservation guard
deploy/Dockerfile builds from a pytorch CUDA base by default and a ROCm base via BASE_IMAGE, installs with uv against deploy/torch-constraints.txt, and fails the build if the GPU-specific torch was replaced by a generic wheel.
- 05
Dependency pins documented inline with issue numbers
pyproject.toml caps setuptools (<80), pyannote-audio (<4) and pedalboard (<0.9.21) and fastapi (<0.137) with comments naming the upstream breakage and the VoiceStudio issue, so pins can be revisited deliberately.
- 06
Provenance watermarking of generated speech
pyproject.toml depends on audioseal to embed an imperceptible neural watermark in AI-generated audio for later detection.
How VoiceStudio is built
How VoiceStudio is structured
VoiceStudio is a Python + JavaScript monorepo. docs/STRUCTURE.md describes the layout.
- Python:
pyproject.toml(packageomnivoice, built with hatchling, managed with uv) covers:backend/: the FastAPI app (backend/main.py) withapi/,services/,engines/,worker/,plugins/,hooks/,schemas/,migrations/and the MCP server (backend/mcp_server.py,backend/mcp_shim),omnivoice/: the bundled TTS model code (models, training, eval, CLI).
- JavaScript: a Bun workspace (
package.json,packageManager: bun@1.4.2) withfrontend/(Vite + React, plus the oldersrc-taurishell) andelectron/(the current desktop app).turbo.jsonorchestrates builds. - Operations:
deploy/(Dockerfile, docker-compose, remote worker installer),bin/(prebuiltomnivoice-ttsbinaries per OS),scripts/(install, release, smoke and benchmark scripts) andnotebooks/(a Colab notebook).
Architecture decision records live in docs/adr. AGENTS.md, CLAUDE.md and .claude/agents give coding agents instructions.
Frontend
Both UIs use React 19. frontend/ is a Vite app with Radix UI primitives, Tailwind CSS 4, TanStack Query, Table and Virtual, i18next and PostHog. electron/ adds TanStack Router, Form, Store and Pacer, Base UI, Vidstack and dash.js for media playback, and electron-updater. It is built with electron-vite and electron-builder, with a packaging contract test (electron/tests/packaging-contract.mjs).
The Tauri shell in frontend/src-tauri is kept for reference. docs/electron-migration.md and electron/PARITY.md track the move to Electron.
Backend & APIs
The backend is FastAPI on uvicorn, with WebSockets for live events and scalar-fastapi for the API reference. Model work uses PyTorch (opens in a new tab), torchaudio and Hugging Face Transformers and Accelerate. The engines cover:
- Speech generation: the default OmniVoice model, KittenTTS for small English voices, and on Apple Silicon the
mlx-audioengines. - Transcription: WhisperX (with wav2vec2 alignment for dubbing timing), faster-whisper,
mlx-whisperandparakeet-mlx. - Audio processing: pyannote-audio for diarization, Demucs for source separation, pedalboard for effects, yt-dlp for downloading media, and audioseal for watermarking.
Remote workers (backend/worker, deploy/install-worker, docs/remote-workers.md) let a desktop client use a GPU on another machine.
Data & persistence
Local state is managed with Alembic migrations (alembic.ini, backend/migrations). Tests cover migration safety, schema reconciliation and database backups. Model weights are cached under the Hugging Face home directory, which the Docker image sets to /app/omnivoice_data/huggingface. Jobs, voices, projects and generated audio are stored on the local disk.
Build, test & deploy
- Python tests are in
tests/: a large pytest suite covering ASR fallbacks, cloning, audiobook generation, authentication and CSRF, plus evals and smoke tests. - Frontend tests use Vitest and several Playwright configs: e2e, production bundle, performance and visual.
- GitHub Actions:
ci.yml,docker.yml,electron-build.yml,electron-release.yml,evals.yml,install-smoke.yml,security.yml,docs-drift.yml,build-omnivoice-tts.ymlandcosyvoice-dependencies.yml. deploy/Dockerfilebuilds the frontend with Bun, then installs the backend on a CUDA or ROCm PyTorch base with FFmpeg. It setsOMNIVOICE_SERVER_MODE=1for headless servers.
Self-hosting notes
Desktop users download releases for macOS, Windows or Linux, or run scripts/install.sh / install.ps1. Server users run deploy/docker-compose.yml, choosing the CUDA or ROCm image. Developers run bun run dev:web, which starts uv sync, the API on 3900 and the Vite frontend together. The code is AGPL-3.0. pyproject.toml notes that a commercial license is available, and that the bundled OmniVoice model remains Apache-2.0 upstream.
What to copy (and what not to)
What to copy
- Platform markers in
pyproject.tomlso each machine installs only the model backends it can run. - A build-time guard in the Dockerfile that fails if a dependency bump replaces the GPU-specific PyTorch.
- Writing the reason and issue number next to every dependency cap.
What not to copy
- Keeping two desktop shells (Tauri and Electron) in the tree during a long migration doubles the surface to test. Set a date to remove the old one.
- Committing prebuilt binaries under
bin/inflates the repository. Release assets are a better home.
Sources & repo audit
- pyproject.toml (Python dependencies and pins)
- deploy/Dockerfile (CUDA/ROCm server image)
- electron/package.json (desktop shell)
- docs/STRUCTURE.md
Independent analysis of repository at github.com/debpalash/VoiceStudio. Spotted an inaccuracy? Use the claim form to request a correction.
Maintainer? Add the architecture badge to your README
[](https://stackitfast.com/project/voicestudio) Scaffold it with your agent
Paste this prompt into Claude Code, Cursor, Windsurf or AGY to start a project with VoiceStudio's architecture.
- 1Copy the promptThe full markdown spec, with every layer and decision.
- 2Open your AI toolClaude Code, Cursor, Windsurf or Copilot, in a new repo.
- 3Paste and scaffoldUse it as the first instruction; review before you ship.
# MISSION: Scaffold "VoiceStudio" Production Architecture
You are an expert Senior Staff Software Architect and Full-Stack Engineer. Your mission is to scaffold and implement a production-grade, highly reliable, and modular codebase following the proven architecture of **VoiceStudio**.
---
## 1. PROJECT SPECIFICATIONS & BENCHMARK
- **Reference Architecture**: VoiceStudio
- **What It Does**: VoiceStudio is an open-source, local-first alternative to ElevenLabs for voice cloning, voice design, video dubbing, dictation, transcription and audiobook creation. It runs a Python/PyTorch model backend behind a local API with a React desktop app.
- **Domain & Category**: Local AI Voice Cloning & Speech Studio
- **Production Scale**: 2-5 people
- **Development Mode**: HYBRID
- **Architectural Rationale**: Heavy model work stays in a Python FastAPI backend with a pluggable, hardware-aware engine registry. Desktop, web, MCP and remote-worker clients all use the same local API, so voice models run on the user's own GPU without a cloud dependency.
- **Live Website Reference**: https://voicestudio.sh
- **Source Repository**: https://github.com/debpalash/VoiceStudio
---
## 2. PRODUCTION TECH STACK
- **Full Stack Array**: Python, FastAPI, PyTorch, Transformers, WhisperX, MLX, Alembic, Model Context Protocol, TypeScript, React, Electron, Tauri, TanStack Query, Tailwind CSS, Radix UI, Vite, Bun, Turborepo, Docker, Playwright
- **Primary Language(s)**: Python, JavaScript, TypeScript, Rust, CSS
- **License of the reference repo**: AGPL-3.0
- **Frontend**: Electron — Current desktop shell (electron/) built with electron-vite and electron-builder; React — React 19 UIs for the Vite web frontend and the Electron renderer; TanStack Query — Server-state fetching against the local API, alongside TanStack Router, Form and Virtual; Tailwind CSS — Styling (Tailwind 4) with shadcn components on Radix UI and Base UI; Tauri — Earlier desktop shell (frontend/src-tauri, Rust) whose migration to Electron is documented in docs/
- **Backend & APIs**: Python — Backend service (backend/main.py) and the bundled omnivoice TTS package, managed with uv and hatchling; FastAPI — Local HTTP and WebSocket API on port 3900 with a Scalar-rendered OpenAPI reference; PyTorch — Runs TTS, cloning, diarization and separation models on CUDA, ROCm, Apple MPS or CPU
- **Data & persistence**: Alembic — Schema migrations for the local app database under backend/migrations
- **Infrastructure & deploy**: Bun — JavaScript package manager and script runner for the frontend/electron workspaces; Turborepo — Build orchestration across the JS workspaces; Docker — CUDA and ROCm server images (deploy/Dockerfile) with a Bun-built frontend stage; Playwright — E2E, production-bundle, performance and visual regression tests for the frontend
- **Tooling, testing & ops**: Transformers — Hugging Face model loading for the default OmniVoice engine and other TTS/ASR backends; WhisperX — Primary transcription with word-level forced alignment used for dubbing lip-sync; MLX — Apple Silicon engines (mlx-whisper, mlx-audio, parakeet-mlx) gated by platform markers; Model Context Protocol — backend/mcp_server.py exposes voice generation and transcription as tools for AI agents
---
## 3. KEY ARCHITECTURAL DECISIONS (audited from https://github.com/debpalash/VoiceStudio @ 47a06c0)
1. **Python model backend behind a local API, UI as a separate client**: backend/ (FastAPI + uvicorn with WebSockets) owns engines, jobs and storage, and the React UIs in frontend/ and electron/ talk to it over HTTP on localhost:3900; the same API powers the MCP server and remote workers.
2. **Pluggable engine registry with hardware-gated dependencies**: backend/engines hosts multiple TTS and ASR engines; pyproject.toml uses platform markers so MLX engines install only on Apple Silicon while WhisperX/faster-whisper, KittenTTS, Demucs and pyannote cover other platforms.
3. **Electron replaced Tauri as the desktop shell**: frontend/src-tauri (Rust) still exists, but electron/ is the default dev and build target in the root package.json, and a series of docs/electron-*.md files plus electron/PARITY.md track feature parity during the migration.
4. **One Docker image per GPU vendor with a torch-preservation guard**: deploy/Dockerfile builds from a pytorch CUDA base by default and a ROCm base via BASE_IMAGE, installs with uv against deploy/torch-constraints.txt, and fails the build if the GPU-specific torch was replaced by a generic wheel.
5. **Dependency pins documented inline with issue numbers**: pyproject.toml caps setuptools (<80), pyannote-audio (<4) and pedalboard (<0.9.21) and fastapi (<0.137) with comments naming the upstream breakage and the VoiceStudio issue, so pins can be revisited deliberately.
6. **Provenance watermarking of generated speech**: pyproject.toml depends on audioseal to embed an imperceptible neural watermark in AI-generated audio for later detection.
---
## 4. NON-NEGOTIABLE ARCHITECTURAL GUARDRAILS
1. **Monorepo & Modular Separation**:
- Structure as a Turborepo monorepo with strict package boundaries:
- `apps/web`: Application UI, routing, layouts, and server endpoints.
- `packages/ui`: Shared design tokens, CSS variables, and Radix UI primitive components.
- `packages/db`: Database schemas, client singleton, declarative migrations, and seed scripts.
- `packages/config`: Shared TypeScript, ESLint, and build configurations.
2. **Strict Type Safety & Zero `any` Policy**:
- Enable `strict: true`, `noImplicitAny: true`, and `strictNullChecks: true`.
- Validate ALL external inputs, API request bodies, and query parameters with **Zod** schemas before execution.
3. **Frontend & Rendering Guidelines**:
- Utilize Vite 6 with TanStack Router for fully type-safe routing. Manage server state and caching via TanStack Query v5 with optimistic UI updates.
4. **Design System & Aesthetics**:
- Keep every color, radius, shadow and font in a single token file (CSS variables) and consume tokens everywhere; never hardcode hex values in components.
- Prefer crisp 1px borders and one subtle shadow scale over blurry default shadows. Pair one sans-serif for body/headings with one monospace for tags, badges, metrics, and code.
5. **Data Layer & Reliability**:
- Write declarative schema definitions with foreign keys, composite indexes on queried filters, and automated timestamp triggers.
- Use connection pooling and prepared statements for serverless database execution.
---
## 5. STEP-BY-STEP SCAFFOLDING ROADMAP
- **Phase 1: Workspace & Root Config**: Initialize package manager, monorepo configuration (`turbo.json`, `tsconfig.base.json`, `package.json`).
- **Phase 2: Database Schema & Client**: Set up the data layer (Alembic): client, connection pool, models, and migration scripts.
- **Phase 3: Design Tokens & UI Primitives**: Build accessible `Button`, `Input`, `Card`, `Badge`, and layout wrappers inside `packages/ui`.
- **Phase 4: Core Application Routes & Handlers**: Implement primary authentication, user session handling, and application routes.
- **Phase 5: Quality Assurance & Build Verification**: Run `tsc --noEmit`, ESLint, Prettier, and smoke test suites to ensure zero compilation or runtime errors.
---
## 6. EXECUTION INSTRUCTIONS
1. Review all specifications, architectural guardrails, and stack choices above.
2. Present the full monorepo directory tree structure.
3. Systematically generate the complete, production-ready codebase according to the 5-phase roadmap above — starting with the root workspace setup, followed by the database schema, UI design system package, and full-stack application routes until the repository is fully scaffolded and ready to run.Frequently asked about VoiceStudio
What is VoiceStudio built with?
VoiceStudio has a Python backend built on FastAPI and PyTorch, with Transformers, WhisperX and MLX-based engines, and a React 19 frontend styled with Tailwind CSS. The desktop app ships as an Electron shell, and Bun and Turborepo manage the JavaScript workspaces.
Does VoiceStudio run fully offline?
Local workflows run on the user's own hardware with models downloaded from Hugging Face. Remote GPU workers and remote services are optional, and according to the README usage analytics require consent.
Which GPUs does VoiceStudio support?
The backend uses PyTorch, so it runs on NVIDIA CUDA, AMD ROCm (a separate Docker image variant), Apple Silicon (with extra MLX engines) or CPU. Hardware-specific notes live in docs/.
Can AI agents use VoiceStudio?
Yes. The backend exposes an MCP server (backend/mcp_server.py) and a local HTTP API, so agents can generate speech, clone voices or transcribe audio as tools.
One email a month: new deep dives and stack trends
New source-audited architectures, head-to-head comparisons and the monthly stack report. No spam, unsubscribe anytime.
Similar architectures
- MindsHub (formerly MindsDB)AI-agentDesktop & Web AI Agent Workspace · 20+ PeopleShares Python · TypeScript · Electron
- MagnitudeHybridLocal LLM Inference Engine & Hardware Profiler · 2-5 peopleShares TypeScript · Bun · Electron
- OpenMAICClassicMulti-Agent Interactive AI Classroom · 6-20 peopleShares TypeScript · React · Tailwind CSS
- CopilotKitAI-agentAI Copilot & Assistant SDK · 2-5 peopleShares TypeScript · React · Python
- LibreChatHybridMulti-Model AI Agent & Chat Workspace · 6-20 PeopleShares TypeScript · React · Vite
- LangfuseAI-agentLLM Observability · 6-20 peopleShares TypeScript · React · Turborepo