Magnitude
Audited from github.com/magnitudedev/magnitude
Magnitude is an open-source local inference engine and desktop app. It profiles the machine, recommends and tunes open models for that hardware, runs them through a Rust llama.cpp-based server, and connects them to coding agents in one click.
- Language
- TypeScript
- License
- Apache-2.0
- Running for
- 3 months
- Team
- 2-5 people
Why this architecture
Performance-critical inference sits in a Rust workspace around vendored llama.cpp, while the app, CLI, agent and model management live in a Bun + Effect TypeScript monorepo. The native engine can be pinned and benchmarked on its own while the product layer changes quickly.
Tech stack by layer
12 technologies · audited Sep 25, 2026- ReactReact 19 UI in web/ (Base UI, shadcn, Monaco, cmdk), shared by the desktop app
- ElectronCross-platform desktop shell built with electron-vite that bundles the CLI and daemon
- ViteBundler for the web UI and, through electron-vite, the desktop renderer
- Tailwind CSSStyling for the web and desktop UIs (Tailwind 4 via @tailwindcss/vite)
- RustInference workspace (inference/crates/icn-*): engine, hardware profiling, catalog, server
- llama.cppModel execution through llama-cpp-2 bindings vendored as a pinned native submodule
- AxumHTTP server for the local inference API in the icn-server crate
- BunRuntime and package manager for the TypeScript workspaces, CLI and scripts
- EffectTyped effect system used across CLI, SDK, agent and ICN packages, including @effect/rpc
- TurborepoOrchestrates build and typecheck across the Bun workspaces
- OpenTelemetryOTLP traces and logs from the Rust inference server
- GitHub ActionsRelease, desktop-native, integration and update-acceptance workflows per platform
- Harness connectionspackages/harness-connections wires local models into agents such as Pi, OpenCode and Hermes
- Hugging Face HubModel downloads via the hf-hub crate for the catalog and model management
Magnitude architecture diagram
Open SVGDiagram as text
- Desktop app (Electron · React) → Daemon & SDK (TypeScript · @effect/rpc): RPC
- magnitude CLI (Bun · Effect) → Daemon & SDK (TypeScript · @effect/rpc): RPC
- Agent harnesses (Pi · OpenCode · Hermes) → Inference server (Rust · Axum): local models
- Daemon & SDK (TypeScript · @effect/rpc) → Inference server (Rust · Axum): HTTP
- Inference server (Rust · Axum) → Engine (llama.cpp · hardware profiling): generate
- Engine (llama.cpp · hardware profiling) → Model catalog (local weights): load
- Inference server (Rust · Axum) → Hugging Face Hub (model downloads): download
Key architectural decisions
5 decisions- 01
TypeScript product layer on top of a separate Rust inference workspace
The root Bun workspace (cli, desktop, web, packages/*) holds the app, agent and client logic, while inference/ is its own Cargo workspace (icn-engine, icn-hardware, icn-speculative, icn-server, …) exposed to TypeScript through packages/icn and packages/icn-protocol.
- 02
llama.cpp vendored as a pinned native dependency
inference/Cargo.toml depends on llama-cpp-2 from native/llama-cpp-rs (a git submodule, see .gitmodules) and inference/native-pin.toml plus verify:native-pin and verify:safe-bindings scripts guard the exact native revision and the unsafe boundary.
- 03
Hardware profiling before model choice
The icn-hardware crate (with an icn-model-assessment binary) and inference/catalog back the app's claim of estimating tokens per second for every catalog model before download; parity and benchmark crates check results against references.
- 04
Effect as the application backbone
CLI, desktop, web, agent and ICN packages all depend on effect 3 with @effect/platform and @effect/rpc, and patches/ carries a local fix for @effect-atom/atom, so services, errors and RPC share one typed runtime model.
- 05
One desktop app ships the CLI and daemon
desktop/ (Electron, electron-vite) runs build:native from packages/daemon-management before bundling, so installing the app also installs the magnitude CLI and the local daemon agents connect to.
How Magnitude is built
How Magnitude is structured
Magnitude is a Bun monorepo (package.json sets packageManager: bun@1.4.2, and turbo.json drives the builds) that contains a separate Rust workspace. Top-level workspaces:
cli/: themagnitudecommand line, built with Commander and Effect.desktop/: the Electron app (electron.vite.config.ts).web/: the React UI, reused by the desktop app as@magnitudedev/web.inference/: a Cargo workspace (inference/Cargo.toml) with the ICN crates, plus a smallpackage.jsonso Bun scripts can drive cargo.integrations/pi: an integration for the Pi agent.packages/*: about 35 internal packages, includingagent,ai,providers,icn,icn-protocol,acn,event-core,storage,tracing,sdk,release,daemon-managementandharness-connections.
Design notes live in design/ (subfolders for inference, harness, storage, release, …) and info/. User documentation is a Mintlify site under docs/ (docs/docs.json). AGENTS.md files at the root and in cli/, design/ and inference/ give coding agents instructions per area.
Frontend
web/ uses React 19 with @base-ui/react, shadcn, cmdk for the command palette, @monaco-editor/react for code views, react-arborist for file trees and react-markdown + shiki for rendered output. State bindings come from @effect-atom/atom-react. Styling is Tailwind CSS 4 through @tailwindcss/vite, and the build uses Vite 6.
desktop/ wraps that UI in Electron via electron-vite, adds @effect/platform-node for the main process, and uses Playwright as a dev dependency for tests.
Backend & APIs
The inference engine is Rust (edition = "2024") split into focused crates: icn-engine, icn-hardware, icn-models, icn-catalog, icn-speculative, icn-reasoning, icn-api, icn-contracts, icn-server, icn-parity and benchmark-runner. Model execution goes through llama-cpp-2 from native/llama-cpp-rs. The HTTP surface is Axum with hyper clients, and model files are fetched with hf-hub. The crate list names speculative decoding and reasoning handling as first-class parts of the design.
On the TypeScript side, packages/icn wraps the engine lifecycle, hardware, catalog and instances. packages/agent implements the agent loop and depends on ripgrep, shell-classifier, vcs, skills and scratchpad packages. Services communicate through Effect (opens in a new tab) and @effect/rpc.
Data & persistence
No server database is visible. Local state is managed by packages/storage. Logs and events are written under ~/.magnitude/logs (see the events and logs scripts in the root package.json). Models are downloaded into the local model store managed by the ICN catalog and model-management code.
Build, test & deploy
bun run buildgenerates the version and runsturbo build.bun run typecheckrunstscacross workspaces.- Rust:
cargo test --workspace, plusparityandbenchmarkcommands defined ininference/package.json.devstartsicn-server serve --fakefor UI work without a GPU. - Changesets (
.changeset/) manage versioning. - GitHub Actions has dedicated workflows for releases (
release.yml,release-build.yml,release-inference-macos.yml,release-candidate-dry-run.yml),desktop-native.yml,integrations.yml,hosted-update-acceptance.ymlandwindows-update-consumer.yml. - Languages in the repo include PowerShell, NSIS, AppleScript and Swift, used for per-OS installers and platform glue.
Self-hosting notes
Users install the desktop app for macOS, Windows or Linux. It includes the CLI. The README lists Apple Silicon, NVIDIA, AMD and CPU-only machines as supported. Contributors run bun run bootstrap and bun run dev. The inference crates need a Rust toolchain (inference/rust-toolchain.toml) and the native submodule.
What to copy (and what not to)
What to copy
- Isolating the native engine in its own Cargo workspace with a pinned upstream revision and explicit safe-binding checks.
- Parity and benchmark harnesses (
icn-parity,benchmark-runner) that compare engine output against reference builds before release. - Per-area
AGENTS.mdfiles that give coding agents scoped instructions in a large monorepo.
What not to copy
- Around 35 internal packages is a lot of structure for a young product. The setup takes more to understand than a flat layout with fewer, larger packages.
- Patching third-party packages under
patches/works, but every patch is a maintenance cost on upgrades.
Sources & repo audit
- Audited
- Sep 25, 2026
- Commit
- fa33238
- License
- Apache-2.0
- Root package.json (Bun workspaces and scripts)
- inference/Cargo.toml (ICN crates, llama-cpp-2, axum, OpenTelemetry)
- desktop/package.json (Electron, electron-vite)
- web/package.json (React 19, Base UI, Monaco, Tailwind 4)
Independent analysis of repository at github.com/magnitudedev/magnitude. Spotted an inaccuracy? Use the claim form to request a correction.
Maintainer? Add the architecture badge to your README
[](https://stackitfast.com/project/magnitude) Scaffold it with your agent
Paste this prompt into Claude Code, Cursor, Windsurf or AGY to start a project with Magnitude's architecture.
- 1Copy the promptThe full markdown spec, with every layer and decision.
- 2Open your AI toolClaude Code, Cursor, Windsurf or Copilot, in a new repo.
- 3Paste and scaffoldUse it as the first instruction; review before you ship.
# MISSION: Scaffold "Magnitude" Production Architecture
You are an expert Senior Staff Software Architect and Full-Stack Engineer. Your mission is to scaffold and implement a production-grade, highly reliable, and modular codebase following the proven architecture of **Magnitude**.
---
## 1. PROJECT SPECIFICATIONS & BENCHMARK
- **Reference Architecture**: Magnitude
- **What It Does**: Magnitude is an open-source local inference engine and desktop app. It profiles the machine, recommends and tunes open models for that hardware, runs them through a Rust llama.cpp-based server, and connects them to coding agents in one click.
- **Domain & Category**: Local LLM Inference Engine & Hardware Profiler
- **Production Scale**: 2-5 people
- **Development Mode**: HYBRID
- **Architectural Rationale**: Performance-critical inference sits in a Rust workspace around vendored llama.cpp, while the app, CLI, agent and model management live in a Bun + Effect TypeScript monorepo. The native engine can be pinned and benchmarked on its own while the product layer changes quickly.
- **Live Website Reference**: https://magnitude.dev
- **Source Repository**: https://github.com/magnitudedev/magnitude
---
## 2. PRODUCTION TECH STACK
- **Full Stack Array**: TypeScript, Rust, Bun, Effect, llama.cpp, Axum, Electron, React, Vite, Tailwind CSS, Turborepo, OpenTelemetry
- **Primary Language(s)**: TypeScript, Rust, C, Svelte, Python
- **License of the reference repo**: Apache-2.0
- **Frontend**: React — React 19 UI in web/ (Base UI, shadcn, Monaco, cmdk), shared by the desktop app; Electron — Cross-platform desktop shell built with electron-vite that bundles the CLI and daemon; Vite — Bundler for the web UI and, through electron-vite, the desktop renderer; Tailwind CSS — Styling for the web and desktop UIs (Tailwind 4 via @tailwindcss/vite)
- **Backend & APIs**: Rust — Inference workspace (inference/crates/icn-*): engine, hardware profiling, catalog, server; llama.cpp — Model execution through llama-cpp-2 bindings vendored as a pinned native submodule; Axum — HTTP server for the local inference API in the icn-server crate; Bun — Runtime and package manager for the TypeScript workspaces, CLI and scripts; Effect — Typed effect system used across CLI, SDK, agent and ICN packages, including @effect/rpc
- **Infrastructure & deploy**: Turborepo — Orchestrates build and typecheck across the Bun workspaces; OpenTelemetry — OTLP traces and logs from the Rust inference server; GitHub Actions — Release, desktop-native, integration and update-acceptance workflows per platform
- **Tooling, testing & ops**: Harness connections — packages/harness-connections wires local models into agents such as Pi, OpenCode and Hermes; Hugging Face Hub — Model downloads via the hf-hub crate for the catalog and model management
---
## 3. KEY ARCHITECTURAL DECISIONS (audited from https://github.com/magnitudedev/magnitude @ fa33238)
1. **TypeScript product layer on top of a separate Rust inference workspace**: The root Bun workspace (cli, desktop, web, packages/*) holds the app, agent and client logic, while inference/ is its own Cargo workspace (icn-engine, icn-hardware, icn-speculative, icn-server, …) exposed to TypeScript through packages/icn and packages/icn-protocol.
2. **llama.cpp vendored as a pinned native dependency**: inference/Cargo.toml depends on llama-cpp-2 from native/llama-cpp-rs (a git submodule, see .gitmodules) and inference/native-pin.toml plus verify:native-pin and verify:safe-bindings scripts guard the exact native revision and the unsafe boundary.
3. **Hardware profiling before model choice**: The icn-hardware crate (with an icn-model-assessment binary) and inference/catalog back the app's claim of estimating tokens per second for every catalog model before download; parity and benchmark crates check results against references.
4. **Effect as the application backbone**: CLI, desktop, web, agent and ICN packages all depend on effect 3 with @effect/platform and @effect/rpc, and patches/ carries a local fix for @effect-atom/atom, so services, errors and RPC share one typed runtime model.
5. **One desktop app ships the CLI and daemon**: desktop/ (Electron, electron-vite) runs build:native from packages/daemon-management before bundling, so installing the app also installs the magnitude CLI and the local daemon agents connect to.
---
## 4. NON-NEGOTIABLE ARCHITECTURAL GUARDRAILS
1. **Monorepo & Modular Separation**:
- Structure as a Turborepo monorepo with strict package boundaries:
- `apps/web`: Application UI, routing, layouts, and server endpoints.
- `packages/ui`: Shared design tokens, CSS variables, and Radix UI primitive components.
- `packages/db`: Database schemas, client singleton, declarative migrations, and seed scripts.
- `packages/config`: Shared TypeScript, ESLint, and build configurations.
2. **Strict Type Safety & Zero `any` Policy**:
- Enable `strict: true`, `noImplicitAny: true`, and `strictNullChecks: true`.
- Validate ALL external inputs, API request bodies, and query parameters with **Zod** schemas before execution.
3. **Frontend & Rendering Guidelines**:
- Utilize Vite 6 with TanStack Router for fully type-safe routing. Manage server state and caching via TanStack Query v5 with optimistic UI updates.
4. **Design System & Aesthetics**:
- Keep every color, radius, shadow and font in a single token file (CSS variables) and consume tokens everywhere; never hardcode hex values in components.
- Prefer crisp 1px borders and one subtle shadow scale over blurry default shadows. Pair one sans-serif for body/headings with one monospace for tags, badges, metrics, and code.
5. **Data Layer & Reliability**:
- Write declarative schema definitions with foreign keys, composite indexes on queried filters, and automated timestamp triggers.
- Use connection pooling and prepared statements for serverless database execution.
---
## 5. STEP-BY-STEP SCAFFOLDING ROADMAP
- **Phase 1: Workspace & Root Config**: Initialize package manager, monorepo configuration (`turbo.json`, `tsconfig.base.json`, `package.json`).
- **Phase 2: Database Schema & Client**: Set up the data layer: client, connection pool, models, and migration scripts.
- **Phase 3: Design Tokens & UI Primitives**: Build accessible `Button`, `Input`, `Card`, `Badge`, and layout wrappers inside `packages/ui`.
- **Phase 4: Core Application Routes & Handlers**: Implement primary authentication, user session handling, and application routes.
- **Phase 5: Quality Assurance & Build Verification**: Run `tsc --noEmit`, ESLint, Prettier, and smoke test suites to ensure zero compilation or runtime errors.
---
## 6. EXECUTION INSTRUCTIONS
1. Review all specifications, architectural guardrails, and stack choices above.
2. Present the full monorepo directory tree structure.
3. Systematically generate the complete, production-ready codebase according to the 5-phase roadmap above — starting with the root workspace setup, followed by the database schema, UI design system package, and full-stack application routes until the repository is fully scaffolded and ready to run.Frequently asked about Magnitude
What is Magnitude built with?
Magnitude pairs a Rust inference workspace, built on llama.cpp bindings and an Axum server, with a TypeScript product layer on Bun and Effect. The desktop app is Electron with a React 19 and Tailwind UI, and Turborepo builds the monorepo.
How is Magnitude different from Ollama or LM Studio?
According to its README, Magnitude profiles the local hardware and estimates performance for every catalog model before download, then tunes the chosen model (context size, speculative decoding) for that machine. Ollama and LM Studio run whichever model the user picks.
Does Magnitude need a cloud server or API keys?
No. Models are downloaded and run locally through the Rust inference server, and prompts and files stay on the machine. A remote-server mode is documented separately in docs/remote-server.mdx.
Which coding agents can connect to Magnitude?
The README lists harnesses such as Pi, OpenCode, Hermes, OpenClaw, Codex and Claude Code, connected through the packages/harness-connections workspace package and an integration under integrations/pi.
One email a month: new deep dives and stack trends
New source-audited architectures, head-to-head comparisons and the monthly stack report. No spam, unsubscribe anytime.
Similar architectures
- VoiceStudioHybridLocal AI Voice Cloning & Speech Studio · 2-5 peopleShares TypeScript · React · Electron
- vLLMAI-agentHigh-Throughput LLM Inference & Serving Engine · 20+ peopleShares Rust
- MindsHub (formerly MindsDB)AI-agentDesktop & Web AI Agent Workspace · 20+ PeopleShares TypeScript · Electron · React
- QdrantClassicVector Similarity Search Engine · 20+ PeopleShares Rust
- AxumClassicRust Web Framework · 1M+ MAUShares Rust · Axum
- DioxusClassicCross-Platform Rust App Framework · 6-20 peopleShares Rust · Axum