Skip to content
STACK IT FAST

WeKnora

Curated OSSClassicLLM Knowledge Base & RAG Platform20+ people

Audited from github.com/Tencent/WeKnora

WeKnora is Tencent's open-source LLM knowledge platform. It turns raw documents into a queryable RAG knowledge base with hybrid retrieval, reranking and an agent layer, and exposes it through a web UI, API, CLI, MCP and chat-app integrations.

Language
Go
Database
PostgreSQL
Hosting
Docker
Running for
1 year 2 months
Team
20+ people

Why this architecture

Retrieval, agents and the API sit in one Go service, and parsing is delegated to a Python gRPC service. Storage backends are pluggable, so the same codebase runs as a lite SQLite binary or a full Kubernetes deployment on existing Postgres, Elasticsearch or vector databases.

Tech stack by layer

24 technologies · audited Sep 25, 2026
Frontend & UI
  • VueVue 3 SPA with TDesign components for knowledge bases, chat, agents and settings
  • PiniaClient state management with vue-router and vue-i18n
  • ViteFrontend build, served in production by an nginx image that proxies the API
Backend & APIs
  • GoMain server (cmd/server, internal/*) plus a Go client SDK (client/) and a Cobra CLI (cli/)
  • GinHTTP router and middleware for the REST API documented with Swagger (docs/swagger.yaml)
  • Pythondocreader service that parses PDF, Office, HTML and ebook files into chunks
  • gRPCProtocol between the Go app and the Python docreader (docreader/proto, connect-go)
Data & Persistence
  • PostgreSQLPrimary database (pgx) with ParadeDB-flavoured migrations for full-text and vector search
  • pgvectorVector storage option inside Postgres for embeddings
  • MySQLAlternative relational backend with its own migration set (migrations/mysql)
  • SQLiteLite single-binary mode with sqlite-vec for local vector search (migrations/sqlite)
  • RedisCache and Asynq task queue backend for asynchronous ingestion jobs
  • ElasticsearchPluggable keyword/vector retrieval backend (v7 and v8 clients), alongside OpenSearch
  • QdrantPluggable vector database backend
  • MilvusPluggable vector database backend, with a migration tool in cmd/milvus-migrate
  • Neo4jGraph store for knowledge-graph style retrieval
  • MinIOObject storage for uploaded files, with S3, Alibaba OSS and KS3 as alternatives
Infrastructure & Deploy
  • DockerCompose stack of frontend, app, docreader and data services; separate sandbox image
  • HelmKubernetes chart under helm/ with a CI workflow
  • SwaggerGenerated OpenAPI docs with a contract test (docs/swagger_contract_test.go)
Tooling, Testing & Ops
  • OllamaLocal model provider (Ollama Go client) next to hosted OpenAI-compatible providers
  • Model Context ProtocolMCP client and server in internal/mcp, internal/mcpserver and a Python mcp-server package

WeKnora architecture diagram

Open SVG
WeKnora architecture diagramWeb app → App server (REST); Agents & IM → App server (MCP); App server → Redis (enqueue); Redis → Ingestion jobs (tasks); Ingestion jobs → docreader (parse files); Ingestion jobs → Retrieval backends (embed · index); App server → PostgreSQL (pgx); App server → Object storage (uploads); App server → Model providers (chat · embed)CLIENTSSERVICESWORKERS & JOBSDATA & STORAGEEXTERNALWeb appVue · TDesignAgents & IMMCP · integrationsApp serverGo · GindocreaderPython · gRPCIngestion jobsAsynqPostgreSQLpgvector · ParadeDBRediscache · queueRetrieval backendsES · Qdrant · Milvus · Neo4jObject storageMinIO · S3 · OSSModel providersOllama · OpenAI-compatibleRESTMCPtasksparse filesembed · indexenqueuepgxuploadschat · embed
How the main components of WeKnora connect, drawn from the audited repository.
Diagram as text
  • Web app (Vue · TDesign) → App server (Go · Gin): REST
  • Agents & IM (MCP · integrations) → App server (Go · Gin): MCP
  • App server (Go · Gin) → Redis (cache · queue): enqueue
  • Redis (cache · queue) → Ingestion jobs (Asynq): tasks
  • Ingestion jobs (Asynq) → docreader (Python · gRPC): parse files
  • Ingestion jobs (Asynq) → Retrieval backends (ES · Qdrant · Milvus · Neo4j): embed · index
  • App server (Go · Gin) → PostgreSQL (pgvector · ParadeDB): pgx
  • App server (Go · Gin) → Object storage (MinIO · S3 · OSS): uploads
  • App server (Go · Gin) → Model providers (Ollama · OpenAI-compatible): chat · embed

Key architectural decisions

5 decisions
  1. 01

    Go core, Python only where document parsing needs it

    The server, agents, retrieval and API are Go (go.mod, internal/), while file parsing lives in a separate Python docreader service (markitdown, pypdf, python-docx, trafilatura) reached over gRPC, so each side uses its strongest ecosystem.

  2. 02

    Every storage layer is pluggable

    go.mod pulls clients for Postgres/pgvector, MySQL, SQLite/sqlite-vec, DuckDB, Elasticsearch 7/8, OpenSearch, Qdrant, Milvus, Neo4j, Redis, MinIO, S3, OSS and KS3, and migrations/ keeps separate mysql, paradedb and sqlite trees, letting deployments pick their existing infrastructure.

  3. 03

    A lite single-binary edition next to the full stack

    docs/LITE.md, .env.lite.example, scripts/package-lite.sh, Formula/weknora-lite.rb and deploy/weknora-lite.service package a SQLite-backed build for Homebrew and systemd, while docker-compose.yml and helm/ cover the full multi-service deployment.

  4. 04

    Asynchronous ingestion through a Redis-backed queue

    hibiken/asynq on Redis, ants goroutine pools and robfig/cron handle document parsing, chunking and embedding as background jobs so uploads return quickly.

  5. 05

    Agents with sandboxed tools and IM integrations

    internal/agent, internal/sandbox and internal/localsandbox run agent tools (with an E2B client and a Docker sandbox image), while Lark, Slack, DingTalk and a WeChat mini program (miniprogram/) expose the knowledge base in chat apps.

How WeKnora is built

How WeKnora is structured

WeKnora's repository holds several Go modules plus a Python service:

Path Role
cmd/server, internal/* Main Go application: handler, router, middleware, application, agent, reranking, datasource, database, stream, mcp, mcpserver, sandbox, im, tracing
client/ Go SDK covering agents, knowledge bases, chunks, sessions, retrieval and evaluation
cli/ Separate Go module: a Cobra + huh CLI that uses the SDK and the MCP Go SDK
docreader/ Python gRPC service for document parsing and splitting
mcp-server/ Standalone Python MCP server
frontend/ Vue 3 SPA
miniprogram/ WeChat mini program client
cmd/desktop Desktop build (Wails is listed in licenses/)
helm/, docker/, deploy/ Deployment
website-docs/ VitePress documentation site

Configuration lives in config/config.yaml, with agent presets and model templates next to it. docs/swagger.yaml is generated from the handlers.

Frontend

frontend/ is a Vue 3 app with TDesign components, Pinia, vue-router and vue-i18n, built with Vite. It renders answers with marked, KaTeX, Mermaid and highlight.js, previews documents with pdfjs-dist, docx-preview and @vue-office/pptx, and streams responses through @microsoft/fetch-event-source. It also embeds @xterm/xterm and noVNC to show agent sandbox sessions. In production, an nginx image (frontend/nginx.conf) serves the SPA and proxies the API.

Backend & APIs

The Go server uses Gin with CORS and JWT (golang-jwt) and Viper configuration. It speaks gRPC/Connect to docreader. For web content it uses chromedp, goquery, go-readability and html-to-markdown, and gofeed for feeds. Background work runs on Asynq with ants worker pools and cron schedules. Agents use sandboxes (a Docker sandbox image and an E2B client) and connect to Lark, Slack and DingTalk through their official SDKs. The Model Context Protocol is implemented with mark3labs/mcp-go.

Data & persistence

Storage is pluggable at every layer:

  • Relational: PostgreSQL via pgx with pgvector, where the migrations target ParadeDB (migrations/paradedb); MySQL; SQLite with sqlite-vec for the lite edition; and DuckDB for tabular analysis.
  • Search and vectors: Elasticsearch 7/8, OpenSearch, Qdrant and Milvus. cmd/milvus-migrate moves data between backends.
  • Graph: Neo4j.
  • Cache and queue: Redis.
  • Files: MinIO, AWS S3, Alibaba OSS or Kingsoft KS3.

Migrations run through golang-migrate (make migrate-up, scripts/migrate.sh). scripts/test_paradedb_upgrade.py covers database upgrades.

Build, test & deploy

  • The Makefile wraps build, test, Swagger generation, lint (.golangci.yml) and image builds. It also builds anydoc, an optional Rust-based parsing library under third_party/anydoc-go.
  • .air.toml provides hot reload in development, with docker-compose.dev.yml and scripts/dev.sh.
  • GitHub Actions has workflows per component: app.yml, frontend.yml, docreader.yml, mcp-server.yml, cli.yml, cli-e2e.yml, dsh-plugin.yml, helm.yml, docker-image.yml, go-lint.yml and release-lite.yml.
  • Images are published as wechatopenai/weknora-app, weknora-ui and weknora-docreader. A Helm chart targets Kubernetes.

Self-hosting notes

The full stack runs with docker compose up after copying .env.example. The app, frontend and docreader containers come up with health checks, and make start-ollama adds local models. The Docker sandbox is off by default because mounting the Docker socket gives the app host-root access, as the compose file warns. For a single machine, the lite edition runs on SQLite and installs with Homebrew (Formula/weknora-lite.rb) or as a systemd service. The GitHub license field reads NOASSERTION. Review LICENSE and THIRD_PARTY_NOTICES.md before redistributing.

What to copy (and what not to)

What to copy

  • Putting document parsing in a separate Python service behind gRPC, which keeps the core service in Go.
  • Shipping a lite single-binary edition on SQLite next to the full stack. It lowers the barrier to trying the product.
  • A Swagger contract test, so the published API docs cannot drift from the handlers.

What not to copy

  • Supporting around ten storage backends multiplies test and migration work. Most products should support one database and one vector store until users demand more.
  • Several independent Go modules (root, client/, cli/) plus Python packages need careful version coordination. Use workspaces or keep the module count small.

Sources & repo audit

Audited
Sep 25, 2026
Commit
4364e61
License
NOASSERTION

Independent analysis of repository at github.com/Tencent/WeKnora. Spotted an inaccuracy? Use the claim form to request a correction.

Maintainer? Add the architecture badge to your README
architecture: stackitfast
[![Architecture on STACK IT FAST](https://stackitfast.com/badge/weknora.svg)](https://stackitfast.com/project/weknora)
Use this stack

Scaffold it with your agent

Paste this prompt into Claude Code, Cursor, Windsurf or AGY to start a project with WeKnora's architecture.

  1. 1Copy the promptThe full markdown spec, with every layer and decision.
  2. 2Open your AI toolClaude Code, Cursor, Windsurf or Copilot, in a new repo.
  3. 3Paste and scaffoldUse it as the first instruction; review before you ship.
use-this-stack.md · 60 lines · 7.8 KB
# MISSION: Scaffold "WeKnora" Production Architecture
You are an expert Senior Staff Software Architect and Full-Stack Engineer. Your mission is to scaffold and implement a production-grade, highly reliable, and modular codebase following the proven architecture of **WeKnora**.
---
## 1. PROJECT SPECIFICATIONS & BENCHMARK
- **Reference Architecture**: WeKnora
- **What It Does**: WeKnora is Tencent's open-source LLM knowledge platform. It turns raw documents into a queryable RAG knowledge base with hybrid retrieval, reranking and an agent layer, and exposes it through a web UI, API, CLI, MCP and chat-app integrations.
- **Domain & Category**: LLM Knowledge Base & RAG Platform
- **Production Scale**: 20+ people
- **Development Mode**: CLASSIC
- **Architectural Rationale**: Retrieval, agents and the API sit in one Go service, and parsing is delegated to a Python gRPC service. Storage backends are pluggable, so the same codebase runs as a lite SQLite binary or a full Kubernetes deployment on existing Postgres, Elasticsearch or vector databases.
- **Live Website Reference**: https://weknora.weixin.qq.com
- **Source Repository**: https://github.com/Tencent/WeKnora
---
## 2. PRODUCTION TECH STACK
- **Full Stack Array**: Go, Gin, Vue, Pinia, Vite, Python, gRPC, PostgreSQL, pgvector, MySQL, SQLite, Redis, Elasticsearch, OpenSearch, Qdrant, Milvus, Neo4j, MinIO, Ollama, Model Context Protocol, Docker, Helm, Kubernetes, Swagger
- **Primary Language(s)**: Go, Vue, TypeScript, Python, JavaScript, PLpgSQL
- **License of the reference repo**: NOASSERTION
- **Frontend**: Vue — Vue 3 SPA with TDesign components for knowledge bases, chat, agents and settings; Pinia — Client state management with vue-router and vue-i18n; Vite — Frontend build, served in production by an nginx image that proxies the API
- **Backend & APIs**: Go — Main server (cmd/server, internal/*) plus a Go client SDK (client/) and a Cobra CLI (cli/); Gin — HTTP router and middleware for the REST API documented with Swagger (docs/swagger.yaml); Python — docreader service that parses PDF, Office, HTML and ebook files into chunks; gRPC — Protocol between the Go app and the Python docreader (docreader/proto, connect-go)
- **Data & persistence**: PostgreSQL — Primary database (pgx) with ParadeDB-flavoured migrations for full-text and vector search; pgvector — Vector storage option inside Postgres for embeddings; MySQL — Alternative relational backend with its own migration set (migrations/mysql); SQLite — Lite single-binary mode with sqlite-vec for local vector search (migrations/sqlite); Redis — Cache and Asynq task queue backend for asynchronous ingestion jobs; Elasticsearch — Pluggable keyword/vector retrieval backend (v7 and v8 clients), alongside OpenSearch; Qdrant — Pluggable vector database backend; Milvus — Pluggable vector database backend, with a migration tool in cmd/milvus-migrate; Neo4j — Graph store for knowledge-graph style retrieval; MinIO — Object storage for uploaded files, with S3, Alibaba OSS and KS3 as alternatives
- **Infrastructure & deploy**: Docker — Compose stack of frontend, app, docreader and data services; separate sandbox image; Helm — Kubernetes chart under helm/ with a CI workflow; Swagger — Generated OpenAPI docs with a contract test (docs/swagger_contract_test.go)
- **Tooling, testing & ops**: Ollama — Local model provider (Ollama Go client) next to hosted OpenAI-compatible providers; Model Context Protocol — MCP client and server in internal/mcp, internal/mcpserver and a Python mcp-server package
---
## 3. KEY ARCHITECTURAL DECISIONS (audited from https://github.com/Tencent/WeKnora @ 4364e61)
1. **Go core, Python only where document parsing needs it**: The server, agents, retrieval and API are Go (go.mod, internal/), while file parsing lives in a separate Python docreader service (markitdown, pypdf, python-docx, trafilatura) reached over gRPC, so each side uses its strongest ecosystem.
2. **Every storage layer is pluggable**: go.mod pulls clients for Postgres/pgvector, MySQL, SQLite/sqlite-vec, DuckDB, Elasticsearch 7/8, OpenSearch, Qdrant, Milvus, Neo4j, Redis, MinIO, S3, OSS and KS3, and migrations/ keeps separate mysql, paradedb and sqlite trees, letting deployments pick their existing infrastructure.
3. **A lite single-binary edition next to the full stack**: docs/LITE.md, .env.lite.example, scripts/package-lite.sh, Formula/weknora-lite.rb and deploy/weknora-lite.service package a SQLite-backed build for Homebrew and systemd, while docker-compose.yml and helm/ cover the full multi-service deployment.
4. **Asynchronous ingestion through a Redis-backed queue**: hibiken/asynq on Redis, ants goroutine pools and robfig/cron handle document parsing, chunking and embedding as background jobs so uploads return quickly.
5. **Agents with sandboxed tools and IM integrations**: internal/agent, internal/sandbox and internal/localsandbox run agent tools (with an E2B client and a Docker sandbox image), while Lark, Slack, DingTalk and a WeChat mini program (miniprogram/) expose the knowledge base in chat apps.
---
## 4. NON-NEGOTIABLE ARCHITECTURAL GUARDRAILS
1. **Monorepo & Modular Separation**:
   - Structure as a Turborepo monorepo with strict package boundaries:
     - `apps/web`: Application UI, routing, layouts, and server endpoints.
     - `packages/ui`: Shared design tokens, CSS variables, and Radix UI primitive components.
     - `packages/db`: Database schemas, client singleton, declarative migrations, and seed scripts.
     - `packages/config`: Shared TypeScript, ESLint, and build configurations.
2. **Strict Type Safety & Zero `any` Policy**:
   - Enable `strict: true`, `noImplicitAny: true`, and `strictNullChecks: true`.
   - Validate ALL external inputs, API request bodies, and query parameters with **Zod** schemas before execution.
3. **Frontend & Rendering Guidelines**:
   - Utilize Vite 6 with TanStack Router for fully type-safe routing. Manage server state and caching via TanStack Query v5 with optimistic UI updates.
4. **Design System & Aesthetics**:
   - Keep every color, radius, shadow and font in a single token file (CSS variables) and consume tokens everywhere; never hardcode hex values in components.
   - Prefer crisp 1px borders and one subtle shadow scale over blurry default shadows. Pair one sans-serif for body/headings with one monospace for tags, badges, metrics, and code.
5. **Data Layer & Reliability**:
   - Write declarative schema definitions with foreign keys, composite indexes on queried filters, and automated timestamp triggers.
   - Use connection pooling and prepared statements for serverless database execution.
---
## 5. STEP-BY-STEP SCAFFOLDING ROADMAP
- **Phase 1: Workspace & Root Config**: Initialize package manager, monorepo configuration (`turbo.json`, `tsconfig.base.json`, `package.json`).
- **Phase 2: Database Schema & Client**: Set up the data layer (PostgreSQL, pgvector, MySQL, SQLite, Redis, Elasticsearch, Qdrant, Milvus, Neo4j, MinIO): client, connection pool, models, and migration scripts.
- **Phase 3: Design Tokens & UI Primitives**: Build accessible `Button`, `Input`, `Card`, `Badge`, and layout wrappers inside `packages/ui`.
- **Phase 4: Core Application Routes & Handlers**: Implement primary authentication, user session handling, and application routes.
- **Phase 5: Quality Assurance & Build Verification**: Run `tsc --noEmit`, ESLint, Prettier, and smoke test suites to ensure zero compilation or runtime errors.
---
## 6. EXECUTION INSTRUCTIONS
1. Review all specifications, architectural guardrails, and stack choices above.
2. Present the full monorepo directory tree structure.
3. Systematically generate the complete, production-ready codebase according to the 5-phase roadmap above — starting with the root workspace setup, followed by the database schema, UI design system package, and full-stack application routes until the repository is fully scaffolded and ready to run.
Scaffolded something with this prompt?
Would you pick this stack for a llm knowledge base & rag platform project?

Frequently asked about WeKnora

What is WeKnora built with?

WeKnora is a Go backend on Gin with a Vue 3 and TDesign frontend built by Vite, plus a Python docreader service reached over gRPC. It stores data in PostgreSQL (with pgvector), MySQL or SQLite and can use Elasticsearch, OpenSearch, Qdrant, Milvus or Neo4j for retrieval.

Can I self-host WeKnora?

Yes. It ships a docker-compose.yml with frontend, app and docreader services, a Helm chart for Kubernetes, and a lite single-binary edition backed by SQLite that installs through Homebrew or runs as a systemd service.

Which LLMs does WeKnora support?

It works with local models through Ollama and with hosted OpenAI-compatible providers configured in config/builtin_models.yaml. A model catalog under scripts/model-catalog is checked against models.dev.

Does WeKnora support MCP?

Yes. The Go server has MCP client and server packages (internal/mcp, internal/mcpserver), and the repository also ships a standalone Python MCP server in mcp-server/ so agents can query knowledge bases.

One email a month: new deep dives and stack trends

New source-audited architectures, head-to-head comparisons and the monthly stack report. No spam, unsubscribe anytime.

use-this-stack.md