A production-ready personal RAG system. Multi-workspace semantic search across tens of thousands of personal knowledge documents directly from Claude Desktop, Claude.ai web, or Claude iOS app — via the Model Context Protocol.
At a glance
- 42,000+ sources / 182,000+ chunks indexed across 6 workspaces (work / consulting / personal / shared research / AI vendor canon / secrets)
- Multilingual semantic search —
bge-m3embeddings +bge-reranker-v2-m3cross-encoder; effective on Vietnamese, English, and code-mixed text - Retrieval quality — Hit@3 = 97.8%, MRR = 0.948 on the held-out personal eval set
- End-to-end p95 latency: 840 ms warm / 2.3 s cold (embed query → ANN search → rerank → tunnel)
- Resource footprint: 388 MB idle / 501 MB active — runs comfortably on a MacBook Pro M2 Max alongside other workloads
- Database size: 3.2 GB Postgres (pgvector + metadata)
- MCP Streamable HTTP server with OAuth 2.0 (PKCE + DCR) + legacy bearer fallback
- Persistent HTTPS endpoint via Cloudflare named tunnel
- S3-native storage — MinIO BlobStore primary with filesystem mirror fallback; dual-scheme URIs (
s3://+file://) - Auto-ingest pipeline — files written to the mount path are automatically chunked, embedded, and indexed via idempotent SHA-256 hash check
- Workspace-scoped retrieval — separate MCP tools per workspace (
kb_search_ll,kb_search_mindx,kb_search_personal,kb_search_shared,kb_search_canon) for trust-tier-aware routing
Stack
Python 3.11 · MCP SDK · FastMCP · Starlette + uvicorn · bge-m3 · bge-reranker-v2-m3 · Postgres 16 + pgvector (HNSW) · MinIO S3 (BlobStore abstraction) · Cloudflare Tunnel · launchd (mount-watcher, S3 hourly mirror, daily pg backup)
Documentation
| Doc | Read this for |
|---|---|
| PRD | What & why — problem framing, goals, scope, milestones, success metrics |
| Architecture | System diagrams, data flows, component responsibilities, failure modes |
| Implementation | Tech stack, code structure, schema, performance numbers, security model, reproducibility steps |
| Notes | Chronological decision log + gotchas + working-session hours |
Quickstart for clients
Claude Desktop (macOS):
// ~/Library/Application Support/Claude/claude_desktop_config.json
{
"mcpServers": {
"Personal-RAG": {
"command": "npx",
"args": ["-y", "mcp-remote", "https://<your-host>/mcp",
"--header", "Authorization:${AUTH_HEADER}"],
"env": { "AUTH_HEADER": "Bearer <your-token>" }
}
}
}
Claude.ai web / iOS app: Settings → Connectors → Add Custom Connector → enter your MCP URL → complete OAuth login.
Project status
| Day | Milestone |
|---|---|
| 1 | MCP server scaffold + Cloudflare tunnel + bearer auth |
| 2 | kb_health / kb_ingest / kb_search / kb_stats tools + ADB schema |
| 3 | Bulk migrate first 5,000+ sources / 45K chunks |
| 3b | Multilingual upgrade: BGE-en → multilingual-e5-small |
| 4-5 | Sync workflow refactor — 4 source types auto-ingest |
| 6 | Persistent named tunnel on a custom domain |
| 7 | Weekly backup + disaster recovery script |
| 8 | OAuth 2.0 — Claude.ai web + iOS app access |
| M2 (May 2026) | Embedder + DB swap: e5-small → bge-m3 + reranker; Oracle ADB → local Postgres + pgvector |
| M2b | Multi-workspace refactor (6 workspaces + scoped MCP tools + client_filter for LL tenants) |
| M3 (May 2026) | S3 migration: MinIO BlobStore primary + FS mirror fallback (s3:// + file:// dual-scheme URIs) |
| M3b | _canon workspace added for AI vendor docs (powers AI-Canon-Crawler) |
Foundation for downstream products
This RAG infrastructure is the shared foundation for 9 downstream production native AI products: Eval-Framework, Knowledge-Audit, Mail-Assistant, AI-Canon-Crawler, Mac-Translator, Diagram-Engine, Voice-Assistant, Health-Coach, Lumi. Each calls scoped kb_search_* tools rather than building its own vector pipeline.