AI · MCP · Developer tools
RepoBrain
An MCP server that gives AI coding assistants a ready-made map of 23 codebases, so they answer from an index instead of re-reading files.
- tokens per question (A/B)
- −59%
- tokens per question (A/B)
- cost per question (A/B)
- −43%
- cost per question (A/B)
- codebases indexed
- 23
- codebases indexed
- nodes in the knowledge graph
- ~308k
- nodes in the knowledge graph
The problem
Our team works across 23 codebases: React web, React Native, Android (Kotlin/Java) and iOS (Swift). AI coding assistants like Claude Code and Cursor answered every question by re-reading the same files, burning tokens. Questions like “what breaks if I change this?” or “how does UPI payment flow across the apps?” needed someone who already knew the whole system.
What I built
An MCP server that every developer’s AI assistant connects to with one command.
Merge to prod branch→Webhook→Incremental indexer→Index + knowledge graph→MCP tools→Claude Code / Cursor / Codex
- One index of every codebase: components, imports, API calls, call graph and git history, mirrored from each repo’s production branch.
- Hybrid search: exact keyword matching for product terms (UPI, OTP, FCM) plus semantic search over precomputed embeddings.
- Cross-repo answers: blast radius (“what breaks if I change X”) and end-to-end feature flow across web, app and native code.
- Team memory: decisions and root causes saved by one developer’s assistant are recalled by everyone else’s.
- Always fresh: every merge re-indexes only the changed files; a 195-file merge takes about 15 seconds. Answers say exactly which commit they reflect.
- Safe by default: secret files are never served, tokens can be limited to specific repos, and the team’s knowledge is stored separately from the rebuildable index.
How I measured it
- A/B runs: the same questions run through headless Claude Code with only the brain vs only local file reads. The brain used 59% fewer tokens and cost 43% less on average.
- Retrieval eval: a 50-question test set with independent answers. Ranking quality (MRR) went from 0.51 to 0.59 after the hybrid-search changes.
What I learned
- Build the eval before tuning ranking. Several changes that “felt better” scored worse.
- Split storage by who writes it. Rebuilding the index can never touch irreplaceable team knowledge.
- Small, precise answers (paths, summaries, short snippets) beat dumping whole files into the model.