← All work

AI · MCP · Developer tools

RepoBrain

An MCP server that gives AI coding assistants a ready-made map of 23 codebases, so they answer from an index instead of re-reading files.

tokens per question (A/B)
−59%
tokens per question (A/B)
cost per question (A/B)
−43%
cost per question (A/B)
codebases indexed
23
codebases indexed
nodes in the knowledge graph
~308k
nodes in the knowledge graph

The problem

Our team works across 23 codebases: React web, React Native, Android (Kotlin/Java) and iOS (Swift). AI coding assistants like Claude Code and Cursor answered every question by re-reading the same files, burning tokens. Questions like “what breaks if I change this?” or “how does UPI payment flow across the apps?” needed someone who already knew the whole system.

What I built

An MCP server that every developer’s AI assistant connects to with one command.

Merge to prod branch→Webhook→Incremental indexer→Index + knowledge graph→MCP tools→Claude Code / Cursor / Codex
  • One index of every codebase: components, imports, API calls, call graph and git history, mirrored from each repo’s production branch.
  • Hybrid search: exact keyword matching for product terms (UPI, OTP, FCM) plus semantic search over precomputed embeddings.
  • Cross-repo answers: blast radius (“what breaks if I change X”) and end-to-end feature flow across web, app and native code.
  • Team memory: decisions and root causes saved by one developer’s assistant are recalled by everyone else’s.
  • Always fresh: every merge re-indexes only the changed files; a 195-file merge takes about 15 seconds. Answers say exactly which commit they reflect.
  • Safe by default: secret files are never served, tokens can be limited to specific repos, and the team’s knowledge is stored separately from the rebuildable index.

How I measured it

  • A/B runs: the same questions run through headless Claude Code with only the brain vs only local file reads. The brain used 59% fewer tokens and cost 43% less on average.
  • Retrieval eval: a 50-question test set with independent answers. Ranking quality (MRR) went from 0.51 to 0.59 after the hybrid-search changes.

What I learned

  • Build the eval before tuning ranking. Several changes that “felt better” scored worse.
  • Split storage by who writes it. Rebuilding the index can never touch irreplaceable team knowledge.
  • Small, precise answers (paths, summaries, short snippets) beat dumping whole files into the model.