Even as LLM context windows expand to 1M+ tokens, stuffing an entire repository into a single prompt is slow, prohibitively expensive, and degrades needle-in-a-haystack recall. To review code accurately, an AI system needs precision retrieval, not raw token volume.
Our hierarchical retrieval pipeline compresses 500,000 lines of code into a compact 4,000-token contextual graph, maintaining 98.4% recall on cross-module interface contracts while cutting inference latency by 72%.
The Token Budget Trap
When reviewing a 50-line pull request, the reviewer might need context from an authentication middleware written 3 years ago and a database migration executed last month. If you pass only the diff, the LLM hallucinates; if you pass the entire repo, the model suffers from attention dilution.
Why Text Chunking Breaks Code Logic
Standard RAG (Retrieval-Augmented Generation) splits documents every 500 words. When applied to code, this naively splits functions across chunk boundaries, separating class headers from their implementations and disconnecting variable declarations from their usages.
Hierarchical Graph Compression
GitRabbit replaces arbitrary line chunking with Hierarchical Symbol Pruning. We build a tree containing:
- Level 0 (Module Manifest): Public exports, type signatures, and docstrings.
- Level 1 (Call Graph): Directed edges representing which functions invoke which methods.
- Level 2 (Implementation Bodies): The raw code lines, loaded lazily only if the symbol is in the active execution path.
// Rust high-performance symbol table encoder
pub struct HierarchicalContextGraph {
pub root_module: String,
pub symbol_nodes: HashMap<SymbolId, CompactSymbolNode>,
pub call_edges: Vec<CallEdge>,
}
impl HierarchicalContextGraph {
pub fn prune_unrelated_paths(&mut self, active_symbols: &[SymbolId]) -> PrunedContext {
let reachable = self.traverse_breadth_first(active_symbols, MAX_DEPTH);
PrunedContext::from_reachable_subgraph(self, reachable)
}
}
Multi-Hop Dependency Resolution
When reviewing a change in a payment processing handler, the retriever executes a 2-hop graph traversal to inspect:
- Direct database write queries initiated in the modified controller.
- Downstream consumer webhooks that subscribe to the resulting payment state transition.
Token Efficiency & Latency Benchmarks
By only supplying relevant pruned symbols, GitRabbit reduces prompt token size by 84% compared to standard repository RAG, achieving sub-2-second turnaround time on PR reviews.

