Why Codebase Context is Non-Negotiable
Modern software development rarely happens within single isolated files. A single modification to a service, database function, or API route cascades across upstream imports, downstream consumers, and system-wide architectural conventions. Traditional review tools and naive AI assistants inspect files in isolation or slice files into arbitrary 500-character text windows. This breaks function signatures, loses syntax hierarchy, and misses critical cross-file logic:Without Codebase Context (Naive Line Analysis)
With Blame’s GraphRAG Engine
Blame’s GraphRAG Architecture
Blame replaces naive chunking with a deterministic Tree-sitter AST engine, a multi-tier relational dependency graph, and a two-stage cross-encoder retrieval pipeline.The 4-Stage Indexing & Graph Pipeline
When a repository is connected or updated via GitHub webhooks, Blame processes files through four stages:1. Smart Filtering & SHA-256 Hash Verification
Before launching heavy analysis, Blame validates whether indexing is required:- Automatically ignores minified bundles, lockfiles (
package-lock.json,yarn.lock,pnpm-lock.yaml,Cargo.lock), sourcemaps, images, binaries, and build directories (node_modules,.next,dist,target,vendor). - Computes a SHA-256 content hash of the file. If the file hasn’t changed since the last commit, database writes and embedding generation are skipped entirely.
2. Universal AST Parsing with Tree-sitter
Blame utilizes language-specific WebAssembly Tree-sitter grammars (supporting TypeScript, JavaScript, Python, Go, Rust, Java, C++, Ruby, PHP, Swift, and more) to extract semantic nodes via declarative S-expressions:@def.fn: Functions, async handlers, generator functions, and class methods.@def.class: Classes, interfaces, structs, traits, enums, and type aliases.@call.direct&@call.method: Function invocations and method calls.@type.ref: Type annotations, interface implementations, and generics.@import.path: Module import sources, named specifiers, and namespace bindings.
3. Multi-Tier Relational Graph Resolution
Extracted references and imports are resolved across files in PostgreSQL using a multi-pass confidence hierarchy:4. Vector Generation (voyage-code-4)
Every parsed syntax node is transformed into a high-dimensional vector using Voyage AI’s code-specialized embedding model (voyage-code-4). Chunks are stored alongside their qualified signatures, line spans, and dependency links in our pgVector DB.
Real-Time Context Retrieval & Cross-Encoder Reranking
When Blame reviews a Pull Request or executes an autonomous agent step, it does not rely on simple keyword matches. It runs a hybrid two-stage retrieval pipeline:- Candidate Retrieval (pgvector): Blame performs cosine similarity search on the repository’s vector index to pull a candidate pool of relevant code chunks.
- Cross-Encoder Reranking (
rerank-3): The candidate pool is evaluated with a cross-encoder model that scores actual contextual relevance against the query prompt. - Graph Impact Traversal (
getDependentChunks): Blame queries the dependency graph to fetch all downstream callers and upstream dependencies connected to the top-ranked symbols.
How Graph Context Powers Blame
1. Cross-File Impact Analysis
When reviewing PR changes, Blame traces downstream callers to verify if signature changes, argument updates, or return types break consuming files.2. Architectural Consistency & Pattern Recognition
Blame compares proposed code changes against existing workspace patterns (e.g., error wrapping, transaction wrappers, ORM idioms, logging standards) and flags anti-patterns.3. Deep Symbol Navigation for AI Agents
When investigating complex tasks or refactorings, Blame’s autonomous agents can traverse from a controller down into database models and out to shared types without hallucinating file locations.Next Steps
- Autonomous Execution Agent — Discover how Blame agents leverage graph context to navigate and modify code.
- Automated PR Reviews — How pull request reviews use graph awareness to catch subtle bugs.
- Connecting Repositories — Add your GitHub repositories and start indexing.