> ## Documentation Index
> Fetch the complete documentation index at: https://docs.getblame.com/llms.txt
> Use this file to discover all available pages before exploring further.

# GraphRAG & Code Intelligence

## Why Codebase Context is Non-Negotiable

Modern software development rarely happens within single isolated files. A single modification to a service, database function, or API route cascades across upstream imports, downstream consumers, and system-wide architectural conventions.

Traditional review tools and naive AI assistants inspect files in isolation or slice files into arbitrary 500-character text windows. This breaks function signatures, loses syntax hierarchy, and misses critical cross-file logic:

### Without Codebase Context (Naive Line Analysis)

```typescript theme={null}
// 1. Without Codebase Context (Naive Line Analysis)

export async function cancelSubscription(
  subscriptionId: string,
  userId: string
) {
  return await db.subscriptions.cancel(subscriptionId);
}

// ❌ Misses: No ownership check verifying the user can cancel this subscription
// ❌ Misses: Missing validation for subscriptions that are already cancelled
// ❌ Misses: No billing event dispatched to the payment processing queue
```

### With Blame's GraphRAG Engine

```typescript theme={null}
// 2. With Blame's GraphRAG Engine

export async function cancelSubscription(
  subscriptionId: string,
  userId: string
) {
  return await db.subscriptions.cancel(subscriptionId);
}

// ✅ Detects: auth/subscriptionGuards.ts requires ownership verification
// ✅ Detects: validation/subscriptions.ts defines `assertCancellableSubscription`
// ✅ Detects: billing/events.ts dispatches `SubscriptionCancelled` after mutations
// ⚠️ Flags: 2 background workers depend on the cancellation event being emitted
```

***

## Blame's GraphRAG Architecture

Blame replaces naive chunking with a deterministic **Tree-sitter AST engine**, a **multi-tier relational dependency graph**, and a **two-stage cross-encoder retrieval pipeline**.

```mermaid theme={null}
flowchart TD
    subgraph Ingestion & Parsing
        Files[Source Code] --> HashCheck{SHA-256 Hash Changed?}
        HashCheck -- No --> Skip[Skip Indexing]
        HashCheck -- Yes --> TS[Tree-sitter WASM Engine]
        TS --> SExpr[S-Expression AST Extraction]
    end

    subgraph Relational Graph Resolution
        SExpr --> Chunks[Syntax Chunks: Functions / Classes / Types]
        SExpr --> Deps[Raw Dependency References]
        Deps --> P1[Pass 2A: Exact Qualified Name Match - Conf 1.0]
        Deps --> P2[Pass 2B: Same-File Symbol Match - Conf 0.85]
        Deps --> P3[Pass 2C: Cross-File Symbol Match - Conf 0.60]
        P1 --> GraphDB[(PostgreSQL + pgvector Graph)]
        P2 --> GraphDB
        P3 --> GraphDB
        Chunks --> GraphDB
    end

    subgraph Hybrid Retrieval & Reranking
        Query[PR Diff / Chat Query] --> VectorSearch[Voyage Code Embeddings + Cosine Search]
        GraphDB --> VectorSearch
        VectorSearch --> Candidates[Candidate Chunk Pool]
        Candidates --> Rerank[Voyage AI Rerank-3 Cross-Encoder]
        Rerank --> GraphTraverse[Downstream / Upstream Graph Traversal]
        GraphTraverse --> Context[Exact High-Relevance Agent Context]
    end
```

***

## The 4-Stage Indexing & Graph Pipeline

When a repository is connected or updated via GitHub webhooks, Blame processes files through four stages:

### 1. Smart Filtering & SHA-256 Hash Verification

Before launching heavy analysis, Blame validates whether indexing is required:

* Automatically ignores minified bundles, lockfiles (`package-lock.json`, `yarn.lock`, `pnpm-lock.yaml`, `Cargo.lock`), sourcemaps, images, binaries, and build directories (`node_modules`, `.next`, `dist`, `target`, `vendor`).
* Computes a SHA-256 content hash of the file. If the file hasn't changed since the last commit, database writes and embedding generation are skipped entirely.

### 2. Universal AST Parsing with Tree-sitter

Blame utilizes language-specific WebAssembly Tree-sitter grammars (supporting TypeScript, JavaScript, Python, Go, Rust, Java, C++, Ruby, PHP, Swift, and more) to extract semantic nodes via declarative S-expressions:

* `@def.fn`: Functions, async handlers, generator functions, and class methods.
* `@def.class`: Classes, interfaces, structs, traits, enums, and type aliases.
* `@call.direct` & `@call.method`: Function invocations and method calls.
* `@type.ref`: Type annotations, interface implementations, and generics.
* `@import.path`: Module import sources, named specifiers, and namespace bindings.

### 3. Multi-Tier Relational Graph Resolution

Extracted references and imports are resolved across files in PostgreSQL using a multi-pass confidence hierarchy:

| **Resolution Pass** | **Target Scope** | **Confidence Score** | **Description** |
| :- | :- | :- | :- |
| **Pass 2A** | `file/path.ts#Class.method` | `1.00` | Exact match against fully-qualified symbol names across files. |
| **Pass 2B** | Same File | `0.85` | Matches local function and variable references defined within the same file. |
| **Pass 2C** | Cross-File Symbol | `0.60` | Resolves bare symbol invocations across separate files throughout the repository. |

### 4. Vector Generation (`voyage-code-4`)

Every parsed syntax node is transformed into a high-dimensional vector using Voyage AI's code-specialized embedding model (`voyage-code-4`). Chunks are stored alongside their qualified signatures, line spans, and dependency links in our pgVector DB.

***

## Real-Time Context Retrieval & Cross-Encoder Reranking

When Blame reviews a Pull Request or executes an autonomous agent step, it does not rely on simple keyword matches. It runs a **hybrid two-stage retrieval pipeline**:

```mermaid theme={null}
graph LR
    Input[Code Change / User Prompt] --> Stage1[Stage 1: pgvector Candidate Search]
    Stage1 --> Pool[12-16 Top Vector Candidates]
    Pool --> Stage2[Stage 2: Voyage Rerank-3 Cross-Encoder]
    Stage2 --> TopPicks[Top 3-4 High-Relevance Chunks]
    TopPicks --> Stage3[Stage 3: Graph Traversal for Dependents]
    Stage3 --> FinalContext[Complete Architectural Context]
```

1. **Candidate Retrieval (pgvector):** Blame performs cosine similarity search on the repository's vector index to pull a candidate pool of relevant code chunks.
2. **Cross-Encoder Reranking (**`rerank-3`**):** The candidate pool is evaluated with a cross-encoder model that scores actual contextual relevance against the query prompt.
3. **Graph Impact Traversal (**`getDependentChunks`**):** Blame queries the dependency graph to fetch all downstream callers and upstream dependencies connected to the top-ranked symbols.

***

## How Graph Context Powers Blame

### 1. Cross-File Impact Analysis

When reviewing PR changes, Blame traces downstream callers to verify if signature changes, argument updates, or return types break consuming files.

### 2. Architectural Consistency & Pattern Recognition

Blame compares proposed code changes against existing workspace patterns (e.g., error wrapping, transaction wrappers, ORM idioms, logging standards) and flags anti-patterns.

### 3. Deep Symbol Navigation for AI Agents

When investigating complex tasks or refactorings, Blame's autonomous agents can traverse from a controller down into database models and out to shared types without hallucinating file locations.

***

## Next Steps

* **Autonomous Execution Agent** — Discover how Blame agents leverage graph context to navigate and modify code.
* **Automated PR Reviews** — How pull request reviews use graph awareness to catch subtle bugs.
* **Connecting Repositories** — Add your GitHub repositories and start indexing.
