Building an Internal AI Assistant That Actually Knows Your Codebase and Documentation
How we built a RAG-based internal AI assistant that answers questions about our codebase, docs, and processes with 91% accuracy and sub-3-second response time.

Every engineering team I've worked with has the same problem: tribal knowledge locked in Slack threads, Notion pages nobody can find, and code comments that reference decisions made two years ago. We built an internal AI assistant that indexes our entire codebase, documentation, and communication history — then answers questions with cited sources in under 3 seconds. After 4 months in production, it handles 800+ queries per day and has become the first place engineers go for answers.
Why Off-the-Shelf Solutions Failed
We evaluated three commercial knowledge assistant products before building our own. All failed for the same reasons:
| Product | Problem | Impact |
|---|---|---|
| Tool A (SaaS) | Couldn't index private Git repos deeply | Missed 60% of code-related answers |
| Tool B (Enterprise) | 48-hour sync delay | Answers referenced stale docs |
| Tool C (Startup) | Generic chunking destroyed code context | Hallucinated function signatures |
The fundamental issue: off-the-shelf RAG systems treat code as text. But code has structure — imports, call graphs, type definitions. A function's meaning depends on its dependencies, not just its surrounding lines.
Architecture: Code-Aware RAG
Our system has four major components:
- Multi-source indexer — Ingests code, docs, Slack, and Notion with source-specific chunking strategies
- Hybrid retrieval — Combines semantic search, keyword search, and graph-based code navigation
- Context assembly — Builds optimal context windows from retrieved chunks with deduplication
- Generation + citation — Produces answers with inline citations to specific files/docs
The Indexing Pipeline
Different content types need different chunking strategies:
// indexer/chunk-strategies.ts
interface ChunkStrategy {
contentType: string;
chunkSize: number;
overlap: number;
splitOn: string;
enrichWith: string[];
}
const STRATEGIES: Record<string, ChunkStrategy> = {
// Code files: chunk by function/class, include imports and types
'source-code': {
contentType: 'code',
chunkSize: 1500, // tokens
overlap: 200,
splitOn: 'ast-node', // Split on function/class boundaries
enrichWith: ['imports', 'type-definitions', 'jsdoc-comments'],
},
// Markdown docs: chunk by heading, preserve hierarchy
'documentation': {
contentType: 'markdown',
chunkSize: 800,
overlap: 100,
splitOn: 'heading', // h2/h3 boundaries
enrichWith: ['parent-heading', 'page-title', 'last-modified'],
},
// Slack messages: chunk by thread, include full context
'slack-thread': {
contentType: 'conversation',
chunkSize: 2000,
overlap: 0, // Threads are atomic
splitOn: 'thread', // One chunk per thread
enrichWith: ['channel-name', 'participants', 'reactions', 'resolution'],
},
// Architecture Decision Records: chunk whole, they're concise
'adr': {
contentType: 'decision',
chunkSize: 3000,
overlap: 0,
splitOn: 'document', // Keep ADRs intact
enrichWith: ['status', 'date', 'stakeholders'],
},
};
For code specifically, we use AST-based chunking instead of naive text splitting:
// indexer/code-chunker.ts
import * as ts from 'typescript';
interface CodeChunk {
content: string;
metadata: {
filePath: string;
symbolName: string;
symbolType: 'function' | 'class' | 'interface' | 'module';
imports: string[];
exportedBy: string[];
calledBy: string[]; // From call graph analysis
dependencies: string[]; // Types and modules this depends on
lastModified: string;
authors: string[]; // From git blame
};
embedding?: number[];
}
function chunkTypeScriptFile(filePath: string, source: string): CodeChunk[] {
const sourceFile = ts.createSourceFile(filePath, source, ts.ScriptTarget.Latest);
const chunks: CodeChunk[] = [];
// Walk AST and extract top-level declarations
ts.forEachChild(sourceFile, (node) => {
if (ts.isFunctionDeclaration(node) || ts.isArrowFunction(node.parent)) {
chunks.push(buildFunctionChunk(node, filePath, source));
} else if (ts.isClassDeclaration(node)) {
chunks.push(buildClassChunk(node, filePath, source));
} else if (ts.isInterfaceDeclaration(node)) {
chunks.push(buildInterfaceChunk(node, filePath, source));
}
});
// Also chunk the file-level context (imports, constants, types)
chunks.push(buildFileContextChunk(sourceFile, filePath, source));
return chunks;
}
Hybrid Retrieval: Three Search Strategies
No single retrieval method works for all queries. We combine three:
// retrieval/hybrid-retriever.ts
interface RetrievalResult {
chunks: ScoredChunk[];
strategy: string;
confidence: number;
}
async function hybridRetrieve(
query: string,
context: QueryContext
): Promise<RetrievalResult> {
// Strategy 1: Semantic search (good for "how does X work?" questions)
const semanticResults = await vectorSearch(query, {
topK: 15,
minScore: 0.72,
filter: context.scopeFilter,
});
// Strategy 2: Keyword/BM25 search (good for exact names, error messages)
const keywordResults = await bm25Search(query, {
topK: 10,
fields: ['content', 'metadata.symbolName', 'metadata.filePath'],
});
// Strategy 3: Graph traversal (good for "what calls X?" or "what depends on Y?")
const graphResults = await codeGraphSearch(query, {
maxDepth: 3,
relationTypes: ['calls', 'imports', 'implements', 'extends'],
});
// Reciprocal Rank Fusion to combine results
const fused = reciprocalRankFusion([
{ results: semanticResults, weight: 0.5 },
{ results: keywordResults, weight: 0.3 },
{ results: graphResults, weight: 0.2 },
]);
return {
chunks: fused.slice(0, 10), // Top 10 for context window
strategy: determineStrategy(semanticResults, keywordResults, graphResults),
confidence: fused[0]?.score || 0,
};
}
Context Assembly: Fitting the Right Information
With retrieved chunks, we need to assemble them into optimal context for the LLM:
// assembly/context-builder.ts
interface AssembledContext {
systemPrompt: string;
contextChunks: string;
totalTokens: number;
}
function assembleContext(
chunks: ScoredChunk[],
query: string,
maxTokens: number = 12000
): AssembledContext {
// Deduplicate overlapping chunks
const deduped = deduplicateChunks(chunks);
// Sort by relevance score
const sorted = deduped.sort((a, b) => b.score - a.score);
// Fit into context window with priority ordering
let contextParts: string[] = [];
let tokenCount = 0;
for (const chunk of sorted) {
const chunkTokens = estimateTokens(chunk.content);
if (tokenCount + chunkTokens > maxTokens) break;
contextParts.push(formatChunkWithCitation(chunk));
tokenCount += chunkTokens;
}
return {
systemPrompt: buildAssistantPrompt(query),
contextChunks: contextParts.join('\n---\n'),
totalTokens: tokenCount,
};
}
function formatChunkWithCitation(chunk: ScoredChunk): string {
const source = chunk.metadata.filePath || chunk.metadata.pageTitle;
return `[Source: ${source}]\n${chunk.content}`;
}
Accuracy Metrics
We evaluate accuracy weekly against a test set of 200 questions with known correct answers:
| Category | Accuracy | Avg Response Time | Citation Accuracy |
|---|---|---|---|
| Code questions ("How does X work?") | 89% | 2.4s | 94% |
| Process questions ("How do I deploy?") | 95% | 1.8s | 97% |
| Architecture questions ("Why did we choose X?") | 87% | 2.9s | 91% |
| Debugging ("Why is X failing?") | 82% | 3.1s | 88% |
| People questions ("Who owns X?") | 93% | 1.5s | 96% |
| Overall | 91% | 2.3s | 93% |
Usage Patterns After 4 Months
The assistant handles 800+ queries per day. Usage patterns reveal what engineers actually need:
| Time of Day | Query Volume | Top Category |
|---|---|---|
| 9-10 AM | 120 queries | "How do I..." (onboarding, setup) |
| 10 AM-12 PM | 200 queries | "How does X work?" (deep dives) |
| 1-3 PM | 180 queries | "Who owns..." and "Where is..." |
| 3-5 PM | 200 queries | "Why is X failing?" (debugging) |
| 5-7 PM | 100 queries | "What's the process for..." (shipping) |
New engineers use the assistant 4× more than tenured engineers in their first month. After month 3, usage normalizes — but never drops below 5 queries/day even for senior engineers.
Keeping the Index Fresh
Stale data kills trust. We implemented real-time and batch indexing:
// indexer/sync-manager.ts
const SYNC_STRATEGIES = {
// Git: webhook-triggered on push
code: {
trigger: 'webhook',
event: 'push',
latency: '< 30 seconds',
strategy: 'incremental-diff',
},
// Notion: polling every 5 minutes
documentation: {
trigger: 'poll',
interval: '5 minutes',
latency: '< 5 minutes',
strategy: 'last-modified-check',
},
// Slack: real-time via Socket Mode
conversations: {
trigger: 'realtime',
event: 'message',
latency: '< 5 seconds',
strategy: 'append-only',
},
// Full re-index: weekly for consistency
fullReindex: {
trigger: 'cron',
schedule: 'Sunday 2 AM',
strategy: 'complete-rebuild',
},
};
Cost Breakdown
| Component | Monthly Cost | Notes |
|---|---|---|
| Embedding generation (OpenAI ada-002) | $340 | Initial index + daily updates |
| Vector database (Pinecone) | $230 | 5M vectors, 1536 dimensions |
| LLM inference (Claude Sonnet) | $1,800 | 800 queries/day × $0.075 avg |
| Compute (indexing pipeline) | $180 | Lambda + ECS tasks |
| Storage (raw content cache) | $45 | S3 |
| Total | $2,595/month |
At 800 queries/day and 80 engineers, that's $1.08 per engineer per day. Against an estimated 15 minutes saved per query (searching Slack, asking colleagues, reading docs), the time savings are worth approximately $32K/month in engineering productivity.
Lessons Learned
1. Code needs AST-aware chunking. Naive text splitting at 500 tokens breaks functions in half, loses import context, and produces chunks that hallucinate when retrieved.
2. Citations build trust. Engineers won't trust an answer they can't verify. Every response links to the specific file, doc page, or Slack thread it drew from.
3. "I don't know" is better than hallucination. When retrieval confidence is below 0.65, the assistant says "I couldn't find a definitive answer, but here's what might be relevant..." This preserves trust.
4. Slack threads are gold. The most valuable knowledge in most organizations exists in Slack threads where decisions were debated and made. Index them.
5. Freshness matters more than completeness. An assistant that gives you last week's API docs is worse than one that admits it doesn't know. Real-time indexing from Git webhooks is essential.
Key Takeaways
-
AST-based code chunking dramatically outperforms naive text splitting — functions need their imports, types, and callers to be understood by the LLM.
-
Hybrid retrieval (semantic + keyword + graph) catches 40% more relevant results than semantic search alone, especially for exact symbol names and error messages.
-
Citations are non-negotiable — without verifiable sources, engineers won't trust AI-generated answers about their own systems.
-
Index freshness determines adoption — stale answers destroy trust faster than anything else. Webhook-triggered indexing for code is the minimum.
-
The ROI is clear at scale — $2,600/month for 80 engineers saving 15 minutes per query equals ~12× return on investment.
-
Start with documentation, add code second — docs are easier to chunk correctly and give faster wins. Add code-aware indexing once you've proven the pattern.
The internal knowledge assistant has become our most-used internal tool after the IDE itself. It didn't replace documentation — it made documentation actually discoverable. That's the real value: not generating answers from nothing, but surfacing the answers that already exist in your organization.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.