Building an Internal AI Assistant That Actually Knows Your Codebase and Documentation

How we built a RAG-based internal AI assistant that answers questions about our codebase, docs, and processes with 91% accuracy and sub-3-second response time.

#ai#knowledge-assistant#rag#internal-tools#enterprise
Cover image for the article: Building an Internal AI Assistant That Actually Knows Your Codebase and Documentation

Every engineering team I've worked with has the same problem: tribal knowledge locked in Slack threads, Notion pages nobody can find, and code comments that reference decisions made two years ago. We built an internal AI assistant that indexes our entire codebase, documentation, and communication history — then answers questions with cited sources in under 3 seconds. After 4 months in production, it handles 800+ queries per day and has become the first place engineers go for answers.

Why Off-the-Shelf Solutions Failed

We evaluated three commercial knowledge assistant products before building our own. All failed for the same reasons:

ProductProblemImpact
Tool A (SaaS)Couldn't index private Git repos deeplyMissed 60% of code-related answers
Tool B (Enterprise)48-hour sync delayAnswers referenced stale docs
Tool C (Startup)Generic chunking destroyed code contextHallucinated function signatures

The fundamental issue: off-the-shelf RAG systems treat code as text. But code has structure — imports, call graphs, type definitions. A function's meaning depends on its dependencies, not just its surrounding lines.

Architecture: Code-Aware RAG

Internal AI Assistant Architecture

Our system has four major components:

  1. Multi-source indexer — Ingests code, docs, Slack, and Notion with source-specific chunking strategies
  2. Hybrid retrieval — Combines semantic search, keyword search, and graph-based code navigation
  3. Context assembly — Builds optimal context windows from retrieved chunks with deduplication
  4. Generation + citation — Produces answers with inline citations to specific files/docs

The Indexing Pipeline

Different content types need different chunking strategies:

// indexer/chunk-strategies.ts
interface ChunkStrategy {
  contentType: string;
  chunkSize: number;
  overlap: number;
  splitOn: string;
  enrichWith: string[];
}

const STRATEGIES: Record<string, ChunkStrategy> = {
  // Code files: chunk by function/class, include imports and types
  'source-code': {
    contentType: 'code',
    chunkSize: 1500,     // tokens
    overlap: 200,
    splitOn: 'ast-node', // Split on function/class boundaries
    enrichWith: ['imports', 'type-definitions', 'jsdoc-comments'],
  },

  // Markdown docs: chunk by heading, preserve hierarchy
  'documentation': {
    contentType: 'markdown',
    chunkSize: 800,
    overlap: 100,
    splitOn: 'heading',  // h2/h3 boundaries
    enrichWith: ['parent-heading', 'page-title', 'last-modified'],
  },

  // Slack messages: chunk by thread, include full context
  'slack-thread': {
    contentType: 'conversation',
    chunkSize: 2000,
    overlap: 0,          // Threads are atomic
    splitOn: 'thread',   // One chunk per thread
    enrichWith: ['channel-name', 'participants', 'reactions', 'resolution'],
  },

  // Architecture Decision Records: chunk whole, they're concise
  'adr': {
    contentType: 'decision',
    chunkSize: 3000,
    overlap: 0,
    splitOn: 'document', // Keep ADRs intact
    enrichWith: ['status', 'date', 'stakeholders'],
  },
};

For code specifically, we use AST-based chunking instead of naive text splitting:

// indexer/code-chunker.ts
import * as ts from 'typescript';

interface CodeChunk {
  content: string;
  metadata: {
    filePath: string;
    symbolName: string;
    symbolType: 'function' | 'class' | 'interface' | 'module';
    imports: string[];
    exportedBy: string[];
    calledBy: string[];      // From call graph analysis
    dependencies: string[];  // Types and modules this depends on
    lastModified: string;
    authors: string[];       // From git blame
  };
  embedding?: number[];
}

function chunkTypeScriptFile(filePath: string, source: string): CodeChunk[] {
  const sourceFile = ts.createSourceFile(filePath, source, ts.ScriptTarget.Latest);
  const chunks: CodeChunk[] = [];

  // Walk AST and extract top-level declarations
  ts.forEachChild(sourceFile, (node) => {
    if (ts.isFunctionDeclaration(node) || ts.isArrowFunction(node.parent)) {
      chunks.push(buildFunctionChunk(node, filePath, source));
    } else if (ts.isClassDeclaration(node)) {
      chunks.push(buildClassChunk(node, filePath, source));
    } else if (ts.isInterfaceDeclaration(node)) {
      chunks.push(buildInterfaceChunk(node, filePath, source));
    }
  });

  // Also chunk the file-level context (imports, constants, types)
  chunks.push(buildFileContextChunk(sourceFile, filePath, source));

  return chunks;
}

Hybrid Retrieval: Three Search Strategies

No single retrieval method works for all queries. We combine three:

// retrieval/hybrid-retriever.ts
interface RetrievalResult {
  chunks: ScoredChunk[];
  strategy: string;
  confidence: number;
}

async function hybridRetrieve(
  query: string,
  context: QueryContext
): Promise<RetrievalResult> {
  // Strategy 1: Semantic search (good for "how does X work?" questions)
  const semanticResults = await vectorSearch(query, {
    topK: 15,
    minScore: 0.72,
    filter: context.scopeFilter,
  });

  // Strategy 2: Keyword/BM25 search (good for exact names, error messages)
  const keywordResults = await bm25Search(query, {
    topK: 10,
    fields: ['content', 'metadata.symbolName', 'metadata.filePath'],
  });

  // Strategy 3: Graph traversal (good for "what calls X?" or "what depends on Y?")
  const graphResults = await codeGraphSearch(query, {
    maxDepth: 3,
    relationTypes: ['calls', 'imports', 'implements', 'extends'],
  });

  // Reciprocal Rank Fusion to combine results
  const fused = reciprocalRankFusion([
    { results: semanticResults, weight: 0.5 },
    { results: keywordResults, weight: 0.3 },
    { results: graphResults, weight: 0.2 },
  ]);

  return {
    chunks: fused.slice(0, 10), // Top 10 for context window
    strategy: determineStrategy(semanticResults, keywordResults, graphResults),
    confidence: fused[0]?.score || 0,
  };
}

Context Assembly: Fitting the Right Information

With retrieved chunks, we need to assemble them into optimal context for the LLM:

// assembly/context-builder.ts
interface AssembledContext {
  systemPrompt: string;
  contextChunks: string;
  totalTokens: number;
}

function assembleContext(
  chunks: ScoredChunk[],
  query: string,
  maxTokens: number = 12000
): AssembledContext {
  // Deduplicate overlapping chunks
  const deduped = deduplicateChunks(chunks);

  // Sort by relevance score
  const sorted = deduped.sort((a, b) => b.score - a.score);

  // Fit into context window with priority ordering
  let contextParts: string[] = [];
  let tokenCount = 0;

  for (const chunk of sorted) {
    const chunkTokens = estimateTokens(chunk.content);
    if (tokenCount + chunkTokens > maxTokens) break;

    contextParts.push(formatChunkWithCitation(chunk));
    tokenCount += chunkTokens;
  }

  return {
    systemPrompt: buildAssistantPrompt(query),
    contextChunks: contextParts.join('\n---\n'),
    totalTokens: tokenCount,
  };
}

function formatChunkWithCitation(chunk: ScoredChunk): string {
  const source = chunk.metadata.filePath || chunk.metadata.pageTitle;
  return `[Source: ${source}]\n${chunk.content}`;
}

Accuracy Metrics

We evaluate accuracy weekly against a test set of 200 questions with known correct answers:

CategoryAccuracyAvg Response TimeCitation Accuracy
Code questions ("How does X work?")89%2.4s94%
Process questions ("How do I deploy?")95%1.8s97%
Architecture questions ("Why did we choose X?")87%2.9s91%
Debugging ("Why is X failing?")82%3.1s88%
People questions ("Who owns X?")93%1.5s96%
Overall91%2.3s93%

Accuracy by Query Type

Usage Patterns After 4 Months

The assistant handles 800+ queries per day. Usage patterns reveal what engineers actually need:

Time of DayQuery VolumeTop Category
9-10 AM120 queries"How do I..." (onboarding, setup)
10 AM-12 PM200 queries"How does X work?" (deep dives)
1-3 PM180 queries"Who owns..." and "Where is..."
3-5 PM200 queries"Why is X failing?" (debugging)
5-7 PM100 queries"What's the process for..." (shipping)

New engineers use the assistant 4× more than tenured engineers in their first month. After month 3, usage normalizes — but never drops below 5 queries/day even for senior engineers.

Keeping the Index Fresh

Stale data kills trust. We implemented real-time and batch indexing:

// indexer/sync-manager.ts
const SYNC_STRATEGIES = {
  // Git: webhook-triggered on push
  code: {
    trigger: 'webhook',
    event: 'push',
    latency: '< 30 seconds',
    strategy: 'incremental-diff',
  },

  // Notion: polling every 5 minutes
  documentation: {
    trigger: 'poll',
    interval: '5 minutes',
    latency: '< 5 minutes',
    strategy: 'last-modified-check',
  },

  // Slack: real-time via Socket Mode
  conversations: {
    trigger: 'realtime',
    event: 'message',
    latency: '< 5 seconds',
    strategy: 'append-only',
  },

  // Full re-index: weekly for consistency
  fullReindex: {
    trigger: 'cron',
    schedule: 'Sunday 2 AM',
    strategy: 'complete-rebuild',
  },
};

Cost Breakdown

ComponentMonthly CostNotes
Embedding generation (OpenAI ada-002)$340Initial index + daily updates
Vector database (Pinecone)$2305M vectors, 1536 dimensions
LLM inference (Claude Sonnet)$1,800800 queries/day × $0.075 avg
Compute (indexing pipeline)$180Lambda + ECS tasks
Storage (raw content cache)$45S3
Total$2,595/month

At 800 queries/day and 80 engineers, that's $1.08 per engineer per day. Against an estimated 15 minutes saved per query (searching Slack, asking colleagues, reading docs), the time savings are worth approximately $32K/month in engineering productivity.

Lessons Learned

1. Code needs AST-aware chunking. Naive text splitting at 500 tokens breaks functions in half, loses import context, and produces chunks that hallucinate when retrieved.

2. Citations build trust. Engineers won't trust an answer they can't verify. Every response links to the specific file, doc page, or Slack thread it drew from.

3. "I don't know" is better than hallucination. When retrieval confidence is below 0.65, the assistant says "I couldn't find a definitive answer, but here's what might be relevant..." This preserves trust.

4. Slack threads are gold. The most valuable knowledge in most organizations exists in Slack threads where decisions were debated and made. Index them.

5. Freshness matters more than completeness. An assistant that gives you last week's API docs is worse than one that admits it doesn't know. Real-time indexing from Git webhooks is essential.

Key Takeaways

  1. AST-based code chunking dramatically outperforms naive text splitting — functions need their imports, types, and callers to be understood by the LLM.

  2. Hybrid retrieval (semantic + keyword + graph) catches 40% more relevant results than semantic search alone, especially for exact symbol names and error messages.

  3. Citations are non-negotiable — without verifiable sources, engineers won't trust AI-generated answers about their own systems.

  4. Index freshness determines adoption — stale answers destroy trust faster than anything else. Webhook-triggered indexing for code is the minimum.

  5. The ROI is clear at scale — $2,600/month for 80 engineers saving 15 minutes per query equals ~12× return on investment.

  6. Start with documentation, add code second — docs are easier to chunk correctly and give faster wins. Add code-aware indexing once you've proven the pattern.

The internal knowledge assistant has become our most-used internal tool after the IDE itself. It didn't replace documentation — it made documentation actually discoverable. That's the real value: not generating answers from nothing, but surfacing the answers that already exist in your organization.

Comments

    No comments yet. Be the first to share your thoughts.