Building Enterprise Knowledge Systems with Claude

How we built a company-wide knowledge base powered by Claude that reduced internal search time by 73% and improved cross-team knowledge sharing.

#claude#enterprise#knowledge-base#ai
Cover image for the article: Building Enterprise Knowledge Systems with Claude

Every enterprise has the same problem: critical knowledge is trapped in Slack threads, Confluence pages nobody reads, Google Docs with broken permissions, and the heads of engineers who left six months ago. We spent three months building a Claude-powered knowledge system that unified 47,000 documents across our organization and made institutional knowledge queryable in natural language.

Here's how we architected it, the mistakes we made, and the benchmarks that convinced leadership to roll it out company-wide.

The Problem

Our engineering org had grown from 30 to 180 people in two years. The symptoms were predictable:

  • New engineers spent 3-4 days per sprint searching for context
  • The same architectural decisions were relitigated every quarter
  • Tribal knowledge left with every departing engineer
  • Support teams couldn't find relevant internal docs to resolve escalations

We measured internal search satisfaction at 23% — meaning 77% of the time, people couldn't find what they needed through existing tools.

Architecture Overview

The system has four main components: ingestion pipeline, vector store, Claude reasoning layer, and feedback loop.

Enterprise Knowledge Base Architecture

Ingestion Pipeline

We built a multi-source ingestion system that handles Confluence, Google Drive, Slack (public channels only), GitHub repos, and Notion databases.

import Anthropic from '@anthropic-ai/sdk';
import { QdrantClient } from '@qdrant/js-client-rest';

interface DocumentChunk {
  id: string;
  content: string;
  metadata: {
    source: string;
    author: string;
    lastModified: Date;
    team: string;
    accessLevel: 'public' | 'team' | 'restricted';
  };
  embedding?: number[];
}

class KnowledgeIngestionPipeline {
  private anthropic: Anthropic;
  private qdrant: QdrantClient;

  constructor() {
    this.anthropic = new Anthropic();
    this.qdrant = new QdrantClient({ url: process.env.QDRANT_URL });
  }

  async ingestDocument(rawContent: string, metadata: DocumentChunk['metadata']): Promise<void> {
    // Step 1: Use Claude to extract structured knowledge
    const extraction = await this.anthropic.messages.create({
      model: 'claude-sonnet-4-20250514',
      max_tokens: 4096,
      messages: [{
        role: 'user',
        content: `Extract structured knowledge from this document. Identify:
        1. Key decisions and their rationale
        2. Technical specifications or constraints
        3. Process descriptions
        4. Named entities (services, teams, people)
        
        Document:\n${rawContent}`
      }]
    });

    // Step 2: Chunk intelligently based on semantic boundaries
    const chunks = await this.semanticChunk(rawContent, extraction.content[0].text);

    // Step 3: Generate embeddings and store
    for (const chunk of chunks) {
      await this.qdrant.upsert('knowledge_base', {
        points: [{
          id: chunk.id,
          vector: chunk.embedding,
          payload: { ...chunk.metadata, content: chunk.content }
        }]
      });
    }
  }

  private async semanticChunk(content: string, structure: string): Promise<DocumentChunk[]> {
    // Claude identifies natural breakpoints rather than fixed-size chunking
    const response = await this.anthropic.messages.create({
      model: 'claude-sonnet-4-20250514',
      max_tokens: 2048,
      messages: [{
        role: 'user',
        content: `Given this document structure:\n${structure}\n\nIdentify optimal chunk boundaries 
        (200-800 tokens each) that preserve semantic coherence. Return as JSON array of 
        {start_marker, end_marker, topic} objects.\n\nDocument:\n${content}`
      }]
    });
    // Process response into chunks...
    return [];
  }
}

Query and Reasoning Layer

The query layer doesn't just do similarity search — it uses Claude to understand intent, reformulate queries, and synthesize answers from multiple sources.

import anthropic
from qdrant_client import QdrantClient
from dataclasses import dataclass

@dataclass
class KnowledgeResponse:
    answer: str
    sources: list[dict]
    confidence: float
    follow_up_questions: list[str]

class EnterpriseKnowledgeQuery:
    def __init__(self):
        self.client = anthropic.Anthropic()
        self.qdrant = QdrantClient(url=os.environ["QDRANT_URL"])

    async def query(self, user_question: str, user_context: dict) -> KnowledgeResponse:
        # Step 1: Query expansion with Claude
        expanded = self.client.messages.create(
            model="claude-sonnet-4-20250514",
            max_tokens=1024,
            messages=[{
                "role": "user",
                "content": f"""Expand this enterprise knowledge query into 3 search variations.
                Consider the user's team ({user_context['team']}) and role ({user_context['role']}).
                
                Original query: {user_question}
                
                Return JSON: {{"queries": ["variation1", "variation2", "variation3"]}}"""
            }]
        )

        search_queries = json.loads(expanded.content[0].text)["queries"]

        # Step 2: Multi-query vector search
        all_results = []
        for q in search_queries:
            results = self.qdrant.search(
                collection_name="knowledge_base",
                query_vector=self.embed(q),
                limit=10,
                query_filter=self._build_access_filter(user_context)
            )
            all_results.extend(results)

        # Step 3: Deduplicate and rank
        unique_results = self._deduplicate(all_results)
        top_results = sorted(unique_results, key=lambda x: x.score, reverse=True)[:15]

        # Step 4: Synthesize answer with Claude
        context_text = "\n---\n".join([
            f"[Source: {r.payload['source']}] {r.payload['content']}"
            for r in top_results
        ])

        synthesis = self.client.messages.create(
            model="claude-sonnet-4-20250514",
            max_tokens=2048,
            messages=[{
                "role": "user",
                "content": f"""Based on these internal knowledge sources, answer the question.
                Cite sources inline. If information conflicts, note the discrepancy.
                If you cannot fully answer, say what's missing.
                
                Question: {user_question}
                
                Sources:\n{context_text}"""
            }]
        )

        return KnowledgeResponse(
            answer=synthesis.content[0].text,
            sources=[{"title": r.payload["source"], "score": r.score} for r in top_results[:5]],
            confidence=self._calculate_confidence(top_results),
            follow_up_questions=self._generate_follow_ups(user_question, synthesis.content[0].text)
        )

Access Control and Security

Enterprise knowledge systems must respect existing permissions. We implemented a layered access model:

  1. Document-level ACLs — synced from source systems every 15 minutes
  2. Team-based filtering — users see their team's content plus public content
  3. PII detection — Claude scans ingested content and flags sensitive data before indexing
  4. Audit logging — every query and response is logged for compliance

Benchmarks

After 90 days of production usage across 180 engineers:

MetricBeforeAfterImprovement
Avg. search time to answer23 min6.2 min73% reduction
Search satisfaction rate23%81%3.5x improvement
New engineer ramp time14 days8 days43% faster
Repeated architecture discussions12/quarter3/quarter75% reduction
Knowledge doc contributions4/week17/week4.25x increase

The last metric surprised us most — when people can find existing docs, they're more motivated to contribute new ones.

Cost Analysis

Running this system for 180 users costs approximately $2,400/month in Claude API calls:

  • Ingestion (daily sync): ~$180/month
  • Query expansion + synthesis: ~$1,800/month (avg. 45 queries/user/month)
  • PII detection and classification: ~$420/month

Compared to the engineering time saved (estimated 2.3 hours/engineer/week at average fully-loaded cost), the ROI is roughly 31x.

Lessons Learned

Semantic chunking beats fixed-size chunking. Our initial implementation used 512-token fixed chunks. Switching to Claude-guided semantic chunking improved answer relevance by 34%.

Freshness matters more than completeness. Users lost trust when the system returned outdated information. We added staleness indicators and prioritize recently-modified documents in ranking.

Feedback loops are essential. We added thumbs-up/down on every response. After 3 months, we had 12,000 feedback signals that we used to fine-tune our retrieval ranking.

Start with one team, then expand. Rolling out to the entire company at once would have been a disaster. We started with the platform team, iterated for 6 weeks, then expanded department by department.

Conclusion

Building a Claude-powered knowledge base isn't just about RAG — it's about understanding how enterprise knowledge flows, who owns it, and how to make it accessible without breaking trust boundaries. The technical implementation matters, but the organizational change management matters more. Start small, measure obsessively, and let the results speak for themselves.

Comments

    No comments yet. Be the first to share your thoughts.