Building Enterprise Knowledge Systems with Claude
How we built a company-wide knowledge base powered by Claude that reduced internal search time by 73% and improved cross-team knowledge sharing.

Every enterprise has the same problem: critical knowledge is trapped in Slack threads, Confluence pages nobody reads, Google Docs with broken permissions, and the heads of engineers who left six months ago. We spent three months building a Claude-powered knowledge system that unified 47,000 documents across our organization and made institutional knowledge queryable in natural language.
Here's how we architected it, the mistakes we made, and the benchmarks that convinced leadership to roll it out company-wide.
The Problem
Our engineering org had grown from 30 to 180 people in two years. The symptoms were predictable:
- New engineers spent 3-4 days per sprint searching for context
- The same architectural decisions were relitigated every quarter
- Tribal knowledge left with every departing engineer
- Support teams couldn't find relevant internal docs to resolve escalations
We measured internal search satisfaction at 23% — meaning 77% of the time, people couldn't find what they needed through existing tools.
Architecture Overview
The system has four main components: ingestion pipeline, vector store, Claude reasoning layer, and feedback loop.
Ingestion Pipeline
We built a multi-source ingestion system that handles Confluence, Google Drive, Slack (public channels only), GitHub repos, and Notion databases.
import Anthropic from '@anthropic-ai/sdk';
import { QdrantClient } from '@qdrant/js-client-rest';
interface DocumentChunk {
id: string;
content: string;
metadata: {
source: string;
author: string;
lastModified: Date;
team: string;
accessLevel: 'public' | 'team' | 'restricted';
};
embedding?: number[];
}
class KnowledgeIngestionPipeline {
private anthropic: Anthropic;
private qdrant: QdrantClient;
constructor() {
this.anthropic = new Anthropic();
this.qdrant = new QdrantClient({ url: process.env.QDRANT_URL });
}
async ingestDocument(rawContent: string, metadata: DocumentChunk['metadata']): Promise<void> {
// Step 1: Use Claude to extract structured knowledge
const extraction = await this.anthropic.messages.create({
model: 'claude-sonnet-4-20250514',
max_tokens: 4096,
messages: [{
role: 'user',
content: `Extract structured knowledge from this document. Identify:
1. Key decisions and their rationale
2. Technical specifications or constraints
3. Process descriptions
4. Named entities (services, teams, people)
Document:\n${rawContent}`
}]
});
// Step 2: Chunk intelligently based on semantic boundaries
const chunks = await this.semanticChunk(rawContent, extraction.content[0].text);
// Step 3: Generate embeddings and store
for (const chunk of chunks) {
await this.qdrant.upsert('knowledge_base', {
points: [{
id: chunk.id,
vector: chunk.embedding,
payload: { ...chunk.metadata, content: chunk.content }
}]
});
}
}
private async semanticChunk(content: string, structure: string): Promise<DocumentChunk[]> {
// Claude identifies natural breakpoints rather than fixed-size chunking
const response = await this.anthropic.messages.create({
model: 'claude-sonnet-4-20250514',
max_tokens: 2048,
messages: [{
role: 'user',
content: `Given this document structure:\n${structure}\n\nIdentify optimal chunk boundaries
(200-800 tokens each) that preserve semantic coherence. Return as JSON array of
{start_marker, end_marker, topic} objects.\n\nDocument:\n${content}`
}]
});
// Process response into chunks...
return [];
}
}
Query and Reasoning Layer
The query layer doesn't just do similarity search — it uses Claude to understand intent, reformulate queries, and synthesize answers from multiple sources.
import anthropic
from qdrant_client import QdrantClient
from dataclasses import dataclass
@dataclass
class KnowledgeResponse:
answer: str
sources: list[dict]
confidence: float
follow_up_questions: list[str]
class EnterpriseKnowledgeQuery:
def __init__(self):
self.client = anthropic.Anthropic()
self.qdrant = QdrantClient(url=os.environ["QDRANT_URL"])
async def query(self, user_question: str, user_context: dict) -> KnowledgeResponse:
# Step 1: Query expansion with Claude
expanded = self.client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=1024,
messages=[{
"role": "user",
"content": f"""Expand this enterprise knowledge query into 3 search variations.
Consider the user's team ({user_context['team']}) and role ({user_context['role']}).
Original query: {user_question}
Return JSON: {{"queries": ["variation1", "variation2", "variation3"]}}"""
}]
)
search_queries = json.loads(expanded.content[0].text)["queries"]
# Step 2: Multi-query vector search
all_results = []
for q in search_queries:
results = self.qdrant.search(
collection_name="knowledge_base",
query_vector=self.embed(q),
limit=10,
query_filter=self._build_access_filter(user_context)
)
all_results.extend(results)
# Step 3: Deduplicate and rank
unique_results = self._deduplicate(all_results)
top_results = sorted(unique_results, key=lambda x: x.score, reverse=True)[:15]
# Step 4: Synthesize answer with Claude
context_text = "\n---\n".join([
f"[Source: {r.payload['source']}] {r.payload['content']}"
for r in top_results
])
synthesis = self.client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=2048,
messages=[{
"role": "user",
"content": f"""Based on these internal knowledge sources, answer the question.
Cite sources inline. If information conflicts, note the discrepancy.
If you cannot fully answer, say what's missing.
Question: {user_question}
Sources:\n{context_text}"""
}]
)
return KnowledgeResponse(
answer=synthesis.content[0].text,
sources=[{"title": r.payload["source"], "score": r.score} for r in top_results[:5]],
confidence=self._calculate_confidence(top_results),
follow_up_questions=self._generate_follow_ups(user_question, synthesis.content[0].text)
)
Access Control and Security
Enterprise knowledge systems must respect existing permissions. We implemented a layered access model:
- Document-level ACLs — synced from source systems every 15 minutes
- Team-based filtering — users see their team's content plus public content
- PII detection — Claude scans ingested content and flags sensitive data before indexing
- Audit logging — every query and response is logged for compliance
Benchmarks
After 90 days of production usage across 180 engineers:
| Metric | Before | After | Improvement |
|---|---|---|---|
| Avg. search time to answer | 23 min | 6.2 min | 73% reduction |
| Search satisfaction rate | 23% | 81% | 3.5x improvement |
| New engineer ramp time | 14 days | 8 days | 43% faster |
| Repeated architecture discussions | 12/quarter | 3/quarter | 75% reduction |
| Knowledge doc contributions | 4/week | 17/week | 4.25x increase |
The last metric surprised us most — when people can find existing docs, they're more motivated to contribute new ones.
Cost Analysis
Running this system for 180 users costs approximately $2,400/month in Claude API calls:
- Ingestion (daily sync): ~$180/month
- Query expansion + synthesis: ~$1,800/month (avg. 45 queries/user/month)
- PII detection and classification: ~$420/month
Compared to the engineering time saved (estimated 2.3 hours/engineer/week at average fully-loaded cost), the ROI is roughly 31x.
Lessons Learned
Semantic chunking beats fixed-size chunking. Our initial implementation used 512-token fixed chunks. Switching to Claude-guided semantic chunking improved answer relevance by 34%.
Freshness matters more than completeness. Users lost trust when the system returned outdated information. We added staleness indicators and prioritize recently-modified documents in ranking.
Feedback loops are essential. We added thumbs-up/down on every response. After 3 months, we had 12,000 feedback signals that we used to fine-tune our retrieval ranking.
Start with one team, then expand. Rolling out to the entire company at once would have been a disaster. We started with the platform team, iterated for 6 weeks, then expanded department by department.
Conclusion
Building a Claude-powered knowledge base isn't just about RAG — it's about understanding how enterprise knowledge flows, who owns it, and how to make it accessible without breaking trust boundaries. The technical implementation matters, but the organizational change management matters more. Start small, measure obsessively, and let the results speak for themselves.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.