Designing an AI Content Moderation System at Scale
Architecture patterns, model cascading strategies, and operational lessons from building content moderation systems processing millions of items daily

Content moderation is one of the most challenging AI engineering problems. The stakes are high - miss harmful content and users are at risk; over-moderate and you destroy the user experience. At scale, you need systems that make millions of decisions per day with sub-second latency while handling adversarial inputs and evolving policy requirements.
This article covers the architecture I recommend for production content moderation, from the model cascade to the human review pipeline.
System Architecture
A modern content moderation system uses a multi-stage cascade to balance speed, cost, and accuracy:
| Stage | Purpose | Latency | Coverage |
|---|---|---|---|
| Rule Engine | Blocklist matching, regex patterns | < 2ms | 15-20% of violations |
| Fast Classifier | Lightweight model (DistilBERT) | < 10ms | 60-70% of violations |
| Deep Classifier | Large model (DeBERTa/GPT) | < 100ms | 90-95% of violations |
| LLM Review | Nuanced policy interpretation | < 2s | Edge cases (< 5%) |
| Human Review | Final arbiter for appeals | Hours | < 1% of all content |
The cascade approach reduces costs by 85% compared to running the most accurate model on all content.
The Rule Engine Layer
Start with deterministic rules. They are fast, explainable, and catch obvious violations:
import re
from typing import Optional
from dataclasses import dataclass
@dataclass
class ModerationResult:
action: str # "allow", "block", "review"
category: str
confidence: float
reason: str
class RuleEngine:
def __init__(self):
self.blocklist = self._load_blocklist()
self.patterns = self._compile_patterns()
def check(self, text: str) -> Optional[ModerationResult]:
normalized = text.lower().strip()
# Exact blocklist match
for term, category in self.blocklist.items():
if term in normalized:
return ModerationResult(
action="block",
category=category,
confidence=0.99,
reason=f"Blocklist match: {category}",
)
# Pattern matching for evasion attempts
for pattern, category in self.patterns:
if pattern.search(normalized):
return ModerationResult(
action="review",
category=category,
confidence=0.85,
reason=f"Pattern match: {category}",
)
return None # Pass to next stage
def _compile_patterns(self):
"""Patterns that catch common evasion techniques."""
return [
(re.compile(r"s[\s._*]+e[\s._*]+l[\s._*]+l"), "spam"),
(re.compile(r"\b(?:k|c)[\W_]*(?:i|1)[\W_]*l[\W_]*l\b"), "violence"),
# Leetspeak and character substitution patterns
(re.compile(r"[h4][a@][t7][e3]"), "hate_speech"),
]
ML Classification Layer
The fast classifier handles the bulk of moderation decisions:
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from torch.nn.functional import softmax
class ModerationClassifier:
CATEGORIES = [
"safe", "hate_speech", "harassment", "violence",
"sexual_content", "spam", "self_harm", "illegal_activity"
]
THRESHOLDS = {
"hate_speech": 0.75,
"harassment": 0.80,
"violence": 0.70,
"sexual_content": 0.75,
"spam": 0.85,
"self_harm": 0.60, # Lower threshold = more sensitive
"illegal_activity": 0.65,
}
def __init__(self, model_path: str):
self.tokenizer = AutoTokenizer.from_pretrained(model_path)
self.model = AutoModelForSequenceClassification.from_pretrained(model_path)
self.model.eval()
@torch.inference_mode()
def classify(self, text: str) -> ModerationResult:
inputs = self.tokenizer(
text, return_tensors="pt", truncation=True, max_length=512
)
outputs = self.model(**inputs)
probs = softmax(outputs.logits, dim=-1)[0]
max_idx = probs.argmax().item()
category = self.CATEGORIES[max_idx]
confidence = probs[max_idx].item()
if category == "safe":
return ModerationResult("allow", "safe", confidence, "ML: safe")
threshold = self.THRESHOLDS.get(category, 0.80)
if confidence >= threshold:
return ModerationResult("block", category, confidence, f"ML: {category}")
elif confidence >= threshold * 0.7:
return ModerationResult("review", category, confidence, f"ML: uncertain")
else:
return ModerationResult("allow", "safe", 1 - confidence, "ML: below threshold")
Training Data Strategy
The quality of moderation models depends heavily on training data composition:
| Data Source | Volume | Quality | Bias Risk |
|---|---|---|---|
| Human-labeled production data | High | High | Medium (annotator bias) |
| Synthetic adversarial examples | Medium | Medium | Low |
| Academic datasets (HateXplain, etc.) | Low | High | High (domain mismatch) |
| Red team outputs | Low | Very High | Low |
| User reports (confirmed) | Medium | High | Medium (reporting bias) |
I recommend this training data mix:
- 50% human-labeled production data
- 20% synthetic adversarial examples
- 15% confirmed user reports
- 10% red team outputs
- 5% academic datasets (for coverage gaps)
Handling Adversarial Inputs
Adversarial users constantly evolve evasion techniques. Build defenses for common attacks:
class TextNormalizer:
"""Normalize adversarial text modifications."""
UNICODE_CONFUSABLES = {
"\u0430": "a", "\u0435": "e", "\u043e": "o", # Cyrillic
"\u0251": "a", "\u0261": "g", # IPA
"\uff41": "a", "\uff42": "b", # Fullwidth
}
def normalize(self, text: str) -> str:
# Step 1: Unicode normalization
text = self._replace_confusables(text)
# Step 2: Remove zero-width characters
text = re.sub(r"[\u200b\u200c\u200d\ufeff]", "", text)
# Step 3: Collapse repeated characters
text = re.sub(r"(.)\1{3,}", r"\1\1", text)
# Step 4: Remove leetspeak
text = self._decode_leetspeak(text)
return text
def _replace_confusables(self, text: str) -> str:
for char, replacement in self.UNICODE_CONFUSABLES.items():
text = text.replace(char, replacement)
return text
def _decode_leetspeak(self, text: str) -> str:
leet_map = {"0": "o", "1": "i", "3": "e", "4": "a", "5": "s", "7": "t"}
return "".join(leet_map.get(c, c) for c in text)
Performance Metrics
Track these metrics for a production moderation system:
| Metric | Target | Why It Matters |
|---|---|---|
| Precision (per category) | > 95% | False positives destroy user trust |
| Recall (per category) | > 90% | Missed content creates safety risk |
| Latency (P95) | < 100ms | User experience on post creation |
| Time to Moderate | < 500ms | Including async pipeline |
| Appeal overturn rate | < 5% | Measures decision quality |
| False positive rate | < 2% | Tracks over-moderation |
Feedback Loop Architecture
The system must continuously improve through human feedback:
class FeedbackLoop:
def __init__(self, db_client, model_registry):
self.db = db_client
self.registry = model_registry
def process_appeal(self, content_id: str, decision: str):
"""Process human review decision."""
original = self.db.get_moderation_decision(content_id)
if decision != original.action:
# Record as training example
self.db.insert_training_example(
text=original.text,
label=decision,
source="appeal_overturn",
priority="high",
)
# Update real-time metrics
self._update_accuracy_metrics(original.category, was_correct=False)
def trigger_retrain(self):
"""Retrain when sufficient new examples accumulate."""
new_examples = self.db.count_new_training_examples()
if new_examples >= 5000:
self.registry.schedule_training(
dataset_version=self.db.get_latest_dataset_version(),
notify_on_complete=True,
)
Scaling Considerations
At 10M+ items per day, you need to think carefully about infrastructure:
| Scale | Architecture | Estimated Cost |
|---|---|---|
| 100K/day | Single GPU server + rule engine | $800/month |
| 1M/day | Auto-scaling GPU cluster + caching | $5,000/month |
| 10M/day | Cascade + dedicated hardware + CDN | $25,000/month |
| 100M/day | Custom inference chips + edge deployment | $150,000/month |
The cascade architecture is the single biggest cost lever. By filtering 70-80% of content through cheap rules and fast models, you reserve expensive deep models for truly ambiguous cases.
Key Takeaways
- Build a cascade, not a monolith. Rule engines catch 15-20% of violations at near-zero cost; save expensive models for hard cases.
- Threshold tuning per category is essential. Self-harm requires higher sensitivity (lower threshold) than spam. One-size-fits-all thresholds lead to either over- or under-moderation.
- Adversarial robustness is ongoing. Budget 20% of your moderation engineering time for evasion defense.
- Human-in-the-loop is not optional. Even the best models need human review for policy edge cases and continuous improvement.
- Measure false positives as aggressively as false negatives. Over-moderation is invisible to safety teams but devastating to user experience.
Content moderation is never solved, only managed. The best systems combine fast deterministic rules, calibrated ML models, and efficient human review into a pipeline that continuously learns and adapts to new threats.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.