Designing an AI Content Moderation System at Scale

Architecture patterns, model cascading strategies, and operational lessons from building content moderation systems processing millions of items daily

#content-moderation#system-design#safety#nlp
Cover image for the article: Designing an AI Content Moderation System at Scale

Content moderation is one of the most challenging AI engineering problems. The stakes are high - miss harmful content and users are at risk; over-moderate and you destroy the user experience. At scale, you need systems that make millions of decisions per day with sub-second latency while handling adversarial inputs and evolving policy requirements.

This article covers the architecture I recommend for production content moderation, from the model cascade to the human review pipeline.

System Architecture

A modern content moderation system uses a multi-stage cascade to balance speed, cost, and accuracy:

Chart

StagePurposeLatencyCoverage
Rule EngineBlocklist matching, regex patterns< 2ms15-20% of violations
Fast ClassifierLightweight model (DistilBERT)< 10ms60-70% of violations
Deep ClassifierLarge model (DeBERTa/GPT)< 100ms90-95% of violations
LLM ReviewNuanced policy interpretation< 2sEdge cases (< 5%)
Human ReviewFinal arbiter for appealsHours< 1% of all content

The cascade approach reduces costs by 85% compared to running the most accurate model on all content.

The Rule Engine Layer

Start with deterministic rules. They are fast, explainable, and catch obvious violations:

import re
from typing import Optional
from dataclasses import dataclass

@dataclass
class ModerationResult:
    action: str  # "allow", "block", "review"
    category: str
    confidence: float
    reason: str

class RuleEngine:
    def __init__(self):
        self.blocklist = self._load_blocklist()
        self.patterns = self._compile_patterns()

    def check(self, text: str) -> Optional[ModerationResult]:
        normalized = text.lower().strip()

        # Exact blocklist match
        for term, category in self.blocklist.items():
            if term in normalized:
                return ModerationResult(
                    action="block",
                    category=category,
                    confidence=0.99,
                    reason=f"Blocklist match: {category}",
                )

        # Pattern matching for evasion attempts
        for pattern, category in self.patterns:
            if pattern.search(normalized):
                return ModerationResult(
                    action="review",
                    category=category,
                    confidence=0.85,
                    reason=f"Pattern match: {category}",
                )

        return None  # Pass to next stage

    def _compile_patterns(self):
        """Patterns that catch common evasion techniques."""
        return [
            (re.compile(r"s[\s._*]+e[\s._*]+l[\s._*]+l"), "spam"),
            (re.compile(r"\b(?:k|c)[\W_]*(?:i|1)[\W_]*l[\W_]*l\b"), "violence"),
            # Leetspeak and character substitution patterns
            (re.compile(r"[h4][a@][t7][e3]"), "hate_speech"),
        ]

ML Classification Layer

The fast classifier handles the bulk of moderation decisions:

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from torch.nn.functional import softmax

class ModerationClassifier:
    CATEGORIES = [
        "safe", "hate_speech", "harassment", "violence",
        "sexual_content", "spam", "self_harm", "illegal_activity"
    ]

    THRESHOLDS = {
        "hate_speech": 0.75,
        "harassment": 0.80,
        "violence": 0.70,
        "sexual_content": 0.75,
        "spam": 0.85,
        "self_harm": 0.60,  # Lower threshold = more sensitive
        "illegal_activity": 0.65,
    }

    def __init__(self, model_path: str):
        self.tokenizer = AutoTokenizer.from_pretrained(model_path)
        self.model = AutoModelForSequenceClassification.from_pretrained(model_path)
        self.model.eval()

    @torch.inference_mode()
    def classify(self, text: str) -> ModerationResult:
        inputs = self.tokenizer(
            text, return_tensors="pt", truncation=True, max_length=512
        )
        outputs = self.model(**inputs)
        probs = softmax(outputs.logits, dim=-1)[0]

        max_idx = probs.argmax().item()
        category = self.CATEGORIES[max_idx]
        confidence = probs[max_idx].item()

        if category == "safe":
            return ModerationResult("allow", "safe", confidence, "ML: safe")

        threshold = self.THRESHOLDS.get(category, 0.80)

        if confidence >= threshold:
            return ModerationResult("block", category, confidence, f"ML: {category}")
        elif confidence >= threshold * 0.7:
            return ModerationResult("review", category, confidence, f"ML: uncertain")
        else:
            return ModerationResult("allow", "safe", 1 - confidence, "ML: below threshold")

Training Data Strategy

The quality of moderation models depends heavily on training data composition:

Data SourceVolumeQualityBias Risk
Human-labeled production dataHighHighMedium (annotator bias)
Synthetic adversarial examplesMediumMediumLow
Academic datasets (HateXplain, etc.)LowHighHigh (domain mismatch)
Red team outputsLowVery HighLow
User reports (confirmed)MediumHighMedium (reporting bias)

I recommend this training data mix:

  • 50% human-labeled production data
  • 20% synthetic adversarial examples
  • 15% confirmed user reports
  • 10% red team outputs
  • 5% academic datasets (for coverage gaps)

Handling Adversarial Inputs

Adversarial users constantly evolve evasion techniques. Build defenses for common attacks:

class TextNormalizer:
    """Normalize adversarial text modifications."""

    UNICODE_CONFUSABLES = {
        "\u0430": "a", "\u0435": "e", "\u043e": "o",  # Cyrillic
        "\u0251": "a", "\u0261": "g",  # IPA
        "\uff41": "a", "\uff42": "b",  # Fullwidth
    }

    def normalize(self, text: str) -> str:
        # Step 1: Unicode normalization
        text = self._replace_confusables(text)
        # Step 2: Remove zero-width characters
        text = re.sub(r"[\u200b\u200c\u200d\ufeff]", "", text)
        # Step 3: Collapse repeated characters
        text = re.sub(r"(.)\1{3,}", r"\1\1", text)
        # Step 4: Remove leetspeak
        text = self._decode_leetspeak(text)
        return text

    def _replace_confusables(self, text: str) -> str:
        for char, replacement in self.UNICODE_CONFUSABLES.items():
            text = text.replace(char, replacement)
        return text

    def _decode_leetspeak(self, text: str) -> str:
        leet_map = {"0": "o", "1": "i", "3": "e", "4": "a", "5": "s", "7": "t"}
        return "".join(leet_map.get(c, c) for c in text)

Performance Metrics

Track these metrics for a production moderation system:

MetricTargetWhy It Matters
Precision (per category)> 95%False positives destroy user trust
Recall (per category)> 90%Missed content creates safety risk
Latency (P95)< 100msUser experience on post creation
Time to Moderate< 500msIncluding async pipeline
Appeal overturn rate< 5%Measures decision quality
False positive rate< 2%Tracks over-moderation

Feedback Loop Architecture

The system must continuously improve through human feedback:

class FeedbackLoop:
    def __init__(self, db_client, model_registry):
        self.db = db_client
        self.registry = model_registry

    def process_appeal(self, content_id: str, decision: str):
        """Process human review decision."""
        original = self.db.get_moderation_decision(content_id)

        if decision != original.action:
            # Record as training example
            self.db.insert_training_example(
                text=original.text,
                label=decision,
                source="appeal_overturn",
                priority="high",
            )

            # Update real-time metrics
            self._update_accuracy_metrics(original.category, was_correct=False)

    def trigger_retrain(self):
        """Retrain when sufficient new examples accumulate."""
        new_examples = self.db.count_new_training_examples()
        if new_examples >= 5000:
            self.registry.schedule_training(
                dataset_version=self.db.get_latest_dataset_version(),
                notify_on_complete=True,
            )

Scaling Considerations

At 10M+ items per day, you need to think carefully about infrastructure:

ScaleArchitectureEstimated Cost
100K/daySingle GPU server + rule engine$800/month
1M/dayAuto-scaling GPU cluster + caching$5,000/month
10M/dayCascade + dedicated hardware + CDN$25,000/month
100M/dayCustom inference chips + edge deployment$150,000/month

The cascade architecture is the single biggest cost lever. By filtering 70-80% of content through cheap rules and fast models, you reserve expensive deep models for truly ambiguous cases.

Key Takeaways

  • Build a cascade, not a monolith. Rule engines catch 15-20% of violations at near-zero cost; save expensive models for hard cases.
  • Threshold tuning per category is essential. Self-harm requires higher sensitivity (lower threshold) than spam. One-size-fits-all thresholds lead to either over- or under-moderation.
  • Adversarial robustness is ongoing. Budget 20% of your moderation engineering time for evasion defense.
  • Human-in-the-loop is not optional. Even the best models need human review for policy edge cases and continuous improvement.
  • Measure false positives as aggressively as false negatives. Over-moderation is invisible to safety teams but devastating to user experience.

Content moderation is never solved, only managed. The best systems combine fast deterministic rules, calibrated ML models, and efficient human review into a pipeline that continuously learns and adapts to new threats.

Comments

    No comments yet. Be the first to share your thoughts.