When to Fine-Tune vs When to RAG: A Data-Driven Framework

A systematic decision framework for choosing between fine-tuning and RAG — based on data characteristics, latency requirements, cost constraints, and maintenance burden.

#fine-tuning#rag#ai#decision-framework
Cover image for the article: When to Fine-Tune vs When to RAG: A Data-Driven Framework

"Should we fine-tune or use RAG?" is the wrong question. The right question is: "What are the characteristics of our knowledge, how often does it change, and what quality bar do we need to hit?" Here's the framework I use to make that decision systematically.

The false dichotomy

Most teams frame this as either/or. In practice, the best systems often combine both — RAG for dynamic knowledge and fine-tuning for behavior, tone, and domain-specific reasoning patterns. But you need to start somewhere, and getting the primary approach wrong wastes months.

The decision matrix

Fine-Tuning vs RAG Decision Framework

FactorFavors RAGFavors Fine-Tuning
Knowledge update frequencyDaily/weeklyQuarterly or less
Data volume>10K documents<1K examples
Output formatFlexible/variedHighly consistent
Latency budget>500ms acceptable<200ms required
Accuracy requirementExact source neededReasoning/style needed
Maintenance teamSmall (1-2 engineers)ML team available
Source attributionRequiredNot needed

When RAG wins clearly

RAG excels when your knowledge base is large, changes frequently, or when users need to trace answers back to source documents.

Ideal RAG scenarios:

  • Customer support over product documentation
  • Internal knowledge bases with daily updates
  • Legal/compliance queries requiring source citation
  • Multi-tenant systems where each tenant has different data

RAG cost model

from dataclasses import dataclass

@dataclass
class RAGCostModel:
    """Estimate monthly costs for a RAG system."""
    num_documents: int
    avg_doc_tokens: int
    queries_per_day: int
    embedding_cost_per_1k_tokens: float = 0.00002
    vector_db_cost_per_1m_vectors: float = 30.0
    llm_cost_per_1k_tokens: float = 0.003
    avg_context_tokens: int = 2000
    avg_output_tokens: int = 500

    def monthly_cost(self) -> dict:
        # One-time embedding cost (amortized over 30 days)
        embedding_cost = (
            self.num_documents * self.avg_doc_tokens / 1000
            * self.embedding_cost_per_1k_tokens
        ) / 30

        # Vector DB storage
        vector_cost = (self.num_documents / 1_000_000) * self.vector_db_cost_per_1m_vectors

        # Per-query costs (embedding + LLM)
        daily_query_embedding = (
            self.queries_per_day * 50 / 1000  # ~50 tokens per query
            * self.embedding_cost_per_1k_tokens
        )
        daily_llm = (
            self.queries_per_day
            * (self.avg_context_tokens + self.avg_output_tokens) / 1000
            * self.llm_cost_per_1k_tokens
        )

        monthly_query_cost = (daily_query_embedding + daily_llm) * 30

        return {
            "embedding_amortized": round(embedding_cost, 2),
            "vector_storage": round(vector_cost, 2),
            "query_costs": round(monthly_query_cost, 2),
            "total_monthly": round(embedding_cost + vector_cost + monthly_query_cost, 2),
        }

# Example: 100K docs, 10K queries/day
rag = RAGCostModel(num_documents=100_000, avg_doc_tokens=500, queries_per_day=10_000)
print(rag.monthly_cost())
# {'embedding_amortized': 0.03, 'vector_storage': 3.0, 'query_costs': 2250.0, 'total_monthly': 2253.03}

When fine-tuning wins clearly

Fine-tuning excels when you need consistent output format, domain-specific reasoning, or when the knowledge is stable and well-bounded.

Ideal fine-tuning scenarios:

  • Domain-specific writing style (medical notes, legal drafting)
  • Consistent structured extraction from domain documents
  • Classification with company-specific taxonomy
  • Behavioral alignment that prompting can't achieve

Fine-tuning cost model

interface FineTuningCostModel {
  trainingExamples: number;
  avgTokensPerExample: number;
  trainingEpochs: number;
  trainingCostPerToken: number;
  inferenceCostPer1kTokens: number;
  queriesPerDay: number;
  retrainingFrequencyDays: number;
}

function calculateFineTuningCosts(config: FineTuningCostModel): {
  trainingCostPerRun: number;
  monthlyInferenceCost: number;
  monthlyTrainingAmortized: number;
  totalMonthly: number;
} {
  const totalTrainingTokens =
    config.trainingExamples * config.avgTokensPerExample * config.trainingEpochs;

  const trainingCostPerRun = (totalTrainingTokens / 1000) * config.trainingCostPerToken;

  const monthlyInferenceCost =
    config.queriesPerDay * 30 * (500 / 1000) * config.inferenceCostPer1kTokens;

  const monthlyTrainingAmortized = trainingCostPerRun * (30 / config.retrainingFrequencyDays);

  return {
    trainingCostPerRun: Math.round(trainingCostPerRun * 100) / 100,
    monthlyInferenceCost: Math.round(monthlyInferenceCost * 100) / 100,
    monthlyTrainingAmortized: Math.round(monthlyTrainingAmortized * 100) / 100,
    totalMonthly: Math.round((monthlyInferenceCost + monthlyTrainingAmortized) * 100) / 100,
  };
}

// Example: 5K training examples, 10K queries/day
const ftCost = calculateFineTuningCosts({
  trainingExamples: 5000,
  avgTokensPerExample: 1000,
  trainingEpochs: 3,
  trainingCostPerToken: 0.000008,
  inferenceCostPer1kTokens: 0.0012,
  queriesPerDay: 10000,
  retrainingFrequencyDays: 90,
});
// { trainingCostPerRun: 120, monthlyInferenceCost: 180, monthlyTrainingAmortized: 40, totalMonthly: 220 }

Head-to-head comparison

Testing both approaches on the same task: answering questions about company policies (500 documents, 2000 test queries).

MetricRAGFine-TuningRAG + Fine-Tuning
Accuracy89%82%93%
Latency (p50)1,800ms450ms2,100ms
Hallucination rate4%12%2%
Source attributionYesNoYes
Monthly cost (10K QPD)$2,250$220$2,400
Update lagMinutesWeeksMinutes
Setup time1 week3 weeks4 weeks

The hybrid approach (RAG for knowledge + fine-tuned model for reasoning) gives the best quality but highest cost and complexity.

The decision flowchart

Answer these five questions:

  1. Does your knowledge change more than monthly? → RAG
  2. Do users need source citations? → RAG
  3. Is consistent output format critical? → Fine-tuning (or structured output)
  4. Is latency under 500ms required? → Fine-tuning (or cached RAG)
  5. Do you have <1000 high-quality training examples? → RAG (fine-tuning needs data)

If you answered "RAG" to questions 1 or 2, start there. Fine-tuning on top of RAG is an optimization, not a starting point.

Common mistakes

Mistake 1: Fine-tuning for knowledge injection. Fine-tuning teaches behavior, not facts. It's unreliable for factual recall and you can't update knowledge without retraining.

Mistake 2: RAG without evaluation. Teams build RAG systems and assume they work. Without measuring retrieval quality (precision@k, recall@k) independently from generation quality, you can't diagnose failures.

Mistake 3: Choosing based on hype. RAG had its hype cycle. Fine-tuning is having one now. Neither is universally better. The data characteristics determine the answer.

Mistake 4: Ignoring the hybrid. A fine-tuned model with RAG context often outperforms either alone. The fine-tuned model is better at reasoning over retrieved context because it's trained on your domain's patterns.

Maintenance burden comparison

Ongoing WorkRAGFine-Tuning
Content updatesIndex new documentsRetrain (hours-days)
Quality monitoringRetrieval metrics + generation metricsOutput quality metrics
Drift detectionEmbedding drift, schema changesModel degradation over time
InfrastructureVector DB, embedding service, chunkingTraining pipeline, model hosting
Team skills neededData engineeringML engineering

Key takeaways

  • RAG for dynamic knowledge and source attribution; fine-tuning for behavior and consistency
  • Start with RAG unless you have a clear fine-tuning use case — it's faster to build and easier to maintain
  • The hybrid approach (RAG + fine-tuned model) gives the best quality but demand it only when the simpler approach falls short
  • Measure retrieval and generation quality independently — they fail for different reasons
  • Cost difference is 10x at scale — factor this into your architecture decision early

The framework isn't about finding the "right" answer. It's about making the trade-offs explicit so you can choose deliberately rather than defaulting to whatever the latest blog post recommends.

Comments

    No comments yet. Be the first to share your thoughts.