Shipping LLMs to Production: Lessons from the Trenches
What actually breaks when you put large language models in front of real users, and the engineering practices that keep AI features reliable.

Prototyping with LLMs takes an afternoon. Shipping them to production, where real users depend on them, is a different discipline entirely. Here's what I've learned leading teams that run AI features in production at scale — serving over 2 million LLM-powered interactions per month.
The demo-to-production gap
An LLM demo that works 9 times out of 10 feels magical. In production, a 10% failure rate means thousands of broken user experiences per day. The gap between "works in the demo" and "works at scale" is where most AI projects die.
At Rafeeq, our first AI feature — an intelligent order routing suggestion — worked flawlessly in testing. The first week in production, it failed on 8% of requests because real-world addresses had formatting the model hadn't seen. That experience shaped everything we built afterward.
Three things close the demo-to-production gap:
1. Treat prompts like code
Prompts are logic. They belong in version control, they need code review, and every change needs regression testing.
- Keep prompts in the repo, not in a dashboard someone edits by hand. We store ours in
src/prompts/with semantic versioning. - Build a golden dataset: 50–200 real examples with expected outputs. We started with 50 edge cases from production logs and grew to 180 over six months.
- Run evals on every prompt change, like unit tests. A one-word edit can silently degrade a distant use case.
- Tag prompts with metadata: model version, temperature, max tokens. When you switch from GPT-4 to Claude, you need to know which prompts were tuned for which model's quirks.
We run our eval suite in CI. A prompt PR doesn't merge unless it passes 95% of the golden dataset with no regressions on previously-passing cases. This has caught silent degradations at least a dozen times.
2. Design for failure, because failure is guaranteed
LLM calls fail in ways traditional APIs don't: timeouts, rate limits, malformed output, hallucinated fields, refusals, and the newest failure mode — the model confidently returning plausible but completely wrong structured data.
- Validate structure: force JSON output with a schema, and validate it. Retry with the validation error in the prompt. We use Zod schemas and retry up to 3 times with increasingly explicit instructions.
- Set timeouts aggressively and always have a fallback path: a simpler model, a cached answer, or a graceful "not available" state. Our timeouts are 8s for GPT-4 class, 4s for smaller models.
- Never let the model write directly to anything important. Human-in-the-loop or constrained tool use for anything with side effects.
- Implement circuit breakers. If a model endpoint fails 5 times in 60 seconds, stop calling it and fall back to a cached response or rule-based logic. This prevents cascade failures during provider outages.
- Version your fallback chain. We maintain a three-tier hierarchy: primary model → smaller fallback model → rule-based logic. Each tier has its own latency budget and quality threshold.
3. Measure quality, not just uptime
Your observability stack needs new dimensions. Traditional monitoring tells you the system is up. LLM monitoring tells you the system is still smart.
| Traditional metric | LLM-era addition |
|---|---|
| Latency, error rate | Token cost per request |
| Throughput | Eval scores over time |
| Availability | Output validation failure rate |
| Response codes | User feedback signal (thumbs, retries, abandons) |
| Error logs | Hallucination detection rate |
If you can't answer "did last week's model upgrade make answers better or worse?", you're flying blind.
We track a composite quality score that combines automated evals (run on 5% of production traffic) with user feedback signals. When the score drops below threshold, an alert fires and we can correlate it with model version changes, prompt updates, or data drift.
4. Cost management is an engineering problem
LLM costs can spiral silently. A developer adding a "helpful" system prompt with 2000 extra tokens doesn't feel expensive until you multiply by 100K daily requests.
- Set per-endpoint token budgets and alert on breaches.
- Cache aggressively. Semantic caching (same question = same answer within TTL) reduced our costs by 40%.
- Use the cheapest model that meets quality requirements for each use case. Not everything needs GPT-4 — 70% of our use cases run on smaller, faster models.
- Monitor cost per user action, not just total spend. When cost-per-action starts climbing, investigate before the bill arrives.
The organizational lesson
The biggest mistake I see is treating AI features as a side project of the data team. LLM features are product engineering: they need the same ownership, on-call, and quality bar as your payments flow. The teams that get this right ship AI that users actually trust.
At Rafeeq, AI features are owned by the product teams that use them, with a shared "AI Platform" layer that handles model routing, caching, observability, and cost allocation. The platform team doesn't own features — they own reliability.
What I'd do differently
If I started again:
- Build the eval framework first, before the first prompt.
- Start with the smallest model that works and scale up only when evals prove you need it.
- Invest in semantic caching from day one — the cost savings compound fast.
- Hire an ML engineer who's done production, not research. The skills are different.
Start small, instrument everything, and earn your way to more autonomy for the model. Reliability is the feature.
Recommended reading

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Comments
No comments yet. Be the first to share your thoughts.