The Rewrite Temptation: A Decision Framework That Saved Us from a 6-Month Mistake

A data-driven framework for the rewrite vs. refactor decision, based on our experience nearly making a catastrophic 6-month rewrite that refactoring solved in 8 weeks.

#rewrite#refactor#technical-debt#startups#decision-making
Cover image for the article: The Rewrite Temptation: A Decision Framework That Saved Us from a 6-Month Mistake

Every startup hits the moment where the original codebase feels like it's holding the company back. Deploys break things. Features take 3× longer than they should. New engineers take months to become productive. The siren call of "let's just rewrite it" grows louder every sprint. We nearly answered that call. Our plan was 6 months. The likely reality was 9-12 months. What we actually needed was 8 weeks of targeted refactoring. This is the framework that helped us make the right call.

Why Rewrites Fail (Data)

The industry data on rewrites is sobering:

OutcomeFrequencySource Pattern
Delivered on time and budget12%Strong architectural vision + small scope
Delivered late but successfully34%Underestimated complexity, but persevered
Abandoned or significantly descoped28%Business needs changed during rewrite
Delivered but worse than original14%Lost institutional knowledge encoded in code
Company failed during rewrite12%Competitor shipped features while team was rebuilding

That's a 54% failure rate (abandoned + worse + company failed). And these aren't incompetent teams — they're experienced engineers who underestimated what the existing system actually did.

The core problem: existing code contains knowledge that nobody remembers acquiring. Every weird conditional, every edge case handler, every "why is this here?" comment represents a lesson learned in production. Rewrites lose that institutional memory.

The Rewrite Trap: Our Near-Miss

Our situation at month 18 of the startup:

  • Monolithic Node.js app: 180K lines of code, 3 years of organic growth
  • Deploy frequency: Down from 3/day to 2/week (fear of breaking things)
  • Onboarding time: New engineers took 6-8 weeks to ship first meaningful PR
  • Incident frequency: 2-3 P2+ incidents per week
  • Feature velocity: Product was frustrating engineering couldn't ship faster

The engineering team (12 people) was unanimous: "We need to rewrite this." I was tempted to agree. The codebase was genuinely painful. But before committing, I insisted on a structured analysis.

The Decision Framework

I developed a 5-factor framework that scores the situation on a rewrite-vs-refactor spectrum:

Rewrite vs Refactor Decision Framework

Factor 1: Where Is the Pain Concentrated?

Score: 1 (Rewrite) ←→ 5 (Refactor)

- Pain is everywhere, all layers: Score 1
- Pain is in 3+ major modules: Score 2
- Pain is in 2 major modules: Score 3
- Pain is concentrated in 1 module: Score 4
- Pain is at boundaries (APIs, integrations): Score 5

Our score: 3 — Pain was concentrated in two areas: the API routing layer (spaghetti middleware) and the payment processing module (impossible to test).

Factor 2: Is the Data Model Sound?

Score: 1 (Rewrite) ←→ 5 (Refactor)

- Data model is fundamentally wrong for the business: Score 1
- Multiple competing models, inconsistent: Score 2
- Model is okay but poorly normalized: Score 3
- Model is solid, schema is reasonable: Score 4
- Model is well-designed, schema is clean: Score 5

Our score: 4 — Surprisingly, our database schema was well-designed. The original engineer had done good data modeling. The mess was in the application layer, not the data layer.

Factor 3: Can You Ship Features During the Transition?

Score: 1 (Rewrite) ←→ 5 (Refactor)

- Must freeze features for 3+ months: Score 1
- Must freeze for 1-3 months: Score 2
- Reduced feature velocity (50%) during transition: Score 3
- Minimal impact on feature work: Score 4
- Transition enables faster feature work immediately: Score 5

Our score: 2 for rewrite, 4 for refactor — A rewrite would require a 4-month feature freeze (we'd need to maintain the old system AND build the new one). Refactoring could happen alongside feature work with 20% velocity reduction.

Factor 4: Team Capacity and Expertise

Score: 1 (Rewrite) ←→ 5 (Refactor)

- Team has deep expertise in target architecture: Score 1 (rewrite is feasible)
- Team has moderate expertise: Score 2
- Team would need to learn during rewrite: Score 3
- Team is junior/mixed, needs guardrails: Score 4
- Team has never done a rewrite: Score 5

Our score: 3 — The team was a mix of senior and mid-level. They could execute a rewrite, but the learning curve during execution would add 30-50% to timeline estimates.

Factor 5: Business Pressure and Runway

Score: 1 (Rewrite) ←→ 5 (Refactor)

- 18+ months runway, no competitive pressure: Score 1
- 12-18 months runway, moderate pressure: Score 2
- 6-12 months runway, significant pressure: Score 3
- 3-6 months runway, intense pressure: Score 4
- <3 months or existential competitive threat: Score 5

Our score: 4 — We had 10 months of runway and a competitor shipping features monthly. A 6-month rewrite would mean 6 months of silence while the competitor advanced.

Scoring Summary

FactorScoreWeightWeighted
Pain concentration325%0.75
Data model health420%0.80
Feature continuity425%1.00
Team capacity315%0.45
Business pressure415%0.60
Total3.60

Interpretation:

  • Score 1.0-2.0: Rewrite is likely the right call
  • Score 2.0-3.0: Significant refactor with possible partial rewrite
  • Score 3.0-4.0: Targeted refactor (our result)
  • Score 4.0-5.0: Incremental improvement, no major refactor needed

What We Actually Did: The 8-Week Refactor

Instead of rewriting 180K lines from scratch, we identified the two highest-pain modules and rebuilt them in place using the Strangler Fig pattern:

Week 1-2: API Routing Layer Refactor

// Before: 2000-line middleware chain with implicit state
app.use(authMiddleware);        // Sets req.user (sometimes)
app.use(tenantMiddleware);      // Sets req.tenant (depends on req.user)
app.use(rateLimitMiddleware);   // Depends on req.tenant
app.use(featureFlagMiddleware); // Depends on req.user AND req.tenant
// ... 14 more middlewares

// After: Explicit, typed, composable handlers
const authenticatedRoute = pipe(
  validateToken,
  loadUser,
  loadTenant,
  checkRateLimit,
  loadFeatureFlags,
);

router.get('/api/v1/orders',
  authenticatedRoute,
  validateQuery(OrderQuerySchema),
  ordersController.list,
);

Impact: Deploy frequency went from 2/week to 5/week within 2 weeks of this change. The explicit routing made it obvious what each endpoint did and what could break.

Week 3-6: Payment Module Rebuild

We rebuilt the payment module behind a feature flag using the Strangler Fig pattern:

// Strangler fig: new implementation alongside old
class PaymentService {
  async processPayment(order: Order): Promise<PaymentResult> {
    if (featureFlags.isEnabled('new-payment-flow', order.tenantId)) {
      return this.newPaymentFlow(order);
    }
    return this.legacyPaymentFlow(order);
  }

  // New: clean, tested, typed
  private async newPaymentFlow(order: Order): Promise<PaymentResult> {
    const validated = PaymentValidator.validate(order);
    const intent = await this.stripe.createPaymentIntent(validated);
    await this.eventBus.emit('payment.initiated', { orderId: order.id, intentId: intent.id });
    return { status: 'pending', intentId: intent.id };
  }

  // Legacy: untouched, running for old tenants
  private async legacyPaymentFlow(order: Order): Promise<PaymentResult> {
    // ... original 800 lines of spaghetti ...
  }
}

We migrated tenants one at a time to the new flow:

  • Week 3: Internal testing (our own tenant)
  • Week 4: 3 low-volume tenants (< 100 transactions/day)
  • Week 5: 50% of tenants by volume
  • Week 6: 100% on new flow, legacy code deleted

Week 7-8: Testing and Documentation

With the two pain points resolved, we invested in:

  • Integration tests for the new routing layer
  • Contract tests for the payment module
  • Architecture Decision Records explaining the new patterns
  • Onboarding documentation for the refactored areas

Results: 8 Weeks vs. Projected 6 Months

MetricBefore RefactorAfter RefactorProjected Rewrite
Timeline-8 weeks6+ months
Feature freeze-0 weeks4+ months
Deploy frequency2/week8/weekUnknown
New engineer onboarding6-8 weeks3-4 weeks2-3 weeks (optimistic)
Incident frequency2-3/week0.5/weekUnknown
Code coverage34%72% (refactored areas)90%+ (theoretical)
Lines of code removed-12,000All 180K (then rewrite)

The rewrite would have achieved better theoretical outcomes (clean architecture everywhere) but at 6-9× the time cost with a 54% historical failure rate. The refactor delivered 80% of the benefit in 13% of the time.

When Rewriting IS the Right Call

Refactoring isn't always the answer. You should genuinely rewrite when:

  1. The data model is wrong — if your fundamental entities and relationships are incorrect, refactoring code on top of a broken model just makes the code prettier while the business logic remains wrong.

  2. Technology is end-of-life — if you're on Python 2, an unsupported framework, or a language nobody can hire for, the ongoing cost of staying exceeds the rewrite cost.

  3. The system is genuinely small — rewriting a 5,000-line service takes weeks, not months. The risk calculus changes entirely for small codebases.

  4. You're pre-product-market-fit — if you haven't found PMF yet, a rewrite can be acceptable because the current code serves no validated purpose anyway.

  5. Security is fundamentally compromised — if the architecture has security flaws that can't be patched (authentication baked into every layer rather than centralized), a rewrite may be the only safe path.

Key Takeaways

  1. The Strangler Fig pattern is your best friend — rebuild modules behind feature flags, migrate gradually, delete legacy code. No big bang required.

  2. Score the decision with the 5-factor framework — don't let emotion or frustration drive a rewrite decision. Pain concentration, data model health, feature continuity, team capacity, and business pressure should be evaluated systematically.

  3. A rewrite has a 54% failure rate — this isn't pessimism, it's historical data. If you decide to rewrite, plan for it taking 2× your estimate.

  4. Most rewrite pain is actually concentrated — in our case, 85% of the daily frustration came from 15% of the codebase. Fixing that 15% gave most of the relief.

  5. Feature velocity during transition is non-negotiable for startups — you cannot afford 4-6 months of silence while competitors ship. Refactoring preserves velocity; rewrites don't.

  6. Measure the result in deployment frequency — the single best proxy for "is the codebase healthy?" is how often you can safely deploy. Track it before, during, and after.

The desire to rewrite is natural. Existing code is frustrating because you see all its flaws. But the desire to rewrite is often the desire for a clean start — and clean starts don't exist in production. The best you can achieve is continuous improvement, one targeted refactor at a time.

Comments

    No comments yet. Be the first to share your thoughts.