5 Technical Decisions I Regret and What They Taught Me About Humility
An honest retrospective on five architectural and technology choices that seemed right at the time but created lasting pain, and the decision frameworks that emerged.

Every CTO has a graveyard of decisions they would reverse if they could. Not the obvious mistakes — nobody regrets choosing a technology that clearly failed. The painful regrets are the decisions that seemed smart, had defensible reasoning, and still turned out wrong in ways that cost years of engineering effort to unwind.
Here are five of mine. Not hypothetical case studies dressed up for a blog post — actual decisions I made, with the reasoning that seemed sound at the time and the consequences that proved otherwise. The point is not self-flagellation. It is to extract the meta-patterns that might help you avoid your own version of these mistakes.
Regret 1: Choosing Microservices at 12 Engineers
The decision (2021): With a team of 12 engineers and a monolith that was growing "too complex," I advocated splitting into 8 microservices. We had read the Amazon two-pizza team papers, seen the Netflix architecture talks, and believed that smaller services would increase developer velocity.
The reasoning: Our deployment pipeline was slow (40 minutes). Engineers were stepping on each other with merge conflicts. Feature velocity was declining. Microservices would let teams deploy independently.
What actually happened:
| Metric | Before (Monolith) | After (8 Services) | Net Effect |
|---|---|---|---|
| Deploy time | 40 min | 12 min per service | Improved |
| Deploys per week | 3 | 2.1 (total across all services) | Worse |
| Incidents per month | 4 | 11 | Much worse |
| Developer onboarding time | 2 weeks | 6 weeks | Much worse |
| Cross-team feature delivery | 1-2 weeks | 4-8 weeks | Much worse |
| Infrastructure cost | $8,400/mo | $23,100/mo | 2.75x |
The deploy time improved, but everything else got worse. With 12 engineers split across 8 services, most services had 1-2 owners. When those people were on vacation or left the company, knowledge was lost. Cross-service features required coordinating 3-4 teams of 1-2 people, which meant meetings, API contracts, and versioning overhead that a monolith handles with a function call.
The lesson: Microservices are an organizational scaling pattern, not a complexity management pattern. They solve coordination problems between large teams (50+ engineers) at the cost of introducing distributed systems complexity. At 12 engineers, the coordination problems are manageable with good git practices and a faster CI pipeline. We should have invested in build optimization, not service decomposition.
The framework I now use:
Service count should roughly equal: team_count / 2
If you have fewer than 4 teams (< 20 engineers):
→ Monolith with good module boundaries
→ Invest in build/deploy speed instead
If you have 4-10 teams (20-60 engineers):
→ 3-6 services aligned to domain boundaries
→ Each service owned by 1 full team minimum
If you have 10+ teams (60+ engineers):
→ Service per team boundary
→ Platform team manages shared infrastructure
Regret 2: Building a Custom Authentication System
The decision (2022): Instead of using Auth0 or Cognito, we built our own JWT-based authentication system. The reasoning was cost (Auth0 was $3/user at our scale) and "we need custom flows our provider cannot support."
The reasoning: We needed passwordless magic links, multi-tenant role hierarchies, and custom MFA flows. Auth0 supported some of these but not all, and the workarounds were hacky. Building our own seemed like a 4-week project that would save $15K/month at scale.
What actually happened:
The initial build took 4 weeks as estimated. Maintaining it took 20% of one senior engineer's time permanently. Over three years:
| Cost Category | Custom Auth (actual) | Auth0 (projected) |
|---|---|---|
| Initial build | $48,000 (eng time) | $0 |
| Ongoing maintenance (3 yr) | $180,000 (eng time) | $0 |
| Security incidents (2 breaches) | $95,000 (response + audit) | $0 (their liability) |
| Compliance audit prep | $35,000/year | Included |
| Monthly SaaS cost | $0 | $15,000/mo |
| 3-year total | $428,000 | $540,000 |
The raw cost comparison looks favorable to custom-built ($428K vs $540K). But the $428K does not account for opportunity cost — that senior engineer could have shipped revenue-generating features instead of patching auth edge cases. And the security incidents were not just financial: they damaged customer trust in ways no dollar figure captures.
The lesson: Never build authentication, payments, or email delivery. These are solved problems where the downside of failure dramatically outweighs the cost savings of custom implementation. The "custom flow" requirement that justifies building your own almost always has a workaround in a mature provider that you just have not found yet.
Regret 3: Premature Multi-Region Deployment
The decision (2022): Before we had product-market fit confirmed, I architected a multi-region active-active deployment. The reasoning was that our target customers (enterprise fintech) would require low-latency access from multiple geographies and we needed to "build it right from the start."
The reasoning: Retrofitting multi-region is expensive. Our competitors were single-region and experiencing latency complaints from international customers. Being multi-region from day one would be a competitive advantage.
What actually happened: We spent 4 months building multi-region infrastructure. During that time:
- A competitor launched a similar product with a single-region monolith
- They acquired 40% of our target market while we were building infrastructure
- Our multi-region architecture added $12,000/month in infrastructure cost before we had paying customers
- When we finally launched, 94% of our traffic came from a single region for the first 18 months
The multi-region capability that we built "for the future" was not needed until 2024. Had we launched single-region in 2022 and migrated in 2024, the total engineering investment would have been similar, but we would have had the product in market 4 months earlier.
The lesson: Infrastructure decisions should serve the current and next business stage, not the stage you hope to reach in 2+ years. Build for the revenue milestone you are approaching, not the one you aspire to. Multi-region is a Series B problem. Do not solve it at pre-seed.
| Stage | Right Infrastructure Investment | Wrong Investment |
|---|---|---|
| Pre-PMF | Single region, simple deployment, fast iteration | Multi-region, complex orchestration |
| PMF confirmed (< $1M ARR) | Reliability basics, monitoring, CI/CD | Custom platforms, service mesh |
| Growth ($1-10M ARR) | Auto-scaling, cost optimization, DR | Kubernetes (for most teams) |
| Scale ($10M+ ARR) | Multi-region, platform team, compliance | Rewriting in a "better" language |
Regret 4: Choosing GraphQL for Internal Service Communication
The decision (2023): We adopted GraphQL not just as our client-facing API but also as the communication protocol between internal services. The reasoning was consistency — one query language everywhere — and the developer experience of typed schemas.
The reasoning: Engineers loved the GraphQL developer experience. Having a consistent query language between frontend-to-backend and backend-to-backend communication meant less context switching. The schema stitching tools would give us a unified graph.
What actually happened:
| Concern | Expected | Actual |
|---|---|---|
| Performance | "Clients fetch exactly what they need" | N+1 query explosions between services |
| Debugging | "Clear schema = clear contracts" | Traces span 6 services through query resolution |
| Caching | "HTTP caching with persisted queries" | Impossible to cache POST-based internal queries |
| Team velocity | "Shared tooling, less context switching" | Schema coordination bottlenecks |
| Latency (p99) | < 200ms | 850ms (resolver chain waterfalls) |
GraphQL is excellent for client-facing APIs where diverse clients need different response shapes. It is terrible for service-to-service communication where you control both sides and need predictable performance characteristics.
The resolver pattern creates implicit dependencies that are invisible until they cascade. Service A resolves a field by calling Service B, which resolves a nested field by calling Service C. A latency spike in Service C manifests as a timeout in Service A with no obvious connection in logs.
The lesson: Match the communication pattern to the relationship type:
- Client to API Gateway: GraphQL (diverse clients, flexible queries)
- Service to service (sync): gRPC or REST (predictable performance, explicit contracts)
- Service to service (async): Events/messages (decoupled, resilient)
We migrated internal communication to gRPC over 6 months. P99 latency dropped from 850ms to 120ms. But that 6-month migration is time we would not have needed if we had made the right choice initially.
Regret 5: Delaying the Data Platform Investment
The decision (2021-2023): For two years, I deprioritized building a proper data platform. Analytics queries ran against production databases. ETL was a collection of cron jobs. Data scientists worked with CSV exports. Every quarter, I said "next quarter we will invest in data infrastructure."
The reasoning: We were a small team with urgent product work. Data infrastructure felt like a luxury. The cron jobs worked. The CSV exports were fine for now. We would build it properly when we had the resources.
What actually happened:
By the time we invested in data infrastructure (late 2023), the cleanup was enormous:
- 47 orphaned cron jobs, 12 of which had been silently failing for months
- Production database performance degraded 30% due to analytical query load
- Three separate "sources of truth" for customer metrics, none of which agreed
- Data scientists spending 60% of their time on data preparation vs analysis
- A critical pricing decision made on incorrect data from a broken ETL job (estimated $400K revenue impact)
The cost of proper data infrastructure in 2021 would have been approximately $80K (engineering time + tooling). The cost of the delayed investment: $80K for the infrastructure plus $200K in engineering cleanup plus $400K in the pricing error. Total cost of delay: approximately $600K.
The lesson: Data infrastructure is not optional once you make decisions based on data. If anyone in the company is querying a database to inform a business decision, you need a proper data platform. The question is not whether, but when — and "next quarter" repeated for two years becomes an expensive compounding mistake.
The Meta-Pattern: Decision Humility
Looking across these five regrets, the meta-pattern is overconfidence in prediction:
- Microservices: Overconfidence that our team would grow faster than it did
- Custom auth: Overconfidence that our custom needs were truly unique
- Multi-region: Overconfidence in our growth timeline
- GraphQL everywhere: Overconfidence that consistency trumps fit-for-purpose
- Delayed data platform: Overconfidence that "later" has no compounding cost
The framework I now apply to every significant technical decision:
Before committing, answer honestly:
1. What am I assuming about the future that might not be true?
2. What is the cost if I am wrong and need to reverse this in 18 months?
3. Is there a simpler option that keeps more futures open?
4. Am I solving today's problem or a problem I expect to have?
5. Who benefits from this decision being complex?
Key Takeaways
- Microservices below 20 engineers create more problems than they solve. Invest in build speed, not service boundaries, until organizational coordination is your actual bottleneck.
- Never build auth, payments, or email delivery. The total cost of ownership for custom-built security-critical infrastructure always exceeds vendor cost when you account for maintenance, incidents, and opportunity cost.
- Build infrastructure for the stage you are in, not the stage you hope to reach. Premature architecture investments delay time-to-market without delivering the value they promise until years later.
- Match communication patterns to relationship types. GraphQL for clients, gRPC for internal sync, events for internal async. Consistency across different relationship types is a false economy.
- Data infrastructure compounds negatively when delayed. Every quarter you postpone proper data engineering, the cleanup cost grows and the risk of decisions made on bad data increases.
The hardest part of admitting these regrets publicly is acknowledging that smarter decisions were available at the time. I was not working with incomplete information — I was working with overconfident interpretation of the information I had. Technical humility is not uncertainty about your skills; it is calibrated confidence about your predictions of the future.
Recommended reading

The Legacy of Leadership: What Remains When You Leave
The thing people remember is not your architecture. It is not your processes. It is how you made them feel. Reflections on what actually endures from engineering leadership.

What 3 A.M. Incidents Taught Me That AWS Certifications Never Did
Twenty production incidents reviewed: why understanding beats fixing, what certifications actually train, and the habits that keep a team calm at 3 a.m.

An Engineering Leader's Sustainable Weekly Rhythm
A realistic weekly rhythm that balances strategy, people, and operational work — without burning out or losing yourself in back-to-back meetings.

Comments
No comments yet. Be the first to share your thoughts.