Blog
Notes from the field: cloud, AI, leadership, and building startups.

Detecting Data Drift and Triggering Automated Model Retraining
Production patterns for detecting statistical drift in model inputs and outputs, with automated retraining pipelines that maintain model freshness without manual intervention.

Recognizing Invisible Work in Engineering Teams
The engineers keeping everything running deserve more than a Slack emoji — a guide to seeing, naming, and rewarding the work that nobody notices until it stops.

Unit and Integration Testing for Terraform Modules
A comprehensive guide to testing Terraform modules at every level—from static analysis to integration tests—ensuring infrastructure code is reliable before it reaches production.

Managing Conflict Between Engineers
When two smart people disagree loudly — a manager's guide to navigating productive conflict and knowing when to intervene.

Quiet Leadership for Introverted Engineering Leaders
You don't have to be loud to lead. A guide for introverted engineering leaders who want to embrace their natural strengths rather than perform extroversion.

A/B Testing ML Models in Production with Statistical Rigor
How to run statistically valid A/B tests on ML models, handle metric sensitivity, and make confident promotion decisions in production.

Auto-Generating API Documentation from Codebases with Claude
How we built a documentation pipeline that generates and maintains API docs directly from source code, reducing doc drift to near-zero and saving 15 hours per sprint.

Having Career Conversations That Matter
Beyond "where do you see yourself in 5 years" — how to have career development talks that actually help your engineers grow and feel supported.

Complete Monitoring Stack for 200-Pod Kubernetes Clusters
How to build a production-grade Prometheus and Grafana monitoring stack for large Kubernetes clusters with custom metrics, intelligent alerting, and capacity planning.

Supporting Engineers Through Layoffs
Leading with humanity when your company is letting people go — a guide for engineering leaders navigating the most painful part of the job.

Creating a Safe Space for Technical Disagreement
How to build a team culture where engineers can disagree constructively about technical approaches without it becoming personal, political, or destructive

The Scope Creep Playbook: Protecting Delivery Without Being a Gatekeeper
A practical framework for managing scope creep in engineering projects — when to say yes, when to push back, and how to protect timelines without killing innovation.

Real-Time Feature Serving with Sub-5ms P99 Latency
Architecture and implementation patterns for feature stores that serve ML features in real-time with consistent sub-5ms p99 latency at scale.

AI-Powered Postmortem Analysis: Finding Systemic Patterns with Kiro
How Kiro analyzes incident postmortems to identify recurring systemic patterns, predict future failure modes, and generate actionable prevention strategies.

Letting Go of Technical Control as a Leader
The hardest lesson in engineering leadership: your team will do it differently than you would, and that's not just okay — it's the whole point.

Shipping LLMs to Production: Lessons from the Trenches
What actually breaks when you put large language models in front of real users, and the engineering practices that keep AI features reliable.

Implementing Error Budgets That Drive Engineering Decisions
A practical guide to implementing SRE error budget policies that align reliability targets with product velocity and empower teams to make data-driven tradeoffs.

First-Time Manager Survival Guide
The identity crisis of becoming a manager — and how to find your footing when everything you were good at no longer applies.

Reducing Alert Volume by 83% While Catching More Incidents
How we eliminated alert fatigue by restructuring our monitoring from symptom-based to SLO-based alerting, reducing pages from 147/week to 25 while improving detection

Building Psychological Safety in Engineering Teams
Creating teams where people feel safe to fail, disagree, and ask for help — the foundation of high-performing engineering cultures.

LLM Function Calling Reliability Patterns for Production
Battle-tested patterns for reliable LLM function calling including retry strategies, parameter validation, timeout handling, and graceful degradation in agentic systems

ML Model Versioning and Registry Patterns for Production
How to implement model versioning and registry patterns that enable reproducible deployments, safe rollbacks, and audit trails for production ML systems.

The Developer Tools Market Opportunity
How to identify, validate, and capture market opportunities in developer tools where technical founders have a natural advantage

Engineering Leader Burnout Recovery
I burned out as a CTO. Here's what I wish someone had told me earlier — the warning signs, the recovery, and the hard truths about sustainable leadership.

Giving Difficult Feedback With Kindness
How to deliver hard feedback without destroying someone's confidence — a practical guide rooted in empathy and respect for the humans we lead.

Freelance Engineering Economics in 2026: How AI Changed Rates, Demand, and Specialization
Data-driven analysis of how AI tools are reshaping freelance engineering: rate changes, in-demand specializations, and strategies for independent engineers.

Measuring AI-Generated Code Quality in Production
A metrics framework for evaluating AI code generation — functional correctness, maintainability, security, and long-term impact on engineering velocity.

AI Carbon Reporting for Enterprises: Calculating and Disclosing Scope 3 AI Emissions
Guide to calculating Scope 3 AI emissions for enterprise carbon reporting. GHG Protocol methodology, SEC/CSRD compliance, and disclosure frameworks.

Building an Internal AI Assistant That Actually Knows Your Codebase and Documentation
How we built a RAG-based internal AI assistant that answers questions about our codebase, docs, and processes with 91% accuracy and sub-3-second response time.

AI Code Review in CI/CD: Catching Architectural Violations Before They Ship
How we integrated Claude into our CI/CD pipeline to catch architectural violations, security issues, and performance anti-patterns, blocking 340 problematic PRs in 6 months.
