Blog
Notes from the field: cloud, AI, leadership, and building startups.

The State of Agentic AI in 2026: Capabilities, Limitations, and Production Readiness
Comprehensive analysis of agentic AI in 2026 covering production capabilities, current limitations, and enterprise readiness benchmarks with real deployment data.

Observability for AI Agents: Tracing Multi-Step Reasoning Chains in Production
How to implement production observability for AI agents including distributed tracing, reasoning chain analysis, and debugging multi-step failures.

Measuring and Reducing AI Workload Carbon Emissions: A Practical Engineering Guide
Building a carbon-aware scheduling system for ML training and inference workloads that reduced our AI infrastructure emissions by 42% while maintaining SLA commitments.

Why the Gulf Will Produce the Next Wave of Logistics Tech Unicorns
Capital, demographics, infrastructure, and regulation are converging in the GCC. A thesis from inside a Qatari delivery platform doing 16M orders a year.

Per-Team Cost Allocation in Shared Kubernetes Clusters: From Chaos to Clarity
Implementing accurate per-namespace cost allocation in multi-tenant Kubernetes clusters, covering request vs. usage attribution, shared resource amortization, and building showback dashboards that drive accountability.

AI-Assisted Capacity Planning: Surviving Black Friday Without Over-Provisioning
How we used ML forecasting models to predict Black Friday traffic patterns, pre-provision infrastructure with surgical precision, and handle 47x normal load without wasting $180K on idle capacity.

The CTO Role in 2027: Managing Agents, Not Just Engineers
AI agents already review our code and run our migrations. My honest projection of the 2027 engineering org, and which CTO skills appreciate or depreciate.

Measuring and Eliminating Toil: From 40% to 12% of Engineering Time
A systematic approach to identifying, measuring, and automating toil—the repetitive operational work that scales linearly with service growth and prevents engineers from doing creative work.

Beyond LangChain: Production AI Orchestration Patterns
Why LangChain struggles in production and the orchestration patterns that work better for reliable, observable, and maintainable AI workflows at scale.

The Legacy of Leadership: What Remains When You Leave
The thing people remember is not your architecture. It is not your processes. It is how you made them feel. Reflections on what actually endures from engineering leadership.

Serverless Postgres in Production: Branching, Scale-to-Zero, and the End of Database Provisioning
Running Neon serverless Postgres in production for 8 months — covering database branching workflows, scale-to-zero economics, connection pooling, and migration from RDS.

Synthetic Monitors From 12 Regions Catching Issues Before Users
How to implement global synthetic monitoring that detects availability and performance degradation from every major region—catching issues minutes before real users are impacted.

What 3 A.M. Incidents Taught Me That AWS Certifications Never Did
Twenty production incidents reviewed: why understanding beats fixing, what certifications actually train, and the habits that keep a team calm at 3 a.m.

ML-Driven Autoscaling: Predicting Traffic 30 Minutes Ahead of the Spike
Building a predictive autoscaling system that uses time-series forecasting to scale infrastructure before traffic arrives, eliminating cold-start latency during demand surges.

Monitoring and Governing Claude API Costs at Scale
How we built a cost governance platform that reduced Claude API spend by 38% while maintaining output quality through intelligent routing, caching, and budget enforcement.

An Engineering Leader's Sustainable Weekly Rhythm
A realistic weekly rhythm that balances strategy, people, and operational work — without burning out or losing yourself in back-to-back meetings.

I Reviewed 400 CVs to Hire 5 Engineers in Qatar. Here''s What Stood Out
One quarter of hiring in Doha: the funnel numbers, the signals that predicted great engineers, the red flags that never failed, and what did not matter.

Post-Acquisition Technical Integration Playbook
How CTOs navigate the technical integration process after an acquisition, from day-one decisions through full platform consolidation

Why We Throw a Party When Things Break — And How It Made Us Better
How celebrating failures transformed our engineering culture from one where people hid mistakes into one where every incident became a gift of collective learning.

Reducing LLM Response Latency from 3.2s to 800ms in Production
Practical techniques for cutting LLM response latency by 75% through streaming, caching, prompt optimization, and model routing without sacrificing output quality.

Building a Blameless Postmortem Culture That Actually Improves Reliability
A practical framework for conducting blameless postmortems that produce actionable improvements—covering facilitation, templates, action item tracking, and cultural patterns that make learning stick.

RDS Proxy in Production: What the Docs Don''t Tell You
A year of RDS Proxy under 16M orders: multiplexing that works, the pinning trap that silently disables it, and the failover win nobody markets.

Software Supply Chain Security with SBOM at Scale: From Compliance to Defense
Implementing automated SBOM generation, vulnerability correlation, and policy enforcement across 200+ microservices to meet regulatory requirements and prevent supply chain attacks.

gRPC vs REST: When the Performance Difference Actually Matters
Real-world benchmarks comparing gRPC and REST in microservices architectures, with guidance on when protocol choice creates meaningful business impact.

Kiro Is the DevOps Engineer I Didn''t Know I Needed
AWS''s agentic IDE became our platform team''s strongest member: Terraform, IAM, CI/CD, runbooks, and real coding skills, with production use cases.

Mentoring Engineers From Different Backgrounds: Adapting Your Style
Mentoring across difference means recognizing that your path is not the only valid path. How to adapt your mentoring style to truly support each person where they are.

When the War Reached Our Cloud: Evacuating an AWS Region in 6 Hours
The attack that took down AWS Bahrain forced an emergency migration: our DR plan under real fire, and how Kiro moved 63 services in 6 hours, not 3 weeks.

Using AI to Right-Size Infrastructure: How We Saved $200K/Year
A deep dive into building an ML-powered infrastructure optimization system that analyzes usage patterns and automatically right-sizes compute, storage, and database resources.

Hybrid Search: Combining Keyword and Semantic Retrieval for Production
Building a production hybrid search system that combines BM25 keyword matching with vector semantic search for superior relevance at scale.

My Team Looks Six Times Bigger, and I Haven''t Hired Anyone
How we delegated code review and verification to an AI agent, grew review capacity from 310 to 1,860 cycles a week, and kept a human on every merge.
