Blog
Notes from the field: cloud, AI, leadership, and building startups.

Choosing Embedding Models: Quality vs Latency vs Cost
A practical guide to selecting embedding models for production — benchmarking OpenAI, Cohere, Voyage AI, and open-source options across quality, speed, and cost.

Refactoring Monorepos Safely with Kiro and AI-Guided Dependency Analysis
How Kiro maps implicit dependencies across monorepo packages and guides safe refactoring without breaking downstream consumers.

LLM Distillation: Building Smaller, Faster Models for Production
Practical guide to distilling large language models into compact, production-ready versions with benchmarks on quality retention, latency improvements, and cost savings

Managing Up: Giving Your CEO the Right Information at the Right Time
A practical guide to managing up as an engineering leader — what your CEO needs to hear, when to escalate, and how to build trust through communication cadence.

Building AI Agents: Kiro vs Claude vs OpenAI — Which Stack for Which Use Case?
A detailed technical comparison of Kiro, Claude, and OpenAI agent stacks covering architecture, performance benchmarks, cost, and ideal use cases for each platform.

Claude Computer Use in Production: Automating Complex GUI Workflows at Scale
Production patterns for Claude computer use automation achieving 91% task success across GUI-dependent workflows with structured error recovery.

AI''s Impact on Software Engineering Jobs in 2026: What the Data Actually Shows
Data-driven analysis of which engineering roles AI is augmenting vs automating in 2026, based on hiring trends, surveys, and productivity metrics.

Shift-Left Container Security: Building a Zero-Vulnerability CI/CD Pipeline
How we implemented multi-layer container image scanning that blocked 340+ vulnerabilities from reaching production in 90 days

RAG Chunking Strategies: Why 512 Tokens Is Almost Never the Right Answer
A data-driven analysis of chunking strategies for RAG pipelines, comparing fixed-size, semantic, recursive, and document-aware approaches with retrieval benchmarks.

CTO vs VP Engineering: When to Split the Role
Understanding when your startup needs both a CTO and a VP Engineering and how to navigate the organizational transition without disrupting delivery

Agentic AI: How Kiro Enables Fully Autonomous Coding Workflows
Kiro agentic architecture delivers autonomous coding workflows with spec-driven development, achieving 87% first-pass acceptance in production teams.

AI Datacenter Energy Consumption in 2026: A Data-Driven Analysis
AI datacenters projected to consume 4.5% of global electricity by 2027. Analysis of IEA data, TWh projections, and infrastructure implications for CTOs.

AWS CloudFormation StackSets for Multi-Region Deployments
Implementing CloudFormation StackSets for consistent multi-region and multi-account infrastructure deployment with drift detection and operational strategies

AWS ElastiCache Redis Cluster Mode: Multi-Tenant Caching at Scale
How we scaled Redis cluster mode to serve 200+ tenants with sub-millisecond latency while reducing cache infrastructure costs by 40%.

Workload Identity Federation: Keyless Authentication from AWS and Azure to GCP
Eliminating service account keys by federating workload identity from AWS, Azure, GitHub Actions, and Kubernetes to GCP using OIDC and SAML protocols.

Running Llama 3 in Production: Cost Comparison with API Providers at 1M Requests/Day
A detailed cost analysis of self-hosting Llama 3 vs using API providers at 1M daily requests, including GPU costs, ops overhead, and break-even calculations.

Infrastructure Playbook for Expanding to 5 New Countries
A technical guide for startup CTOs planning international expansion, covering multi-region architecture, data residency, compliance, and latency optimization.

Making AI Agents Reliable Enough for Production
Battle-tested patterns for building AI agents that handle failures gracefully — retry strategies, state machines, human-in-the-loop fallbacks, and observability.

Systematic Model Evaluation Frameworks for Claude in Production
Building comprehensive evaluation and benchmarking systems for Claude deployments with custom metrics, regression detection, and continuous monitoring.

AI-Generated Test Strategies: How Kiro Covers Edge Cases Humans Miss
Kiro generates comprehensive testing strategies from specs, consistently finding edge cases and failure modes that manual test planning overlooks.

Cutting Our Kubernetes Bill: A Practical Playbook
Concrete, ordered steps to reduce Kubernetes infrastructure costs: right-sizing, spot instances, autoscaling, and the observability to keep it that way.

Your Pricing Model Is an Engineering Decision. Here Is Why Most CTOs Ignore It.
How pricing model complexity creates hidden engineering costs, and a framework for evaluating pricing structures by their implementation and maintenance burden.

Building an Engineering Mentorship Program
How to design a mentorship program that accelerates engineer growth, builds organizational knowledge, and creates a culture of continuous development

SOC2 Compliance Through Infrastructure Automation: Evidence Collection at Scale
How we automated SOC2 Type II evidence collection using AWS Config, reducing audit prep from 6 weeks to 3 days

The Rewrite Temptation: A Decision Framework That Saved Us from a 6-Month Mistake
A data-driven framework for the rewrite vs. refactor decision, based on our experience nearly making a catastrophic 6-month rewrite that refactoring solved in 8 weeks.

Building an AI Copilot for Enterprise Applications
Architecture patterns for embedding AI copilots into enterprise software including context management, action execution, permission models, and evaluation frameworks

Supply Chain Security with GCP Artifact Registry: Vulnerability Scanning and Binary Authorization
Implementing end-to-end container supply chain security using Artifact Registry vulnerability scanning, Binary Authorization, and attestation workflows.

1:1 Meetings That Develop Engineers Into Leaders
How to structure one-on-one meetings that go beyond status updates and actively develop your engineers leadership capabilities.

Pinecone vs Weaviate vs pgvector: Production Benchmarks
Head-to-head comparison of vector databases under production load — latency, throughput, cost, and operational complexity with real benchmark data.

Multi-Account Strategy for 50+ AWS Accounts with Automated Guardrails
How we designed and implemented an AWS Organizations structure for 50+ accounts with SCPs, automated provisioning, and centralized governance at scale.
