Why We Throw a Party When Things Break — And How It Made Us Better
How celebrating failures transformed our engineering culture from one where people hid mistakes into one where every incident became a gift of collective learning.

The incident that changed our culture happened on a Tuesday at 2:13 PM. A junior engineer — let's call her Noor — pushed a database migration that took down our production API for 47 minutes. Customers were affected. Revenue was lost. The on-call engineer scrambled to rollback while the leadership Slack channel lit up with questions.
When it was over, Noor was shaking. Literally, physically trembling. She found me in a meeting room and said, with tears in her eyes: "I'm so sorry. I understand if you need to let me go."
Let me go. For a production incident. At a company that claims to have "blameless culture."
That moment broke something in me — because I realized that despite all our postmortem processes and "blameless" documentation, our culture wasn't actually safe. Noor believed, genuinely believed, that a mistake would end her employment. And if she believed it, others probably did too.
That's when I decided: we weren't going to just tolerate failures. We were going to celebrate them.
The Problem With "Blameless" Culture
Most engineering organizations claim to be blameless. They have postmortem templates. They have documents that say "focus on systems, not people." And yet, somehow, everyone still knows who caused the last outage. People still carry shame. Engineers still hesitate to make changes because they're afraid of being the one who breaks production.
The issue is that "blameless" is passive. It says what you won't do (blame people). It doesn't say what you will do (actively celebrate learning from failure). It removes punishment without replacing it with something positive.
And human psychology doesn't work in the absence of signals. If you neither punish nor celebrate failure, people will still assume punishment exists — because that's what every other system in their life (school, sports, previous jobs) has taught them.
You have to replace the absence with a presence. You have to make the safety visible, loud, and repeated.
What We Built: The Failure Celebration System
After the Noor incident, I introduced three connected practices. Together, they transformed how our team related to mistakes.
1. The "Oops Award" (Weekly)
Every Friday in standup, we give a lighthearted award to someone who made an interesting mistake that week. The "winner" shares what happened, what they learned, and gets a round of applause. I bought a cheap plastic trophy from a dollar store — it lives on the winner's desk (or virtual desk icon) for a week.
Rules:
- The person self-nominates. Nobody is ever nominated by someone else.
- The mistake must include a learning. "I broke staging" alone isn't enough. "I broke staging because I assumed the migration was idempotent, and now I always test that" qualifies.
- The response is always: curiosity, laughter, and applause. Never judgment.
- Leadership (me) goes first when we launched this. For weeks, I was the primary self-nominator until others felt safe enough to join.
2. The Monthly "Failure Retro" (Team Meeting)
Once a month, we dedicate a full team meeting (45 minutes) to collectively examining our failures. Not incidents specifically — all failures. Failed experiments. Wrong predictions. Bad decisions. Missed opportunities.
The format:
- Each person brings one "failure" (loosely defined)
- They share it in 3 minutes: what happened, what they'd do differently
- The team asks curious questions
- We identify patterns across failures
Some of the most valuable discussions we've ever had came from these meetings. We discovered that 40% of our production issues traced back to the same root cause (insufficient staging environment parity). We found that most "bad decisions" happened when we skipped our own design review process under time pressure. Patterns became visible that no single postmortem could reveal.
3. The "First Failure" Welcome (Onboarding)
When a new engineer joins, we tell them explicitly: "You will break something in production within your first month. That's not a warning — it's an expectation. When it happens, we're going to celebrate that you were moving fast enough to hit the edge cases. Please don't suffer in silence. Tell us immediately and we'll fix it together."
This sets the cultural tone from day one. New hires learn that speed and risk-taking are valued, and that incidents are collective responsibilities, not individual shames.
| Practice | Frequency | Duration | What It Builds |
|---|---|---|---|
| Oops Award | Weekly | 5 minutes | Normalizes small mistakes |
| Failure Retro | Monthly | 45 minutes | Pattern recognition, collective learning |
| First Failure Welcome | On join | 5 minutes in onboarding | Psychological safety from day one |
| Incident celebrations | On occurrence | Post-resolution | Reframes incidents as learning |
| Learning logs | Continuous | 5 min per entry | Individual growth tracking |
What Changed (The Data)
I want to share what actually happened after we implemented these practices, because "celebrating failure" can sound like fluffy culture stuff that doesn't actually move metrics. It does.
Mean time to detection dropped by 60%. When people aren't afraid of incidents, they report issues immediately instead of trying to fix them quietly and hoping nobody notices.
Incident frequency initially increased, then decreased significantly. In the first quarter, reported incidents went up — because people started reporting things they would have previously hidden. By the third quarter, total incidents dropped because we were catching and fixing root causes that had been invisible.
Deployment frequency increased by 40%. When the fear of breaking production decreases, people ship more frequently, in smaller increments. Smaller deployments are inherently safer. A virtuous cycle.
Noor became one of our strongest engineers. Within a year of "the incident," she was leading our infrastructure reliability efforts. The same person who thought she'd be fired became the person who made our systems more resilient for everyone.
How to Start (Even if Your Culture Isn't There Yet)
If your team currently hides mistakes and you want to shift toward celebration, here's the path:
Week 1-4: Lead by Example
Share your own failures first. Publicly. In team meetings. In Slack. Make it clear that the most senior person in the room makes mistakes too. This is non-negotiable — if leadership doesn't go first, nobody will follow.
What I shared in our first week: "Last Thursday, I gave incorrect guidance in a design review that sent the team in the wrong direction for two days. I should have said 'I'm not sure' instead of guessing. Lesson: it's better to say 'let me look into this' than to give bad direction confidently."
Week 4-8: Invite Participation Gently
Start the Oops Award. Make it lighthearted. Emphasize that it's voluntary. Don't force it. Some people will participate immediately. Others will watch for weeks before they feel safe. That's okay. Safety builds gradually.
Month 2-3: Formalize the Failure Retro
Once the weekly Oops Award is normalized, introduce the monthly failure retro. Frame it as: "We're so good at learning from individual incidents. What if we looked for patterns across all our stumbles?"
Month 3-6: Watch for Resistance
Some people will resist this cultural shift. Common resistances:
- "We're celebrating mediocrity." No — we're celebrating learning. The failure itself isn't celebrated. The learning is.
- "Some failures are unacceptable." True. We celebrate learning from good-faith mistakes, not negligence or repeated careless errors. There's a line.
- "My team won't go for this." They might not — at first. Trust builds slowly. Give it six months before deciding it doesn't work.
The Boundary: What We Don't Celebrate
I want to be clear about what this isn't. We don't celebrate:
- Repeated identical mistakes. Making a mistake once is learning. Making the same mistake five times is a performance issue.
- Negligence. Forgetting to test because you were careless is different from a test not catching an edge case you couldn't foresee.
- Violations of safety practices. If we have guardrails and someone bypasses them deliberately, that's not a "celebration" moment.
The line is: good-faith effort + unexpected outcome = learning opportunity. Anything that falls on the "bad faith" or "known risk ignored" side needs a different conversation — a compassionate one, but a different one.
The Deeper Why
Celebrating failures isn't ultimately about metrics or deployment frequency. It's about something more fundamental: it's about creating an environment where people can be fully human at work.
Humans make mistakes. That's not a bug in our species — it's a feature. Every mistake carries information. Every failure reveals something about the system that success hides. When you create a culture that treats failures as gifts rather than crimes, you unlock the full innovative potential of your team.
People who aren't afraid to fail try things. They experiment. They push boundaries. They propose wild ideas. They take the kind of risks that lead to breakthroughs. And when those risks don't pan out? They bring the learning back to the team, openly and immediately, so everyone benefits.
That's the culture Noor helped us build. Not because she broke production — but because her fear made the gap between our stated values and our lived values painfully visible. And once we saw it, we couldn't unsee it.
Start This Week
If you take one thing from this article, let it be this: share one of your own failures with your team before Friday. Not a humble-brag. A real mistake. Something that made you wince. Say what you learned. Watch what happens.
That's how it starts. One leader. One vulnerability. One moment of "I'm human too." The rest follows from there.
Break things. Learn things. Celebrate both. Your team — and your systems — will be stronger for it.
Recommended reading

The Legacy of Leadership: What Remains When You Leave
The thing people remember is not your architecture. It is not your processes. It is how you made them feel. Reflections on what actually endures from engineering leadership.

What 3 A.M. Incidents Taught Me That AWS Certifications Never Did
Twenty production incidents reviewed: why understanding beats fixing, what certifications actually train, and the habits that keep a team calm at 3 a.m.

An Engineering Leader's Sustainable Weekly Rhythm
A realistic weekly rhythm that balances strategy, people, and operational work — without burning out or losing yourself in back-to-back meetings.

Comments
No comments yet. Be the first to share your thoughts.