Building Documentation Habits That Actually Stick
Why most documentation initiatives fail and a practical system for creating a self-sustaining documentation culture in engineering teams.

Every engineering organization I have joined has had the same problem: documentation that is either nonexistent or so outdated it is actively harmful. The pattern is predictable — someone launches a documentation initiative, the team writes docs for 2-3 weeks with enthusiasm, then maintenance drops off, docs rot, and within six months everyone is back to Slack questions and tribal knowledge. As teams grow past the informal communication threshold described in scaling engineering teams, this problem becomes existential.
I have finally cracked this problem, not with better tooling or more documentation mandates, but by redesigning the incentive structure around documentation. The key insight: documentation fails when it is a separate activity. It succeeds when it is embedded in existing workflows.
Why Documentation Initiatives Fail
The fundamental problem is that writing documentation has no immediate reward for the author. The person who benefits most from documentation is the future reader — a new hire in 3 months, a teammate debugging at 2 AM, or your future self who has forgotten why this design decision was made.
This creates a classic tragedy of the commons: everyone benefits from documentation, but nobody is incentivized to produce it. Add to this:
- The freshness problem: Documentation is most accurate when the code is written but least needed at that moment. By the time someone needs it, it is stale.
- The discovery problem: Even when good docs exist, people cannot find them. Documentation scattered across wikis, README files, Notion pages, and Confluence spaces is documentation that might as well not exist.
- The ownership problem: When everyone owns documentation, nobody maintains it. Shared ownership means shared neglect.
The Three-Layer Documentation System
After iterating across three organizations, I have settled on a three-layer system that addresses each failure mode:
┌─────────────────────────────────────────────────────────┐
│ Documentation Layers │
├─────────────────────────────────────────────────────────┤
│ │
│ Layer 3: Knowledge Base (How things work) │
│ ├── Architecture overviews │
│ ├── System design documents │
│ ├── Onboarding guides │
│ └── Updated: Quarterly review cycle │
│ │
│ Layer 2: Decision Records (Why we chose this) │
│ ├── Architecture Decision Records (ADRs) │
│ ├── RFC documents │
│ ├── Post-mortem reports │
│ └── Updated: At decision time (immutable after) │
│ │
│ Layer 1: Code-Adjacent Docs (What this does) │
│ ├── README files per service/module │
│ ├── API documentation (auto-generated where possible) │
│ ├── Inline code comments for non-obvious logic │
│ └── Updated: With every code change (PR requirement) │
│ │
└─────────────────────────────────────────────────────────┘
Each layer has a different maintenance cadence and different ownership model. This is critical — trying to maintain all documentation at the same cadence is what causes burnout.
Layer 1: Code-Adjacent Documentation
This is documentation that lives next to the code and is maintained as part of code changes. It includes:
- A
README.mdin every service and significant module directory - OpenAPI specs or GraphQL schemas that auto-generate API docs
The key principle: Layer 1 documentation updates must be a code review requirement. If a PR changes behavior, the reviewer checks that the README reflects the new state.
- Inline comments explaining non-obvious business logic
- Configuration files with comments explaining each option
The enforcement mechanism: PR reviews check that code-adjacent docs are updated when the code changes. This is a review checklist item, not a suggestion. If you change a service's behavior, you update its README. If you add an API endpoint, it appears in the OpenAPI spec.
Why this works: documentation updates happen in the same PR as code changes, when the author has full context. The marginal effort is 5-10 minutes per PR, not a separate writing session.
Layer 2: Decision Records
Decision records capture why something was done, not what it does. They are written at decision time and never modified afterward (they can be superseded by new decisions). This includes:
- Architecture Decision Records (ADRs): Why we chose this database, this framework, this deployment strategy
- RFCs: Proposals for significant changes, with discussion and resolution
- Post-mortems: What happened during incidents, what we learned, what we changed
The enforcement mechanism: Significant technical decisions cannot be merged without an ADR. The threshold: any change that affects more than one team, introduces a new technology, or changes a data model requires a written decision record.
Why this works: decision records are point-in-time artifacts. They never go stale because they capture a moment. The context they provide is invaluable months later when someone asks "why did we use Kafka instead of SQS?"
Layer 3: Knowledge Base
The knowledge base contains high-level understanding documents: system architecture overviews, onboarding guides, operational runbooks. These are the most expensive to maintain and the most likely to go stale.
The enforcement mechanism: Quarterly documentation review cycle. Every team spends one sprint day per quarter reviewing and updating their knowledge base section. The review is tracked and the output is visible.
Why this works: accepting that knowledge base docs will lag reality by up to a quarter and explicitly scheduling maintenance is more sustainable than expecting continuous maintenance that never happens.
Making Documentation Discoverable
Documentation that cannot be found does not exist. We solved discovery with three principles:
Single search. All documentation is searchable from one place. Whether it is in a README, an ADR, or the knowledge base, engineers search one tool and find it. We chose to index everything into our documentation platform rather than spreading across tools.
Consistent structure. Every service README follows the same template. Every ADR uses the same format. Every runbook has the same sections. Consistency means engineers know where to look without thinking.
Link from code. Critical code paths include comments linking to relevant documentation. A database query that looks suboptimal might link to the ADR explaining why it is written that way. A configuration value links to the runbook that explains when to change it.
The Documentation Champions Model
Instead of making everyone equally responsible for documentation (which means nobody is responsible), we appointed documentation champions:
- One champion per team (rotates quarterly)
- 10% of their time is explicitly allocated to documentation
- They review PRs for documentation completeness
- They run the quarterly knowledge base review for their team
- They triage documentation requests from other teams
Champions are not the only people who write docs — they are the people who ensure docs get written. It is a coordination role, not a production role.
Measuring Documentation Health
You cannot improve what you do not measure. We track:
| Metric | Target | How Measured |
|---|---|---|
| README coverage | 100% of services have current README | Automated scan |
| ADR coverage | Every significant decision has an ADR | Manual review quarterly |
| Documentation freshness | <90 days since last update per service | Git timestamps |
| New hire onboarding time | <2 weeks to first PR | HR tracking |
| Documentation search success | >80% queries return useful results | Search analytics |
The onboarding metric is the most meaningful business outcome. When documentation is working, new engineers become productive faster because they can self-serve answers instead of interrupting teammates.
Overcoming Resistance
Engineers resist documentation for legitimate reasons. Address them directly:
"It will be outdated tomorrow." That is why we tier documentation by maintenance cadence and embed Layer 1 docs in the code change workflow.
"Nobody reads the docs anyway." That is a discovery problem, not a production problem. Fix search and linking, and reading will follow.
"I would rather write code." Layer 1 docs take 5-10 minutes per PR. The time you save in not answering Slack questions pays this back within a week.
"Documentation is someone else's job." No, it is not. But it is the champion's job to make sure it happens and to keep the barrier low.
Practical Implementation Steps
If you are starting from scratch:
Month 1: Establish the template for Layer 1 (service READMEs) and mandate it for all new services. Retrofit the 3 most critical existing services.
Month 2: Introduce ADRs for all new technical decisions. Write ADRs retroactively for the 5 most important existing decisions that new hires always ask about.
Month 3: Appoint documentation champions. Run the first quarterly knowledge base review to establish baselines.
Month 4-6: Measure, iterate, and expand. Add more services to Layer 1 coverage. Build the searchability infrastructure. Celebrate teams with high documentation health scores.
Key Takeaways
Documentation culture that sticks requires systemic design, not motivational speeches:
- Structure documentation in layers with different maintenance cadences
- Embed documentation updates in existing code change workflows (Layer 1)
- Make decision records immutable point-in-time artifacts that never go stale (Layer 2)
- Accept knowledge base staleness and schedule explicit quarterly maintenance (Layer 3)
- Solve discovery through single search, consistent structure, and code-to-doc links
- Appoint documentation champions with explicit time allocation
- Measure documentation health through coverage, freshness, and onboarding time
The teams that sustain documentation culture are not the ones with the best writers or the most discipline. They are the ones that designed the system so documentation happens as a natural byproduct of work, not as an additional burden on top of it. Tools like Kiro for DevOps can now automate parts of Layer 1 documentation by generating docs from code changes — turning AI agents into documentation co-authors.
Recommended reading

The Legacy of Leadership: What Remains When You Leave
The thing people remember is not your architecture. It is not your processes. It is how you made them feel. Reflections on what actually endures from engineering leadership.

What 3 A.M. Incidents Taught Me That AWS Certifications Never Did
Twenty production incidents reviewed: why understanding beats fixing, what certifications actually train, and the habits that keep a team calm at 3 a.m.

An Engineering Leader's Sustainable Weekly Rhythm
A realistic weekly rhythm that balances strategy, people, and operational work — without burning out or losing yourself in back-to-back meetings.

Comments
No comments yet. Be the first to share your thoughts.