9 Engineering Rules That Production Outages Teach the Hard Way

| | 4 min read

Summary: nine engineering rules at a glance

  • Prove it in production. Staging is a guess, so test failure on purpose and use an error budget to decide how much risk to take.
  • Make reliability a shared decision. SLOs only work when engineering and the business agree on them, and on-call should be boring, not heroic.
  • Treat platforms as products and keep architecture honest. Fix what developers work around, and don’t distribute a problem you haven’t solved.
  • Run incidents for clarity. Blameless postmortems record what really happened, and a good incident commander asks questions instead of diving into logs.
  • Design for the next person. Complexity is a debt that your successor pays.

Tools change every year. Kubernetes replaced the scripts that came before it, and something will replace Kubernetes. Yet the same lessons keep coming back after every serious outage. Here are nine engineering rules that have outlasted each wave of tooling, and why each one holds up.

Engineering rules for proving reliability

1. If it hasn’t failed in production, you don’t know it works

Staging is a hypothesis about production. It never has the same traffic, data shape or odd dependencies. The only real test is production, so the choice is between breaking things on purpose while the stakes are low (game days, fault injection) and finding out on the system’s own schedule.

2. An error budget tells you what you can do next

Uptime reports history. An error budget turns it into a decision. A 99.9% target allows roughly 8 hours 45 minutes of downtime a year, and that time will be spent somehow. The useful question is who gets to decide: risky launches, migrations or experiments. When the budget is healthy, ship faster. When it is spent, slow down and fix reliability.

3. On-call should be boring

If engineers regularly save the day at 3am, that is a sign the system depends on individuals. Heroics make a good story and a bad design. Boring on-call means alerts that matter, runbooks that work and failures that heal themselves.

4. SLOs are a conversation, not a dashboard

The exact number matters less than the agreement behind it. A service level objective forces engineering and the business to say out loud how reliable something needs to be and what it is worth. If the SLO only lives on a team dashboard, nobody outside the team is making trade-offs with it.

Engineering rules for platforms and architecture

5. Your platform is a product

Internal tools rarely fail loudly. People work around them. If engineers bypass your platform, treat that as a bug report: the platform is slower, harder or less clear than the alternative. Developers are customers, so talk to them, measure adoption and fix the friction.

6. Don’t distribute a problem you haven’t solved

Orchestration does not repair a weak design, it scales it. Microservices without clear domain boundaries do not remove coupling, they move it into network calls where it is harder to see and debug. Get the boundaries right in a simpler setup first. (Related: Clean Architecture and the Command Bus: When It Helps and When It Adds Complexity.)

Engineering rules for running an incident

7. The postmortem is your most valuable engineering document

A good postmortem puts what people assumed next to what actually happened, and that gap is where the learning is. It only works in a blameless culture. People share the real sequence of events when they know honesty will not be punished. For a related habit, see Systematic Debugging: Reproduce, Isolate, Test and Verify.

8. The best incident commander asks more than they explain

On a bridge call, the commander’s job is clarity, not showing expertise. Ask who owns the next step, what is known, what is a guess and what changes if we wait. If the commander is deep in the logs, nobody is coordinating, so hand the investigating to someone else.

Engineering rules for the long term

9. Complexity is a debt your successor pays

An abstraction that feels obvious to its author is often a puzzle to the next reader. Someone will debug your code at 2am, possibly long after you have left the team, and they will not have your context. Choose the simpler design, name things plainly, and write down why decisions were made.

What these engineering rules have in common

Most of these rules are about people and incentives rather than technology. Error budgets and SLOs make trade-offs explicit. Boring on-call, blameless postmortems and calm incident commanders remove the need for heroes. Restraint with architecture and complexity protects whoever comes next.

If you only act on one this week, pick the one your team skips most often, and ask when you last tested that part of your system for real. Which of the nine would you add to or challenge?

Subscribe to Our Newsletter

We don’t spam! Read our privacy policy for more info.