Skip to content
aviral gupta

// Lesson 1 of 1 · ~25 min · Intermediate

Blameless incidents and DORA metrics

After this lesson you can decide when an incident needs a postmortem, write one without blame, and read your team’s delivery through the DORA metrics.

You will be able to

  • Apply the SRE book’s postmortem triggers
  • Turn blame into system-focused findings
  • Match each DORA metric to what it measures
  1. Warm-up · Activity 1 of 7

    Warm-up: after a bad deploy, what is usually the fastest way to restore service?

  2. Predict · Activity 2 of 7

    Decide: which of these incidents meet the postmortem triggers listed in the Google SRE book? Pick all that apply.

    Select all that apply.

  3. Practice · Activity 3 of 7

    Which sentence belongs in a blameless postmortem?

  4. Practice · Activity 4 of 7

    Match each DORA metric to what it measures.

  5. Practice · Activity 5 of 7

    Last month: 40 deployments, 6 of which needed a rollback or hotfix. What is the change fail rate, as a percentage?

    %
  6. Brain teaser · Activity 6 of 7

    Brain teaser. A manager sets a target: “Cut change fail rate in half this quarter.” The team hits it by deploying once a month instead of daily. What happened to the other metrics?

  7. Apply · Activity 7 of 7

    Mini-task. Write a one-page postmortem outline for: “A config typo took the API down for 25 minutes; on-call rolled back.” Use headings.

    Check your work against this list

Exit ticket

5 questions, no hints. Score 80% or more to complete the lesson.

Finish every activity above to unlock the exit ticket.

Report a problem

Spotted something wrong or unclear? Say what, and it will be checked and fixed.

#

At least 20 characters.

Only if you want a reply.

Key ideas

Blameless means fixing systems, not people

The Google SRE book describes a blameless postmortem as one that focuses on the contributing causes of an incident without pointing at any individual or team for bad or inappropriate behaviour. It assumes everyone acted with good intentions on the information they had. People who fear punishment hide the details you need.

Agree the triggers in advance

Common triggers listed in the SRE book include user-visible downtime or degradation beyond a threshold, data loss of any kind, on-call engineer intervention such as a rollback or rerouting traffic, a resolution time above a threshold, and a monitoring failure, which usually means a person found the incident rather than an alert. Any one is enough. Deciding these beforehand removes the awkward “does this one count?” debate, and any stakeholder can still ask for a postmortem.

Measure the delivery system

DORA’s software delivery metrics are change lead time, deployment frequency, failed deployment recovery time, change fail rate and deployment rework rate. They describe the team’s delivery system, not individuals, so use them to find bottlenecks, never to rank people.

Sources

Last reviewed September 28, 2026