Severity levels
Fill this in: replace the descriptions above with thresholds somebody can apply without a meeting — a percentage of users, a named flow, a revenue path.
Grading rules#
- Grade on impact, not cause. A one-line config error that logs everybody out is a SEV1.
- When in doubt, grade up. Downgrading is easy and free; upgrading late costs the time nobody spent responding.
- Data risk overrides everything. Possible exposure, loss or corruption of customer data is your highest severity whether or not anything looks broken.
Severity drives the response, not the other way around. If a SEV2 needs the whole team, it was a SEV1 — change the grade rather than quietly running a bigger response than the grade implies.
What each level obliges#
| Severity | Response | Communication | Write-up |
|---|---|---|---|
| SEV1 | Page immediately, all hands available | Status page, updates every 30 minutes | Always |
| SEV2 | Page the primary | Internal channel, hourly | Always |
| SEV3 | Working hours | Ticket | If it recurs |
| SEV4 | Backlog | None | No |