Checkout Observability & Reliability
A cross-team operational-excellence initiative that cut P1–P2 incidents ~85% (~15–20/wk → ~2–3/wk) and mitigation time from ~2 hours to ~15 minutes, by fixing ownership first and tooling second.
Context
Checkout was running across multiple new services and migrations, and we were drowning in incidents with no clear ownership. Dashboards mixed signal and noise, errors were classified inconsistently between services, and the person who could fix a given alert was usually the person who had written it.
Problem
Nobody could say with confidence which endpoints and components our team owned. That gap is not something a better threshold fixes: an alert that reaches no specific owner gets ignored, however well it is tuned.
What I did
- Started and led a cross-team operational-excellence initiative, beginning with a map of domain ownership across every endpoint and component.
- Mined postmortem action items for the failures we had already paid for but never systematically fixed.
- Built one consolidated set of Datadog monitors with monitor-based SLOs, replacing scattered per-service dashboards.
- Made a per-monitor triage guide a hard acceptance criterion, so any on-call engineer could self-mitigate instead of escalating to the author.
- Set 3 SLOs across every endpoint the team owned: 99.8% availability, 98.5% success rate, and per-endpoint p95 latency.
- Closed the loop on the roughly monthly incidents the monitors missed by adding custom metrics and guides for those too.
Reliability is an ownership problem before it's a tooling problem
We already had monitoring. What we did not have was agreement about who owned what, so alerts fired into a gap where no specific team was accountable for answering them.
The first work was a map of domain ownership across endpoints and components: which team owned each surface, who was accountable when it broke, and where the ambiguous seams were. Most of that was negotiation between teams rather than engineering, and everything after it depended on the result.
The second input was postmortem action items. Each represented a failure the organization had already suffered and never systematically fixed, so mining them told us what to monitor without guessing.
The triage guide as an acceptance criterion
A monitor was not considered done until it had a triage guide. A monitor without one did not ship.
Before, a page during an incident routed implicitly to whoever understood that subsystem, which meant either an escalation or an hour of an on-call engineer reverse-engineering someone else's alert under pressure. Afterwards the page carried its own answer: what the monitor means, how to confirm it, what to do first, and when to escalate.
monitor:
name: checkout-place-order-success-rate
owner: shop-platform # from the ownership map, never blank
slo:
objective: 98.5 # success rate
window: 30d
triage: # required: no guide, no merge
means: "Place-order success below SLO for the rolling window."
confirm: "Break down by payment provider and country."
first_action: "Check provider status; if isolated, route the funnel to legacy."
escalate_when: "Error budget burn > 2x after mitigation."Consolidating onto one set of monitors with monitor-based SLOs followed from the same problem. Scattered per-service dashboards let every team define "healthy" differently, so nobody could compare anything or trust a number they hadn't built themselves.
Closing the loop on what the monitors missed
A monitor set built from known failures catches known failures. Roughly once a month an incident got through that no monitor had anticipated.
Each one produced a custom metric and its own triage guide, so the blind spots shrank on a schedule instead of accumulating. Proactive alerting on invalid products entering the cart came out of this loop, after a pricing problem originating in undocumented legacy code had to be found the hard way.
The same signals doubled as the go/no-go view during platform cutovers. The migrations were running concurrently, and a shared error taxonomy let a cutover decision come from a small set of numbers everyone already trusted.
Result
- P1–P2 incidents cut ~85%, from ~15–20/week to ~2–3/week, from ownership mapping plus monitors built on real postmortem history.
- MTTR (to mitigation) cut from ~2 hours to ~15 minutes, because the triage guide requirement meant the on-call engineer who got paged could act instead of escalate.
- 3 SLOs across every endpoint the team owned: 99.8% availability, 98.5% success rate, per-endpoint p95 latency, all enforced as monitor-based objectives.
- Safer cutovers: go/no-go decisions backed by a small set of trusted, consistently classified metrics.
- Observability and SLO standards other commerce services adopted, which is the part that outlasted the initiative.
Tech
- OpenTelemetry
- Prometheus
- Micrometer
- Datadog
- PagerDuty
- Sentry
- AWS
- EKS