Tags: web-dev concept

Service Level Objectives

Date: 2026-08-17


A published target for reliability, and an explicit budget for failing to meet it. The technical content is small; the point is political — it converts “is the site reliable enough” from an argument about feelings into a number both engineering and the business agreed to in advance.


A service level objective (SLO) is a target value for a measured aspect of reliability over a time window — for example, 99.5% of checkout requests succeeding in under a second over 28 days.

The three terms

What it isExample
SLI — service level indicatorThe measurementProportion of checkout requests served in under 1s without a 5xx
SLO — service level objectiveYour internal target for it99.5% over a rolling 28 days
SLA — service level agreementA contract with money attached99.0%, or the customer gets a credit

Set the SLO tighter than any SLA, so you breach your own target well before you breach a contract. SLAs are a legal artefact; SLOs are the operational one, and most teams should only have the second.

The error budget

The inversion is what makes it useful.

SLO 99.5% over 28 days
    → allowed failure = 0.5%
    → 28 days × 24h × 60m = 40,320 minutes
    → budget = 201 minutes of failure per 28 days

spent so far this period      48 min   ┃████░░░░░░░░░░░░░░░░┃  24%
remaining                    153 min

    ↑ this number is the decision-making instrument

The budget converts reliability into a resource that gets spent. Budget remaining means ship — take the risky migration, run the experiment, roll out faster. Budget exhausted means stop feature work and fix reliability, by prior agreement rather than by argument.

100% is the wrong target and saying so is the point. Every nine costs disproportionately more, the user’s own network is less reliable than that anyway, and a team with no error budget can never ship anything. An SLO that’s never been breached is set too loose to be informative.

Choosing an SLI

The measurement is where this succeeds or fails.

  • Measure what the customer experiences, not what the server did. Server-side latency excludes DNS, TLS, the network and rendering — and those are most of what a customer perceives as slow — Real User Monitoring, Field vs Lab Data
  • Measure at the right place. Load balancer logs catch failures your application never saw; application metrics can’t report an outage that stopped the application
  • Per critical journey, not per service. “Checkout availability” is meaningful; “the pricing service’s availability” only matters through its effect on the first
  • Use good-events / valid-events, a ratio, and be explicit about what’s excluded. Bots, health checks and 4xx caused by the client usually don’t count — write the exclusions down or they become the argument later
  • Percentiles, never averages. A mean of 400ms with a p99 of 9s is a broken service for one user in a hundred, and the mean conceals it — Percentiles in Performance, Percentiles and Quantiles

Two or three SLIs per critical journey — typically availability, latency, and sometimes correctness or freshness. More than that and nobody tracks any of them.

Setting the number

Not aspirationally. Measure current performance for a month, then set the SLO at or slightly above what you already achieve, and tighten later if it’s genuinely insufficient. An SLO adopted at 99.9% by a service currently doing 99.2% is breached on day one, which teaches everyone that breaching it is normal.

Sanity checks on the arithmetic:

99%      → 7h 12m of failure per 28 days      generous
99.5%    → 3h 21m                             realistic for most commerce
99.9%    → 40m                                expensive; needs real redundancy
99.99%   → 4m                                 a serious engineering programme

Ask what the customer would actually notice. If checkout being down for ten minutes at 3am costs almost nothing, the SLO shouldn’t be priced as though it costs a lot.

What it’s for politically

This is the part worth being explicit about:

  • It ends the “is it stable enough” argument by settling it once, numerically, in advance
  • It gives engineering a defensible reason to say no to a release — not a preference, an agreed threshold that has been crossed
  • It gives the business a defensible reason to say yes to speed while the budget holds, which is the half engineers tend to forget and the half that gets the whole thing adopted
  • It prices reliability work. “Nine more minutes of downtime allowed this month” is a fundable statement; “the system feels fragile” is not

None of this works if the budget is ignored when exhausted. The first time feature work continues through a blown budget, the mechanism is decoration — the commitment has to be made before it costs anything.

Where it interacts

  • Alerting — burn-rate alerts are the mature form: page on a fast burn, ticket on a slow one, rather than on a static threshold
  • Progressive Delivery — remaining budget is a sensible gate on how aggressively to ramp a rollout
  • Incident Response — incident severity maps naturally to budget consumed, and the postmortem records the spend
  • Multi-Tenancy — an aggregate SLO can look healthy while one tenant is permanently broken; measure per tenant where isolation is promised
  • Guardrail Metrics — the experimentation equivalent: an agreed line that stops a rollout regardless of what the primary metric says