Ingest API erroring since 09:14 UTC Webhook delivery paused 7 of 9 services normal Next update 11:30 UTC

Current Updated 10:52 UTC, 6 minutes ago

Two things are broken

The ingest API has been returning errors since 09:14 UTC. Webhook delivery is paused because it depends on ingest. Seven other services are working normally. Nobody has fixed it yet.

  • Started 09:14 UTC
  • Elapsed 1h 38m
  • Cause known
  • Fix in progress

Every service, right now

Nine services, one row each, no summary number at the top. A single figure covering nine services is an average, and an average is how a page reports 78% while the thing you use is down.

Service status, 10:52 UTC
Service State Since 30d uptime What that means
Ingest API Erroring 09:14 UTC 97.412% Returns 500 on about one write in three.
Webhook delivery Paused 09:21 UTC 98.006% Queued, not dropped. They will arrive late.
Query API Normal n/a 99.987% Reads are unaffected by the ingest fault.
Dashboard Normal n/a 99.951% Charts will show a gap for the outage window.
Auth Normal n/a 99.999% Sign in works.
Exports Normal n/a 99.872% Scheduled exports ran at 08:00 as usual.
Alerting Normal n/a 99.944% Alerts on ingest data are late by the queue depth.
Billing Normal n/a 100.000% Nothing here has broken in 30 days.
Status page Normal n/a 99.998% Hosted away from everything above it.

What we refuse to call it

Incident language drifts toward whatever is least alarming to write. These are the phrases we have banned from this page and the sentence each one gets replaced with.

Banned phrasing and its replacement
Phrase What we write instead
Elevated error rates About one write in three returns a 500.
Degraded performance Queries take 4 seconds. They usually take 90ms.
We are aware of an issue It broke at 09:14. We saw it at 09:14.
A small number of customers 1,204 accounts. Yours is one of them.
We apologise for any inconvenience Nothing. The apology goes in the postmortem.

Three rules this page runs on

These three blocks sit on nine of the twelve columns. The three they leave empty are drawn rather than implied, so the grid is visible in the space it has left over. That is the same argument the rest of the page makes about its own structure.

Green is not the default

A service is normal when a check says so within the last 60 seconds. A check that has not reported is unknown, and unknown is written on the page as unknown.

The clock starts when it broke

Elapsed time counts from the first failed request, not from the moment somebody opened the incident. The gap between those two is published in the postmortem.

Uptime keeps its decimals

99.9% and 99.94% are different promises and get written differently. Rounding up to the nearest comfortable figure is how a number stops being a measurement.

Be told when it breaks

One message when a service changes state, and one when it is fixed. No weekly digest, no product news, no reactivation mail six months after you unsubscribe.