Current Updated 10:52 UTC, 6 minutes ago
Two things are broken
The ingest API has been returning errors since 09:14 UTC. Webhook delivery is paused because it depends on ingest. Seven other services are working normally. Nobody has fixed it yet.
Every service, right now
Nine services, one row each, no summary number at the top. A single figure covering nine services is an average, and an average is how a page reports 78% while the thing you use is down.
| Service | State | Since | 30d uptime | What that means |
|---|---|---|---|---|
| Ingest API | Erroring | 09:14 UTC | 97.412% | Returns 500 on about one write in three. |
| Webhook delivery | Paused | 09:21 UTC | 98.006% | Queued, not dropped. They will arrive late. |
| Query API | Normal | n/a | 99.987% | Reads are unaffected by the ingest fault. |
| Dashboard | Normal | n/a | 99.951% | Charts will show a gap for the outage window. |
| Auth | Normal | n/a | 99.999% | Sign in works. |
| Exports | Normal | n/a | 99.872% | Scheduled exports ran at 08:00 as usual. |
| Alerting | Normal | n/a | 99.944% | Alerts on ingest data are late by the queue depth. |
| Billing | Normal | n/a | 100.000% | Nothing here has broken in 30 days. |
| Status page | Normal | n/a | 99.998% | Hosted away from everything above it. |
What we refuse to call it
Incident language drifts toward whatever is least alarming to write. These are the phrases we have banned from this page and the sentence each one gets replaced with.
| Phrase | What we write instead |
|---|---|
| Elevated error rates | About one write in three returns a 500. |
| Degraded performance | Queries take 4 seconds. They usually take 90ms. |
| We are aware of an issue | It broke at 09:14. We saw it at 09:14. |
| A small number of customers | 1,204 accounts. Yours is one of them. |
| We apologise for any inconvenience | Nothing. The apology goes in the postmortem. |
Three rules this page runs on
These three blocks sit on nine of the twelve columns. The three they leave empty are drawn rather than implied, so the grid is visible in the space it has left over. That is the same argument the rest of the page makes about its own structure.
Green is not the default
A service is normal when a check says so within the last 60 seconds. A check that has not reported is unknown, and unknown is written on the page as unknown.
The clock starts when it broke
Elapsed time counts from the first failed request, not from the moment somebody opened the incident. The gap between those two is published in the postmortem.
Uptime keeps its decimals
99.9% and 99.94% are different promises and get written differently. Rounding up to the nearest comfortable figure is how a number stops being a measurement.
Be told when it breaks
One message when a service changes state, and one when it is fixed. No weekly digest, no product news, no reactivation mail six months after you unsubscribe.