# The p99 that was a cron job

- Published: 2026-07-28
- Reading time: 3 min
- Tags: PostgreSQL, Observability, Latency

For three weeks one endpoint’s p99 went to nine seconds every night at 02:10 and recovered on its own by 02:40. Nobody was awake to see it and the dashboard averaged it away. The fix was one word of SQL; the finding was the three weeks.

The alert was never loud enough to wake anybody. A single endpoint — the one the mobile app calls on launch — went from a 180 ms p99 to about nine seconds, every night, and was back to normal before anyone in Tehran had breakfast. It did this for three weeks.

First mistake

### The dashboard was averaging it away

Our latency panel was a one-hour rolling mean . Thirty minutes of nine-second responses inside a day of 180 ms responses moves a daily mean by a few milliseconds. It moves the p99 by nine seconds. We had the percentile in the metrics store the whole time; the panel simply did not draw it. The first change was not a fix. It was replacing every mean on that board with p50, p95 and p99 on the same axis, so a tail that separates from the median is visible as a shape rather than as a number somebody has to compare against memory.

What it was

### A nightly job holding a lock it did not need

At 02:10 a job rebuilt a materialised view used by the app-launch query. It ran REFRESH MATERIALIZED VIEW without CONCURRENTLY , which takes an ACCESS EXCLUSIVE lock — and every read of that view queues behind it. The refresh took about half an hour because the view had grown by a factor of forty since the day it was written. The job was two years old. It had been correct on the day it was written and had been quietly wrong ever since the table it read got big.

```
-- before: takes ACCESS EXCLUSIVE, every reader queues behind it
REFRESH MATERIALIZED VIEW app_launch_summary;

-- after: readers keep the previous snapshot while the new one builds.
-- needs a UNIQUE index on the view, which is the whole cost of the fix.
REFRESH MATERIALIZED VIEW CONCURRENTLY app_launch_summary;
```

One line of SQL was the bug. The thing worth writing down is that a nightly nine-second p99 survived three weeks of a team looking at its own dashboards every day — because the panel answered a question nobody was asking.


---

Source: https://larsima.com/nl/insights/p99-was-a-cron-job
Organisation: LARSIMA (شرکت فن‌آوران توسعه لار سیما), registered in Iran, no. 318515, since 2007-12-02.
