Engineering

Production incident

Playwright hung for 15 hours and the worker still looked Up

97/98 companies done · scheduler blocked 15+ hours · no join timeout

On 10 July 2026 the scheduled Germany fetch started at 00:11 UTC. One generic career page needed a browser. Playwright started at 00:11:58 and never returned. At 00:13:37 the worker logged 97 of 98 companies done. The last one was still in flight.

The container stayed “Up.” The scheduler was blocked on wait_for_fetch_thread() with no timeout. Country fetches did not run again until a manual restart at 15:50 — more than 15 hours later. Chromium children from 00:12 were still running until that restart.

What we saw

Germany run 341, 10 July 2026, UTC
00:11:53Country fetch started, 98 companies
00:11:58Generic Playwright scrape started — never returned
00:13:37Last log: 97/98 done
00:36:13Postgres: failed, “server restarted” (there was no restart)
15:50:30docker restart of the fetch worker; scheduler resumed

The 00:36 failure message was the panel. Panel and worker are separate processes sharing fetch_runs. The panel reaped the row as an orphan while the worker thread was still alive. The database said failed. The worker kept waiting forever.

What it was not

HTTP timeouts were already 15s. Playwright page.goto had a timeout. A stuck renderer does not care. The process watchdog did not exist. The country join did not exist as a ceiling either. One company owned the six-hour cycle.

Layered timeouts

Every blocking boundary gets a hard ceiling. Inner layers must be strictly shorter than outer layers.

HTTP client                         15s
Playwright process (whole fallback) 90s
Per-company wait_for               300s
Scheduler join on the country     2700s
Scheduler sleep                      6h  (only after join returns)

On timeout the company is skipped and logged. Remaining companies in that country continue. The scheduler joins, finalizes, and sleeps. It does not wait for a browser that will never come back.

The panel no longer reaps running fetches it does not own. Orphan reap stays on worker bootstrap, after a real restart.

What timeouts do not bound

Timeouts bound how long a unit of work may run. They do not bound how many OS threads and browsers you create while it runs. In September the worker ran out of threads anyway. That write-up is can't start new thread.

The rule

Nothing in the fetch path may block without a bounded timeout. One hung career page must not block the rest of the country, the rest of the cycle, or the scheduler loop. “Up” is not a health signal when the join has no deadline.

This worker is why the catalog of visa-sponsored software jobs refreshes about every six hours. How Kuchup finds roles · More engineering notes