On 10 July 2026 the scheduled Germany fetch started at 00:11 UTC. One generic career page needed a browser. Playwright started at 00:11:58 and never returned. At 00:13:37 the worker logged 97 of 98 companies done. The last one was still in flight.
The container stayed “Up.” The scheduler was blocked on wait_for_fetch_thread() with no timeout. Country fetches did not run again until a manual restart at 15:50 — more than 15 hours later. Chromium children from 00:12 were still running until that restart.
What we saw
| 00:11:53 | Country fetch started, 98 companies |
|---|---|
| 00:11:58 | Generic Playwright scrape started — never returned |
| 00:13:37 | Last log: 97/98 done |
| 00:36:13 | Postgres: failed, “server restarted” (there was no restart) |
| 15:50:30 | docker restart of the fetch worker; scheduler resumed |
The 00:36 failure message was the panel. Panel and worker are separate processes sharing fetch_runs. The panel reaped the row as an orphan while the worker thread was still alive. The database said failed. The worker kept waiting forever.
What it was not
HTTP timeouts were already 15s. Playwright page.goto had a timeout. A stuck renderer does not care. The process watchdog did not exist. The country join did not exist as a ceiling either. One company owned the six-hour cycle.
Layered timeouts
Every blocking boundary gets a hard ceiling. Inner layers must be strictly shorter than outer layers.
HTTP client 15s
Playwright process (whole fallback) 90s
Per-company wait_for 300s
Scheduler join on the country 2700s
Scheduler sleep 6h (only after join returns)On timeout the company is skipped and logged. Remaining companies in that country continue. The scheduler joins, finalizes, and sleeps. It does not wait for a browser that will never come back.
The panel no longer reaps running fetches it does not own. Orphan reap stays on worker bootstrap, after a real restart.
What timeouts do not bound
Timeouts bound how long a unit of work may run. They do not bound how many OS threads and browsers you create while it runs. In September the worker ran out of threads anyway. That write-up is can't start new thread.
The rule
Nothing in the fetch path may block without a bounded timeout. One hung career page must not block the rest of the country, the rest of the cycle, or the scheduler loop. “Up” is not a health signal when the join has no deadline.