Engineering

Production measurement

One event loop: the thread error is gone

357 thread errors in the outage · 0 in 3,375 attempts after

On 2 September country fetch stopped starting an OS thread per company. Companies now run on one event loop, bounded by asyncio.Semaphore. The outage was RuntimeError: can't start new thread. That is the number this change is allowed to claim.

A first version of this note compared Germany’s median duration before and after. That comparison is wrong. The healthy week ran at concurrency 4 on the thread pool. The after window runs at concurrency 2 on the event loop. Two knobs moved. Failed outage runs lasted about a second with new_jobs = 0; those are not a speed baseline. Duration is not scored here.

The same error, before and after

Grafana Cloud on this box is host health: disk, available RAM, a probe of /api/health. It does not store per-company fetch outcomes. Counts are Postgres — company_fetch_attempts and fetch_runs — queried on 6 September 2026.

Production picked up FETCH_SCHEDULE_CONCURRENCY=2 at 16:09 UTC on 3 September (first run: Armenia). That is the after cut for “is the thread error gone,” not a claim that two is faster or slower than four.

Same failure string. Outage window is 28 August through 1 September. After starts 3 September 16:09 UTC.
Outage attempt errors357 / 358 were can't start new thread
Attempts since 3 September 16:09 UTC3,375 / 3,375 ok
Thread errors in that window0
Attempt errors in that window0
Board flags, 1 September186 / 394 companies
Board flags, 6 September3 / 401 companies

The three flags left are today’s ATS problems — arculus in Germany, bol and Channable in the Netherlands — dated 6 September. They are not leftover “cannot start a thread” stamps. Successful fetches cleared those.

Did real catalogs finish

Empty scheduler countries (austria, joblet, mauritius, uae, united-state) still fail every cycle with “No catalog.” They failed in August too. They are excluded. What remains is whether a country that has a catalog completed.

Real catalogs only. Collapse is 30–31 August, when every scheduled country run failed in about a second. After is the same question with the new loop.
WindowRunsOKNew jobs
Healthy, 20–27 August352345 (98.0%)383
Collapse, 30–31 August9100
After, from 3 September 16:09 UTC154153 (99.4%)160

The one miss after the cut is Germany run 3991, 5 September 11:00 UTC: 90/111 done, exit 1, no duration. A scrape that stopped, not can't start new thread. The next Germany cycle was 111/111 again. New jobs per day are back in the pre-outage band, about 37–51, against zero on 30 and 31 August.

Completing versus not completing is the same question in all three rows. How many seconds Germany took at concurrency 4 versus 2 is not.

Still true

Five scheduler countries still have no catalog. Drop them from the schedule. One event loop did not invent that, and it did not fix it.

Concurrency 2 and one Chromium are RAM knobs on a t4g.micro. They shipped in the same deploy. They are not evidence about event-loop speed. Judge the model on whether can't start new thread returns.

This worker is why the catalog of visa-sponsored software jobs refreshes about every six hours. How Kuchup finds roles · More engineering notes