On 2 September country fetch stopped starting an OS thread per company. Companies now run on one event loop, bounded by asyncio.Semaphore. The outage was RuntimeError: can't start new thread. That is the number this change is allowed to claim.
A first version of this note compared Germany’s median duration before and after. That comparison is wrong. The healthy week ran at concurrency 4 on the thread pool. The after window runs at concurrency 2 on the event loop. Two knobs moved. Failed outage runs lasted about a second with new_jobs = 0; those are not a speed baseline. Duration is not scored here.
The same error, before and after
Grafana Cloud on this box is host health: disk, available RAM, a probe of /api/health. It does not store per-company fetch outcomes. Counts are Postgres — company_fetch_attempts and fetch_runs — queried on 6 September 2026.
Production picked up FETCH_SCHEDULE_CONCURRENCY=2 at 16:09 UTC on 3 September (first run: Armenia). That is the after cut for “is the thread error gone,” not a claim that two is faster or slower than four.
| Outage attempt errors | 357 / 358 were can't start new thread |
|---|---|
| Attempts since 3 September 16:09 UTC | 3,375 / 3,375 ok |
| Thread errors in that window | 0 |
| Attempt errors in that window | 0 |
| Board flags, 1 September | 186 / 394 companies |
| Board flags, 6 September | 3 / 401 companies |
The three flags left are today’s ATS problems — arculus in Germany, bol and Channable in the Netherlands — dated 6 September. They are not leftover “cannot start a thread” stamps. Successful fetches cleared those.
Did real catalogs finish
Empty scheduler countries (austria, joblet, mauritius, uae, united-state) still fail every cycle with “No catalog.” They failed in August too. They are excluded. What remains is whether a country that has a catalog completed.
| Window | Runs | OK | New jobs |
|---|---|---|---|
| Healthy, 20–27 August | 352 | 345 (98.0%) | 383 |
| Collapse, 30–31 August | 91 | 0 | 0 |
| After, from 3 September 16:09 UTC | 154 | 153 (99.4%) | 160 |
The one miss after the cut is Germany run 3991, 5 September 11:00 UTC: 90/111 done, exit 1, no duration. A scrape that stopped, not can't start new thread. The next Germany cycle was 111/111 again. New jobs per day are back in the pre-outage band, about 37–51, against zero on 30 and 31 August.
Completing versus not completing is the same question in all three rows. How many seconds Germany took at concurrency 4 versus 2 is not.
Still true
Five scheduler countries still have no catalog. Drop them from the schedule. One event loop did not invent that, and it did not fix it.
Concurrency 2 and one Chromium are RAM knobs on a t4g.micro. They shipped in the same deploy. They are not evidence about event-loop speed. Judge the model on whether can't start new thread returns.