mirror of
https://github.com/permissionlesstech/bitchat.git
synced 2026-07-25 01:45:20 +00:00
The two stress tests I added (test_concurrentUpsertsAndForceSaveDoNotRaceOrHang and test_manyManagersDeinitDoNotWedgeTeardown) spawned units of work that the test could not deterministically join before returning: - test_manyManagersDeinitDoNotWedgeTeardown created 200 managers that each scheduled a fire-and-forget queue.async(.barrier) save on the manager's own private queue; the test has no handle to await that per-manager barrier work, so a large backlog of it could still be executing after the test returned. - test_concurrentUpsertsAndForceSaveDoNotRaceOrHang spawned 32 DispatchQueue.global().async workers, each also scheduling per-manager barrier saves; the DispatchGroup only joined the worker loops, not the manager's internal barrier work. Under --enable-code-coverage (CI only), LLVM writes .profraw from an atexit handler; if instrumented worker threads are still live during that dump the process deadlocks at exit — matching the CI signature exactly: all 145 tests start, then a ~5-minute wedge, then Killed:9. It reproduced only on the constrained CI runner, not locally (18 cores drained the backlog before exit), which is why earlier local runs looked clean. The production fix is already verified: ThreadSanitizer is clean on the identity + announce-handler suites (no data race, no re-entrant deadlock), and the deterministic unit tests cover the signing-key pin refusal, persistence across re-init, and the persisted-pin fallback. A stress test that destabilizes CI is worse than no stress test, so both are removed along with the now-unused LockedKeychain double. Verified: `time swift test --parallel --enable-code-coverage --skip PerformanceBaselineTests` is green 6x and the process exits ~2.9s after the last test (tests run in ~1.5s); no teardown wedge under coverage even with LIBDISPATCH_COOPERATIVE_POOL_STRICT=1 and a single worker. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>