mirror of
https://github.com/permissionlesstech/bitchat.git
synced 2026-07-27 06:45:22 +00:00
* Deflake two iOS-sim CI tests with unbounded timing assumptions Both failed on main-adjacent CI runs during this session's PRs. Neither was a product bug; both asserted things about real time that a loaded runner is under no obligation to honour. **NetworkReachabilityGateTests: a wall-clock upper bound.** test_monitor_duplicateUpdatesDoNotPostponeOfflineCommit slept 500ms for real, then asserted total elapsed time was under 1.4s to prove a duplicate mid-window had not restarted the 1.0s debounce. One CI run took 3.75s. No wall-clock bound can separate "deadline preserved" from "runner is slow", because Task.sleep and the asyncAfter flush are both real time and neither is bounded above. The deadline property was already covered deterministically one level down: test_debounce_duplicateObservationsPreservePendingDeadline drives ReachabilityDebounce with injected timestamps and checks pendingRemaining directly. So the monitor test now asserts only what needs a real monitor — that a duplicate still yields exactly one committed false through the debounce — with an injected clock for the arithmetic and a generous liveness budget. Renamed to say what it actually checks. No coverage lost, and it runs in 0.14s instead of ~3.8s because the real sleep is gone. **NoiseEncryptionServiceTests: injected timeouts that also arm during setup.** #1483 diagnosed and fixed exactly this in the quarantine-restore test, but two sibling tests kept the shape. Their injected ordinaryResponderHandshakeTimeout (0.04 and 0.06) also arms during the establishSessions setup handshake, where bob is the responder — so a preempted runner fires it mid-setup, tears down the half-open responder, and message 3 gets answered as a fresh initiation. The CI failure named setup's own `#expect(finalMessage == nil)` seeing a 96-byte message 2, which is precisely the signature #1483 recorded. Raised both to 1.0s, matching #1483's remedy: the scenarios still need the responder timeout to fire, and it still does, just with room for setup to complete first. Thin 1-second waitUntil budgets in the file now use the shared TestConstants.longTimeout; every one of them backs a positive assertion, so waitUntil still returns the moment the condition holds and nothing gets slower in the passing case. No assertion weakened in either file. Verified 3x sequentially and 3x with all 18 cores saturated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Raise two more starvation-prone test deadlines Both surfaced on the CI runs for this PR and #1487, both in tests neither PR touches, both the same shape as the two already fixed here: a deadline sized for the work rather than for a runner executing many suites at once. VoiceNotePlaybackControllerTests waited 5s for a @MainActor Task that playback schedules for the session acquire and its failure path. When that Task is not scheduled in time the helper reports *two* failures — the wait, and the `!isPlaying` the un-run failure path has not reset yet — which reads like a playback bug rather than a starved scheduler. That is exactly what CI showed. GeoRelayDirectoryTests waited 10s for work the directory runs in Task.detached(priority: .utility); utility priority competes with every other suite. The retry-scheduling case timed out at exactly 10.06s with the retry never scheduled, reading like a missing retry rather than a starved background task. Raised to 30s each with the reasoning recorded at the helper. Both helpers return as soon as their condition holds, so nothing slows down when tests pass — only the genuine-failure case takes longer to report. Verified 3x with all cores saturated: both suites clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Make the flake class unrepeatable, not just fixed Raising four deadlines fixed the four tests that happened to fire. The class was still there: fourteen separate waitUntil helpers, most defaulting to 1.0s, plus wait call sites with literal budgets. The fifth instance would have landed the same way, on someone else's unrelated PR. The rule, now written down in TestConstants.settleTimeout: **a wait deadline is not a latency budget.** It exists so a genuine hang eventually fails the suite, so size it for the worst-case scheduler, never for how long the operation should take. Waits return as soon as their condition holds, so a generous deadline is free in the passing case and only extends genuine failures. - Every wait helper now defaults to TestConstants.settleTimeout (30s), and the literal wait call sites below the floor were converted too — sixteen sites across ten files. - TestTimingHygieneTests enforces it by scanning the test sources: wait defaults and wait call sites must be at least minimumSettleTimeout, and no test may assert an upper bound on elapsed wall-clock time (the assertion that started this, which cannot separate correct behaviour from a slow machine). - Both rules waive per line with "test-timing-ok: <reason>", accepted on the line or in the comment block above it so the reason has room to be a sentence. One legitimate use so far: a NEGATIVE wait in NoiseCoverageTests asserting a promotion has *not* completed within 50ms, where a long deadline would only make the suite slow while still passing. Injected production timeouts are deliberately not matched — the Noise handshake timeouts are the behaviour under test, and short values are correct there. Verified the guard actually fails: a canary file with both banned shapes was flagged with file and line, and the suite went green again once it was removed. Full suite passes with all 18 cores saturated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Close the guard's blind spot: named short timeouts A fifth flake landed on CI while the first guard was in review — GossipSyncBoardTests timing out at 1.03s — and the guard did not catch it, because the deadline was `TestConstants.shortTimeout` rather than a literal. A literals-only scan cannot see a short value behind a symbol, which is the more common way it is written: 81 wait sites used shortTimeout (1s) or defaultTimeout (5s), every one of them below the floor. Rather than guess which of those were safe to raise, all 81 were converted and the suite timed. Runtime went 15s -> 72s, which located the genuine negative waits precisely: ten tests that assert something does *not* happen and therefore always run their deadline out. Measurement instead of a heuristic, since a mis-guess in either direction is invisible — too short reintroduces the flake, too long silently costs a minute a run. Those 28 sites now use `TestConstants.negativeWaitWindow`, a named constant whose doc explains the inverted reasoning: for a negative wait, starvation can only make the assertion *more* likely to hold, so short is correct, and the name states the polarity instead of leaving a bare literal that reads like the mistake. Suite is back to 15.3s. The guard now also rejects `timeout: TestConstants.shortTimeout` and `defaultTimeout`, while accepting `negativeWaitWindow` by name. Re-verified with a canary carrying every banned shape — literal default, both named constants, and an elapsed-time upper bound. All four were flagged with file and line; the suite went green again on removal. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Delete shortTimeout now that nothing may use it The hygiene guard bans `TestConstants.shortTimeout` at every wait site and the last users were converted, so Periphery correctly flagged the constant itself as dead and failed CI. Remove it from both TestConstants copies; the banned-name entry stays so the symbol cannot quietly return. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: jack <jackjackbits@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
272 lines
11 KiB
Swift
272 lines
11 KiB
Swift
import Combine
|
|
import XCTest
|
|
@testable import bitchat
|
|
|
|
/// Covers the reachability-gate decision logic (pure debounce) and the
|
|
/// `NetworkActivationService` wiring that suppresses Tor/relay startup when
|
|
/// there is provably no network.
|
|
@MainActor
|
|
final class NetworkReachabilityGateTests: XCTestCase {
|
|
|
|
// MARK: - Pure debounce logic
|
|
|
|
func test_debounce_satisfiedStaysReachable() {
|
|
var d = ReachabilityDebounce(interval: 2.5, initial: true)
|
|
let t0 = Date()
|
|
// An interface remains present: no change, no pending.
|
|
XCTAssertNil(d.observe(reachable: true, at: t0))
|
|
XCTAssertTrue(d.committed)
|
|
XCTAssertFalse(d.hasPendingChange)
|
|
}
|
|
|
|
func test_debounce_unsatisfiedSuppressesAfterInterval() {
|
|
var d = ReachabilityDebounce(interval: 2.5, initial: true)
|
|
let t0 = Date()
|
|
// Path drops: not committed immediately (within debounce window).
|
|
XCTAssertNil(d.observe(reachable: false, at: t0))
|
|
XCTAssertTrue(d.committed)
|
|
XCTAssertTrue(d.hasPendingChange)
|
|
// Still within window.
|
|
XCTAssertNil(d.flush(at: t0.addingTimeInterval(1.0)))
|
|
XCTAssertTrue(d.committed)
|
|
// Past the window: commit unreachable.
|
|
XCTAssertEqual(d.flush(at: t0.addingTimeInterval(2.5)), false)
|
|
XCTAssertFalse(d.committed)
|
|
XCTAssertFalse(d.hasPendingChange)
|
|
}
|
|
|
|
func test_debounce_flapIsIgnored() {
|
|
var d = ReachabilityDebounce(interval: 2.5, initial: true)
|
|
let t0 = Date()
|
|
// Drop then recover well within the window — must never commit a change.
|
|
XCTAssertNil(d.observe(reachable: false, at: t0))
|
|
XCTAssertTrue(d.hasPendingChange)
|
|
XCTAssertNil(d.observe(reachable: true, at: t0.addingTimeInterval(0.5)))
|
|
XCTAssertFalse(d.hasPendingChange, "recovery should cancel the pending drop")
|
|
// A late flush after the original deadline is a no-op (nothing pending).
|
|
XCTAssertNil(d.flush(at: t0.addingTimeInterval(3.0)))
|
|
XCTAssertTrue(d.committed)
|
|
}
|
|
|
|
func test_debounce_recoverAfterOutageCommitsAfterInterval() {
|
|
var d = ReachabilityDebounce(interval: 2.5, initial: false)
|
|
let t0 = Date()
|
|
XCTAssertNil(d.observe(reachable: true, at: t0))
|
|
XCTAssertTrue(d.hasPendingChange)
|
|
XCTAssertEqual(d.flush(at: t0.addingTimeInterval(2.5)), true)
|
|
XCTAssertTrue(d.committed)
|
|
}
|
|
|
|
func test_debounce_duplicateObservationsPreservePendingDeadline() {
|
|
var d = ReachabilityDebounce(interval: 2.5, initial: true)
|
|
let t0 = Date()
|
|
XCTAssertNil(d.observe(reachable: false, at: t0))
|
|
// Duplicate unsatisfied updates mid-window keep the original deadline.
|
|
XCTAssertNil(d.observe(reachable: false, at: t0.addingTimeInterval(1.0)))
|
|
XCTAssertEqual(d.pendingRemaining(at: t0.addingTimeInterval(1.0)), 1.5)
|
|
// A duplicate arriving past the deadline commits immediately.
|
|
XCTAssertEqual(d.observe(reachable: false, at: t0.addingTimeInterval(2.5)), false)
|
|
XCTAssertNil(d.pendingRemaining(at: t0.addingTimeInterval(2.5)))
|
|
}
|
|
|
|
/// Wiring only: a duplicate mid-window still yields exactly one committed
|
|
/// `false`, published through the monitor's debounce.
|
|
///
|
|
/// This deliberately makes no assertion about *when* the flush fires. It
|
|
/// used to bound elapsed wall-clock time at 1.4 s to prove the deadline was
|
|
/// not restarted, which flaked on loaded CI runners — one observed run took
|
|
/// 3.75 s, because `Task.sleep` and the `asyncAfter` flush are both real
|
|
/// time and neither is bounded above on a busy machine. No wall-clock bound
|
|
/// can distinguish "deadline preserved" from "runner is slow", so the timing
|
|
/// property is asserted where it is computable instead:
|
|
/// `test_debounce_duplicateObservationsPreservePendingDeadline` drives
|
|
/// `ReachabilityDebounce` with injected timestamps and checks
|
|
/// `pendingRemaining` directly.
|
|
///
|
|
/// The clock is injected here so the debounce arithmetic is deterministic
|
|
/// even though the flush itself is scheduled in real time.
|
|
func test_monitor_duplicateUpdatesCommitOnceThroughTheDebounce() async {
|
|
let clock = MutableDate(now: Date(timeIntervalSince1970: 1_784_000_000))
|
|
let monitor = NWPathReachabilityMonitor(
|
|
debounceInterval: 0.2,
|
|
now: { clock.now }
|
|
)
|
|
var received: [Bool] = []
|
|
let cancellable = monitor.reachabilityPublisher.sink { received.append($0) }
|
|
defer { cancellable.cancel() }
|
|
|
|
monitor.ingest(reachable: false)
|
|
// Duplicate unsatisfied update mid-window (e.g. an interface detail
|
|
// change while still offline).
|
|
clock.now = clock.now.addingTimeInterval(0.1)
|
|
monitor.ingest(reachable: false)
|
|
// Past the original deadline, so the scheduled flush commits.
|
|
clock.now = clock.now.addingTimeInterval(0.2)
|
|
|
|
// Generous: this is a liveness check, not a latency bound. A real
|
|
// regression — never committing — still fails, just later.
|
|
let committed = await waitUntil(timeout: 10.0) { !received.isEmpty }
|
|
XCTAssertTrue(committed)
|
|
XCTAssertEqual(received, [false])
|
|
}
|
|
|
|
// MARK: - Service gating
|
|
|
|
func test_start_whenUnreachable_suppressesTorAndRelays() {
|
|
let ctx = makeService(permission: .authorized, reachable: false)
|
|
ctx.service.start()
|
|
|
|
XCTAssertFalse(ctx.service.activationAllowed)
|
|
XCTAssertFalse(ctx.service.isNetworkReachable)
|
|
XCTAssertTrue(ctx.reachability.startCalled)
|
|
XCTAssertEqual(ctx.torController.startIfNeededCallCount, 0)
|
|
XCTAssertEqual(ctx.torController.autoStartAllowedValues, [false])
|
|
XCTAssertEqual(ctx.relayController.connectCallCount, 0)
|
|
XCTAssertEqual(ctx.relayController.disconnectCallCount, 1)
|
|
}
|
|
|
|
func test_start_whenReachable_allowsTorAndRelays() {
|
|
let ctx = makeService(permission: .authorized, reachable: true)
|
|
ctx.service.start()
|
|
|
|
XCTAssertTrue(ctx.service.activationAllowed)
|
|
XCTAssertEqual(ctx.torController.startIfNeededCallCount, 1)
|
|
XCTAssertEqual(ctx.relayController.connectCallCount, 1)
|
|
}
|
|
|
|
func test_reachabilityRecovery_resumesTorAndRelays() async {
|
|
let ctx = makeService(permission: .authorized, reachable: false)
|
|
ctx.service.start()
|
|
XCTAssertFalse(ctx.service.activationAllowed)
|
|
|
|
ctx.reachability.set(true)
|
|
let resumed = await waitUntil { ctx.service.activationAllowed }
|
|
XCTAssertTrue(resumed)
|
|
XCTAssertTrue(ctx.service.isNetworkReachable)
|
|
XCTAssertGreaterThanOrEqual(ctx.torController.startIfNeededCallCount, 1)
|
|
XCTAssertGreaterThanOrEqual(ctx.relayController.connectCallCount, 1)
|
|
}
|
|
|
|
func test_reachabilityLoss_disconnectsRelaysAndStopsTor() async {
|
|
let ctx = makeService(permission: .authorized, reachable: true)
|
|
ctx.service.start()
|
|
XCTAssertTrue(ctx.service.activationAllowed)
|
|
let disconnectsBefore = ctx.relayController.disconnectCallCount
|
|
|
|
ctx.reachability.set(false)
|
|
let suppressed = await waitUntil { !ctx.service.activationAllowed }
|
|
XCTAssertTrue(suppressed)
|
|
XCTAssertFalse(ctx.service.isNetworkReachable)
|
|
XCTAssertGreaterThan(ctx.relayController.disconnectCallCount, disconnectsBefore)
|
|
XCTAssertTrue(ctx.torController.autoStartAllowedValues.contains(false))
|
|
XCTAssertGreaterThanOrEqual(ctx.torController.shutdownCompletelyCallCount, 1)
|
|
}
|
|
|
|
// MARK: - Harness
|
|
|
|
private func makeService(
|
|
permission: LocationChannelManager.PermissionState,
|
|
reachable: Bool
|
|
) -> Context {
|
|
let suiteName = "NetworkReachabilityGateTests-\(UUID().uuidString)"
|
|
let storage = UserDefaults(suiteName: suiteName)!
|
|
storage.removePersistentDomain(forName: suiteName)
|
|
|
|
let permissionSubject = CurrentValueSubject<LocationChannelManager.PermissionState, Never>(permission)
|
|
let favoritesSubject = CurrentValueSubject<Set<Data>, Never>([])
|
|
let reachability = ControllableReachabilityMonitor(initial: reachable)
|
|
let torController = GateMockTorController()
|
|
let relayController = GateMockRelayController()
|
|
let proxyController = GateMockProxyController()
|
|
let service = NetworkActivationService(
|
|
storage: storage,
|
|
locationPermissionPublisher: permissionSubject.eraseToAnyPublisher(),
|
|
mutualFavoritesPublisher: favoritesSubject.eraseToAnyPublisher(),
|
|
permissionProvider: { permissionSubject.value },
|
|
mutualFavoritesProvider: { favoritesSubject.value },
|
|
reachabilityMonitor: reachability,
|
|
torController: torController,
|
|
relayController: relayController,
|
|
proxyController: proxyController,
|
|
notificationCenter: NotificationCenter()
|
|
)
|
|
return Context(
|
|
service: service,
|
|
reachability: reachability,
|
|
torController: torController,
|
|
relayController: relayController
|
|
)
|
|
}
|
|
|
|
private func waitUntil(
|
|
timeout: TimeInterval = TestConstants.settleTimeout,
|
|
condition: @escaping @MainActor () -> Bool
|
|
) async -> Bool {
|
|
let deadline = Date().addingTimeInterval(timeout)
|
|
while Date() < deadline {
|
|
if condition() { return true }
|
|
try? await Task.sleep(nanoseconds: 10_000_000)
|
|
}
|
|
return condition()
|
|
}
|
|
}
|
|
|
|
@MainActor
|
|
private struct Context {
|
|
let service: NetworkActivationService
|
|
let reachability: ControllableReachabilityMonitor
|
|
let torController: GateMockTorController
|
|
let relayController: GateMockRelayController
|
|
}
|
|
|
|
@MainActor
|
|
private final class ControllableReachabilityMonitor: NetworkReachabilityMonitoring {
|
|
private let subject: CurrentValueSubject<Bool, Never>
|
|
private(set) var startCalled = false
|
|
|
|
init(initial: Bool) {
|
|
subject = CurrentValueSubject(initial)
|
|
}
|
|
|
|
var isReachable: Bool { subject.value }
|
|
var reachabilityPublisher: AnyPublisher<Bool, Never> {
|
|
subject.removeDuplicates().dropFirst().eraseToAnyPublisher()
|
|
}
|
|
func start() { startCalled = true }
|
|
func stop() { startCalled = false }
|
|
func set(_ reachable: Bool) { subject.send(reachable) }
|
|
}
|
|
|
|
@MainActor
|
|
private final class GateMockTorController: NetworkActivationTorControlling {
|
|
private(set) var autoStartAllowedValues: [Bool] = []
|
|
private(set) var startIfNeededCallCount = 0
|
|
private(set) var shutdownCompletelyCallCount = 0
|
|
func setAutoStartAllowed(_ allowed: Bool) { autoStartAllowedValues.append(allowed) }
|
|
func startIfNeeded() { startIfNeededCallCount += 1 }
|
|
func shutdownCompletely() { shutdownCompletelyCallCount += 1 }
|
|
}
|
|
|
|
@MainActor
|
|
private final class GateMockRelayController: NetworkActivationRelayControlling {
|
|
private(set) var connectCallCount = 0
|
|
private(set) var disconnectCallCount = 0
|
|
func connect() { connectCallCount += 1 }
|
|
func disconnect() { disconnectCallCount += 1 }
|
|
}
|
|
|
|
private final class GateMockProxyController: NetworkActivationProxyControlling {
|
|
private(set) var proxyModes: [Bool] = []
|
|
func setProxyMode(useTor: Bool) { proxyModes.append(useTor) }
|
|
}
|
|
|
|
/// Controllable clock, so debounce arithmetic is deterministic even where the
|
|
/// flush itself is scheduled in real time.
|
|
private final class MutableDate: @unchecked Sendable {
|
|
var now: Date
|
|
|
|
init(now: Date) {
|
|
self.now = now
|
|
}
|
|
}
|