Files
bitchat/bitchat/Services/BLE/BLEConnectionScheduler.swift
T
266827ceff Extend BLE mesh range: relax RSSI gates, lift sparse TTL clamps, faster walk-back reconnects (#1338)
* Extend mesh range: relax RSSI gates, lift sparse TTL clamps, drain connection queue

Range improvements to the BLE mesh, all policy-level (no wire/protocol
changes):

- Drain the connection candidate queue from the maintenance tick.
  Weak-RSSI discoveries are enqueued rather than connected, but the
  queue was only drained on disconnect/failure/timeout events — an
  isolated node surrounded only by weak (distant) peers queued them
  all and never connected to anyone.
- Relax isolated RSSI floors from -90/-92 to -95/-100 and relax after
  30s instead of 60s. When isolated, a fringe connection beats no
  connection; CoreBluetooth rarely reports below -100 so prolonged
  isolation now effectively accepts any decodable peer.
- Drop the global high-timeout RSSI escalation (-80 after 3 timeouts
  in 60s). One flaky distant peer could blind the node to every other
  edge-of-range peer; per-peripheral cooldown, the discovery ignore
  window, and score bias already contain flaky links individually.
- Relay at full incoming TTL in thin chains (degree <= 2). Sparse
  line topologies are exactly where every hop counts and where flood
  cost is minimal; previously messages lost a hop to the clamp.
- Raise the fragment relay TTL cap from 5 to 7 in sparse graphs so
  media reaches as far as text; dense graphs keep the 5-hop clamp to
  contain full-fanout fragment floods.
- Extend peer reachability retention from 21s to 60s (verified) /
  45s (unverified) so duty-cycled nodes (worst-case dense announce
  interval 38s) don't forget peers between announces.
- Extend the directed store-and-forward spool window from 15s to 60s
  so brief link gaps heal via the periodic flush.

957 tests pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Reconnect quickly after walk-away disconnects

Field test (walk away + return between two devices) showed reconnect
landing exactly 15.0s after the supervision-timeout disconnect: the
scheduler records a dropped established connection via
recordDisconnectError into the same map as connect timeouts, and
handleDiscovery hard-ignores rediscoveries for 15s.

Those are different situations. A connect attempt that timed out means
the peer likely isn't reachable, so backing off is right. A dropped
established connection usually means the peer walked out of range and
will return — track it separately and only ignore rediscoveries for 3s
(enough for CoreBluetooth to settle), so walking back into range
reconnects ~12s sooner.

Disconnect errors also no longer feed the weak-link cooldown or the
candidate-score timeout bias; those penalties now apply only to peers
that never answered a connect attempt.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Honor the disconnect settle window on the queue drain path

Codex review caught that the 3s settle window was only enforced in
handleDiscovery. A candidate can already be sitting in the queue when
its peripheral drops (weak-RSSI adverts are enqueued even while
connected, since the RSSI check precedes the existing-state check),
and didDisconnectPeripheral immediately drains the queue — so the
stale entry could reconnect right through the window, recreating the
reconnect/cancel thrash it exists to prevent.

nextCandidate now defers such candidates with retryAfter for the
window's remainder, mirroring the weak-link cooldown pattern.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Relay on one link per bound peer instead of both dual-role links

Three-device field test (star topology) showed every relayed fragment
arriving twice at the leaf: dual-role pairs hold two live links (we as
central writing to their peripheral, they as central subscribed to
ours) and broadcast/relay fanout sent the same packet down both — 2x
airtime on exactly the pairs that talk most, with the receiver just
discarding the duplicate.

The fanout selector now collapses link selection to one link per bound
peer, preferring the peripheral (write) side since it has per-link
flow control via canSendWriteWithoutResponse, while notifications
share the peripheral manager's update queue across all centrals.
Links with no bound peer yet (pre-announce) pass through untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Only notify "bitchatters nearby" on the empty-to-populated transition

Devices sitting idle and connected kept re-firing the notification.
Two bugs in handleNetworkAvailability:

- Peers first sighted during the 5-minute cooldown were never added to
  recentlySeenPeers (the formUnion only ran when a notification
  fired), so they stayed "new" forever and re-triggered on the next
  routine peer-list event once the cooldown lapsed.
- There was no went-from-zero gate at all: any unseen peer notified,
  even while already meshed with others who are visible in the app.

Every sighted peer is now recorded regardless of cooldown, and the
notification only fires when the mesh transitions from confirmed-empty
to populated with genuinely new peers. meshWasEmpty resets only via
the existing confirmed-empty paths (30s empty confirmation, 10-minute
quiet reset), so brief link flaps stay silent. The cooldown becomes
injectable so tests can prove the transition gate independently.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Bump version to 1.5.2; Xcode 26.5 project settings update

Marketing version 1.5.1 -> 1.5.2 (pbxproj + Release.xcconfig).
Project settings refresh from Xcode 26.5: upgrade-check stamp, drop
redundant DEVELOPMENT_TEAM self-references, scheme version stamps.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Disable string catalog symbol generation

The Xcode 26.5 settings refresh enabled STRING_CATALOG_GENERATE_SYMBOLS
(the new default), which fails on the literal "%@" key in
Localizable.xcstrings — a pure format placeholder can't become a Swift
identifier. Nothing in the codebase references generated catalog
symbols, so turn the feature off rather than renaming keys around it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Guard _PreviewHelpers references for archive builds

Archiving for TestFlight failed: _PreviewHelpers is a development
asset, so its sources (PreviewKeychainManager, BitchatMessage.preview)
are excluded from Release/archive builds, and two call sites
referenced them unconditionally:

- TextMessageView's #Preview block — now wrapped in #if DEBUG
- FavoritesPersistenceService.makeDefaultKeychain's test branch — the
  in-memory-keychain-under-test path is now #if DEBUG; tests always
  run Debug so behavior is unchanged, and Release always gets the real
  KeychainManager

Verified with an iOS Release arm64 build (the archive configuration).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: jack <jackjackbits@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-12 10:38:39 +02:00

294 lines
10 KiB
Swift

import Foundation
struct BLEConnectionCandidate<Peripheral> {
let peripheral: Peripheral
let peripheralID: String
let rssi: Int
let name: String
let isConnectable: Bool
let discoveredAt: Date
}
struct BLEExistingConnectionState {
let isConnecting: Bool
let isConnected: Bool
let lastConnectionAttempt: Date?
}
enum BLEPeripheralConnectionState {
case disconnected
case connecting
case connected
}
enum BLEDiscoveryDecision: Equatable {
case ignore
case queued
case scheduleRetry(after: TimeInterval)
case cancelStaleConnection
case connectNow
}
enum BLEConnectionQueueDecision<Peripheral> {
case none
case retryAfter(TimeInterval)
case connect(BLEConnectionCandidate<Peripheral>)
}
final class BLEConnectionScheduler<Peripheral> {
private let maxCentralLinks: Int
private let connectRateLimitInterval: TimeInterval
private let candidateCap: Int
private let weakLinkCooldownSeconds: TimeInterval
private let weakLinkRSSICutoff: Int
private var lastGlobalConnectAttempt: Date = .distantPast
private var candidates: [BLEConnectionCandidate<Peripheral>] = []
private var failureCounts: [String: Int] = [:]
private var recentConnectTimeouts: [String: Date] = [:]
// Tracked separately from connect timeouts: a peer we held a connection
// with and lost (walked out of range) usually comes back, so it only gets
// a brief rediscovery ignore — not the timeout backoff/cooldown treatment
// reserved for peers that never answered a connect attempt.
private var recentDisconnects: [String: Date] = [:]
private var lastIsolatedAt: Date?
private let initialDynamicRSSIThreshold: Int
private(set) var dynamicRSSIThreshold: Int
var candidateCount: Int {
candidates.count
}
init(
maxCentralLinks: Int = TransportConfig.bleMaxCentralLinks,
connectRateLimitInterval: TimeInterval = TransportConfig.bleConnectRateLimitInterval,
candidateCap: Int = TransportConfig.bleConnectionCandidatesMax,
weakLinkCooldownSeconds: TimeInterval = TransportConfig.bleWeakLinkCooldownSeconds,
weakLinkRSSICutoff: Int = TransportConfig.bleWeakLinkRSSICutoff,
dynamicRSSIThreshold: Int = TransportConfig.bleDynamicRSSIThresholdDefault
) {
self.maxCentralLinks = maxCentralLinks
self.connectRateLimitInterval = connectRateLimitInterval
self.candidateCap = candidateCap
self.weakLinkCooldownSeconds = weakLinkCooldownSeconds
self.weakLinkRSSICutoff = weakLinkRSSICutoff
self.initialDynamicRSSIThreshold = dynamicRSSIThreshold
self.dynamicRSSIThreshold = dynamicRSSIThreshold
}
func handleDiscovery(
_ candidate: BLEConnectionCandidate<Peripheral>,
connectedOrConnectingCount: Int,
existingState: BLEExistingConnectionState?,
peripheralState: BLEPeripheralConnectionState,
now: Date
) -> BLEDiscoveryDecision {
guard candidate.isConnectable else { return .ignore }
if candidate.rssi <= dynamicRSSIThreshold {
enqueue(candidate)
return .queued
}
if connectedOrConnectingCount >= maxCentralLinks {
enqueue(candidate)
return .queued
}
if let retryDelay = rateLimitRetryDelay(now: now) {
enqueue(candidate)
return .scheduleRetry(after: retryDelay)
}
if let existingState {
if existingState.isConnected || existingState.isConnecting {
return .ignore
}
if let lastAttempt = existingState.lastConnectionAttempt,
now.timeIntervalSince(lastAttempt) < 2.0 {
return .ignore
}
}
if let lastTimeout = recentConnectTimeouts[candidate.peripheralID],
now.timeIntervalSince(lastTimeout) < TransportConfig.bleTimeoutDiscoveryIgnoreSeconds {
return .ignore
}
if let lastDisconnect = recentDisconnects[candidate.peripheralID],
now.timeIntervalSince(lastDisconnect) < TransportConfig.bleDisconnectDiscoveryIgnoreSeconds {
return .ignore
}
switch peripheralState {
case .disconnected:
return .connectNow
case .connecting, .connected:
return .cancelStaleConnection
}
}
func enqueue(_ candidate: BLEConnectionCandidate<Peripheral>) {
if let existingIndex = candidates.firstIndex(where: { $0.peripheralID == candidate.peripheralID }) {
candidates[existingIndex] = candidate
} else {
candidates.append(candidate)
}
candidates.sort {
if $0.rssi != $1.rssi { return $0.rssi > $1.rssi }
return $0.discoveredAt < $1.discoveredAt
}
if candidates.count > candidateCap {
candidates.removeLast(candidates.count - candidateCap)
}
}
func nextCandidate(
connectedOrConnectingCount: Int,
isAlreadyConnectingOrConnected: (String) -> Bool,
now: Date
) -> BLEConnectionQueueDecision<Peripheral> {
guard connectedOrConnectingCount < maxCentralLinks else { return .none }
if let retryDelay = rateLimitRetryDelay(now: now) {
return .retryAfter(retryDelay)
}
while !candidates.isEmpty {
candidates.sort { score($0, now: now) > score($1, now: now) }
let candidate = candidates.removeFirst()
guard candidate.isConnectable else { continue }
if let delay = weakLinkRetryDelay(for: candidate, now: now) {
enqueue(candidate)
return .retryAfter(delay)
}
if let delay = disconnectSettleDelay(for: candidate, now: now) {
enqueue(candidate)
return .retryAfter(delay)
}
if isAlreadyConnectingOrConnected(candidate.peripheralID) {
continue
}
return .connect(candidate)
}
return .none
}
func recordConnectionAttempt(at now: Date) {
lastGlobalConnectAttempt = now
}
func recordConnectionSuccess(peripheralID: String) {
failureCounts[peripheralID] = 0
recentConnectTimeouts.removeValue(forKey: peripheralID)
recentDisconnects.removeValue(forKey: peripheralID)
}
func recordConnectionFailure(peripheralID: String) {
failureCounts[peripheralID, default: 0] += 1
}
func recordDisconnectError(peripheralID: String, at now: Date) {
recentDisconnects[peripheralID] = now
}
func recordConnectionTimeout(peripheralID: String, at now: Date) {
recentConnectTimeouts[peripheralID] = now
recordConnectionFailure(peripheralID: peripheralID)
}
func pruneConnectionTimeouts(before cutoff: Date) {
recentConnectTimeouts = recentConnectTimeouts.filter { $0.value >= cutoff }
recentDisconnects = recentDisconnects.filter { $0.value >= cutoff }
}
func reset() {
lastGlobalConnectAttempt = .distantPast
candidates.removeAll()
failureCounts.removeAll()
recentConnectTimeouts.removeAll()
recentDisconnects.removeAll()
lastIsolatedAt = nil
dynamicRSSIThreshold = initialDynamicRSSIThreshold
}
@discardableResult
func updateRSSIThreshold(
connectedCount: Int,
connectedOrConnectingLinkCount: Int,
now: Date
) -> Int {
if connectedCount == 0 {
if lastIsolatedAt == nil { lastIsolatedAt = now }
let isolatedAt = lastIsolatedAt ?? now
let elapsed = now.timeIntervalSince(isolatedAt)
dynamicRSSIThreshold = elapsed > TransportConfig.bleIsolationRelaxThresholdSeconds
? TransportConfig.bleRSSIIsolatedRelaxed
: TransportConfig.bleRSSIIsolatedBase
return dynamicRSSIThreshold
}
lastIsolatedAt = nil
// Flaky links are handled per-peripheral (weak-link cooldown, discovery
// ignore window, score bias) — never globally, so one flaky distant peer
// can't blind us to every other edge-of-range peer.
var threshold = TransportConfig.bleDynamicRSSIThresholdDefault
if connectedOrConnectingLinkCount >= maxCentralLinks || candidates.count >= candidateCap {
threshold = TransportConfig.bleRSSIConnectedThreshold
}
dynamicRSSIThreshold = threshold
return threshold
}
private func rateLimitRetryDelay(now: Date) -> TimeInterval? {
let elapsed = now.timeIntervalSince(lastGlobalConnectAttempt)
guard elapsed < connectRateLimitInterval else { return nil }
return connectRateLimitInterval - elapsed + 0.05
}
private func weakLinkRetryDelay(
for candidate: BLEConnectionCandidate<Peripheral>,
now: Date
) -> TimeInterval? {
guard let lastTimeout = recentConnectTimeouts[candidate.peripheralID] else { return nil }
let elapsed = now.timeIntervalSince(lastTimeout)
guard elapsed < weakLinkCooldownSeconds && candidate.rssi <= weakLinkRSSICutoff else { return nil }
let remaining = weakLinkCooldownSeconds - elapsed
return min(max(2.0, remaining), 15.0)
}
// The disconnect settle window must hold on the queue path too: a stale
// candidate enqueued while the peripheral was still connected would
// otherwise reconnect immediately via the post-disconnect queue drain,
// bypassing the window and recreating reconnect/cancel thrash.
private func disconnectSettleDelay(
for candidate: BLEConnectionCandidate<Peripheral>,
now: Date
) -> TimeInterval? {
guard let lastDisconnect = recentDisconnects[candidate.peripheralID] else { return nil }
let remaining = TransportConfig.bleDisconnectDiscoveryIgnoreSeconds - now.timeIntervalSince(lastDisconnect)
guard remaining > 0 else { return nil }
return remaining + 0.05
}
private func score(_ candidate: BLEConnectionCandidate<Peripheral>, now: Date) -> Int {
let failures = failureCounts[candidate.peripheralID] ?? 0
let penalty = min(20, 1 << min(4, failures))
let timeoutBias = recentConnectTimeouts[candidate.peripheralID].map {
now.timeIntervalSince($0) < 60 ? 10 : 0
} ?? 0
let base = (candidate.isConnectable ? 1000 : 0) + (candidate.rssi + 100) * 2
let recency = -Int(now.timeIntervalSince(candidate.discoveredAt) * 10)
return base + recency - penalty - timeoutBias
}
}