mirror of
https://github.com/permissionlesstech/bitchat.git
synced 2026-07-26 19:25:23 +00:00
Live push-to-talk voice for DMs (streams while you talk, voice note as fallback) (#1403)
* Live push-to-talk voice for DMs: stream while you talk, voice note as fallback Holding the mic in a DM now streams AAC frames live over the Noise session (walkie-talkie style, ~0.5s mouth-to-ear at one hop) while recording the same audio as a normal voice note. On release the note ships through the existing fileTransfer pipeline; receivers that heard the live stream absorb it silently into the same bubble (matched by the burst ID embedded in the file name), so reliability comes for free and nobody sees duplicates. Protocol: - NoisePayloadType.voiceFrame = 0x08 carrying VoiceBurstPacket (burstID + seq + START/data/END/CANCELED, length-prefixed AAC frames) - 210-byte burst-content budget keeps each Noise packet inside the 256-byte padding bucket: one BLE frame, never the fragment scheduler - fire-and-forget: frames are dropped (never queued) without an established session; live is only offered when the peer is mesh-reachable Receive: - ChatLiveVoiceCoordinator assembles bursts (jitter-ordered, 0.5s gap skip, 3s idle end, flood/size caps), persists progressively as ADTS .aac so even a partial burst is a replayable bubble - live autoplay only when the conversation is on screen, app active, and the new app-info "live voice messages" toggle is on (also gates live sending) - one-playback-at-a-time via a shared ExclusivePlayback slot Capture: - PTTCaptureEngine taps AVAudioEngine, dual-encodes: live AAC frames + the finalized .m4a (same 16kHz/mono/16kbps settings as VoiceRecorder) - VoiceRecordingViewModel now drives a pluggable VoiceCaptureSession; the composer HUD shows a pulsing LIVE treatment when streaming Includes the push-to-talk design doc, 6 new localization keys across all 29 locales, and unit tests for framing, packetizer budget, ADTS output, codec round-trip, and the assembly/absorb lifecycle. Public-mesh PTT (MessageType 0x29) lands separately on top of this. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * PTT follow-ups from review + field test: peer-ID normalization, toggle gates inbound, drop-path diagnostics Codex review fixes (#1403): - makeVoiceCaptureSession normalizes the selected peer with toShort() before the reachability/session checks and binds the send target to that same routing ID — a conversation selected under the stable 64-hex Noise key no longer silently falls back to a classic note while the short-ID session is established - the live-voice toggle now gates inbound bursts too: off means classic-notes-only in both directions (no live bubble, partial file, or early notification; the finalized note still arrives), with a test Field-test diagnostics (first device run: DM frames decrypted but no bubble appeared, with no log evidence of which guard dropped them): - coordinator logs undecodable frames (size + hex prefix) and blocked drops - makeAssembly logs directory/file-handle failures instead of returning nil silently - PTTLiveVoiceSession logs capture start and finish (packet/frame/duration counts); PTTCaptureEngine logs engine start success/failure with the input format; BLEService.sendVoiceFrame logs no-session drops Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Fix iPhone live-capture failure: dead input unit (AURemoteIO -10851, 0 Hz) Field testing showed the phone's live capture failing at mic enable with AURemoteIO -10851 and an input format of 0 Hz / 2 ch — an input unit bound to an earlier (playback-only or settling) audio session. The Mac, which has no session lifecycle, captured fine, which is why public bursts from the Mac worked while phone-side sends degraded from working (first hold) to sporadic to dead across holds. Three layers of defense: - PTTCaptureEngine recreates its AVAudioEngine on every start(), after the session is configured, so the input unit binds to the session that is active now; a dead input (0 Hz or 0 channels) is now a distinct, logged error instead of a silent setup failure - PTTLiveVoiceSession retries the capture start once after a 150 ms route-settle pause - VoiceRecordingViewModel falls back to the classic VoiceRecorder within the same hold if the live engine still cannot start — a route glitch now costs the live stream, never the voice note Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: jack <jackjackbits@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
jack
Claude Fable 5
parent
6886035632
commit
eacd8f0750
@@ -0,0 +1,175 @@
|
||||
//
|
||||
// PTTAudioCodec.swift
|
||||
// bitchat
|
||||
//
|
||||
// This is free and unencumbered software released into the public domain.
|
||||
// For more information, see <https://unlicense.org>
|
||||
//
|
||||
|
||||
import AVFoundation
|
||||
import BitLogger
|
||||
import Foundation
|
||||
|
||||
/// Streaming PCM -> AAC-LC encoder for live voice. Stateful (the AAC encoder
|
||||
/// carries a bit reservoir across frames); one instance per burst.
|
||||
/// Not thread-safe — confine to one queue.
|
||||
final class PTTFrameEncoder {
|
||||
private let converter: AVAudioConverter
|
||||
private var pendingInput: [AVAudioPCMBuffer] = []
|
||||
|
||||
init?() {
|
||||
guard let pcm = PTTAudioFormat.pcmFormat,
|
||||
let aac = PTTAudioFormat.aacFormat,
|
||||
let converter = AVAudioConverter(from: pcm, to: aac)
|
||||
else { return nil }
|
||||
converter.bitRate = PTTAudioFormat.bitRate
|
||||
self.converter = converter
|
||||
}
|
||||
|
||||
/// Feeds PCM (16 kHz mono float) and returns every complete AAC frame the
|
||||
/// encoder produced. Frames come out ~130 bytes each at 16 kbps.
|
||||
func encode(_ buffer: AVAudioPCMBuffer) -> [Data] {
|
||||
pendingInput.append(buffer)
|
||||
return drainConverter()
|
||||
}
|
||||
|
||||
private func drainConverter() -> [Data] {
|
||||
var frames: [Data] = []
|
||||
while true {
|
||||
let output = AVAudioCompressedBuffer(
|
||||
format: converter.outputFormat,
|
||||
packetCapacity: 8,
|
||||
maximumPacketSize: max(converter.maximumOutputPacketSize, 1)
|
||||
)
|
||||
var error: NSError?
|
||||
let status = converter.convert(to: output, error: &error) { [weak self] _, outStatus in
|
||||
guard let self, let next = self.pendingInput.first else {
|
||||
outStatus.pointee = .noDataNow
|
||||
return nil
|
||||
}
|
||||
self.pendingInput.removeFirst()
|
||||
outStatus.pointee = .haveData
|
||||
return next
|
||||
}
|
||||
if status == .error {
|
||||
SecureLogger.error("PTT encode failed: \(error?.localizedDescription ?? "unknown")", category: .session)
|
||||
return frames
|
||||
}
|
||||
frames.append(contentsOf: Self.extractPackets(from: output))
|
||||
// .haveData means the output buffer filled and more may be ready;
|
||||
// anything else means the converter wants more input.
|
||||
if status != .haveData { return frames }
|
||||
}
|
||||
}
|
||||
|
||||
private static func extractPackets(from buffer: AVAudioCompressedBuffer) -> [Data] {
|
||||
guard buffer.packetCount > 0, let descriptions = buffer.packetDescriptions else { return [] }
|
||||
var frames: [Data] = []
|
||||
frames.reserveCapacity(Int(buffer.packetCount))
|
||||
for index in 0..<Int(buffer.packetCount) {
|
||||
let description = descriptions[index]
|
||||
guard description.mDataByteSize > 0 else { continue }
|
||||
let start = buffer.data.advanced(by: Int(description.mStartOffset))
|
||||
frames.append(Data(bytes: start, count: Int(description.mDataByteSize)))
|
||||
}
|
||||
return frames
|
||||
}
|
||||
}
|
||||
|
||||
/// Streaming AAC-LC -> PCM decoder for live voice. Stateful; one instance per
|
||||
/// inbound burst. Not thread-safe — confine to one queue/actor.
|
||||
final class PTTFrameDecoder {
|
||||
private let converter: AVAudioConverter
|
||||
private let pcmFormat: AVAudioFormat
|
||||
private let aacFormat: AVAudioFormat
|
||||
|
||||
init?() {
|
||||
guard let pcm = PTTAudioFormat.pcmFormat,
|
||||
let aac = PTTAudioFormat.aacFormat,
|
||||
let converter = AVAudioConverter(from: aac, to: pcm)
|
||||
else { return nil }
|
||||
self.converter = converter
|
||||
self.pcmFormat = pcm
|
||||
self.aacFormat = aac
|
||||
}
|
||||
|
||||
/// Decodes one raw AAC frame to PCM. Returns nil for malformed input or
|
||||
/// while the decoder is still priming (the first frame of a stream).
|
||||
func decode(_ frame: Data) -> AVAudioPCMBuffer? {
|
||||
guard !frame.isEmpty, frame.count <= 8 * 1024 else { return nil }
|
||||
|
||||
let input = AVAudioCompressedBuffer(format: aacFormat, packetCapacity: 1, maximumPacketSize: frame.count)
|
||||
frame.withUnsafeBytes { raw in
|
||||
guard let base = raw.baseAddress else { return }
|
||||
input.data.copyMemory(from: base, byteCount: frame.count)
|
||||
}
|
||||
input.byteLength = UInt32(frame.count)
|
||||
input.packetCount = 1
|
||||
input.packetDescriptions?.pointee = AudioStreamPacketDescription(
|
||||
mStartOffset: 0,
|
||||
mVariableFramesInPacket: 0,
|
||||
mDataByteSize: UInt32(frame.count)
|
||||
)
|
||||
|
||||
guard let output = AVAudioPCMBuffer(
|
||||
pcmFormat: pcmFormat,
|
||||
frameCapacity: PTTAudioFormat.samplesPerFrame * 2
|
||||
) else { return nil }
|
||||
|
||||
var consumed = false
|
||||
var error: NSError?
|
||||
let status = converter.convert(to: output, error: &error) { _, outStatus in
|
||||
if consumed {
|
||||
outStatus.pointee = .noDataNow
|
||||
return nil
|
||||
}
|
||||
consumed = true
|
||||
outStatus.pointee = .haveData
|
||||
return input
|
||||
}
|
||||
guard status != .error else {
|
||||
SecureLogger.debug("PTT decode failed: \(error?.localizedDescription ?? "unknown")", category: .session)
|
||||
return nil
|
||||
}
|
||||
return output.frameLength > 0 ? output : nil
|
||||
}
|
||||
}
|
||||
|
||||
/// Sample-rate/channel converter from the microphone's native format to the
|
||||
/// 16 kHz mono processing format. Stateful; not thread-safe.
|
||||
final class PTTInputResampler {
|
||||
private let converter: AVAudioConverter
|
||||
private let outputFormat: AVAudioFormat
|
||||
private let ratio: Double
|
||||
|
||||
init?(inputFormat: AVAudioFormat) {
|
||||
guard let pcm = PTTAudioFormat.pcmFormat,
|
||||
let converter = AVAudioConverter(from: inputFormat, to: pcm)
|
||||
else { return nil }
|
||||
self.converter = converter
|
||||
self.outputFormat = pcm
|
||||
self.ratio = PTTAudioFormat.sampleRate / inputFormat.sampleRate
|
||||
}
|
||||
|
||||
func resample(_ buffer: AVAudioPCMBuffer) -> AVAudioPCMBuffer? {
|
||||
let capacity = AVAudioFrameCount(Double(buffer.frameLength) * ratio) + 64
|
||||
guard let output = AVAudioPCMBuffer(pcmFormat: outputFormat, frameCapacity: capacity) else { return nil }
|
||||
|
||||
var consumed = false
|
||||
var error: NSError?
|
||||
let status = converter.convert(to: output, error: &error) { _, outStatus in
|
||||
if consumed {
|
||||
outStatus.pointee = .noDataNow
|
||||
return nil
|
||||
}
|
||||
consumed = true
|
||||
outStatus.pointee = .haveData
|
||||
return buffer
|
||||
}
|
||||
guard status != .error else {
|
||||
SecureLogger.debug("PTT resample failed: \(error?.localizedDescription ?? "unknown")", category: .session)
|
||||
return nil
|
||||
}
|
||||
return output.frameLength > 0 ? output : nil
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,84 @@
|
||||
//
|
||||
// PTTAudioFormat.swift
|
||||
// bitchat
|
||||
//
|
||||
// This is free and unencumbered software released into the public domain.
|
||||
// For more information, see <https://unlicense.org>
|
||||
//
|
||||
|
||||
import AVFoundation
|
||||
import Foundation
|
||||
|
||||
/// Shared audio parameters for live push-to-talk: AAC-LC, 16 kHz, mono,
|
||||
/// ~16 kbps — deliberately identical to `VoiceRecorder`'s voice-note settings
|
||||
/// so a burst's finalized `.m4a` and its live frames sound the same.
|
||||
enum PTTAudioFormat {
|
||||
static let sampleRate: Double = 16_000
|
||||
static let channelCount: AVAudioChannelCount = 1
|
||||
static let bitRate = 16_000
|
||||
/// AAC-LC frame size is fixed by the codec: 1024 samples = 64 ms at 16 kHz.
|
||||
static let samplesPerFrame: AVAudioFrameCount = 1024
|
||||
static var frameDuration: TimeInterval { Double(samplesPerFrame) / sampleRate }
|
||||
|
||||
/// Uncompressed processing format (deinterleaved float PCM).
|
||||
static var pcmFormat: AVAudioFormat? {
|
||||
AVAudioFormat(standardFormatWithSampleRate: sampleRate, channels: channelCount)
|
||||
}
|
||||
|
||||
/// Compressed wire format.
|
||||
static var aacFormat: AVAudioFormat? {
|
||||
var description = AudioStreamBasicDescription(
|
||||
mSampleRate: sampleRate,
|
||||
mFormatID: kAudioFormatMPEG4AAC,
|
||||
mFormatFlags: 0,
|
||||
mBytesPerPacket: 0,
|
||||
mFramesPerPacket: samplesPerFrame,
|
||||
mBytesPerFrame: 0,
|
||||
mChannelsPerFrame: channelCount,
|
||||
mBitsPerChannel: 0,
|
||||
mReserved: 0
|
||||
)
|
||||
return AVAudioFormat(streamDescription: &description)
|
||||
}
|
||||
|
||||
/// Voice-note container settings for the finalized `.m4a`, mirroring
|
||||
/// `VoiceRecorder.startRecording()`.
|
||||
static var voiceNoteFileSettings: [String: Any] {
|
||||
[
|
||||
AVFormatIDKey: kAudioFormatMPEG4AAC,
|
||||
AVSampleRateKey: sampleRate,
|
||||
AVNumberOfChannelsKey: Int(channelCount),
|
||||
AVEncoderBitRateKey: bitRate
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
/// Builds ADTS-framed AAC so a receiver can persist a burst progressively:
|
||||
/// unlike `.m4a` (whose moov atom only exists after close), an ADTS `.aac`
|
||||
/// stream is playable at any prefix — a partially received burst is still a
|
||||
/// replayable voice note.
|
||||
enum ADTSFramer {
|
||||
private static let headerSize = 7
|
||||
/// MPEG-4 sampling frequency index for 16 kHz.
|
||||
private static let samplingFrequencyIndex: UInt8 = 8
|
||||
private static let channelConfiguration: UInt8 = 1
|
||||
|
||||
/// Wraps one raw AAC-LC frame in an ADTS header.
|
||||
static func frame(_ aacFrame: Data) -> Data {
|
||||
let frameLength = aacFrame.count + headerSize
|
||||
var data = Data(capacity: frameLength)
|
||||
// Syncword 0xFFF, MPEG-4, layer 00, no CRC.
|
||||
data.append(0xFF)
|
||||
data.append(0xF1)
|
||||
// Profile AAC-LC (audio object type 2 -> bits 01), frequency index,
|
||||
// private bit 0, channel config high bit.
|
||||
data.append((0b01 << 6) | (samplingFrequencyIndex << 2) | ((channelConfiguration >> 2) & 0x1))
|
||||
data.append(((channelConfiguration & 0x3) << 6) | UInt8((frameLength >> 11) & 0x3))
|
||||
data.append(UInt8((frameLength >> 3) & 0xFF))
|
||||
data.append(UInt8((frameLength & 0x7) << 5) | 0x1F)
|
||||
// Buffer fullness 0x7FF (VBR), one AAC frame per ADTS frame.
|
||||
data.append(0xFC)
|
||||
data.append(aacFrame)
|
||||
return data
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,144 @@
|
||||
//
|
||||
// PTTBurstPlayer.swift
|
||||
// bitchat
|
||||
//
|
||||
// This is free and unencumbered software released into the public domain.
|
||||
// For more information, see <https://unlicense.org>
|
||||
//
|
||||
|
||||
import AVFoundation
|
||||
import BitLogger
|
||||
import Foundation
|
||||
|
||||
/// Plays one inbound live voice burst with a small jitter buffer.
|
||||
///
|
||||
/// Frames are decoded and scheduled back-to-back on an `AVAudioPlayerNode`;
|
||||
/// an underrun (missing/late packets) simply pauses output until the next
|
||||
/// buffer arrives, which self-heals timing without explicit silence
|
||||
/// insertion. Playback starts once `TransportConfig.pttJitterBufferSeconds`
|
||||
/// of audio is queued or `pttJitterDeadlineSeconds` has elapsed.
|
||||
@MainActor
|
||||
final class PTTBurstPlayer {
|
||||
private let engine = AVAudioEngine()
|
||||
private let node = AVAudioPlayerNode()
|
||||
private let decoder: PTTFrameDecoder
|
||||
|
||||
private var queuedBuffers: [AVAudioPCMBuffer] = []
|
||||
private var queuedDuration: TimeInterval = 0
|
||||
private var scheduledCount = 0
|
||||
private var engineStarted = false
|
||||
private var finished = false
|
||||
private var stopped = false
|
||||
private var deadlineTask: Task<Void, Never>?
|
||||
|
||||
private(set) var isPlaying = false
|
||||
|
||||
init?() {
|
||||
guard let format = PTTAudioFormat.pcmFormat, let decoder = PTTFrameDecoder() else { return nil }
|
||||
self.decoder = decoder
|
||||
engine.attach(node)
|
||||
engine.connect(node, to: engine.mainMixerNode, format: format)
|
||||
|
||||
deadlineTask = Task { [weak self] in
|
||||
try? await Task.sleep(nanoseconds: UInt64(TransportConfig.pttJitterDeadlineSeconds * 1_000_000_000))
|
||||
self?.startIfReady(force: true)
|
||||
}
|
||||
}
|
||||
|
||||
/// Decodes and queues frames (in burst order). Starts playback when the
|
||||
/// jitter buffer fills.
|
||||
func enqueue(_ frames: [Data]) {
|
||||
guard !stopped else { return }
|
||||
for frame in frames {
|
||||
guard let pcm = decoder.decode(frame) else { continue }
|
||||
if engineStarted {
|
||||
schedule(pcm)
|
||||
} else {
|
||||
queuedBuffers.append(pcm)
|
||||
queuedDuration += Double(pcm.frameLength) / PTTAudioFormat.sampleRate
|
||||
}
|
||||
}
|
||||
startIfReady(force: false)
|
||||
}
|
||||
|
||||
/// The burst ended: stop once everything scheduled has played out.
|
||||
func finishAfterDrain() {
|
||||
finished = true
|
||||
stopIfDrained()
|
||||
}
|
||||
|
||||
/// Immediate stop (cancel, another playback taking over, teardown).
|
||||
func stop() {
|
||||
guard !stopped else { return }
|
||||
stopped = true
|
||||
deadlineTask?.cancel()
|
||||
queuedBuffers = []
|
||||
if engineStarted {
|
||||
node.stop()
|
||||
engine.stop()
|
||||
}
|
||||
isPlaying = false
|
||||
VoiceNotePlaybackCoordinator.shared.deactivate(self)
|
||||
}
|
||||
|
||||
private func startIfReady(force: Bool) {
|
||||
guard !engineStarted, !stopped, !queuedBuffers.isEmpty else { return }
|
||||
guard force || queuedDuration >= TransportConfig.pttJitterBufferSeconds else { return }
|
||||
|
||||
#if os(iOS)
|
||||
do {
|
||||
let session = AVAudioSession.sharedInstance()
|
||||
try session.setCategory(.playback, mode: .spokenAudio, options: [.mixWithOthers])
|
||||
try session.setActive(true, options: [])
|
||||
} catch {
|
||||
SecureLogger.error("PTT playback session activation failed: \(error)", category: .session)
|
||||
}
|
||||
#endif
|
||||
|
||||
engine.prepare()
|
||||
do {
|
||||
try engine.start()
|
||||
} catch {
|
||||
SecureLogger.error("PTT playback engine failed to start: \(error)", category: .session)
|
||||
stopped = true
|
||||
return
|
||||
}
|
||||
engineStarted = true
|
||||
isPlaying = true
|
||||
VoiceNotePlaybackCoordinator.shared.activate(self)
|
||||
node.play()
|
||||
|
||||
let buffered = queuedBuffers
|
||||
queuedBuffers = []
|
||||
queuedDuration = 0
|
||||
for buffer in buffered {
|
||||
schedule(buffer)
|
||||
}
|
||||
}
|
||||
|
||||
private func schedule(_ buffer: AVAudioPCMBuffer) {
|
||||
scheduledCount += 1
|
||||
node.scheduleBuffer(buffer) { [weak self] in
|
||||
Task { @MainActor [weak self] in
|
||||
guard let self else { return }
|
||||
self.scheduledCount -= 1
|
||||
self.stopIfDrained()
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
private func stopIfDrained() {
|
||||
guard finished, scheduledCount <= 0 else { return }
|
||||
stop()
|
||||
}
|
||||
}
|
||||
|
||||
extension PTTBurstPlayer: ExclusivePlayback {
|
||||
/// A live stream can't meaningfully pause; yielding the floor stops it.
|
||||
/// The burst keeps assembling to file, so nothing is lost.
|
||||
nonisolated func pauseForExclusivity() {
|
||||
Task { @MainActor [weak self] in
|
||||
self?.stop()
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,171 @@
|
||||
//
|
||||
// PTTCaptureEngine.swift
|
||||
// bitchat
|
||||
//
|
||||
// This is free and unencumbered software released into the public domain.
|
||||
// For more information, see <https://unlicense.org>
|
||||
//
|
||||
|
||||
import AVFoundation
|
||||
import BitLogger
|
||||
import Foundation
|
||||
|
||||
/// Captures microphone audio for a live push-to-talk burst, producing both:
|
||||
/// - live AAC frames via `onFrames` (called on the capture queue), and
|
||||
/// - a finalized `.m4a` voice note on `stop()` — the same artifact
|
||||
/// `VoiceRecorder` produces, so the existing voice-note send pipeline
|
||||
/// handles delivery to receivers that missed the live stream.
|
||||
final class PTTCaptureEngine {
|
||||
/// Hard cap matching `VoiceRecorder.maxRecordingDuration`: past it the
|
||||
/// engine keeps running (the UI owns the gesture) but stops encoding.
|
||||
private static let maxCaptureDuration: TimeInterval = 120
|
||||
|
||||
/// Recreated on every `start()`: an engine whose input unit was
|
||||
/// instantiated against an earlier (playback-only or inactive) audio
|
||||
/// session keeps reporting a dead 0 Hz / 2 ch input format and fails to
|
||||
/// enable the mic (AURemoteIO -10851, observed on iPhone field tests).
|
||||
private var engine = AVAudioEngine()
|
||||
private let queue = DispatchQueue(label: "chat.bitchat.ptt.capture", qos: .userInitiated)
|
||||
|
||||
// Capture-queue-confined state.
|
||||
private var resampler: PTTInputResampler?
|
||||
private var encoder: PTTFrameEncoder?
|
||||
private var file: AVAudioFile?
|
||||
private var fileURL: URL?
|
||||
private var encodedFrameCount = 0
|
||||
private var running = false
|
||||
private var captureStart = Date()
|
||||
|
||||
/// Called on the capture queue with each batch of encoded AAC frames.
|
||||
var onFrames: (([Data]) -> Void)?
|
||||
|
||||
enum CaptureError: Error {
|
||||
case inputUnavailable
|
||||
case audioSetupFailed
|
||||
}
|
||||
|
||||
func start(outputURL: URL) throws {
|
||||
#if os(iOS)
|
||||
try Self.configureAudioSession()
|
||||
#endif
|
||||
|
||||
// Fresh engine per capture so its input unit binds to the session
|
||||
// that is active *now* (see `engine` doc comment).
|
||||
engine = AVAudioEngine()
|
||||
let inputFormat = engine.inputNode.outputFormat(forBus: 0)
|
||||
guard inputFormat.sampleRate > 0, inputFormat.channelCount > 0 else {
|
||||
SecureLogger.error("PTT: capture input unavailable (input reports \(Int(inputFormat.sampleRate)) Hz, \(inputFormat.channelCount) ch)", category: .session)
|
||||
throw CaptureError.inputUnavailable
|
||||
}
|
||||
guard let resampler = PTTInputResampler(inputFormat: inputFormat),
|
||||
let encoder = PTTFrameEncoder(),
|
||||
let pcmFormat = PTTAudioFormat.pcmFormat
|
||||
else { throw CaptureError.audioSetupFailed }
|
||||
|
||||
let file = try AVAudioFile(
|
||||
forWriting: outputURL,
|
||||
settings: PTTAudioFormat.voiceNoteFileSettings,
|
||||
commonFormat: pcmFormat.commonFormat,
|
||||
interleaved: pcmFormat.isInterleaved
|
||||
)
|
||||
|
||||
queue.sync {
|
||||
self.resampler = resampler
|
||||
self.encoder = encoder
|
||||
self.file = file
|
||||
self.fileURL = outputURL
|
||||
self.encodedFrameCount = 0
|
||||
self.captureStart = Date()
|
||||
self.running = true
|
||||
}
|
||||
|
||||
engine.inputNode.installTap(onBus: 0, bufferSize: 4096, format: inputFormat) { [weak self] buffer, _ in
|
||||
self?.queue.async { self?.process(buffer) }
|
||||
}
|
||||
engine.prepare()
|
||||
do {
|
||||
try engine.start()
|
||||
} catch {
|
||||
SecureLogger.error("PTT: capture engine failed to start (input: \(Int(inputFormat.sampleRate)) Hz, \(inputFormat.channelCount) ch): \(error)", category: .session)
|
||||
engine.inputNode.removeTap(onBus: 0)
|
||||
queue.sync { self.teardown(deleteFile: true) }
|
||||
throw error
|
||||
}
|
||||
SecureLogger.info("PTT: capture engine running (input: \(Int(inputFormat.sampleRate)) Hz, \(inputFormat.channelCount) ch)", category: .session)
|
||||
}
|
||||
|
||||
/// Stops capture and finalizes the `.m4a`. Returns the file URL and the
|
||||
/// number of encoded AAC frames (each `PTTAudioFormat.frameDuration` long).
|
||||
func stop() -> (url: URL?, encodedFrames: Int) {
|
||||
engine.inputNode.removeTap(onBus: 0)
|
||||
engine.stop()
|
||||
let result: (URL?, Int) = queue.sync {
|
||||
let url = fileURL
|
||||
let frames = encodedFrameCount
|
||||
teardown(deleteFile: false)
|
||||
return (url, frames)
|
||||
}
|
||||
#if os(iOS)
|
||||
Self.deactivateAudioSession()
|
||||
#endif
|
||||
return result
|
||||
}
|
||||
|
||||
func cancel() {
|
||||
engine.inputNode.removeTap(onBus: 0)
|
||||
engine.stop()
|
||||
queue.sync { teardown(deleteFile: true) }
|
||||
#if os(iOS)
|
||||
Self.deactivateAudioSession()
|
||||
#endif
|
||||
}
|
||||
|
||||
// MARK: - Capture queue
|
||||
|
||||
private func process(_ buffer: AVAudioPCMBuffer) {
|
||||
guard running,
|
||||
Date().timeIntervalSince(captureStart) < Self.maxCaptureDuration,
|
||||
let resampled = resampler?.resample(buffer)
|
||||
else { return }
|
||||
|
||||
do {
|
||||
try file?.write(from: resampled)
|
||||
} catch {
|
||||
SecureLogger.error("PTT capture file write failed: \(error)", category: .session)
|
||||
}
|
||||
|
||||
guard let frames = encoder?.encode(resampled), !frames.isEmpty else { return }
|
||||
encodedFrameCount += frames.count
|
||||
onFrames?(frames)
|
||||
}
|
||||
|
||||
private func teardown(deleteFile: Bool) {
|
||||
running = false
|
||||
// Releasing the AVAudioFile finalizes the .m4a container.
|
||||
file = nil
|
||||
encoder = nil
|
||||
resampler = nil
|
||||
if deleteFile, let url = fileURL {
|
||||
try? FileManager.default.removeItem(at: url)
|
||||
}
|
||||
fileURL = nil
|
||||
}
|
||||
|
||||
// MARK: - Audio session (iOS)
|
||||
|
||||
#if os(iOS)
|
||||
private static func configureAudioSession() throws {
|
||||
let session = AVAudioSession.sharedInstance()
|
||||
#if targetEnvironment(simulator)
|
||||
try session.setCategory(.playAndRecord, mode: .default, options: [.defaultToSpeaker, .allowBluetoothA2DP])
|
||||
#else
|
||||
try session.setCategory(.playAndRecord, mode: .default, options: [.defaultToSpeaker, .allowBluetoothA2DP, .allowBluetoothHFP])
|
||||
#endif
|
||||
try session.setActive(true, options: .notifyOthersOnDeactivation)
|
||||
}
|
||||
|
||||
private static func deactivateAudioSession() {
|
||||
try? AVAudioSession.sharedInstance().setActive(false, options: .notifyOthersOnDeactivation)
|
||||
}
|
||||
#endif
|
||||
}
|
||||
@@ -0,0 +1,39 @@
|
||||
//
|
||||
// PTTSettings.swift
|
||||
// bitchat
|
||||
//
|
||||
// This is free and unencumbered software released into the public domain.
|
||||
// For more information, see <https://unlicense.org>
|
||||
//
|
||||
|
||||
import Foundation
|
||||
#if os(iOS)
|
||||
import UIKit
|
||||
#elseif os(macOS)
|
||||
import AppKit
|
||||
#endif
|
||||
|
||||
/// User preference for live push-to-talk voice. One switch controls both
|
||||
/// directions: streaming your holds live, and auto-playing inbound bursts.
|
||||
/// Off means voice messages behave exactly like classic voice notes.
|
||||
enum PTTSettings {
|
||||
private static let liveVoiceEnabledKey = "ptt.liveVoiceEnabled"
|
||||
|
||||
static var liveVoiceEnabled: Bool {
|
||||
get { UserDefaults.standard.object(forKey: liveVoiceEnabledKey) as? Bool ?? true }
|
||||
set { UserDefaults.standard.set(newValue, forKey: liveVoiceEnabledKey) }
|
||||
}
|
||||
|
||||
/// Autoplay is foreground-only: audio must never start from the
|
||||
/// background.
|
||||
@MainActor
|
||||
static var isAppActive: Bool {
|
||||
#if os(iOS)
|
||||
return UIApplication.shared.applicationState == .active
|
||||
#elseif os(macOS)
|
||||
return NSApplication.shared.isActive
|
||||
#else
|
||||
return true
|
||||
#endif
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,183 @@
|
||||
//
|
||||
// VoiceCaptureSession.swift
|
||||
// bitchat
|
||||
//
|
||||
// This is free and unencumbered software released into the public domain.
|
||||
// For more information, see <https://unlicense.org>
|
||||
//
|
||||
|
||||
import BitFoundation
|
||||
import BitLogger
|
||||
import Foundation
|
||||
|
||||
/// Capture backend behind the composer's hold-to-record gesture.
|
||||
/// `VoiceRecordingViewModel` drives one session per press; the concrete type
|
||||
/// decides *how* audio leaves the device: `VoiceNoteCaptureSession` records a
|
||||
/// note delivered on release (today's behavior), `PTTLiveVoiceSession`
|
||||
/// additionally streams frames live while the button is held.
|
||||
@MainActor
|
||||
protocol VoiceCaptureSession: AnyObject {
|
||||
/// Whether audio is leaving the device in real time while recording —
|
||||
/// drives the composer's LIVE treatment.
|
||||
var isLive: Bool { get }
|
||||
func requestPermission() async -> Bool
|
||||
func start() async throws
|
||||
/// Stops capture and returns the finalized voice-note file, or nil when
|
||||
/// nothing valid was captured.
|
||||
func finish() async -> URL?
|
||||
func cancel() async
|
||||
}
|
||||
|
||||
/// The classic record-then-send backend, wrapping the shared `VoiceRecorder`.
|
||||
@MainActor
|
||||
final class VoiceNoteCaptureSession: VoiceCaptureSession {
|
||||
var isLive: Bool { false }
|
||||
|
||||
func requestPermission() async -> Bool {
|
||||
await VoiceRecorder.shared.requestPermission()
|
||||
}
|
||||
|
||||
func start() async throws {
|
||||
try await VoiceRecorder.shared.startRecording()
|
||||
}
|
||||
|
||||
func finish() async -> URL? {
|
||||
await VoiceRecorder.shared.stopRecording()
|
||||
}
|
||||
|
||||
func cancel() async {
|
||||
await VoiceRecorder.shared.cancelRecording()
|
||||
}
|
||||
}
|
||||
|
||||
/// Live push-to-talk backend: streams `VoiceBurstPacket`s to one peer while
|
||||
/// recording, then finalizes the same audio as a standard voice note whose
|
||||
/// file name carries the burst ID (`voice_<burstID>.m4a`) so receivers that
|
||||
/// heard the live stream absorb the note silently instead of seeing a
|
||||
/// duplicate.
|
||||
@MainActor
|
||||
final class PTTLiveVoiceSession: VoiceCaptureSession {
|
||||
let burstID: Data
|
||||
|
||||
private let sendPacket: (Data) -> Void
|
||||
private let capture = PTTCaptureEngine()
|
||||
/// Capture-queue-confined stream state: packetizes frames and lazily
|
||||
/// emits START so packet order is guaranteed by queue serialization.
|
||||
private final class StreamState {
|
||||
var packetizer: VoiceBurstPacketizer
|
||||
var sentStart = false
|
||||
init(burstID: Data) {
|
||||
packetizer = VoiceBurstPacketizer(burstID: burstID)
|
||||
}
|
||||
}
|
||||
private let stream: StreamState
|
||||
private var startDate: Date?
|
||||
private var completed = false
|
||||
|
||||
var isLive: Bool { true }
|
||||
|
||||
/// - Parameter sendPacket: delivers one encoded `VoiceBurstPacket` to the
|
||||
/// target peer; must be safe to call from any queue (BLEService hops to
|
||||
/// its own message queue internally).
|
||||
init(sendPacket: @escaping (Data) -> Void) {
|
||||
self.burstID = VoiceBurstPacket.makeBurstID()
|
||||
self.sendPacket = sendPacket
|
||||
self.stream = StreamState(burstID: burstID)
|
||||
}
|
||||
|
||||
func requestPermission() async -> Bool {
|
||||
await VoiceRecorder.shared.requestPermission()
|
||||
}
|
||||
|
||||
func start() async throws {
|
||||
let outputURL = try Self.makeOutputURL(burstID: burstID)
|
||||
let sendPacket = sendPacket
|
||||
let stream = stream
|
||||
capture.onFrames = { frames in
|
||||
if !stream.sentStart {
|
||||
stream.sentStart = true
|
||||
if let start = VoiceBurstPacket(
|
||||
burstID: stream.packetizer.burstID,
|
||||
seq: 0,
|
||||
kind: .start(codec: .aacLC16kMono)
|
||||
) {
|
||||
sendPacket(start.encode())
|
||||
}
|
||||
}
|
||||
for frame in frames {
|
||||
for packet in stream.packetizer.add(frame) {
|
||||
sendPacket(packet)
|
||||
}
|
||||
}
|
||||
// Flush per callback batch: at ~130-byte frames the budget fits
|
||||
// one frame per packet anyway, and holding residue would add
|
||||
// ~100 ms of avoidable latency.
|
||||
for packet in stream.packetizer.flush() {
|
||||
sendPacket(packet)
|
||||
}
|
||||
}
|
||||
do {
|
||||
try capture.start(outputURL: outputURL)
|
||||
} catch {
|
||||
// The HAL can briefly report a dead input right after the audio
|
||||
// session (re)activates while the route settles; one retry after
|
||||
// a short pause covers it (observed on iPhone field tests).
|
||||
SecureLogger.warning("PTT: capture start failed (\(error)) — retrying once after route settle", category: .session)
|
||||
try? await Task.sleep(nanoseconds: 150_000_000)
|
||||
try capture.start(outputURL: outputURL)
|
||||
}
|
||||
startDate = Date()
|
||||
SecureLogger.info("PTT: live burst \(burstID.hexEncodedString()) capture started", category: .session)
|
||||
}
|
||||
|
||||
func finish() async -> URL? {
|
||||
guard !completed else { return nil }
|
||||
completed = true
|
||||
|
||||
let elapsed = startDate.map { Date().timeIntervalSince($0) } ?? 0
|
||||
let (url, encodedFrames) = capture.stop()
|
||||
// stop() drained the capture queue, so touching `stream` is safe now.
|
||||
|
||||
guard elapsed >= VoiceRecorder.minRecordingDuration, let url else {
|
||||
sendControlPacket(.canceled)
|
||||
if let url {
|
||||
try? FileManager.default.removeItem(at: url)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
for packet in stream.packetizer.flush() {
|
||||
sendPacket(packet)
|
||||
}
|
||||
let durationMs = UInt32((Double(encodedFrames) * PTTAudioFormat.frameDuration * 1000).rounded())
|
||||
sendControlPacket(.end(totalDataPackets: stream.packetizer.dataPacketCount, durationMs: durationMs))
|
||||
SecureLogger.info("PTT: live burst \(burstID.hexEncodedString()) finished — \(stream.packetizer.dataPacketCount) data packets, \(encodedFrames) frames, \(durationMs) ms", category: .session)
|
||||
return url
|
||||
}
|
||||
|
||||
func cancel() async {
|
||||
guard !completed else { return }
|
||||
completed = true
|
||||
capture.cancel()
|
||||
sendControlPacket(.canceled)
|
||||
}
|
||||
|
||||
private func sendControlPacket(_ kind: VoiceBurstPacket.Kind) {
|
||||
guard let packet = VoiceBurstPacket(burstID: burstID, seq: stream.packetizer.nextSeq, kind: kind) else { return }
|
||||
sendPacket(packet.encode())
|
||||
}
|
||||
|
||||
private static func makeOutputURL(burstID: Data) throws -> URL {
|
||||
let base = try FileManager.default.url(
|
||||
for: .applicationSupportDirectory,
|
||||
in: .userDomainMask,
|
||||
appropriateFor: nil,
|
||||
create: true
|
||||
)
|
||||
let directory = base
|
||||
.appendingPathComponent("files", isDirectory: true)
|
||||
.appendingPathComponent("voicenotes/outgoing", isDirectory: true)
|
||||
try FileManager.default.createDirectory(at: directory, withIntermediateDirectories: true, attributes: nil)
|
||||
return directory.appendingPathComponent("voice_\(burstID.hexEncodedString()).m4a")
|
||||
}
|
||||
}
|
||||
@@ -181,23 +181,35 @@ final class VoiceNotePlaybackController: NSObject, ObservableObject, AVAudioPlay
|
||||
}
|
||||
}
|
||||
|
||||
/// Ensures only one voice note plays at a time.
|
||||
/// Something that can hold the app's single audio-playback slot and yield it
|
||||
/// when another playback starts (voice notes pause; live bursts stop).
|
||||
protocol ExclusivePlayback: AnyObject {
|
||||
func pauseForExclusivity()
|
||||
}
|
||||
|
||||
extension VoiceNotePlaybackController: ExclusivePlayback {
|
||||
func pauseForExclusivity() {
|
||||
pause()
|
||||
}
|
||||
}
|
||||
|
||||
/// Ensures only one voice playback (note or live burst) runs at a time.
|
||||
final class VoiceNotePlaybackCoordinator {
|
||||
static let shared = VoiceNotePlaybackCoordinator()
|
||||
|
||||
private weak var activeController: VoiceNotePlaybackController?
|
||||
private weak var activeController: (any ExclusivePlayback)?
|
||||
|
||||
private init() {}
|
||||
|
||||
func activate(_ controller: VoiceNotePlaybackController) {
|
||||
func activate(_ controller: any ExclusivePlayback) {
|
||||
if activeController === controller {
|
||||
return
|
||||
}
|
||||
activeController?.pause()
|
||||
activeController?.pauseForExclusivity()
|
||||
activeController = controller
|
||||
}
|
||||
|
||||
func deactivate(_ controller: VoiceNotePlaybackController) {
|
||||
func deactivate(_ controller: any ExclusivePlayback) {
|
||||
if activeController === controller {
|
||||
activeController = nil
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user