Skip to content

feat(audio): low-latency engine: device-paced rendering, adaptive queues, whole-pipeline tests - #59

Open
Horuse wants to merge 43 commits into
devfrom
feat/low-latency
Open

Horuse wants to merge 43 commits into
devfrom
feat/low-latency

Conversation

@Horuse

@Horuse Horuse commented Oct 2, 2026

Copy link
Copy Markdown
Owner

What does this PR do?

Takes the engine from fixed, guessed buffering to latency that follows the engine buffer setting (32 to 2048 frames) and what is measured at runtime. Then it fixes what that exposed, and adds end-to-end tests that run whole pipelines with no sound hardware.

Rendering and buffers

  • Speakers render the graph in the device callback at the configured buffer size. A new buffer-size setting, a round-trip latency readout and a working-block badge come with it.
  • Inputs on the speaker's clock are locked and read one block per block. Every other input is absorbed by a drift-steered resampler instead of splices.
  • Capture buffers are sized to the engine block. The capture normalizer is woken by its capture instead of polling every 1 ms, and resamples one device block instead of 256 frames.
  • An adaptive queue-depth estimator sizes each live queue from measured jitter. It starts from the observed delivery size instead of a fixed 10 ms prior. Silence from a quiet source no longer deepens any buffer.
  • Graph edits keep unchanged nodes running across swaps. A swap that changes what is heard dips (fades out and in) instead of stepping. Edits reach the engine at once: the 400 ms debounce is gone.

Network

  • Wire packets carry one engine block (Opus from 2.5 ms; PCM up to the MTU). The receiver reads the packet size from the packet, so the protocol is unchanged.
  • The receive buffer is steered by the shared drift loop and sized from the link. It counts in the source's own rate. It has no 20 ms floor and no 60 ms start.
  • WebRTC send graphs run at 48 kHz. Joining waits on connection events and gives up on a silent signaling server.

Effects

  • Offloaded effects gather their own blocks, and their pad comes from measured processing time instead of a fixed 5 ms.
  • Gain and mute ramps, splices and fades are set in time, not blocks. The delay time glides like tape. The spectrum keeps the same resolution at any rate. Meter RMS covers the time between ticks.
  • Mono (and the last channel of an odd width) fills both sides of a stereo pair, so linked detectors hear it at its own level.

Correctness fixes found on the way

  • Allocations on the audio thread: input staging, network receive buffer, LUFS history (now histogram-gated), noise suppressor queues.
  • A speaker that stalls (a device overload skips callbacks) no longer keeps the input audio that piled up meanwhile as permanent latency.
  • File seeks are sample-accurate and never play the old position. File read-ahead is not reported as latency, and the playhead shows what is heard.
  • The monitor graph ticks at its own block. It used to run at 3% of real time, which starved every speaker fed by the same file.
  • A slow plugin load is waited for. A main-thread call that never started never runs late.
  • The virtual driver ring is sized in time and refuses stale reads (driver version 6).
  • Engine limits (buffer options, sample-rate bounds, defaults, EQ crossovers, and more) are generated into generated/engine.ts instead of copied in the frontend.

Tests

  • 99 whole-pipeline scenarios (src-tauri/tests/pipelines.rs), run through the app's own ActivePipeline on virtual devices. They cover every input kind × topology, every effect, effect + analyzer on the monitor, recordings, network loopback (PCM and Opus) and edits while playing. Each one checks that every worker keeps real time, no source runs dry and the tone comes out unbroken.
  • Per-effect mono, stereo and six-channel tests; seek, monitor-rate and stall regression tests.

Why is this the right approach?

  • Pull model. Rendering in the device callback removes the worker-to-device ring and its fixed depth. Latency then follows the buffer the user picks, instead of a ring target that ignored it.
  • Measured, not guessed. Queue depths, packet sizes, offload pads and resampler chunks are derived from the engine block, the source rate or runtime measurement. A nine-part audit of hard-coded constants drove most of the follow-up commits.
  • Separate devices are bridged, not merged. Separate devices stay bridged by rings plus resampling. On an output stall the queue realigns under that dropout, which is the usual way device bridges resync after an xrun. A private aggregate device on macOS (one IO cycle for same-clock devices, as DAWs use) was considered and left out as too large for this PR.
  • A Host seam instead of a parallel test harness. The pipeline takes a small Host (UI events plus system or virtual devices) instead of AppHandle and cpal. The tests drive the exact code the app runs. A test-only copy of the wiring would have tested itself; the monitor bug this PR fixes is one the old graph-level tests could not see.

Dependencies: cpal 0.17 → 0.18.2, so device xruns no longer end a stream. No new crates. ebur128 is now built optimized in dev, like rubato and tract, because unoptimized it cannot keep a 64-frame speaker in real time.

Checklist

  • Diff is limited to the change — no unrelated edits (large branch; each commit is one topic)
  • bun run test passes
  • bun run check passes
  • bun run format leaves the tree clean
  • Generated TS types are committed with the Rust change (LatencyReport.sampleRate, engine.ts)
  • No new dependency without a reason in the PR description
  • I read the RT audio path section of docs/CONCEPT.md and confirmed this change adds no allocations, locks, or syscalls to cpal / SCK callbacks or DspWorker::run. One documented exception: Doorbell::ring (Thread::unpark, non-blocking) wakes the capture normalizer and the net/WebRTC send threads instead of letting them poll.

Platform coverage

  • Developed on: macOS (Apple Silicon)
  • Tested on: macOS
  • What I did to test:
    • cargo test: 590 unit tests and 99 whole-pipeline scenarios on virtual devices; bun run test; bun run check.
    • By ear on MacBook built-in mic and speakers, and a Chrome process tap, at 32-frame buffers.
    • Network loopback on 127.0.0.1.
    • Long runs (hours) with the device-overload log.
  • Not tested: Linux, Windows. Not compiled locally; CI is the first build.

Per-OS files touched:

  • capture/linux.rs: PipeWire capture asks for node.latency equal to the engine block and promotes its thread with the real period.
  • capture/windows.rs: loopback is drained once per engine block (minimum 1 ms) instead of every 10 ms. The idle threshold stays at 10 ms, the shared-mode period.
  • pipeline/input/{linux,windows,macos}.rs and pipeline/output/{linux,windows,macos}.rs: the Host seam (AppHandle → Host), file-seek flush plumbing, and a Virtual arm that cannot be reached.
  • native/CATapCapture.swift (macOS): reacts to the output device starting; polls processes at 200 ms instead of 1 s.
  • native/virtual_driver/* (macOS): ring sized in time, masked indexing, stale-read check; driver version 6.

The Windows buffer setting still has no effect on the device period: cpal's WASAPI BufferSize::Fixed only deepens the ring. A lower period needs an IAudioClient3 path, deferred.

DepthEstimator sizes a queue from the jitter it measures (window mean minus
window low, 99th percentile with ~45 s half-life forgetting), so the target
no longer follows the level its own corrections move. Underruns raise it at
once; unused depth is released over about a minute.

rt_guard is a test-only counting allocator that proves a real-time path
never touches the heap.
… buffer size

- Speaker graphs render inside the device callback: no output ring, no
  sleeping worker. The engine block is a new buffer setting (32-2048
  frames, default 256); macOS requests the same device IO buffer via HAL.
- Live sources hold their queue at the measured headroom (cushion):
  startup backlog dropped, drift spliced out or stretched in with
  crossfades, an underrun re-primes once instead of clicking per block.
- Network receive sizes its jitter buffer with the same estimator and
  follows the engine block instead of a fixed 1024.
- Offload pad is max(block, effect floor): plugins 5 ms, noise suppressor
  one model hop. Offload threads run at real-time priority.
- latency_report replaces output_latency_ms: device buffers and hardware
  latency both ways, input queue, processing, output adapter, DSP load,
  overloads and per-node working block.
- GraphSpec reads camelCase, so sampleRate from the frontend is applied
  (it was silently ignored) and bufferFrames arrives.
- Full-width effects process at their real channel count; a mono source
  into an analyzer panicked.
…-block badge

- Settings: buffer size presets with their duration at the pipeline rate.
- Header: real round-trip latency with a per-stage breakdown, DSP load,
  and a warning when audio was dropped.
- Nodes that run at a larger block than the engine buffer show an
  hourglass next to the resampling icon.
- Storybook: mocked latency_report, bufferFrames/workingBlock controls,
  Larger block stories for Noise Suppressor and Plugin.
…ilding them

Monitor sources are keyed like their bridges, so they get their delivery counters.
Carried sources follow the new graph's clock lock, short splices end on the incoming audio, a speaker retired early or mid fade-in hands back cleanly.
…omes back from a dead device

The renderer is built before the stream opens; a plugin no longer reports its offload pad as a working block.
…rmat

The speaker renderer is built before its stream opens.
… engine limits with the frontend

- Wire packets carry one engine block (Opus from 2.5 ms); the receiver
  reads packet size from the packet and sizes its buffer from the link, in
  the source's own rate.
- Capture normalizer is woken by its capture and resamples one device
  block; monitor graph runs the engine block; offload pads from measured
  processing time.
- RT safety: input staging, network, LUFS and noise suppressor buffers are
  sized up front; LUFS uses histogram gating.
- Mono and odd channels fill both sides of a stereo pair; tests run every
  effect mono, stereo and six channels wide.
- Gain/mute ramps, splices, fades, meter RMS and gauges work in time, not
  blocks; declick delay follows its configured width.
- Edits reach the engine at once; hot swaps wait for fresh bridges instead
  of a fixed sleep; xrun thread wakes on stop.
- macOS tap reacts to the output device starting and polls at 200 ms.
- Virtual driver ring is sized in time, masked, and refuses stale reads.
- Engine limits are generated into generated/engine.ts for the frontend.
… block

The read position follows the set time frame by frame, read between
samples, at most 25% faster or slower than the audio, so moving the
slider bends the pitch briefly instead of clicking at every block.
The window scales with the rate (4096 frames at 48 kHz, 16384 at
192 kHz), so a bin stays about 11 Hz and low tones stay apart; the UI
sizes its FFT from the window it receives.
…on, WebRTC graph at 48 kHz

- A seek asks every graph reading the file to fade the old position out
  over 5 ms and drop it, waits for each to answer, then hands over the new
  position, which fades in. Dropping on the command raced the reader.
- Seeks land on the exact frame asked for, not the start of its packet.
- WebRTC send graphs run at 48 kHz like Opus net senders, so the engine
  converts the rate block by block and the encode thread's 256-frame
  resampler is gone; one rule picks every wire sender's rate.
…never started never runs

A main-thread call that has not started within 5 s is withdrawn, so it
cannot run later and leave an instance nobody owns. One that has started
is waited for, up to 60 s, so a large plugin that takes a while to load
opens instead of failing at 5 s.
…ent server

The guest hears each peer connection state as it happens instead of
polling every 250 ms. Reaching the signaling server and hearing the
host's offer are bounded, so a dead or silent server ends the attempt
instead of hanging the join.
…ot latency

- The monitor graph ran 32-frame blocks on a ticker still paced for
  1024, so it played about 3% of real time; a file feeding it and a
  speaker then stopped decoding for the speaker too, and the sound broke.
  The monitor is back on the large timer block, and every timer worker
  now ticks at the block its graph was built for. A test runs the real
  worker at both blocks and checks it keeps real time.
- A file's queue is audio decoded ahead, not held back: the latency
  report no longer counts it, and the playhead shows the position being
  heard rather than the one handed over.
…ual devices

The pipeline no longer takes Tauri's AppHandle and cpal devices directly
but a Host: where UI events go, and whether devices are the system's or
virtual ones. The app passes its window and the system's devices, as
before. Virtual devices exist only in memory: an input plays a tone on
its own clock through the same capture broadcast a device callback uses,
and a speaker pulls the real speaker renderer on its own clock and keeps
what it played. A testkit drives ActivePipeline on them for tests that
run on machines with no sound hardware.

ebur128 is built optimized in dev: unoptimized, a LUFS meter could not
keep a 64-frame speaker in real time.
Every input kind (WAV/FLAC/MP3 files, microphones at 44.1/48/96 kHz, a
four-channel interface, system and app audio) on one speaker at 256 and
32 frames, through a chain with a meter, on two speakers and on a
speaker at another rate; every effect alone; files through effects with
an analyzer on the monitor; recordings; network loopback over PCM and
Opus; edits while playing; mixes and multichannel speakers. Each runs in
real time through the app's engine and checks that every worker keeps
real time, no source runs dry and the sound comes out unbroken. They
run one at a time and need no sound hardware, so they run on CI.
A device overload makes the speaker skip callbacks while its inputs, on
their own devices, deliver on. That audio stayed queued: a mic locked to
the speaker's clock has no resampler to take it back down, so every
overload added its length to the latency for good, until the input ring
overflowed. The renderer now tells a gap of several periods from jitter,
and every live source realigns under that dropout, as after any silence,
without deepening its queue. Counted as OUTPUT_STALLS.
…ny rate

- A paused file reader is woken on resume and seek instead of noticing
  within 10 ms; dropping it no longer waits out its poll.
- The scope ring holds at least 85 ms at any rate (it was 43 ms at
  384 kHz, less than a late tick).
- Declick's end-of-click hold is a time, not three samples at any rate.
A plugin can be an end point of its own (one that sends its input over
the network, as Apple's AUNetSend does), but validation kept only what
reaches an output or an analyzer, so a plugin wired to nothing after it
was dropped before the engine saw it and its node waited forever. A
plugin now counts as an end point like an analyzer: it is validated,
and the monitor graph runs it when no output hears it.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant