Windows audio (WASAPI)
native/windows_audio/wasapi.cpp, bound from Python by tonesphere/native/wasapi.py.
The Python PortAudio host (tonesphere/engine/host.py) remains for Linux and macOS; on
Windows the native backend replaces it for the audio path (the switch-over in
AudioEngine is migration step M8 — until then the native backend is driven directly).
Endpoints
wasapi.endpoints() enumerates active render and capture endpoints through
IMMDeviceEnumerator. Each is identified by its MMDevice endpoint ID
({0.0.0.00000000}.{guid}) — the string Windows itself uses, stable across reboots,
re-enumeration and unplug/replug — rather than a PortAudio index, which shifts whenever a
device comes or goes. Reported per endpoint: friendly name, flow, which default roles it
holds (console, multimedia, communications), the shared-mode mix format, the default and
minimum device periods, the IAudioClient3 low-latency shared-mode period range, and
whether it supports raw processing.
wasapi.watch() registers an IMMNotificationClient; wasapi.poll_events() returns
arrival, removal, state and default-device changes. Notifications arrive on Windows’
threads and are queued there, never on the audio thread.
Streams and threads
NativeEngine.start_wasapi([StreamSpec, ...], master=i) opens every stream on its own
thread (MTA, MMCSS “Pro Audio”). The master stream’s device clock drives the engine:
on each device event its thread runs exactly as many frames as the device wants, cut into
equal blocks of at most the engine’s block size — a 144-frame exclusive period with a
128-frame engine becomes 2 × 72, never 128 + 16, because load is judged per block against
that block’s own period and a 16-frame remainder’s 0.33 ms budget made fixed per-block
costs read as a 152 % overload that never happened. Every other stream is a satellite
on its own clock: it meets the master through a wait-free ring, read through the native
drift resampler on the consuming side.
The drift resampler primes the ring to two device periods, then spends 64 blocks at
ratio 1 learning the fill this pair of streams actually settles at (sampled before each
read, a packet-fed ring averages below its primed level; steering at the primed level
corrects a drift that is not there, and on the AI-04 that warped the first seconds of
every stream). After that a proportional-integral loop steers the fill back to that
setpoint within ±0.5 % of real time, with a ±5 % dead band so two streams on one clock
run at a ratio of exactly 1 (native/engine/resampler.h).
Stream kinds: render, capture, loopback (everything a render endpoint plays) and
process_loopback (one process’s audio, Windows 10 build 20348+).
Shared mode. When the device’s mix format is float32 at the engine rate and width,
the stream uses IAudioClient3::InitializeSharedAudioStream at the smallest period that
covers one engine block (on the development machine, 480 frames = 10 ms). Otherwise it
asks Windows to convert (AUTOCONVERTPCM).
Exclusive mode. Formats are tried in order — float32, int32 with 24 valid bits, int32,
packed int24, int16 — at the engine rate, and the stream reports which it got. A refusal
falls back to shared mode only if the spec allows it, and the status then says
exclusive: False with the refusal’s reason; it is never silent.
Raw mode, on by default. A shared stream asks for raw processing
(AUDCLNT_STREAMOPTIONS_RAW), which bypasses the driver’s enhancement effects on that
stream. This matters: on the development machine, a stream without raw mode came back
+9.7 dB louder, rising over the first second — the Realtek driver’s loudness
processing — measured by playing a tone through PortAudio (not ToneSphere) and capturing
it, so it is the driver, not this code. Whether the endpoint honoured raw mode is in the
stream status.
What is proven, on what
Development machine, 2026-09-29: Windows 11 Pro 25H2 (build 26200), Realtek ALC257
(Realtek HD Audio 2nd output, the default render endpoint; Microphone), Intel Smart
Sound Technology microphone array. All figures are from tests/hardware/test_wasapi.py.
| Claim | Level | Evidence |
|---|---|---|
| Enumeration: stable endpoint IDs, names, formats, periods, default roles, raw support | HARDWARE VERIFIED | test_endpoints_have_stable_ids_names_and_formats — 4 endpoints, all raw-capable |
| Shared render → process loopback, bit-exact | HARDWARE VERIFIED | test_raw_render_is_bit_exact_through_process_loopback: 1 s of white noise, max sample error 0.00e+00. Skips itself on a run where the drift resampler engaged (interpolated samples are correct audio, not identical samples) |
| Shared render → whole-endpoint loopback, level exact at 1 kHz, no gaps | HARDWARE VERIFIED | test_raw_render_comes_back_whole_at_the_level_sent, both loopback kinds |
| Render-to-capture delay | measured | 3521 frames = 73.4 ms (process loopback), 2993 frames = 62.4 ms (endpoint loopback), by cross-correlation. This is the digital path through the Windows audio engine, not an acoustic round trip |
| Endpoint effects after the mix | recorded | whole-endpoint loopback shows a high shelf on this machine: 0.00 dB to 1 kHz, +0.08 (1–4 kHz), +0.61 (4–12 kHz), +0.92 dB (12–20 kHz) — the endpoint’s own processing, which raw mode on one stream cannot bypass |
| Exclusive mode | HARDWARE VERIFIED (runs; output not captured) | Realtek granted exclusive: int32 container / 24 valid bits, 480-frame buffer, reported latency 10.0 ms, 0 glitches, 0 xruns. Loopback cannot capture an exclusive stream, so the audio itself was not verified this way |
| Three clocks: microphone master + render and loopback satellites | HARDWARE VERIFIED | test_a_capture_master_drives_render_and_loopback_satellites: tone arrives at the sent level (±1 %), no gaps, no satellite underrun |
| Audio thread on the device: callback time, allocations | measured | render master at 480 frames: mean 15–19 µs, max 26–34 µs per block of 10 ms; 0 xruns; 0 audio-thread allocations |
| Start/stop cycles | HARDWARE VERIFIED | 5 cycles, no threads left behind |
| A missing endpoint fails with a reason | HARDWARE VERIFIED | test_a_bad_endpoint_fails_with_a_reason_not_silence |
| Device notifications registered and polled | IMPLEMENTED | test_device_notifications_can_be_watched |
| Device removal/arrival while streaming | NOT VERIFIED | no device was unplugged during a test; a stream whose device disappears is designed to fail with AUDCLNT_E_DEVICE_INVALIDATED in its status, but that has not been exercised |
| Capture of a real microphone signal’s content | NOT VERIFIED | the capture master test proves the microphone clock drives the engine; nothing asserted what the microphone heard |
On a USB interface: Audio Array AI-04
2026-09-30, the same machine with an Audio Array AI-04 attached (USB, C-Media VID 0D8C
PID 0269, Windows’ class driver — the maker publishes no driver) and made the default
device: a guitar on input 1, earphones on the output. tests/hardware/test_interface.py
(set TONESPHERE_TEST_INTERFACE for another interface), 20 s per run:
| Mode | Device period | Callback mean / p99 / max | Worst load | Dropouts |
|---|---|---|---|---|
| Exclusive 48 kHz, 24-bit | 144 frames (3.0 ms reported) | 5.8 / 16.0 / 209.4 µs | 14.0 % of a 72-frame (1.5 ms) block | 0 glitches, 0 xruns, 0 underruns |
| Exclusive 44.1 kHz, 24-bit | 132 frames (3.0 ms reported) | 4.7 / 16.0 / 152.7 µs | 10.2 % of a 66-frame block | 0 xruns; input underruns 0–65 frames per run, at start (below) |
| Shared 48 kHz, float32 | 480 frames (10 ms) | 25.1 / 45.3 / 107.5 µs | 1.1 % | 0 xruns; 1 capture discontinuity flag at stream start (WASAPI’s own) |
In every mode, 0 audio-thread allocations. What the input delivered is a real signal from
a physical source: input 1, the guitar, carried its pickup’s mains hum at 50.0 Hz
(Kolkata mains is 50 Hz), −19.9 dBFS rms; input 2, empty, its noise floor at −63.5 dBFS.
Monitoring input → gain bus → output at −40 dB gave the output exactly the input × 0.01,
sample for sample (test_the_monitoring_path_carries_the_input_at_the_gain_set). The
render → loopback digital path measures 33.35 or 43.35 ms (1601 or 2081 frames — one
480-frame period apart, the phase of each start), identical within a run, confidence 34.3.
The AI-04’s endpoint applies no processing of its own: loopback came back at +0.00 dB in
every band, raw or not (the Realtek applies a high shelf).
Its 44.1 kHz input clock. At 44.1 kHz the AI-04’s input delivers 0.2–0.3 % fewer frames than its output consumes (792 frames short in 308,748 at the first measurement), and the shortfall wanders by ±0.15 % over seconds; at 48 kHz the two sides agree exactly. With the first drift resampler (a ±0.1 % limit) that was 355 frames lost every 20 s. Now the resampler follows it: no ongoing loss, but a run can lose a few frames (0–65 measured) while it converges, and the ratio follows the input’s wander — a pitch change of at most ±2.6 cents, over seconds. That case is recorded as a known failure in the test, not hidden. Use 48 kHz on this interface.
Measured round trip through the interface. Later on 2026-09-30 a 6.35 mm cable joined
the AI-04’s headphone output to its input (test_roundtrip.py), a −18 dBFS sweep each time,
captured back at a −8.0 dBFS peak (the path gains about 10 dB; nothing clips):
| Path | Measured round trip | Runs | Nominal / reported |
|---|---|---|---|
| Native WASAPI exclusive, 144-frame period | 17.92 ms (860 frames); 15.94 ms (765) in 2 of 10 starts | 10, confidence 16–24 | 3.0 ms / 6.0 ms |
| Native WASAPI exclusive, engine block 96 / 128 / 256 | 17.92 / 15.94 / 25.15 ms | 1 each | |
| Native WASAPI shared, 480-frame block | 76.9 ms (3688–3692 frames) | 7, confidence 13–19 | 10 ms / – |
| ASIO4ALL 2.22 at 64 frames | 15.58 ms (748 frames) | 3, identical | ASIO.md |
| FlexASIO defaults (shared, 882 frames) | 119.75 ms | 1 |
Within one start the exclusive figure is fixed to the frame; between starts it lands on one of two values 96 frames (two USB packet groups of 1 ms) apart, set by the phase the capture stream starts in. The same tone through the cable came back at the same level through native WASAPI, ASIO4ALL (+0.00 dB) and FlexASIO (−0.03 dB), which also establishes that sound leaves the output jack.
Whole-system loopback, as a source
Every WASAPI output is offered in the routing as “
Devices that come and go
engine/device_monitor.py drains Windows’ endpoint notifications on a control thread and,
300 ms after the last of a burst, has the engine re-enumerate. If a device the routing uses
left, arrived, or has a failed stream, the engine restarts with what is present: other
devices carry on after a gap of a few blocks (the backend opens and closes its streams
together), and a route to the missing device is kept, so when it returns it is reopened
under the same id with no user action. What is proven without hardware is the decision
(tests/test_device_monitor.py); a real device disappearing and returning is exercised on
the ToneSphere virtual cable in the driver VM.
Latency figures and what they mean
Every stream reports reported_latency_ms (IAudioClient::GetStreamLatency) — what
Windows says, not a measurement. The render-to-loopback delay above is a measurement,
but of a digital path that stops before the DAC. The acoustic round trip (output → air or
cable → input) needs a loopback cable or a microphone near a speaker, and is measured by
the round-trip tool of migration step M6; until then it is --.