Sound card: about 0.2 % of packets lost on a 48 kHz I2S wire, 0 % on a 24 kHz wire #49
Description
Activity
Evaluated tonight against the two leads from CircuitPython. Neither fixes this as it stands.
- CircuitPython 10.3.1's adaptive speaker endpoint answers a different problem. Its speaker has no feedback endpoint at all, and declaring it adaptive is what lets Windows' UAC2 driver start without one. Its source still says "async feedback for true clock matching is a later step". usbif's card already starts on Windows, and the clock half of this (Sound card: no feedback, so host and codec clocks drift (a tiny click about once a second at 44.1 kHz) #40) is now handled by explicit feedback in Sound card: send explicit feedback so the host follows the codec's clock #59. Adaptive wouldn't recover lost packets.
- The DWC2 isochronous fixes named in the recon ("keep isochronous transfers on DATA0", "complete periodic IN channels immediately") are both in
hcd_dwc2.c, TinyUSB's host side. The sound card is the device side, and usbif's host uses ESP-IDF's host library, not TinyUSB's.
What might apply. Since 0.21, Espressif added a "Periodic Transfer Interrupt" mode to TinyUSB's DWC2 device driver (
CFG_TUD_DWC2_PTI_ENABLE, commit1d73cd4). It setsDCTL.IgnrFrmNum, so an isochronous OUT endpoint takes the next packet whatever the frame parity. Todaydcd_dwc2.carms an iso OUT endpoint for the next frame's parity, read fromDSTSat re-arm time. A re-arm that lands just after a microframe starts targets the frame after, and that microframe's packet is dropped. That fits a loss that grows with interrupt load, here the I2S DMA at twice the rate. But PTI only works in Buffer DMA mode, and the P4 build runs the DWC2 in slave mode (CFG_TUD_DWC2_DMA_ENABLEdefaults to 0 for the P4 in Espressif's TinyUSB 0.21.0.1). Moving to DMA mode is a larger change, with the P4's cache maintenance (the trap CircuitPython hit on P4 host), so I didn't prototype it tonight.Next step before any of that: confirm the mechanism. On a 48 kHz wire, count the core's incomplete-isochronous-OUT and packet-drop status (
GINTSTS.incompISOOUT/DOEPINT.PktDrpSts) againstuac_pump_clock()'s packet count. If the drops line up with late re-arms, it's this, and DMA plus PTI is the fix. Still open.Measured on the Waveshare P4 panel today (MicroPython 1.29.0, usbif main with explicit feedback from #59, the receive path in IRAM since #50), with a board script that counts the DWC2's own drop flags beside
uac_pump_clock()'s packet count:GINTSTS.incompISOOUTand the iso OUT endpoint'sDOEPINT.PktDrpSts, polled and cleared in a loop, in 10 s windows. Windows played a 440 Hz tone in WASAPI exclusive mode at 48 kHz (uac_play_tone.py), and the codec sat at volume 0.At rest, the 0.2 % is gone on both wires. With the interpreter busy polling (about 84,000 polls a second, so no idle time), nine windows each:
wire feedback pkts/s pump % incompISOOUT PktDrpSts 48 kHz on 8000.2-8000.3 99.997-100.000 0 0 48 kHz off 8000.2-8000.3 100.002-100.004 0 0 Under PSRAM traffic, both wires lose packets, in bursts. The same loop with a 64 KB
bytearraycopy between polls (a memcpy through PSRAM, about 2,200 polls a second):wire lost packets per window, nine windows incompISOOUT seen in the lossy windows 48 kHz 0, 0, 0, 229, 133, 0, 0, 0, 0 79, 48 24 kHz 61, 0, 524, 0, 0, 0, 0, 0, 0 22, 180 A 1 MB copy between polls (46 polls a second) lost 53 packets in one window of nine.
PktDrpStsnever set. Every lossy window hasincompISOOUT, at about one flag per three lost packets, which is what a poll every 450 us catches of drops that come in runs. So the controller reports the endpoint as not armed for those microframes. That's the late re-arm from the last comment, and it isn't specific to the 48 kHz wire: the 24 kHz wire loses just as much once PSRAM is busy, and neither loses anything when only the CPU is busy. The 09-25 measurement predates #50, which moved TinyUSB's interrupt path out of flash for exactly this reason. That's the likeliest reason the at-rest loss is gone.Two more facts for whoever picks this up. The USB OTG interrupt is routed to core 1 (interrupt matrix
USB_OTG_INT_MAPreads CPU interrupt 21 on core 1, nothing on core 0). TinyUSB allocates it withESP_INTR_FLAG_LOWMED(dwc2_esp32.h), so it usually gets level 1, the same level as most other handlers, and any of them can hold it off.Next lead: find what delays the ISR while PSRAM is busy. The cheapest experiment is one variable in the TinyUSB component: allocate the USB interrupt at
ESP_INTR_FLAG_LEVEL3and check that its data (xfer_status, the audio FIFO) sits in internal RAM. Then rerun this stress A/B, several runs each, because the loss comes in bursts. If a higher level doesn't clear it, the structural fix is still the one already named: Buffer DMA mode with Espressif's periodic transfer interrupt (DCTL.IgnrFrmNum), so a late re-arm no longer costs a parity mismatch. Leaving this open.Tried the interrupt-priority lead today. Raising the USB interrupt to level 3 doesn't fix the loss. Bursts still come under PSRAM load, and level 1 loses packets at rest as well, so the loss isn't really tied to the load.
What I changed. I force-included a small header into the TinyUSB component only, which redefines
ESP_INTR_FLAG_LOWMEDasESP_INTR_FLAG_LEVEL3on the P4. Theesp_intr_alloc()indwc2_esp32.his the only use of that flag in the component. The builtdcd_int_enablepasses flags8instead of14, and on the board the USB OTG line (CLIC 21, CPU interrupt 5 on core 1) reads level 3. Everything else enabled on that core is at level 1 (UART0, I2C, the core-1 systimer tick, AHB PDMA OUT ch0, GPIO, USB-Serial-JTAG, the cross-core interrupt) or level 4. Level 3 is the highest a C handler that calls FreeRTOSFromISRAPIs can have, because critical sections mask levels 1 to 3.It was legal. I walked the call graph from
dwc2_int_handler_wrap,audiod_xfer_isr,audiod_sof_israndtud_audio_rx_done_isrin the built ELF. All 38 functions on the isochronous path are in IRAM, andmemsetcomes from ROM. The only flash call is MicroPython'stud_event_hook_cb, which runs when an event is queued fortud_task()(CDC, control), not for each audio packet._dcd_data,_usbd_devand_audiod_fctare in internal RAM, since BSS isn't allowed in PSRAM on this build. The allocation doesn't ask forESP_INTR_FLAG_IRAM, so the level doesn't need it anyway.Results. I used the same probe and setup as the last comment: 10 s windows, nine per run, Windows playing 440 Hz at 48 kHz in exclusive mode, and a 64 KB
bytearraycopy between polls for "load". The host reported 0 underflows in every run. A window counts as lossy when it hasincompISOOUTor at least two missing packets, and a burst means 10 or more lost in one window.interrupt level image wire condition windows lossy bursts packets lost (largest first) 1 (before) MicroPython 1.29, usbif at #63 48 kHz rest 18 2 1 200, 8 1 (before) same 24 kHz rest 9 0 0 0 1 (before) same 48 kHz load 18 1 1 224 1 (before) same 24 kHz load 18 1 1 80 3 usbif main + level 3, micropython-pydevices main 48 kHz rest 18 0 0 0 3 same 24 kHz rest 18 0 0 0 3 same 48 kHz load 36 11 1 48, then 1 to 3 each 3 same 24 kHz load 36 8 1 151, then 1 to 3 each 1 (CLIC set back at runtime) the level-3 image 48 kHz rest 9 1 1 16 1 (CLIC set back at runtime) same 24 kHz rest 9 2 2 120, 24 1 (CLIC set back at runtime) same 48 kHz load 18 4 1 144, then single packets 1 (CLIC set back at runtime) same 24 kHz load 18 3 3 168, 80, 31 In total, level 1 had a burst in 10 of 117 windows and level 3 in 2 of 108. That may be a real reduction, but it isn't a fix: level 3 still lost 151 packets in one 10 s window under load. It also showed more isolated one-to-three-packet drops than level 1 did. A burst shows up anywhere in a run (the first flag came between 10 and 93 s in), and at level 1 it happens at rest too. That contradicts the earlier "at rest, 100.00 % with zero drops": that reading came from too few windows.
What this rules out. If another handler holding off a level-1 interrupt were the cause, level 3 would have cleared it, because nothing else on that core runs at 2 or 3. What remains is something that masks level 3 too, such as a critical section or a spinlock wait, or a re-arm that's on time but still lands on the wrong frame parity. One caveat about the probe:
incompISOOUTalso sets if the host simply doesn't send in a microframe, so the counter alone can't separate a late re-arm from a skipped host transfer.Next step, still the structural one: Buffer DMA mode with Espressif's periodic transfer interrupt (
CFG_TUD_DWC2_PTI_ENABLE,DCTL.IgnrFrmNum). An iso OUT endpoint then accepts the next packet whatever the frame parity, so a re-arm doesn't have to win a race against the microframe boundary. It means moving the P4's DWC2 out of slave mode, along with the cache maintenance that brings. That's a bigger change, to be proposed on its own, so this priority experiment is closed. Leaving the issue open.Measured on the full-speed controller today, not the high-speed one where this was found. That tells us whether the 48 kHz loss follows the I2S wire or the USB path.
The board was the Waveshare ESP32-P4-WIFI6-DEV-KIT with the device on its USB-C marked "USB": micropython-pydevices' new
WIFI6_DEV_KIT_H2variant (micropython-pydevices#83), usbif at 452c99f, MicroPython 1.29.0, TinyUSB in slave mode on the full-speed DWC2. Same setup as before:soundcard.py's pump withcdc+uac, Windows playing a 440 Hz tone at 48 kHz in WASAPI exclusive mode (0 underflows), the codec at volume 0, and 10 s windows ofuac_pump_clock().GINTSTS.incompISOOUTon the full-speed core (0x50040000) was polled and cleared in the same loop. A full-speed host sends 1000 packets a second, not 8000. "Load" is a 64 KBbytearraycopy between polls, about 2,500 polls a second, against about 71,000 at rest.wire condition windows packets/s pump % incompISOOUT starved 48 kHz rest 9 1000.0 in every window 99.997–100.000 0 0 48 kHz load 9 1000.0 in every window 99.991–100.010 0 0 24 kHz rest 9 1000.0 in every window 99.996–100.000 0 0 24 kHz load 9 999.9–1000.1 99.998–100.003 0 0 On the full-speed path the 48 kHz wire loses nothing, and neither does the 24 kHz wire, at rest or under PSRAM load. In 36 windows, no packet went missing and no
incompISOOUTwas raised. Since the I2S wire is the same and only the USB controller changed, this points the high-speed loss at the USB path (the high-speed controller's iso OUT re-arm at 8000 microframes a second) rather than the 48 kHz wire. This is 36 windows, against 10 bursts in 117 at level 1 on the high-speed path, so it's strong evidence but not proof.One trap, for anyone repeating this on any controller. The first run logged each window to a file, and every window after the first lost exactly 12 packets, with one
incompISOOUT. A littlefs write stalls the USB interrupt for about 12 ms. The second run kept the log in RAM and added a control condition, one 100-byte file append per window, at 48 kHz:wire condition windows packets/s lost per window incompISOOUT 48 kHz one file write per window 9 998.8 (the first window, before any write: 1000.0) 11.7–12.5 1 per window So a probe that writes its log between windows creates its own loss. Earlier probes on this issue should be checked for that before their numbers are trusted.
Leaving this open.
Found during the audio meter spike on the P4 (2026-09-25), on current main (TinyUSB v0.21.0.1-micropython1, the ISR re-arm from #47).
With the meter off, delivery from Windows WASAPI at 48 kHz settles near 99.8 % a window or two into any stream when
soundcard.pyruns the I2S wire at the host's 48 kHz (#46).uac_stream_check.pyon a 24 kHz wire holds 100.00 %. So the 48 kHz wire itself costs about one packet in 500. It passes the 99.5 % gate, but it isn't the 100 % the ISR re-arm reaches elsewhere.Not investigated. The suspects: the I2S DMA writes at double the rate competing with the USB interrupt, or the pump's per-block work at 48 kHz. Measure it the #37/#39 way (
uac_pump_clock()against 8,000 packets/s, 10 s windows).