Skip to content

Sound card: about 0.2 % of packets lost on a 48 kHz I2S wire, 0 % on a 24 kHz wire #49

Description

@bdbarnett

Found during the audio meter spike on the P4 (2026-09-25), on current main (TinyUSB v0.21.0.1-micropython1, the ISR re-arm from #47).

With the meter off, delivery from Windows WASAPI at 48 kHz settles near 99.8 % a window or two into any stream when soundcard.py runs the I2S wire at the host's 48 kHz (#46). uac_stream_check.py on a 24 kHz wire holds 100.00 %. So the 48 kHz wire itself costs about one packet in 500. It passes the 99.5 % gate, but it isn't the 100 % the ISR re-arm reaches elsewhere.

Not investigated. The suspects: the I2S DMA writes at double the rate competing with the USB interrupt, or the pump's per-block work at 48 kHz. Measure it the #37/#39 way (uac_pump_clock() against 8,000 packets/s, 10 s windows).

Activity

  1. bdbarnett commented on Oct 8, 2026

    @bdbarnett
    ContributorAuthor

    Evaluated tonight against the two leads from CircuitPython. Neither fixes this as it stands.

    What might apply. Since 0.21, Espressif added a "Periodic Transfer Interrupt" mode to TinyUSB's DWC2 device driver (CFG_TUD_DWC2_PTI_ENABLE, commit 1d73cd4). It sets DCTL.IgnrFrmNum, so an isochronous OUT endpoint takes the next packet whatever the frame parity. Today dcd_dwc2.c arms an iso OUT endpoint for the next frame's parity, read from DSTS at re-arm time. A re-arm that lands just after a microframe starts targets the frame after, and that microframe's packet is dropped. That fits a loss that grows with interrupt load, here the I2S DMA at twice the rate. But PTI only works in Buffer DMA mode, and the P4 build runs the DWC2 in slave mode (CFG_TUD_DWC2_DMA_ENABLE defaults to 0 for the P4 in Espressif's TinyUSB 0.21.0.1). Moving to DMA mode is a larger change, with the P4's cache maintenance (the trap CircuitPython hit on P4 host), so I didn't prototype it tonight.

    Next step before any of that: confirm the mechanism. On a 48 kHz wire, count the core's incomplete-isochronous-OUT and packet-drop status (GINTSTS.incompISOOUT / DOEPINT.PktDrpSts) against uac_pump_clock()'s packet count. If the drops line up with late re-arms, it's this, and DMA plus PTI is the fix. Still open.

  2. bdbarnett commented on Oct 9, 2026

    @bdbarnett
    ContributorAuthor

    Measured on the Waveshare P4 panel today (MicroPython 1.29.0, usbif main with explicit feedback from #59, the receive path in IRAM since #50), with a board script that counts the DWC2's own drop flags beside uac_pump_clock()'s packet count: GINTSTS.incompISOOUT and the iso OUT endpoint's DOEPINT.PktDrpSts, polled and cleared in a loop, in 10 s windows. Windows played a 440 Hz tone in WASAPI exclusive mode at 48 kHz (uac_play_tone.py), and the codec sat at volume 0.

    At rest, the 0.2 % is gone on both wires. With the interpreter busy polling (about 84,000 polls a second, so no idle time), nine windows each:

    wire feedback pkts/s pump % incompISOOUT PktDrpSts
    48 kHz on 8000.2-8000.3 99.997-100.000 0 0
    48 kHz off 8000.2-8000.3 100.002-100.004 0 0

    Under PSRAM traffic, both wires lose packets, in bursts. The same loop with a 64 KB bytearray copy between polls (a memcpy through PSRAM, about 2,200 polls a second):

    wire lost packets per window, nine windows incompISOOUT seen in the lossy windows
    48 kHz 0, 0, 0, 229, 133, 0, 0, 0, 0 79, 48
    24 kHz 61, 0, 524, 0, 0, 0, 0, 0, 0 22, 180

    A 1 MB copy between polls (46 polls a second) lost 53 packets in one window of nine. PktDrpSts never set. Every lossy window has incompISOOUT, at about one flag per three lost packets, which is what a poll every 450 us catches of drops that come in runs. So the controller reports the endpoint as not armed for those microframes. That's the late re-arm from the last comment, and it isn't specific to the 48 kHz wire: the 24 kHz wire loses just as much once PSRAM is busy, and neither loses anything when only the CPU is busy. The 09-25 measurement predates #50, which moved TinyUSB's interrupt path out of flash for exactly this reason. That's the likeliest reason the at-rest loss is gone.

    Two more facts for whoever picks this up. The USB OTG interrupt is routed to core 1 (interrupt matrix USB_OTG_INT_MAP reads CPU interrupt 21 on core 1, nothing on core 0). TinyUSB allocates it with ESP_INTR_FLAG_LOWMED (dwc2_esp32.h), so it usually gets level 1, the same level as most other handlers, and any of them can hold it off.

    Next lead: find what delays the ISR while PSRAM is busy. The cheapest experiment is one variable in the TinyUSB component: allocate the USB interrupt at ESP_INTR_FLAG_LEVEL3 and check that its data (xfer_status, the audio FIFO) sits in internal RAM. Then rerun this stress A/B, several runs each, because the loss comes in bursts. If a higher level doesn't clear it, the structural fix is still the one already named: Buffer DMA mode with Espressif's periodic transfer interrupt (DCTL.IgnrFrmNum), so a late re-arm no longer costs a parity mismatch. Leaving this open.

  3. bdbarnett commented on Oct 9, 2026

    @bdbarnett
    ContributorAuthor

    Tried the interrupt-priority lead today. Raising the USB interrupt to level 3 doesn't fix the loss. Bursts still come under PSRAM load, and level 1 loses packets at rest as well, so the loss isn't really tied to the load.

    What I changed. I force-included a small header into the TinyUSB component only, which redefines ESP_INTR_FLAG_LOWMED as ESP_INTR_FLAG_LEVEL3 on the P4. The esp_intr_alloc() in dwc2_esp32.h is the only use of that flag in the component. The built dcd_int_enable passes flags 8 instead of 14, and on the board the USB OTG line (CLIC 21, CPU interrupt 5 on core 1) reads level 3. Everything else enabled on that core is at level 1 (UART0, I2C, the core-1 systimer tick, AHB PDMA OUT ch0, GPIO, USB-Serial-JTAG, the cross-core interrupt) or level 4. Level 3 is the highest a C handler that calls FreeRTOS FromISR APIs can have, because critical sections mask levels 1 to 3.

    It was legal. I walked the call graph from dwc2_int_handler_wrap, audiod_xfer_isr, audiod_sof_isr and tud_audio_rx_done_isr in the built ELF. All 38 functions on the isochronous path are in IRAM, and memset comes from ROM. The only flash call is MicroPython's tud_event_hook_cb, which runs when an event is queued for tud_task() (CDC, control), not for each audio packet. _dcd_data, _usbd_dev and _audiod_fct are in internal RAM, since BSS isn't allowed in PSRAM on this build. The allocation doesn't ask for ESP_INTR_FLAG_IRAM, so the level doesn't need it anyway.

    Results. I used the same probe and setup as the last comment: 10 s windows, nine per run, Windows playing 440 Hz at 48 kHz in exclusive mode, and a 64 KB bytearray copy between polls for "load". The host reported 0 underflows in every run. A window counts as lossy when it has incompISOOUT or at least two missing packets, and a burst means 10 or more lost in one window.

    interrupt level image wire condition windows lossy bursts packets lost (largest first)
    1 (before) MicroPython 1.29, usbif at #63 48 kHz rest 18 2 1 200, 8
    1 (before) same 24 kHz rest 9 0 0 0
    1 (before) same 48 kHz load 18 1 1 224
    1 (before) same 24 kHz load 18 1 1 80
    3 usbif main + level 3, micropython-pydevices main 48 kHz rest 18 0 0 0
    3 same 24 kHz rest 18 0 0 0
    3 same 48 kHz load 36 11 1 48, then 1 to 3 each
    3 same 24 kHz load 36 8 1 151, then 1 to 3 each
    1 (CLIC set back at runtime) the level-3 image 48 kHz rest 9 1 1 16
    1 (CLIC set back at runtime) same 24 kHz rest 9 2 2 120, 24
    1 (CLIC set back at runtime) same 48 kHz load 18 4 1 144, then single packets
    1 (CLIC set back at runtime) same 24 kHz load 18 3 3 168, 80, 31

    In total, level 1 had a burst in 10 of 117 windows and level 3 in 2 of 108. That may be a real reduction, but it isn't a fix: level 3 still lost 151 packets in one 10 s window under load. It also showed more isolated one-to-three-packet drops than level 1 did. A burst shows up anywhere in a run (the first flag came between 10 and 93 s in), and at level 1 it happens at rest too. That contradicts the earlier "at rest, 100.00 % with zero drops": that reading came from too few windows.

    What this rules out. If another handler holding off a level-1 interrupt were the cause, level 3 would have cleared it, because nothing else on that core runs at 2 or 3. What remains is something that masks level 3 too, such as a critical section or a spinlock wait, or a re-arm that's on time but still lands on the wrong frame parity. One caveat about the probe: incompISOOUT also sets if the host simply doesn't send in a microframe, so the counter alone can't separate a late re-arm from a skipped host transfer.

    Next step, still the structural one: Buffer DMA mode with Espressif's periodic transfer interrupt (CFG_TUD_DWC2_PTI_ENABLE, DCTL.IgnrFrmNum). An iso OUT endpoint then accepts the next packet whatever the frame parity, so a re-arm doesn't have to win a race against the microframe boundary. It means moving the P4's DWC2 out of slave mode, along with the cache maintenance that brings. That's a bigger change, to be proposed on its own, so this priority experiment is closed. Leaving the issue open.

  4. bdbarnett commented on Oct 10, 2026

    @bdbarnett
    ContributorAuthor

    Measured on the full-speed controller today, not the high-speed one where this was found. That tells us whether the 48 kHz loss follows the I2S wire or the USB path.

    The board was the Waveshare ESP32-P4-WIFI6-DEV-KIT with the device on its USB-C marked "USB": micropython-pydevices' new WIFI6_DEV_KIT_H2 variant (micropython-pydevices#83), usbif at 452c99f, MicroPython 1.29.0, TinyUSB in slave mode on the full-speed DWC2. Same setup as before: soundcard.py's pump with cdc+uac, Windows playing a 440 Hz tone at 48 kHz in WASAPI exclusive mode (0 underflows), the codec at volume 0, and 10 s windows of uac_pump_clock(). GINTSTS.incompISOOUT on the full-speed core (0x50040000) was polled and cleared in the same loop. A full-speed host sends 1000 packets a second, not 8000. "Load" is a 64 KB bytearray copy between polls, about 2,500 polls a second, against about 71,000 at rest.

    wire condition windows packets/s pump % incompISOOUT starved
    48 kHz rest 9 1000.0 in every window 99.997–100.000 0 0
    48 kHz load 9 1000.0 in every window 99.991–100.010 0 0
    24 kHz rest 9 1000.0 in every window 99.996–100.000 0 0
    24 kHz load 9 999.9–1000.1 99.998–100.003 0 0

    On the full-speed path the 48 kHz wire loses nothing, and neither does the 24 kHz wire, at rest or under PSRAM load. In 36 windows, no packet went missing and no incompISOOUT was raised. Since the I2S wire is the same and only the USB controller changed, this points the high-speed loss at the USB path (the high-speed controller's iso OUT re-arm at 8000 microframes a second) rather than the 48 kHz wire. This is 36 windows, against 10 bursts in 117 at level 1 on the high-speed path, so it's strong evidence but not proof.

    One trap, for anyone repeating this on any controller. The first run logged each window to a file, and every window after the first lost exactly 12 packets, with one incompISOOUT. A littlefs write stalls the USB interrupt for about 12 ms. The second run kept the log in RAM and added a control condition, one 100-byte file append per window, at 48 kHz:

    wire condition windows packets/s lost per window incompISOOUT
    48 kHz one file write per window 9 998.8 (the first window, before any write: 1000.0) 11.7–12.5 1 per window

    So a probe that writes its log between windows creates its own loss. Earlier probes on this issue should be checked for that before their numbers are trusted.

    Leaving this open.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions