Skip to content

multimer: one Timer, one dispatcher, one wake source per host (the timing redesign) - #101

Merged
bdbarnett merged 10 commits into
mainfrom
timing-redesign
Sep 26, 2026
Merged

bdbarnett merged 10 commits into
mainfrom
timing-redesign

Conversation

@bdbarnett

@bdbarnett bdbarnett commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

The timing layer redesigned from first principles, on Brad's charter of 2026-09-25: one Timer with machine.Timer's shape, one clock, one dispatcher, and one wake source per host, with callbacks delivered on the main thread at a safe point on every interpreter PyDevices runs on. appdev.App loses its timer machinery (the service tick, each display's refresh and every subscription are ordinary timers you can see in multimer.report()), displaydev gains a frame clock, and the docs follow. A script that ends without app.run() keeps running; multimer.report() says what is running, from the REPL or wherever there is a prompt.

What it is and why it is better, the alternatives that lost, the numbers and the ledger of what was measured where: docs/timing-design.md. What changes for user code: docs/multimer-migration.md. The hardware runs and what each saw: docs/timing-hardware-tests.md.

Designed and prototyped in a cloud session (the first four commits, applied with git am on the base it recorded, which was still main); landed and being proven on the bench by a local session. Draft until the hardware phases are in the ledger.

Companions, which need this one merged first: PyDevices/lvgl-bindings#22 (the LVGL driver on one multimer timer), PyDevices/pydevices-examples#149 (examples, the REPL proofs, the benchmarks) and PyDevices/micropython-pydevices#24 (patch 0015 for micropython.exe). lvgl-micropython, lvgl-python and lvgl-circuitpython sync display_driver.py after the lvgl-bindings merge.

Also fixes the Windows REPL window hang Brad found on 2026-09-27: python.exe -i -m examples.roku_remote leaves its WinDisplay window hung (IsHungAppWindow) on the current layer for as long as the prompt waits, because WinDisplay pumps its message queue only from the App's service tick and nothing delivers that tick at _pyrepl's non-alertable wait. The input hook here delivers it. Reproduced and proven in a real console with tools/prove_repl/win_window_alive.py (pydevices-examples): 3 of 5 checks fail on the installed 0.5.5, 0 fail on this branch, source=pending.

bdbarnett and others added 5 commits September 26, 2026 01:35
Timers are deadlines in one list; the host's wake source (a signal, a
pending call from a worker thread, the asyncio loop, one machine.Timer, the
browser's timer, or nothing) only says when to look. Callbacks run on the
main thread at a safe point on every host; a callback never interrupts
another; an overrun lowers the timer's rate instead of taking the thread;
hold() masks delivery for a critical section; report() shows everything
from the REPL, and repl() is the prompt where a host has none.

appdev.App loses its timer machinery: every() returns a multimer.Timer,
the service tick and each display's refresh are timers, and staying alive
past the script body is multimer.keepalive. displaydev gains
refresh_period_ms and a frame_clock per display; the async-timer flag goes.
Tests cover the contract, the wake sources, and the failure classes that
cost time before (re-entrancy, thread affinity, catch-up bursts, the frame
gate).
LVGL reads its indevs from its own timers, so the App's service tick must
stand down while a GUI owns the devices (app.pause_polling), or it consumes
the events first. docs/multimer.md and multimer-internals.md describe the
new layer.
appdev, app-and-board-config, board-configs, displaydev, jupyter, android,
architecture and the READMEs stop describing timer_async, AsyncTimer and
the provider imports; AGENTS.md names the wake sources as private.
…ns, beside multimer.md

The design doc's front answers what the layer is and why it is better;
the alternatives, the numbers and the ledger of what was measured where
follow. The hardware page carries the plans the cloud session wrote and
records what each run on the bench saw.
…ask Windows for 1 ms timer resolution

wake_from_source(safe=True) from the machine and native sources: their
callback is already a bytecode boundary of the main thread (the port
handed it through micropython.schedule), so a second trip through the
scheduler queue only added a wait slice per delivery. On micropython.exe
that slice was the console wait, and a 10 ms timer was served 20 times a
second at an idle prompt; it is about 100 now, measured with
prove_windows.py in a real console.

The pending source asks Windows for its 1 ms timer resolution while it
runs (timeBeginPeriod, as SDL and pygame do). Measured on Windows 11,
python.exe 3.14, a 10 ms timer with a busy main thread: 327 of 500
delivered with 23 ms p99 lateness at the default 15.6 ms, 501 of 500
with 8 ms (the GIL switch interval) at 1 ms.
Phase 5 is measured now, not bench: Windows python.exe and micropython.exe
(overlay 0015) in a real console, the P4 and T-Embed on MicroPython, the
S21 on the pending source, and CircuitPython on the T-Embed. The Numbers on
hardware section carries the old-vs-new tables per host. Two bench findings
fed back into the code: Windows' 15.6 ms default timer resolution and a
redundant scheduler hop for sources the port already schedules.
@bdbarnett

Copy link
Copy Markdown
Contributor Author

A data point from the P4 casting work (2026-09-26): on the ESP32-P4, while examples.roku_remote's LVGL loop runs, the REPL is starved completely, so mpftp can't reach the board until a hard reset. That's on the current released stack, not this branch. Worth checking during review whether the redesign's bytecode-safe-point delivery leaves the REPL reachable on the MCU the way it does at the Windows prompt.

@bdbarnett
bdbarnett marked this pull request as ready for review September 26, 2026 11:09
@bdbarnett
bdbarnett merged commit b6590c2 into main Sep 26, 2026
3 checks passed
@bdbarnett
bdbarnett deleted the timing-redesign branch September 26, 2026 11:09
@bdbarnett

Copy link
Copy Markdown
Contributor Author

Correction to my earlier comment. The P4 does NOT starve the REPL while examples.roku_remote runs. The casting session retested it: with roku_remote up, mpftp exec reached the board 3 of 3 times on the old frozen multimer 0.5.0 and 3 of 3 on the 0.6.1 layer (a /timing overlay). The earlier failures were two other things: the mpftp sidecar RPC needing a reconnect ("Expecting value" is the sidecar, not the board), and the busy cast-pump loops in its own p4_cast_* scripts, a tight Python loop Ctrl-C can't break. roku_remote is timer-driven and returns to the REPL. The T-Embed gave the same answer on 0.6.1 from MIP: REPL median 32 ms while the knob app ran.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

1 participant