Skip to content

[Miniflare][Workflows] Completed instances retain file handles, eventually crashing workerd with “Too many open files” #15809

Description

@alexminza

What versions & operating system are you using?

Collected on 2026-09-23 with:

npx envinfo --system --npmPackages '{wrangler,create-cloudflare,miniflare,@cloudflare/*}' --binaries

Output, with binary installation paths omitted:

  System:
    OS: macOS 27.0
    CPU: (10) arm64 Apple M1 Pro
    Memory: 87.27 MB / 16.00 GB
    Shell: 5.9
  Binaries:
    Node: 26.9.0
    Yarn: 1.22.22
    npm: 12.1.0
    pnpm: 12.5.1
    bun: 1.4.2

The command emitted no npmPackages section in the workspace checkout. npm ls --json separately confirmed Miniflare 5.20260918.0-alpha, workerd 1.20260918.1, Wrangler 4.135.0, @cloudflare/vitest-plugin 1.1.13 and @cloudflare/workers-types 5.20260920.1; create-cloudflare is not installed. Only Miniflare and workerd are needed by the standalone reproduction.

The host meets the documented Wrangler system requirements: macOS 13.5+ and a supported Node.js release; Node.js 26 is Current. No kernel file limits were raised. The reproduction uses a disposable subshell with an 8,192-file process limit.

Please provide a link to a minimal reproduction

https://gist.github.com/alexminza/53755e0a5b27825f455414ed1aec64a4

Describe the Bug

Completed local Workflow instances retain database file handles. Repeatedly creating and completing small batches eventually crashes workerd with accept(): Too many open files (os error 24).

The standalone reproduction has one Workflow class, one immediate step.do() per instance and at most 100 instances in flight. Every batch must complete with the correct output before the next starts. There are no application dependencies, D1/R2 bindings, external requests, Vitest introspection or child Workflow chains. It does not retain earlier batches' handles in an array.

Expected: idle completed engines release runtime resources while their saved status and output remain available. Resource usage should not grow indefinitely with the number of already-completed instances.

Actual: the unmodified runtime retains approximately three file descriptors per completed instance. With an 8,192-file process limit it crashes after 2,600 completed instances. With the host's 61,440 ceiling, an instrumented run reaches 60,131 descriptors after 20,000 completions and loses the runtime connection before 21,000. This is not 25,000 simultaneous executions.

Reproduction

  1. Save miniflare-workflow-files.mjs from the linked reproduction into an empty directory and follow its header's pinned dependency installation command.
  2. Run its subshell command with the 8,192-file limit. The script requests 5,000 instances, logs completed counts and exits nonzero on failure.
  3. Run the same command with --delete-completed. This positive control deletes only its own completed synthetic instances before starting the next batch.

The deletion control is diagnostic, not a proposed history-retention policy. The exact crash count depends on the process limit and other open handles.

Source findings and controlled comparison

The pinned Miniflare Workflow plugin sets preventEviction: true on the SQLite-backed Engine namespace. Workerd's actor lifecycle skips idle shutdown for pinned actors. The local engine also retains uncancelled background waits and does not dispose the object result returned by its internal USER_WORKFLOW.run() RPC.

These experiments used an isolated copy of the same dependency, not an application rewrite:

Variant Result
Unmodified baseline, 8,192 process limit Crashes after 2,600 completions
Unmodified runtime with native deletion after each batch, same limit All 5,000 complete in 54.8 s; exit 0
Allow eviction alone, with/without cancellation of successful step timeouts Still 3,029 open handles after 1,000 completions and 90 idle seconds
Allow eviction and forcibly abort the engine after persisting completion Completes 25,000, but breaks an active event subscription; rejected
Allow eviction and cancel finished background waits, without disposing the run result Still crashes after 2,600 completions at the 8,192 limit
Same non-forced cleanup plus run-result disposal, 8,192 process limit Completes 25,000 in 174.1 s; maximum sampled 4,846 handles, falling to 29 after 15 idle seconds

The final experiment made these changes:

  • Allow normal idle eviction for the Workflow engine namespace.
  • Pass the existing combined step/engine abort signal to the step timeout wait.
  • Pass the engine signal to the grace-period wait and abort that signal when execution finishes; do not call ctx.abort() on successful completion.
  • Dispose the object returned by the internal USER_WORKFLOW.run() RPC after persisting its output. Primitive returns do not require disposal.

First/last saved results remain readable after the 25,000-instance experiment. Separate lifecycle checks pass duplicate creation, step-targeted restart, retry, sleep/pause/resume, event waiting, terminal errors, timeout expiry and event subscriptions. The baseline passes those same lifecycle checks; forced termination introduces the subscription regression.

Descriptor counts came from a separate macOS proc_pidinfo(PROC_PIDLISTFDS) diagnostic. The published script intentionally needs no custom native counter to reproduce the crash. Its deletion control tests release through the existing API, not the experimental patch. The baseline and deletion rows were rerun with this single-file reproduction on 2026-09-23; the instrumented patch comparisons are from earlier that day.

These results locate the resource-lifecycle problem; they are not a fully validated fix. The combined experiment does not isolate every change's individual contribution. Workflow-wide introspection, all upstream regressions and the full application workload still need validation. No deployed Cloudflare Workflows failure has been demonstrated, and no dependency patch is retained.

Related reports

  • #15788 covers retained timers exhausting the active-timer quota within one instance. This report covers file-handle exhaustion across completed instances.
  • PR #15708 addresses step callback results and introspection modifiers. At inspected head 9c13670657fcc109e4fe632d2bf3660c64ff8aaa, it does not change engine eviction or dispose the final USER_WORKFLOW.run() result; it is related, not a verified complete fix for this reproduction.
  • #14839 and PR #14840 concern retained pause listeners, not this file-descriptor failure.

References

Please provide any relevant error logs

Observed in the unmodified 8,192-limit run after 2,600 completions:

*** Fatal uncaught kj::Exception:
rust/cxx/kj-rs-io/ffi.rs:184: overloaded:
accept(): Too many open files (os error 24)

The Node caller then loses its connection to the local runtime. The process exits nonzero; completed batches before the crash had the expected output.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions