Skip to content

Fix empty unsigned sum/prod failing on Metal - #4546

Merged
zcbenz merged 2 commits into
ml-explore:mainfrom
CodeWithMoin:fix-init-reduce-unsigned
Oct 1, 2026
Merged

zcbenz merged 2 commits into
ml-explore:mainfrom
CodeWithMoin:fix-init-reduce-unsigned

Conversation

@CodeWithMoin

Copy link
Copy Markdown
Contributor

mlx runs operations like sum, prod, max on GPUs. The problem in the existing code is related to the init kernel. It's that we don't have unsigned types defined for sum and prod:

#define instantiate_init_sum_prod(name, op)                 \
  instantiate_init_reduce(name, int32, int32_t, op)         \
  instantiate_init_reduce(name, int64, int64_t, op)         \
  instantiate_init_reduce(name, float16, float16_t, op)     \
  instantiate_init_reduce(name, bfloat16, bfloat16_t, op)   \
  instantiate_init_reduce(name, float32, float, op)         \
  instantiate_init_reduce(name, complex64, complex64_t, op)

instantiate_init_min_max, right below it in the same file, lists all 13 types including uint8 to uint64, so min/max already work on empty unsigned arrays.

Most of the time it goes unnoticed, but when the array is empty, there's nothing to add up or multiply, so the answer is only that starting value. init_reduce is reached only there:

// Nothing to reduce just initialize the output
else {
  init_reduce(out, op_name, compute_encoder, d, s);
}

mlx goes looking for the init kernel that was never built, and we get:

>>> mx.sum(mx.array(np.zeros((0,), dtype=np.uint8)))
RuntimeError: [metal::Device] Unable to load kernel init_reduce_sumuint32
>>> mx.prod(mx.array(np.zeros((0,), dtype=np.uint64)))
RuntimeError: [metal::Device] Unable to load kernel init_reduce_produint64

Any ordinary way of ending up with an empty unsigned array does the same:

>>> counts = mx.zeros((8,), dtype=mx.uint8)
>>> mx.sum(mx.zeros((8, 8), dtype=mx.uint8)[:0])             # slice to nothing
>>> mx.sum(mx.take(counts, mx.array([], dtype=mx.int32)))     # empty take
>>> mx.sum(mx.zeros((0, 3), dtype=mx.uint8))                  # empty batch
RuntimeError: [metal::Device] Unable to load kernel init_reduce_sumuint32

This hits uint8, uint16, uint32 and uint64, for both sum and prod. Signed, float and bool types are fine.

On the released 0.32.2 version it can be worse than an error: mx.sum(...).item() raises a catchable RuntimeError, but np.array(mx.sum(...)) aborts the process (libc++abi: terminating) even inside a try/except.

The fix

We add uint32 and uint64 to the sum/prod list. These are the two types missing, because summing any unsigned type produces uint32 (from uint8, uint16, uint32) or uint64 (from uint64). CUDA is unaffected: init_reduce.cu uses dispatch_all_types, so it instantiates every type already.

mlx/backend/metal/kernels/reduce.metal |  2 ++
python/tests/test_reduce.py            | 30 ++++++++++++++++++++++++++++++

How I found this

I ran mlx against numpy across many shapes, types and edge cases, comparing answers. The empty-unsigned case crashed the test script, and narrowing it down gave the exact types and the missing kernels.

How I proved it

I built mlx from source on my MacBook, confirmed this failure exists on unmodified code, applied the fix, rebuilt it and confirmed all 4 unsigned types now return 0 for sum and 1 for prod:

sum uint8 -> 0 / prod uint8 -> 1
sum uint16 -> 0 / prod uint16 -> 1
sum uint32 -> 0 / prod uint32 -> 1
sum uint64 -> 0 / prod uint64 -> 1

The added test test_zero_size_sum_prod_all_dtypes covers 11 dtypes and both empty shapes; it fails on main and passes with the fix.

python -m unittest test_reduce -v → Ran 14 tests, OK
test_ops 163 OK · test_array 102 OK (19 skipped) · test_nn 73 OK (1 skipped)

Built at commit 59d600b5e with Xcode 27 on an M4 Pro.

  • ☑️ I understand it is strictly prohibited to use AI to write PR description
  • AI usage disclosure: I directed an AI assistant for finding the bug, the code and the tests. I verified the fix.

@zcbenz
zcbenz marked this pull request as ready for review October 1, 2026 03:02
@zcbenz
zcbenz force-pushed the fix-init-reduce-unsigned branch from 7af0a79 to c5adcb8 Compare October 1, 2026 03:05
@zcbenz
zcbenz merged commit 49028fb into ml-explore:main Oct 1, 2026
30 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants