Summary
A proposal, after the wasm page target (#866) and the array fork of 2.0.22: a WGSL lane, so bend x.bend -o x.html runs its ! on WebGPU.
Metal gets the GPU for free because Apple's compiler takes the emitted C as it is, 64-bit terms included. A browser takes only WGSL: no 64-bit integers, no pointers, no recursion, no tail calls. So the device code has to be emitted in WGSL from the compiled form, the way the JS lane emits JS from it.
What is measured
A prototype runtime of about 300 lines of JS (rt.js) runs the fork tree on WebGPU with the dispatch boundary as the only fence: terms as pairs of u32, a task is a node on the heap, one indirect dispatch per fork level, and the leaves work their subtree sequentially on a lane. With the leaves translated by hand it draws the same frames as the Bend program (same checksums) at 0.3 ms for Mandelbrot 512², 1.8 ms for the voxel raycaster 512² and 2.4 ms for app_ray_tracer_3d at 1024×768, against 4.5 / 42 / 53 ms on ten wasm threads in the same browser (Chrome 153, Apple M5): the pages. Bendcraft's editable world at 512² was 1.3 ms against 18 in wasm and 4 on Metal, before its hand port was retired for not following the program.
The shape
The hand-written leaves are the missing piece. The Metal kernel is already a switch(fid) inside a loop over the segments; the lane would emit that loop in WGSL, over the same segments, with every term operation on two 32-bit words and the heap as a storage buffer. The host stays the wasm page: it uploads the program, dispatches the fork levels, reads the result back; IO stays on the host, as on Metal.
The question
Is this something you want in Bend, and if so, how: a lane next to JS, or a device target of the C lane? I would rather build it your way than twice.
Written by Claude (Anthropic) with Adriel Santana driving; the numbers are from his M5.
Summary
A proposal, after the wasm page target (#866) and the array fork of 2.0.22: a WGSL lane, so
bend x.bend -o x.htmlruns its!on WebGPU.Metal gets the GPU for free because Apple's compiler takes the emitted C as it is, 64-bit terms included. A browser takes only WGSL: no 64-bit integers, no pointers, no recursion, no tail calls. So the device code has to be emitted in WGSL from the compiled form, the way the JS lane emits JS from it.
What is measured
A prototype runtime of about 300 lines of JS (rt.js) runs the fork tree on WebGPU with the dispatch boundary as the only fence: terms as pairs of u32, a task is a node on the heap, one indirect dispatch per fork level, and the leaves work their subtree sequentially on a lane. With the leaves translated by hand it draws the same frames as the Bend program (same checksums) at 0.3 ms for Mandelbrot 512², 1.8 ms for the voxel raycaster 512² and 2.4 ms for
app_ray_tracer_3dat 1024×768, against 4.5 / 42 / 53 ms on ten wasm threads in the same browser (Chrome 153, Apple M5): the pages. Bendcraft's editable world at 512² was 1.3 ms against 18 in wasm and 4 on Metal, before its hand port was retired for not following the program.The shape
The hand-written leaves are the missing piece. The Metal kernel is already a
switch(fid)inside a loop over the segments; the lane would emit that loop in WGSL, over the same segments, with every term operation on two 32-bit words and the heap as a storage buffer. The host stays the wasm page: it uploads the program, dispatches the fork levels, reads the result back; IO stays on the host, as on Metal.The question
Is this something you want in Bend, and if so, how: a lane next to JS, or a device target of the C lane? I would rather build it your way than twice.
Written by Claude (Anthropic) with Adriel Santana driving; the numbers are from his M5.