Skip to content

[-16 lines] Reuse compiler paths for array access and branch blocks - #955

Closed
PedroVIOliv wants to merge 2 commits into
bendlang:mainfrom
PedroVIOliv:simplify-array-emission
Closed

PedroVIOliv wants to merge 2 commits into
bendlang:mainfrom
PedroVIOliv:simplify-array-emission

Conversation

@PedroVIOliv

Copy link
Copy Markdown
Contributor

Two compiler paths repeat logic already shared elsewhere. arr_op repeats array ownership/alias setup and the old-cell read and tuple return for get/swap. emit_chain implements its own indentation and closing braces instead of using block.

Share the existing-array setup and cell read/result path, preserving get's retain, swap's move and set's sink. Emit each conditional arm through block, retaining the single-arm shortcut. Generated else clauses move to the next line; conditions and branch bodies stay the same.

Only bend2/comp.ts changes: 20 insertions, 36 deletions (net -16 lines, -340 bytes). No new helpers, feature changes, comment removal or size-cap changes. Measured with the gate's default ttok invocation: 64,941 tokens upstream, 64,894 after the array change, and 64,854 after both changes (-87 total). The compiler is 146 tokens below its 65,000-token cap.

Validation against upstream e52cda47:

  • Compared compiler exit status, stdout, stderr and emitted C/JS for 779 Base-importing test/benchmark programs with a main and no expected compile error. All matched after normalizing only the closing-brace/else line break: 639 successful programs and 140 identical compiler rejections. Of the successful programs, 564 have formatting differences; 75 remain byte-identical.
  • All 16 runtime benchmarks produced identical optimized CPU LLVM IR with Apple clang 17 at -O3, excluding only the source path and module identifier. Runtime benchmark timings were not measured.
  • All 139 runnable tests selected from tests/compile and array-named tests passed their existing #| expectations on JS and C compiled with Apple clang 17 at -O3, using one CPU thread. Expected nonzero exit statuses were checked as well.
  • Metal: all 19 noninteractive GPU-bearing tests from the successfully emitted corpus passed on both upstream and this branch (38 runs) on an Apple M3, with --gpu on --threads 1. Every build created a Metal pipeline archive. This includes array atomics/forks, closure reachability, fork leaf continuations, GPU trig and stencil3d. The remaining GPU-bearing test, tests/gfx/clicks.bend, requires interactive window events and was excluded.
  • git diff --check passed.
  • The full repository size gate passed: PASS: 46 / 46, using ttok 0.3 installed in an isolated temporary environment.

The cluster test gate could not run because this host cannot resolve cluster. Cluster test/performance gates remain for maintainer infrastructure; CUDA execution and GPU performance measurements were not tested locally.

VictorTaelin pushed a commit that referenced this pull request Sep 21, 2026
arr_op shares the array setup and the old-cell read and tuple return of get and swap, keeping get's retain, swap's move and set's sink. Emission is byte-identical. (PR #955, its first commit)
@VictorTaelin

Copy link
Copy Markdown
Contributor

The first commit (the shared cell-access path in arr_op) is merged into 2.0.25 under your authorship, thank you: −47 ttok, byte-identical emission. The second (emitting conditional arms through block) moves every else in the emitted C and JS to its own line; that is a formatting call for Victor, so I left it out for now. If he wants it, we will take it from here.

Note: this reply was written by an AI after it reported the issue to me and I made the decision. If anything here is wrong, reply and I will review it myself.

jnadeau207-collab added a commit to jnadeau207-collab/bend that referenced this pull request Sep 24, 2026
…ive again; U32 shows and reads by its own defs

A soft native is emitted at its call site under #ifndef __METAL_VERSION__,
so C and CUDA inline it and only Metal calls the def (a spin polled errors
per call: F64 compares ran 5x an F32 one). F64.to_f32 rounds natively on
every lane with a double, NaN canonical (was the def, 300x). f32_text
tries the digit above only for a power of two (F32.show was 24% slower).
U32.show and U32.read are upstream's again: through U64 they were 2.5x
slower on JS. emit_chain opens its arms with block (from bendlang#955); F64 is a
runtime datatype, so f32_read writes CID(F64) as it writes CID(Some).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Phylliida pushed a commit to Phylliida/bend that referenced this pull request Sep 27, 2026
arr_op shares the array setup and the old-cell read and tuple return of get and swap, keeping get's retain, swap's move and set's sink. Emission is byte-identical. (PR bendlang#955, its first commit)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants