Skip to content

Preserve Python and Node symbols in memory flamegraphs - #556

Merged
not-matthias merged 8 commits into
mainfrom
cod-3654-support-memory-flamegraphs-for-pythonnode
Oct 8, 2026
Merged

not-matthias merged 8 commits into
mainfrom
cod-3654-support-memory-flamegraphs-for-pythonnode

Conversation

@not-matthias

@not-matthias not-matthias commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Capture per-process perf maps and Python JIT unwind data alongside native memory artifacts.
  • Enable V8 perf maps for memory runs while preserving existing NODE_OPTIONS.
  • Document native-allocation coverage and runtime limitations.
  • One weird edge case where binaries can be mapped twice into a single address space

Verification

  • cargo test --release --bin codspeed writes_keyed_artifacts_and_metadata_for_a_streamed_mapping
  • cargo fmt --all --check
  • Local Python 3.12 perf-trampoline and Node 22 memory runs produced perf maps and resolved language frames when parsed offline.

@codspeed

codspeed Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Merging this PR will not alter performance

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

✅ 31 untouched benchmarks
⏩ 6 skipped benchmarks1


Comparing cod-3654-support-memory-flamegraphs-for-pythonnode (8368147) with main (9187e1a)

Open in CodSpeed

Footnotes

  1. 6 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@not-matthias
not-matthias force-pushed the cod-3654-support-memory-flamegraphs-for-pythonnode branch 2 times, most recently from 9108e6c to 1b0d6fd Compare October 5, 2026 14:27
@not-matthias
not-matthias marked this pull request as ready for review October 5, 2026 14:46
@greptile-apps

greptile-apps Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[Medium risk] Adds Python and Node.js symbol support to memory profiling.

The PR appears safe to merge, with a non-blocking gap in CI coverage for overlapping main-branch runs.

Summary

The PR adds Python and Node runtime symbols and JIT unwind artifacts to memory profiles, preserves multiple placements of a mapped module, and expands interpreter coverage in CI. The latest changes add CI concurrency and clarify the two Node wrappers.

Reviews (6) · Last reviewed commit: "fix(runtime-env): emit the Node perf map..." · Reviewed by Greptile

Comment thread crates/exec-harness/src/analysis/mod.rs Outdated
Comment thread src/executor/memory/module_artifacts.rs Outdated
Comment thread crates/memtrack/tests/interpreter_tests.rs
@not-matthias
not-matthias force-pushed the cod-3654-support-memory-flamegraphs-for-pythonnode branch 2 times, most recently from 366917f to 25472c3 Compare October 5, 2026 15:58
Comment thread crates/exec-harness/src/node.rs Outdated
Comment thread src/executor/memory/module_artifacts.rs
Comment thread README.md
@not-matthias
not-matthias force-pushed the cod-3654-support-memory-flamegraphs-for-pythonnode branch from a01c776 to 5e00d30 Compare October 6, 2026 10:12
@not-matthias
not-matthias changed the base branch from main to cod-3746-investigate-unresolved-symbols-and-truncated-memory-call October 6, 2026 10:12
@not-matthias
not-matthias added this pull request to stack #568 October 6, 2026 10:20

@GuillaumeLagrange GuillaumeLagrange left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

olgtm, second round will be quick

Comment thread crates/exec-harness/src/analysis/mod.rs Outdated
Comment thread crates/exec-harness/src/node.rs Outdated
Comment thread crates/exec-harness/src/walltime/benchmark_loop.rs Outdated
Comment thread crates/memtrack/tests/interpreter_tests.rs Outdated
Comment thread src/executor/shared/module_artifacts/loaded_module.rs
Comment thread src/executor/wall_time/profiler/perf/jit_dump.rs Outdated
Base automatically changed from cod-3746-investigate-unresolved-symbols-and-truncated-memory-call to main October 6, 2026 13:06
@not-matthias
not-matthias force-pushed the cod-3654-support-memory-flamegraphs-for-pythonnode branch 2 times, most recently from 1f73202 to 2b88350 Compare October 6, 2026 15:39
@not-matthias
not-matthias force-pushed the cod-3654-support-memory-flamegraphs-for-pythonnode branch from 2b88350 to be99440 Compare October 7, 2026 07:58
Comment thread crates/runner-shared/src/runtime_env/node.rs Outdated
Comment thread crates/runner-shared/tests/node_wrapper.rs Outdated

@GuillaumeLagrange GuillaumeLagrange left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

olgtm

Comment thread crates/runner-shared/src/runtime_env/mod.rs Outdated
Comment thread crates/runner-shared/src/runtime_env/mod.rs
Comment thread crates/runner-shared/src/runtime_env/mod.rs Outdated
Comment thread src/executor/memory/module_artifacts.rs
@not-matthias
not-matthias force-pushed the cod-3654-support-memory-flamegraphs-for-pythonnode branch 2 times, most recently from a414a0a to 78361d2 Compare October 7, 2026 14:14
Comment thread crates/exec-harness/src/runtime_env/node/mod.rs
@not-matthias
not-matthias force-pushed the cod-3654-support-memory-flamegraphs-for-pythonnode branch from 78361d2 to 3140211 Compare October 7, 2026 14:22
@not-matthias
not-matthias force-pushed the cod-3654-support-memory-flamegraphs-for-pythonnode branch from 3140211 to 4818d34 Compare October 8, 2026 09:09

@GuillaumeLagrange GuillaumeLagrange left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

Comment thread crates/exec-harness/src/runtime_env/node/mod.rs
Python and Node write runtime symbols to /tmp/perf-<pid>.map, and Python JIT dumps carry the unwind data needed to walk through interpreter trampolines. Collect both for the benchmark processes before saving the memtrack metadata, reusing the walltime artifact pipeline, so offline allocation stacks keep their runtime frames.
Memory stacks can only name Python and Node frames from the runtime perf maps. Set PYTHONPERFSUPPORT and the Node perf options per benchmark command, as simulation mode already does. Memory mode also passes --interpreted-frames-native-stack, since interpreted JS frames otherwise all resolve to the shared V8 interpreter trampoline.
Track a native allocation made from a Python function (perf trampoline) and a Node function (V8 perf-basic-prof). Assert the allocation carries a captured stack and that the runtime perf map names the allocating function, which offline attribution needs.
A process can map the same file at several addresses at once. V8 remaps its
embedded builtins out of the node binary into its code range, so node runs
from both its original text and that copy. Each placement has its own load
bias, but only the last mapping per (path, pid) was kept, so frames in every
earlier placement lost their symbols and unwind data. In a Node memory
profile, all native node frames showed up as unresolved addresses.

Record every distinct load bias and the unwind data of every executable
mapping for each process, and emit them all in the artifact metadata. The
walltime perf path shares this bookkeeping and gets the same fix.

Add a sample that maps its own text a second time and allocates through
both copies. The test feeds that process's real mappings into the artifact
pipeline and checks that both copies resolve.

Refs COD-1377
…rapper

The runner (`helpers/env.rs`) and exec-harness each defined the environment
a benchmark process needs: `CODSPEED_RUNNER_MODE`, the Python hash seed and
perf-map switches, and the Java tool options. Move them into one
`runner_shared::runtime_env` module, keyed on `MeasurementMode`, which moves
to runner-shared as well. The runner converts its `RunnerMode` and extends
its injected env from the shared list; the values are unchanged except that
memory mode now also sets `PYTHONPERFSUPPORT=1`, so memory flamegraphs can
name Python frames.

Add a `node` wrapper script that the module installs on `PATH`. It runs the
real node with the V8 flags codspeed-node's `getV8Flags()` would request for
the current `CODSPEED_RUNNER_MODE` and node major version: the deterministic
analysis set for simulation and memory, and the perf-prof/log-code set for
walltime. Most of these flags are rejected in `NODE_OPTIONS`, so exec-harness
targets without a codspeed-node integration had no way to get them before.

The wrapper may run under valgrind with the simulation preload library in
`LD_PRELOAD`, which reports a benchmark result from every process that loads
it. The script therefore drops `LD_PRELOAD` for its only subprocess
(`node --version`), restores it for the final `exec`, and avoids command
substitutions whose forked subshells would report on exit. Under callgrind
only the real node process emits the benchmark dump.

The install writes a staging file and renames it into place, once per
process: writing over an executing script, or a second in-process writer
racing with a concurrent fork+exec, fails with `ETXTBSY`.
Set the benchmark environment once in `execute_benchmarks`, before the mode
dispatch, through `runner_shared::runtime_env::apply_to_process`. Every
spawned command inherits it, so the per-command `set_perf_map_env` and
`set_node_options` calls in the memory, simulation and walltime loops go
away, together with the local `node.rs` and its `NODE_OPTIONS` subset.

Node targets now go through the shared `node` wrapper and receive the full
codspeed-node V8 flag set for the mode, including indirect launches through
npm, npx or `#!/usr/bin/env node` scripts. `MeasurementMode` is re-exported
from runner-shared, so the CLI is unchanged.
codspeed-node's analysis flag set has no `--perf-basic-prof`, so the node
wrapper in memory mode produced no `/tmp/perf-<pid>.map` and memory
flamegraphs lost their JS frames. Add the flag for memory mode only.

Switch the memtrack interpreter tests to the shared runtime env instead of
hand-picked interpreter flags, so they exercise the same Python env and
node wrapper the runner injects.
@not-matthias
not-matthias force-pushed the cod-3654-support-memory-flamegraphs-for-pythonnode branch from 4818d34 to 8368147 Compare October 8, 2026 09:44
@greptile-apps

greptile-apps Bot commented Oct 8, 2026

Copy link
Copy Markdown

Comments Outside Diff

These findings could not be posted inline.

  • P2 Queued CI runs can disappear .github/workflows/ci.yml:9 ▶

    If several runs overlap on main, this group permits only one pending run, even though cancellation is disabled for pushes. A third push or manually dispatched run can replace a queued push run before it starts, leaving that commit without CI results. This is a non-blocking gap in main-branch CI coverage.

@not-matthias
not-matthias merged commit 8368147 into main Oct 8, 2026
62 checks passed
@not-matthias
not-matthias deleted the cod-3654-support-memory-flamegraphs-for-pythonnode branch October 8, 2026 10:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants