Conversation
Add set_executed_benchmark_for_pid, which passes an explicit pid to instrument_hooks_set_executed_benchmark instead of the calling process' own. set_executed_benchmark keeps its behavior and delegates to it. Bump instrument-hooks, whose valgrind instrument now writes a "Benchmark pid: <pid>" desc line in the dump part when that pid is not the calling process'. This also brings thread-safe C API exports and the callgrind_toggle_collect helper. Refs COD-3722 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
exec-harness measures a command from its own process, so the dumped part also holds the harness's cost of spawning the command and waiting for it, a fixed overhead added to every benchmark. Spawn the command and pass its pid to set_executed_benchmark_for_pid, so the profile states which process ran the benchmark and the harness's own cost can be told apart. Under valgrind this needs a valgrind build supporting CALLGRIND_REGISTER_DESC; older builds ignore it. Closes COD-3722 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Merging this PR will improve performance by 14.28%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ⚡ | WallTime | memtrack track tar |
9.6 s | 8.4 s | +14.28% |
Tip
Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.
Comparing cod-3722-ignore-process-spawning-overhead-in-exec-harness-simulation (485598a) with spike/cod-3440-memtrack-musl (ef0764f)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Make exec-harness declare the pid of the command it benchmarks, so its own spawning cost can be left out of simulation results.
Since #531, exec-harness turns instrumentation on in its own process and spawns the command, which inherits that state across fork and exec. The dumped part therefore also holds the harness's cost of spawning the command and waiting for it. That is a fixed overhead on every benchmark, about 207k instructions per command locally.
Changes:
set_executed_benchmark_for_pid(pid, uri).set_executed_benchmarkkeeps its behavior and delegates to it.Command::status(). In memory mode the pid only travels to the runner's FIFO, where nothing reads it for exec-harness.desc: Benchmark pid: <pid>in the part when the pid is not the caller's. This bump also brings the thread-safe C API exports and thecallgrind_toggle_collecthelper.Checked locally under the patched valgrind, with the runner's simulation flags on
sh -c '/bin/true; /bin/true; :':Still to do before this is ready:
0codspeed8once feat(callgrind): add CALLGRIND_REGISTER_DESC client request valgrind-codspeed#43 is released. Until then the request is ignored with a warning and results keep the harness's cost.Closes COD-3722