Benchmark: solve each problem N times, take per-problem median - #130
Merged
Merged
Conversation
…hmark execution - Use median instead of mean for timing in benchmark.jl to reduce fluctuation sensitivity - Increase default num_runs from 5 to 11 for a more reliable median - Update CSV header and output columns from *_mean_s to *_median_s - Update compare_benchmarks.jl to use median column names for comparison - Enable CTest when TEST OR BENCHMARK so -DBENCHMARK=ON works without -DTEST=ON - Guard the Julia normal test suite (Pkg.test()) with if(TEST) so it only runs when TEST=ON Co-authored-by: darnstrom <55484604+darnstrom@users.noreply.github.com> Agent-Logs-Url: https://github.com/darnstrom/daqp/sessions/45b32e7d-23fd-4da1-ab7a-c292cea072bd
Copilot
AI
changed the title
[WIP] Improve benchmarking reliability by taking median time
Benchmark: use median timing and decouple from Julia test suite
Mar 21, 2026
Replace `if(TEST OR BENCHMARK)` with `if(TEST) / elseif(BENCHMARK)` to avoid the conflict with CMake's built-in TEST operator in if() expressions (CMP0064). Co-authored-by: darnstrom <55484604+darnstrom@users.noreply.github.com> Agent-Logs-Url: https://github.com/darnstrom/daqp/sessions/00446e09-47ee-43e3-8735-4b5bd6ecd13c
Copilot
AI
changed the title
Benchmark: use median timing and decouple from Julia test suite
Improve benchmark reliability and decouple from Julia test suite
Mar 21, 2026
…lem median Previously, benchmark_qp/lp generated a new problem on each of 11 runs and took the median across those different problems. The intent was instead to generate N distinct problems and solve each M times, taking the median of the M solves to smooth out OS/scheduling jitter. - Add num_problems=10 and num_repeats=5 parameters - Inner loop solves the same problem num_repeats times; take per-problem median - Outer loop collects per-problem medians; report median of those medians - CSV: replace num_runs column with num_problems and num_repeats Co-authored-by: darnstrom <55484604+darnstrom@users.noreply.github.com> Agent-Logs-Url: https://github.com/darnstrom/daqp/sessions/3c83b141-cf4f-4947-9546-5fc52ca5cd7c
Copilot
AI
changed the title
Improve benchmark reliability and decouple from Julia test suite
Benchmark: solve each problem N times, take per-problem median
Mar 21, 2026
darnstrom
marked this pull request as ready for review
March 21, 2026 13:40
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The previous benchmark took the median over 11 distinct problem instances (one solve each), so timing variance came from problem difficulty differences rather than OS/scheduler noise.
Changes
Two-level timing loop in
benchmark_qp/benchmark_lp:num_problems(default 10) distinct problem instancesnum_repeats(default 5) times; per-problem timing = median of those repeatssolve_median= median of the 10 per-problem mediansCSV schema:
num_runsreplaced withnum_problems+num_repeats✨ Let Copilot coding agent set things up for you — coding agent works faster and does higher quality work when set up for your repo.