Skip to content

Benchmark: solve each problem N times, take per-problem median - #130

Merged
darnstrom merged 4 commits into
masterfrom
copilot/improve-benchmark-reliability
Mar 21, 2026
Merged

darnstrom merged 4 commits into
masterfrom
copilot/improve-benchmark-reliability

Conversation

Copilot AI commented Mar 21, 2026 •

Copy link
Copy Markdown
Contributor

The previous benchmark took the median over 11 distinct problem instances (one solve each), so timing variance came from problem difficulty differences rather than OS/scheduler noise.

Changes

  • Two-level timing loop in benchmark_qp / benchmark_lp:

    • Outer loop generates num_problems (default 10) distinct problem instances
    • Inner loop solves each instance num_repeats (default 5) times; per-problem timing = median of those repeats
    • Reported solve_median = median of the 10 per-problem medians
  • CSV schema: num_runs replaced with num_problems + num_repeats

for _ in 1:num_problems
    xref, H, f, A, bupper, blower, sense = generate_test_QP(...)
    prob_solve = Float64[]
    for _ in 1:num_repeats
        x, fval, exitflag, info = quadprog(H, f, A, bupper, blower, sense)
        push!(prob_solve, info.solve_time)
    end
    push!(solve_medians, median(prob_solve))  # smooth out scheduler jitter
end
# reported value: median(solve_medians)

✨ Let Copilot coding agent set things up for you — coding agent works faster and does higher quality work when set up for your repo.

…hmark execution

- Use median instead of mean for timing in benchmark.jl to reduce fluctuation sensitivity
- Increase default num_runs from 5 to 11 for a more reliable median
- Update CSV header and output columns from *_mean_s to *_median_s
- Update compare_benchmarks.jl to use median column names for comparison
- Enable CTest when TEST OR BENCHMARK so -DBENCHMARK=ON works without -DTEST=ON
- Guard the Julia normal test suite (Pkg.test()) with if(TEST) so it only runs when TEST=ON

Co-authored-by: darnstrom <55484604+darnstrom@users.noreply.github.com>
Agent-Logs-Url: https://github.com/darnstrom/daqp/sessions/45b32e7d-23fd-4da1-ab7a-c292cea072bd
Copilot AI changed the title [WIP] Improve benchmarking reliability by taking median time Benchmark: use median timing and decouple from Julia test suite Mar 21, 2026
Copilot AI requested a review from darnstrom March 21, 2026 10:13
Replace `if(TEST OR BENCHMARK)` with `if(TEST) / elseif(BENCHMARK)` to avoid
the conflict with CMake's built-in TEST operator in if() expressions (CMP0064).

Co-authored-by: darnstrom <55484604+darnstrom@users.noreply.github.com>
Agent-Logs-Url: https://github.com/darnstrom/daqp/sessions/00446e09-47ee-43e3-8735-4b5bd6ecd13c
Copilot AI changed the title Benchmark: use median timing and decouple from Julia test suite Improve benchmark reliability and decouple from Julia test suite Mar 21, 2026
…lem median

Previously, benchmark_qp/lp generated a new problem on each of 11 runs
and took the median across those different problems. The intent was instead
to generate N distinct problems and solve each M times, taking the median
of the M solves to smooth out OS/scheduling jitter.

- Add num_problems=10 and num_repeats=5 parameters
- Inner loop solves the same problem num_repeats times; take per-problem median
- Outer loop collects per-problem medians; report median of those medians
- CSV: replace num_runs column with num_problems and num_repeats

Co-authored-by: darnstrom <55484604+darnstrom@users.noreply.github.com>
Agent-Logs-Url: https://github.com/darnstrom/daqp/sessions/3c83b141-cf4f-4947-9546-5fc52ca5cd7c
Copilot AI changed the title Improve benchmark reliability and decouple from Julia test suite Benchmark: solve each problem N times, take per-problem median Mar 21, 2026
@darnstrom
darnstrom marked this pull request as ready for review March 21, 2026 13:40
@darnstrom
darnstrom merged commit 0c8ef74 into master Mar 21, 2026
14 checks passed
@darnstrom
darnstrom deleted the copilot/improve-benchmark-reliability branch July 19, 2026 11:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants