Skip to content
This repository was archived by the owner on Oct 8, 2026. It is now read-only.

test(perf): add manual-only extended coverage and DuckDB baseline - #1241

Merged
bill-ph merged 13 commits into
mainfrom
codex/perf-workload-coverage
Sep 29, 2026
Merged

bill-ph merged 13 commits into
mainfrom
codex/perf-workload-coverage

Conversation

@bill-ph

@bill-ph bill-ph commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

The frozen benchmarks leave common analytical query shapes unmeasured. Add 12 versioned workloads covering filtered uniques/trends, two person joins, event and actor retrieval, ordered funnels, sampled-week retention, two cohorts, and daily-person/session modeling SELECTs.

Manually select scenario=posthog_frozen_perf_extended to run all five configurations, or set coverage_target for one. Shared image builds feed isolated, sequential target jobs with separate artifacts and a pinned Trino digest. Each target gets a four-hour scenario / 270-minute job budget; failures do not cancel remaining targets. Extended queries stay out of nightly runs.

Each query runs one warmup and four measurements. Fixed populated windows, tenant-aware person deduplication, and fixture checks avoid misleadingly empty workloads. Scenario Job names respect Kubernetes limits.

Comparable dev-sized worker resources

Standard frozen perf and extended perf use three Trino workers at 7 CPU / 28 GiB each and one DuckDB worker at 21 CPU / 84 GiB. Worker requests equal limits. Trino uses a 20 GiB heap, 10 GiB per-worker query memory and 30 GiB cluster query memory. DuckDB sizing admission caps rise with its requested profile; the current headroom policy derives a 63 GiB engine memory limit. Aggregate container resources match; coordinator/support overhead and engine memory policies differ.

Extended Trino jobs deploy only the selected comparison cache mode. Bootstrap/general E2E resources retain their small defaults. Nightly keeps its original query set with the larger comparison resources.

Deployment prerequisite: charts #16434 allows node sizes that can accommodate the larger pods. Apply it before running this profile. No larger-profile benchmark has run yet.

Historical measurement

The uncached DuckDB baseline run, commit 9c9f5f305f87b7f0d41c0b54d64e5648aeb8ef83, completed 12 warmups and 48 measurements without errors: 87m06s benchmark phase, 1h42m45s workflow. Sanitized samples and methodology remain in tests/perf/baselines/coverage-v1-duckdb-uncached-2026-09-25.

That run used the old 3 CPU / 12 GiB aggregate worker budget. It is historical data, not a like-for-like comparison base for the new profile. Its runtime extrapolations are also historical; a fresh measured run is required to establish the larger-profile baseline and forecast. The subsequent all-target run hit Trino's old 2 GiB per-node query-memory limit on joins and session modeling; this sizing change addresses that constraint but does not claim those queries now pass.

Validation and limits

  • Existing scenario checks, workflow actionlint, shell syntax, and rendered sizing checks pass. Existing fixture tests pass except the previously reproduced shared-catalog fixture failure. No new perftest unit tests; only existing raw-template substitutions were adapted.
  • just lint still reports six pre-existing SA4023 diagnostics at three unchanged sites.
  • The dependent charts change passes rendering, kubeconform and kube-linter for all nine environments.
  • Sampled-day retention is a performance shape, not a production rate; modeling queries measure computation, not writes. Uncached settings do not guarantee cold OS/storage caches. Four repetitions do not establish tail latency.

@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Test Impact Plan

Deterministic summary of how this PR changes tests, CI runners, and coverage-risk signals.

Summary

Area Added Changed Deleted
Test files 8 8 0
E2E/journey files 0 0 0
Workflow files 0 1 0

Signals

  • Test cases: +0 / -0
  • Assertions: +1 / -1
  • Skips or known failures added: 0
  • Workflow continue-on-error added: 0
  • Workflow path filters added: 0
  • Test commands removed from justfile: 0
  • E2E/journey retry lines added: 0

Coverage risk: neutral or increased

No coverage-reduction warnings detected.

@bill-ph
bill-ph marked this pull request as ready for review September 25, 2026 17:58
@bill-ph bill-ph changed the title test(perf): add analytical workload coverage and focused DuckDB baseline test(perf): add manual-only extended coverage and DuckDB baseline Sep 25, 2026
@bill-ph
bill-ph enabled auto-merge (squash) September 29, 2026 01:39
@bill-ph
bill-ph merged commit a1becc8 into main Sep 29, 2026
22 checks passed
@bill-ph
bill-ph deleted the codex/perf-workload-coverage branch September 29, 2026 01:40
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant