Skip to content

Release 2.0.0: physiology review - #97

Merged
0jonjo merged 38 commits into
mainfrom
release/v2.0.0
Oct 2, 2026
Merged

0jonjo merged 38 commits into
mainfrom
release/v2.0.0

Conversation

@0jonjo

@0jonjo 0jonjo commented Oct 2, 2026

Copy link
Copy Markdown
Owner

calcpace 2.0.0 — physiology review

Combines the reviewed phases into one release (supersedes #94, #95, #96 and two local-only branches). Full details, before/after tables and sources in CHANGELOG.md → [2.0.0].

Breaking

  • Cameron predictor uses Dave Cameron's original formula (the old constants were wrong); CAMERON_A/B/C removed; distances above 100 km raise ArgumentError.
  • Age grading uses the 2025 road tables (Alan Jones / USATF MLDR) instead of the 2023 track tables; WMA_DATA keys/ages changed (18–100).
  • Clock strings are validated everywhere (Checker::CLOCK_FORMAT): seconds/minutes ≥ 60 and garbage now raise InvalidTimeFormatError instead of returning wrong seconds or 0. Every clock the gem writes still parses (signed, padded, day-prefixed).
  • check_positive rejects non-finite numbers.

Changed (numbers)

  • Heat: base curve and duration factor fitted jointly to El Helou et al. 2012 (Table S3), flat from 35 °C; long efforts get much smaller penalties than 1.18.1.
  • Altitude: no step at 914 m (starts at 300 m), extrapolated to 4000 m.
  • Negative/positive splits: ±1% per half (was ±4%).
  • Marathon training band ends at the VDOT-predicted marathon pace.

Added

  • Optional humidity: / dew_point: for heat (WBGT-based effective temperature).
  • PersonalizedPredictor: marathon from training volume (Tanda 2011), personal Riegel exponent.
  • Grade-adjusted pace (Minetti 2002) and track_grade_adjusted_splits.
  • VO2max percentile and label by age and sex (FRIEND, Kaminsky et al. 2015).

Verification

  • 833 runs, 0 failures; rubocop clean; gem builds with the new data files.
  • Per-phase adversarial reviews + a final review of the combined branch + a delta review of the last parser commits.
  • calcpace_web suite run against this branch: only expected number changes (the site stays on ~> 1.18 until its own upgrade PR).

0jonjo added 30 commits October 2, 2026 06:20
- Cameron: replace the made-up exponential constants with Dave Cameron's
  velocity-ratio model, f(d) = 13.49681 - 0.000030363*d + 835.7114/d^0.7905
  (d in metres). The old constants were far more optimistic than Riegel for
  the marathon; the real model is more conservative.
- Altitude: threshold 914.4 m -> 300 m with a linear ramp to the first NCAA
  point (no more 0% -> 1.41% jump at 915 m) and quadratic-fit extrapolation
  to 4000 m instead of a 5.90% cap from 2438 m.
- Splits: negative/positive strategies use +-1% pace per half instead of +-4%.
- Heat: extrapolate 35 C (8.7%) and 40 C (10.9%) from the 25->30 C slope.
Replace the age-grading data with Alan Jones' 2025 road tables (approved
2025-01-10 by the USATF Masters Long Distance Running Council), taken from
MaleRoadStd2025.xlsx / FemaleRoadStd2025.xlsx in
github.com/AlanLyttonJones/Age-Grade-Tables.

The old file, despite its "road" name, held the WMA 2023 track and field
factors and track open standards, with no factors under age 30, so road
age grades were off by about 1-3%. Factors now cover every age from 18 to
100 for 5K, 10K, half marathon and marathon.

Data files renamed to mldr_2025_road*.yml; table_version becomes
MLDR_2025_ROAD_ONE_YEAR_FACTORS_V1. Public constants and category labels
are unchanged.
All review phases will ship together in a single later release.
Cameron's f(d) crosses zero near 445 km, so long distances produced negative
or absurd times. Add CAMERON_MAX_DISTANCE_KM (100 km, keeps '100k' usable)
and raise ArgumentError when either distance exceeds it.

Also strengthen the altitude continuity and YAML structure tests, and move
the changelog entry under Unreleased, marking the removed CAMERON_A/B/C
constants and the new range limit as breaking.
Add PersonalizedPredictor with predictions built from the runner's own data:

- predict_marathon_from_training: Tanda (2011) regression on mean weekly
  distance and mean training pace over the 8 weeks before the race. Inputs
  or a predicted time outside the paper's sample are flagged in
  :out_of_range instead of raising.
- riegel_exponent: personal k fitted to two performances.
- predict_time_personal: Riegel from the performance closer to the target
  (log-distance) with k clamped to [1.01, 1.20].
Link the 2025 road tables at the commit the data came from, correct the
age-clamp note (old table ran to 110, new ends at 100), list the dropped
WMA_DATA track keys under Breaking, and replace the vacuous interpolation
test with one that drives a sparse stubbed table.
Add grade_adjustment_factor, grade_adjusted_pace(_clock) and
track_grade_adjusted_splits, which returns the track_splits fields plus a
per-split :gap. Grades are measured over segments of at least 100 m so GPS
elevation noise does not turn into fake climbing; points without :ele are
flat. track_splits and estimate_detailed_vo2max are unchanged.
Infinity passed the positive check, so an infinite distance or time reached
the formulas: an infinite weekly distance produced a finite marathon
prediction, an infinite pace a FloatDomainError far from the input. It now
raises NonPositiveInputError, like zero, negatives and NaN.
- predict_time_personal: a target between the two known races is now
  interpolated along the Riegel curve through both, with the raw exponent
  and no clamping, so the result agrees with the runner's own data and does
  not depend on argument order. Extrapolation outside the pair keeps the
  clamped exponent from the closer race.
- Time strings go through check_time, like Calculator, AgeGrading and
  Vo2maxEstimator: malformed strings and symbols raise
  InvalidTimeFormatError instead of NonPositiveInputError.
- weekly_distance accepts numeric strings, read the way race distances are.
- Private helpers prefixed with personal_ to avoid mixin name collisions.
- README: training pace range given in seconds (253.3-330.6 s/km).
Add vo2max_percentile(value, age:, sex:) and optional age:/sex: keywords
on vo2max_label, read against the FRIEND registry percentiles of measured
treadmill VO2max (Kaminsky et al., Mayo Clin Proc 2015;90:1515, Table 3),
stored as YAML. Labels map to the same six strings on published
percentile cuts. Without age and sex, vo2max_label is unchanged.
The piecewise duration_factor becomes HEAT_DURATION_FACTORS, a list of
[minutes, factor] points joined by straight lines and flat outside them.
Values are bit-identical to the previous branches (checked on 0-30000 s),
and duration_factor keeps its name and signature for the site mirror.
calculate_penalty (and adjust_time, normalize_time and the *_adjusted
predictions that forward their options) accepts humidity: (relative
humidity, 0-100 %) or dew_point: (in temperature_unit). The temperature
is replaced by the effective temperature that has the same simplified
WBGT (Australian Bureau of Meteorology: 0.567*Ta + 0.393*e + 3.94) at
REFERENCE_HUMIDITY = 50 %, so the existing heat curve and the
base(temp) x duration_factor(seconds) shape stay as they are.

50 % is the humidity at which the simplified WBGT equals the air
temperature for 20-35 C (51-56 %), which is how the temperature-only
points (calibrated on Ely's WBGT figures) already read; humidity: 50
reproduces today's numbers exactly. The equation is solved by bisection
instead of adding dWBGT/0.567, which would double the humidity effect
at 30 C by ignoring the reference air's own vapour pressure.

factors gains :effective_temperature_celsius only when humidity or dew
point is given. Invalid input raises ArgumentError: humidity outside
0-100 or non-numeric, dew point above the temperature, both given, or
either without a temperature.
The factor rose from 3.0x at 3 h to 4.5x at 4 h, i.e. +50% heat penalty
for one more hour. El Helou et al. (2012, PLoS One, 1.8 M finishers of
six majors, Table S3) give, against the optimum temperature, for the
men's median (~3:58) 8.45% at 20 C and 16.9% at 25 C - 3.0x and 3.9x
the 60-minute base - and the men's Q3 (~4:28) no more (3.0x / 4.1x).
4.5x sat above every group at both temperatures; the women's groups
were lower still (1.7-2.5x).

The 3 h anchor (3.0x, Ely et al. 2007: ~9% at 20 C WBGT for a 3 h
runner) is kept; the segment now ends at 3.5x at 4 h and stays flat
after. 35 C / 4 h goes from 39.15% to 30.45%, 40 C / 4 h from 49.05%
to 38.15%. Efforts of 3 h or less are unchanged.
training_paces(50)[:marathon] ran from 4:50 to 4:25/km (75-84% VO2max),
but the VDOT-predicted marathon for VO2max 50 is 3:10:39, 4:31/km, and
Daniels' M pace is exactly that predicted race pace - so the fast end
was 6 s/km quicker than the runner's own marathon.

The fast end now comes from predict_time_from_vo2max(vo2max,
'marathon') (about 80-83% VO2max across 30-70); the slow end stays at
75%. TRAINING_INTENSITIES[:marathon][:high] becomes :race_pace. The
prediction covers VO2max 10-100; beyond it the race-pace fraction of
the nearest bound (0.800 / 0.849) is used, so training_paces keeps
accepting any positive VO2max and stays continuous at the bounds.
…ELOG

Adds the humidity/dew point keywords, the duration factor points, the
heat tables (before/after and 30 C by humidity) and the new marathon
band to the README and to the Unreleased changelog. The adjust_time and
Cameron-adjusted README examples move with the 4 h factor.
- Stretches with elevation shorter than a grade segment and with no full
  segment before them (between missing fixes, or a whole short track)
  are flat instead of graded over a few noisy metres.
- NaN/infinite :ele counts as missing in grade segments instead of
  raising a misleading grade error.
- track_splits skips the GAP bookkeeping, restoring its original cost.
- FRIEND citation: 7,783 tests on adults free of known cardiovascular
  disease, not apparently healthy adults; document truncated float ages,
  the horizontal-distance assumption, and fix README order and the
  CHANGELOG signature.
- The heat curve is now read at the unrounded effective temperature;
  only factors[:effective_temperature_celsius] is rounded. Rounding to
  2 decimals before interpolating made humidity: 50 differ from the
  temperature-only penalty for ~3% of random float temperatures
  (30.42271454 C: 6.69 vs 6.68). humidity: REFERENCE_HUMIDITY now
  returns the air temperature itself instead of a bisection result an
  ulp away. Randomized test over 5000 inputs.
- Dew points below -100 C raise ArgumentError (the Magnus formula has a
  pole at -237.7 C; -240 C used to give T_eff 70 C and the capped
  penalty).
- Complex humidity / dew point and a non-finite temperature combined
  with humidity or dew point raise ArgumentError instead of
  NoMethodError / comparison errors.
The previous 3.0x (3 h) / 3.5x (4 h) points mixed two baselines: they
used penalties measured from the ~6 C optimum, while the gem's heat
curve is zero at 15 C, and they leaned on Ely 2007 percentages for a
3 h runner (~9% at 20 C, ~12% at 25 C) that the paper's abstract does
not contain.

The 3 h and 4 h points are now a weighted least-squares fit to Table
S3: for eight groups (men/women P1, Q1, median, Q3; 2:41-4:54) the
time penalty against 15 C at 20 and 25 C, interpolated linearly between
the published points, divided by the 60-minute base (2.8 / 4.3). 1.0x
at 60 min is kept, 180 and 240 min are free, flat after 240, each sex
carries half the weight; men P1 at 25 C is beyond the table and left
out (15 observations). Result: 1.239 and 2.180 -> 1.24x and 2.18x.
2 h is the straight line between 60 and 180 min (1.12x); no group
finishes under 2:41.

The derivation table lives in environmental_factors.yml and a test
recomputes the fit from the table. The Ely percentages are gone from
the comments; Ely 2007 stays as a qualitative source with its
abstract's numbers. 25 C: 3 h 12.9% -> 5.33%, 4 h 19.35% -> 9.37%.
TRAINING_INTENSITIES[:marathon][:high] goes back to 0.84 (now the
nominal upper bound), so consumers doing arithmetic on the table keep
working. Which zones take their fast end from the VDOT-predicted race
pace is now a separate constant, PREDICTED_RACE_PACE_ZONES
(%i[marathon]). The race-pace intensity is 0.800-0.849 of VO2max over
10-100, so the marathon band stays slower than the threshold band
(from 0.83) below VO2max ~69.5, as in Daniels; tested for 30-69.
- README and CHANGELOG: the El Helou Table S3 derivation (method,
  weighting, per-group table), the new duration points, the 1.18.1 ->
  now heat grid and the 30 C humidity table; recomputed adjust_time and
  Cameron-adjusted examples; Ely 2007 kept only as a qualitative source
  with its abstract's numbers.
- predict_time_adjusted / predict_time_cameron_adjusted list humidity:
  and dew_point: among the forwarded options; the predict_time_adjusted
  example moves to 12214.84 / 6.13% and gains checked single-line
  examples.
- Marathon band: real intensity range 0.800-0.849, PREDICTED_RACE_PACE_ZONES,
  and the marathon/threshold bands no longer overlapping below VO2max
  ~69.5.
- REFERENCE_HUMIDITY no longer claims the heat points were calibrated
  on Ely's WBGT figures.
Against 15 C, El Helou et al. (2012) Table S3 penalties grow 2.70-2.97x
from 20 to 25 C in every finisher group, while the base grew only 1.54x
(2.8 -> 4.3), so no single duration factor could fit both temperatures.

- Base: P proportional to (T - 15)^p with p = log2(P25/P20), the
  sex-weighted mean being 1.497 -> 1.5. base(T) = 4.3 * ((T - 15)/10)^1.5,
  anchored at the original 25 C / 60 min value (the only one available
  for short efforts), stored every 2.5 C from 15 to 40 C (linear
  interpolation within 0.08 points). 20 C: 2.8 -> 1.52; 30 C: 6.5 -> 7.9;
  35 / 40 C (extrapolated): 12.16 / 17.0.
- Duration factor refitted against the new base by the same weighted
  least squares: 1.76x at 3 h, 2.81x at 4 h (flat after); 0.5x / 1.0x
  kept. Free 150 / 210 min points were tried and rejected (1% gain;
  non-monotonic).
- Weighted residual in penalty points: 18.3 -> 8.8; per group the ratio
  is now nearly equal at 20 and 25 C. The rest is a sex effect the model
  cannot see.
- The fit test recomputes both p and the factor points from Table S3;
  the derivation table is in environmental_factors.yml, README and the
  changelog. Examples and predictor tests updated to the new numbers.

40 C for 4 h is now 47.77% (extrapolated); no cap point applied.
The 37.5 C and 40 C points become 12.16, the 35 C value, so 35-40 C
(and anything hotter, clamped to the last point) reads like 35 C. This
is a deliberate choice, not data: everything above 25.2 C, El Helou's
hottest race, is extrapolation, and the uncapped law gave 47.77% for
4 h at 40 C. 40 C now gives 12.16% for 60 min and 34.17% for 4 h; 30 C
at 90% RH (effective 35.94 C) reaches the cap.

The fit test still checks the power law up to 35 C and the cap above
it. README/CHANGELOG grids and examples updated, and the changelog gains
a known limitation: the heat model has no sex term (men read low and
women high at 25 C, per the Table S3 residuals).
check_time accepted any two digits per field, so '05:99' or '1:60:00'
passed validation and were silently converted to the wrong number of
seconds. Seconds must now be below 60, and so must minutes when an hour
field is present. MM:SS still counts minutes past the hour ('75:00' is
75 minutes), matching the padded paces track_splits emits ('66:33').

BREAKING CHANGE: invalid clocks now raise InvalidTimeFormatError.
Document the stricter clock validation and non-finite rejection in the
errors section, show the real equivalent_performance output, list the new
capabilities in the intro, and state the Ruby versions CI actually tests.
0jonjo added 8 commits October 2, 2026 07:32
Bump the version, fold the merged phases' changelog entries into one
deduplicated 2.0.0 section, and list the new capabilities in the gemspec
summary and description.
Only a few entry points called check_time; race_splits, the pace
converters, the Riegel and Cameron predictors and race_time/race_pace
passed strings straight to convert_to_seconds, so '05:99' became 359 s
and 'abc' became 0. convert_to_seconds now validates, which puts the
2.0.0 clock rule on every path. '75:00' stays a valid 75 minutes.

BREAKING CHANGE: invalid or malformed time strings raise
InvalidTimeFormatError in every method, including convert_to_seconds.
lib/calcpace.rb never loaded version.rb, so the constant only existed
when the gemspec had been evaluated.
The stricter parser rejected the gem's own output: signed track_splits
paces ('-0:40'), padded paces past 100 minutes ('123:45'), compact
durations past 100 hours ('400:00:00') and convert_to_clocktime's day
prefix ('1 03:46:40'). Clocks are now read with one grammar,
Checker::CLOCK_FORMAT, that matches exactly what the gem formats: an
optional '-', an optional 'D HH:MM:SS' day prefix, any number of digits
in the leading field, and two-digit minutes and seconds below 60.
convert_to_seconds keeps the sign; methods that need a positive time
still reject negatives through check_positive.

Round-trip tests cover convert_to_clocktime, track_splits paces in both
formats and race_splits.
List every public constant added since 1.18.1 (checked by diffing the
constants of origin/main and this branch) and add the [2.0.0] compare
link, pointing [Unreleased] at v2.0.0...HEAD.
…lper

Vo2maxNorms called AgeGrading's private normalize_age/normalize_sex and
only worked inside Calcpace. Both validators now live in Checker, which
AgeGrading and Vo2maxNorms include, so Vo2maxNorms works on its own.
TrainingZones' copy of vo2_at_velocity duplicated Vo2maxEstimator's; it is
gone and the module documents the siblings it relies on. A composition
test fails if two included modules ever define the same method again.
A clock with hundreds of hour digits parsed to a finite Integer that
check_positive accepted and the formulas turned into Infinity or a
FloatDomainError; check_positive now also requires to_f to be finite.
clock_match only matches valid, ASCII-compatible strings, so UTF-16 or
broken UTF-8 input raises InvalidTimeFormatError instead of an encoding
error from the regexp engine. The docs now say the grammar covers every
clock the gem writes, not exactly those.
@0jonjo
0jonjo marked this pull request as ready for review October 2, 2026 11:30
Copilot AI balanced review requested due to automatic review settings October 2, 2026 11:31
@0jonjo
0jonjo merged commit 3df166c into main Oct 2, 2026
7 checks passed

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants