Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
582 changes: 582 additions & 0 deletions docs/data/operating_point_parity_51.json

Large diffs are not rendered by default.

524 changes: 524 additions & 0 deletions docs/data/yolo_geometry_51.json

Large diffs are not rendered by default.

58 changes: 58 additions & 0 deletions docs/data/yolo_geometry_51/annapolis_pano.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
Bundle: benchmark/annapolis (125 scored panos) match radius 0.022 ground truth: reviewer-confirmed ramps + missed marks
Detection cache: /homes/gws/jonf/RampNet/.model_cache

[y11x_pano_h200] all 125 panos already cached; model load skipped
Operating point: predictions with confidence < 0.25 dropped (models without confidences are unaffected).
model P 95% CI R 95% CI F1 AP tp/fp/fn/ign
-------------------------------------------------------------------------------------------------------------------
y11x_pano 0.956 (0.892-0.983) 0.296 (0.247-0.350) 0.452 0.678 87/4/207/0
y11x_pano_h200 0.986 (0.927-0.998) 0.248 (0.202-0.301) 0.397 0.662 73/1/221/0
AP: all-point interpolated, over the recall-confirmed panos, from the full confidence range (--op-threshold does not truncate it); '-' = no calibrated per-box score.

[y11x_pano] threshold sweep (re-scored from cached detections)
thr P R F1 tp/fp/fn
0.05 0.596 0.731 0.656 215/146/79
0.10 0.856 0.626 0.723 184/31/110 <- best F1
0.15 0.928 0.524 0.670 154/12/140
0.20 0.946 0.415 0.577 122/7/172
0.25 0.956 0.296 0.452 87/4/207
0.30 0.971 0.231 0.374 68/2/226
0.40 0.977 0.143 0.249 42/1/252
0.50 1.000 0.078 0.145 23/0/271
0.60 1.000 0.048 0.091 14/0/280
0.70 1.000 0.014 0.027 4/0/290

[y11x_pano_h200] threshold sweep (re-scored from cached detections)
thr P R F1 tp/fp/fn
0.05 0.676 0.704 0.690 207/99/87 <- best F1
0.10 0.901 0.558 0.689 164/18/130
0.15 0.963 0.446 0.609 131/5/163
0.20 0.963 0.354 0.517 104/4/190
0.25 0.986 0.248 0.397 73/1/221
0.30 1.000 0.173 0.296 51/0/243
0.40 1.000 0.092 0.168 27/0/267
0.50 1.000 0.051 0.097 15/0/279
0.60 1.000 0.020 0.040 6/0/288
0.70 1.000 0.003 0.007 1/0/293

PR curves written to /homes/gws/jonf/RampNet/yolo_eval_results_geometry_51/pr_annapolis_pano (JSON + pr_curves.png)

--- RampNet verdict-based cross-check ---
Panos fully judged: 125 (of 125 seen)
Detections judged: 222 (correct 214, incorrect 8)
Detections duplicate: 3 (redundant hits on an already-counted ramp; folded into the numbers above per this run's duplicate scoring)
Detections unsure: 5 (abstained — not in precision or recall)
Missed ramps marked: 80 (+34 unsure, abstained)
Precision: 0.964 (95% CI 0.931-0.982)
Recall: 0.728 (95% CI 0.674-0.776) [vs ramps visible in the 125 recall-pool panos]

threshold kept precision recall
0.55 222 0.964 0.728
0.60 208 0.976 0.690
0.65 193 0.990 0.650
0.70 181 0.989 0.609
0.75 164 0.994 0.554
0.80 145 0.993 0.490
0.85 117 0.991 0.395
0.90 68 1.000 0.231
0.95 26 1.000 0.088
43 changes: 43 additions & 0 deletions docs/data/yolo_geometry_51/annapolis_tiles.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,43 @@
Bundle: benchmark/annapolis (125 scored panos) match radius 0.022 ground truth: reviewer-confirmed ramps + missed marks
Detection cache: /homes/gws/jonf/RampNet/.model_cache

Operating point: predictions with confidence < 0.25 dropped (models without confidences are unaffected).
model P 95% CI R 95% CI F1 AP tp/fp/fn/ign
-------------------------------------------------------------------------------------------------------------------
y11x_tiles 0.923 (0.870-0.955) 0.490 (0.433-0.547) 0.640 0.738 144/12/150/2
AP: all-point interpolated, over the recall-confirmed panos, from the full confidence range (--op-threshold does not truncate it); '-' = no calibrated per-box score.

[y11x_tiles] threshold sweep (re-scored from cached detections)
thr P R F1 tp/fp/fn
0.05 0.823 0.776 0.799 228/49/66 <- best F1
0.10 0.881 0.704 0.783 207/28/87
0.15 0.914 0.650 0.759 191/18/103
0.20 0.914 0.578 0.708 170/16/124
0.25 0.923 0.490 0.640 144/12/150
0.30 0.953 0.415 0.578 122/6/172
0.40 1.000 0.299 0.461 88/0/206
0.50 1.000 0.167 0.286 49/0/245
0.60 1.000 0.095 0.174 28/0/266
0.70 1.000 0.003 0.007 1/0/293

PR curves written to /homes/gws/jonf/RampNet/yolo_eval_results_geometry_51/pr_annapolis_tiles (JSON + pr_curves.png)

--- RampNet verdict-based cross-check ---
Panos fully judged: 125 (of 125 seen)
Detections judged: 222 (correct 214, incorrect 8)
Detections duplicate: 3 (redundant hits on an already-counted ramp; folded into the numbers above per this run's duplicate scoring)
Detections unsure: 5 (abstained — not in precision or recall)
Missed ramps marked: 80 (+34 unsure, abstained)
Precision: 0.964 (95% CI 0.931-0.982)
Recall: 0.728 (95% CI 0.674-0.776) [vs ramps visible in the 125 recall-pool panos]

threshold kept precision recall
0.55 222 0.964 0.728
0.60 208 0.976 0.690
0.65 193 0.990 0.650
0.70 181 0.989 0.609
0.75 164 0.994 0.554
0.80 145 0.993 0.490
0.85 117 0.991 0.395
0.90 68 1.000 0.231
0.95 26 1.000 0.088
60 changes: 60 additions & 0 deletions docs/data/yolo_geometry_51/bend_pano.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,60 @@
Bundle: benchmark/bend (110 scored panos) match radius 0.022 ground truth: reviewer-confirmed ramps + missed marks
Detection cache: /homes/gws/jonf/RampNet/.model_cache

[y11x_pano_h200] all 110 panos already cached; model load skipped
Operating point: predictions with confidence < 0.25 dropped (models without confidences are unaffected).
model P 95% CI R 95% CI F1 AP tp/fp/fn/ign
-------------------------------------------------------------------------------------------------------------------
y11x_pano 0.974 (0.941-0.989) 0.581 (0.527-0.633) 0.728 0.775 190/5/137/2
y11x_pano_h200 0.989 (0.961-0.997) 0.554 (0.499-0.606) 0.710 0.781 181/2/146/2
AP: all-point interpolated, over the recall-confirmed panos, from the full confidence range (--op-threshold does not truncate it); '-' = no calibrated per-box score.

[y11x_pano] threshold sweep (re-scored from cached detections)
thr P R F1 tp/fp/fn
0.05 0.650 0.810 0.721 265/143/62
0.10 0.819 0.761 0.789 249/55/78 <- best F1
0.15 0.898 0.703 0.789 230/26/97
0.20 0.932 0.633 0.754 207/15/120
0.25 0.974 0.581 0.728 190/5/137
0.30 0.987 0.468 0.635 153/2/174
0.40 0.990 0.306 0.467 100/1/227
0.50 1.000 0.205 0.340 67/0/260
0.60 1.000 0.147 0.256 48/0/279
0.70 1.000 0.101 0.183 33/0/294
0.80 1.000 0.003 0.006 1/0/326

[y11x_pano_h200] threshold sweep (re-scored from cached detections)
thr P R F1 tp/fp/fn
0.05 0.688 0.807 0.743 264/120/63
0.10 0.844 0.758 0.799 248/46/79
0.15 0.939 0.700 0.802 229/15/98 <- best F1
0.20 0.964 0.648 0.775 212/8/115
0.25 0.989 0.554 0.710 181/2/146
0.30 0.987 0.456 0.623 149/2/178
0.40 0.990 0.315 0.478 103/1/224
0.50 0.985 0.199 0.331 65/1/262
0.60 1.000 0.138 0.242 45/0/282
0.70 1.000 0.089 0.163 29/0/298
0.80 1.000 0.037 0.071 12/0/315

PR curves written to /homes/gws/jonf/RampNet/yolo_eval_results_geometry_51/pr_bend_pano (JSON + pr_curves.png)

--- RampNet verdict-based cross-check ---
Panos fully judged: 110 (of 110 seen)
Detections judged: 260 (correct 248, incorrect 12)
Detections duplicate: 1 (redundant hits on an already-counted ramp; folded into the numbers above per this run's duplicate scoring)
Detections unsure: 5 (abstained — not in precision or recall)
Missed ramps marked: 79 (+72 unsure, abstained)
Precision: 0.954 (95% CI 0.921-0.973)
Recall: 0.758 (95% CI 0.709-0.802) [vs ramps visible in the 110 recall-pool panos]

threshold kept precision recall
0.55 260 0.954 0.758
0.60 252 0.952 0.734
0.65 244 0.959 0.716
0.70 221 0.973 0.657
0.75 195 0.974 0.581
0.80 171 0.988 0.517
0.85 138 0.986 0.416
0.90 81 0.975 0.242
0.95 31 0.968 0.092
43 changes: 43 additions & 0 deletions docs/data/yolo_geometry_51/bend_tiles.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,43 @@
Bundle: benchmark/bend (110 scored panos) match radius 0.022 ground truth: reviewer-confirmed ramps + missed marks
Detection cache: /homes/gws/jonf/RampNet/.model_cache

Operating point: predictions with confidence < 0.25 dropped (models without confidences are unaffected).
model P 95% CI R 95% CI F1 AP tp/fp/fn/ign
-------------------------------------------------------------------------------------------------------------------
y11x_tiles 0.975 (0.944-0.989) 0.609 (0.555-0.660) 0.750 0.844 199/5/128/2
AP: all-point interpolated, over the recall-confirmed panos, from the full confidence range (--op-threshold does not truncate it); '-' = no calibrated per-box score.

[y11x_tiles] threshold sweep (re-scored from cached detections)
thr P R F1 tp/fp/fn
0.05 0.870 0.859 0.865 281/42/46
0.10 0.918 0.826 0.870 270/24/57 <- best F1
0.15 0.958 0.758 0.846 248/11/79
0.20 0.961 0.679 0.796 222/9/105
0.25 0.975 0.609 0.750 199/5/128
0.30 0.994 0.505 0.669 165/1/162
0.40 1.000 0.355 0.524 116/0/211
0.50 1.000 0.232 0.377 76/0/251
0.60 1.000 0.107 0.193 35/0/292
0.70 1.000 0.018 0.036 6/0/321

PR curves written to /homes/gws/jonf/RampNet/yolo_eval_results_geometry_51/pr_bend_tiles (JSON + pr_curves.png)

--- RampNet verdict-based cross-check ---
Panos fully judged: 110 (of 110 seen)
Detections judged: 260 (correct 248, incorrect 12)
Detections duplicate: 1 (redundant hits on an already-counted ramp; folded into the numbers above per this run's duplicate scoring)
Detections unsure: 5 (abstained — not in precision or recall)
Missed ramps marked: 79 (+72 unsure, abstained)
Precision: 0.954 (95% CI 0.921-0.973)
Recall: 0.758 (95% CI 0.709-0.802) [vs ramps visible in the 110 recall-pool panos]

threshold kept precision recall
0.55 260 0.954 0.758
0.60 252 0.952 0.734
0.65 244 0.959 0.716
0.70 221 0.973 0.657
0.75 195 0.974 0.581
0.80 171 0.988 0.517
0.85 138 0.986 0.416
0.90 81 0.975 0.242
0.95 31 0.968 0.092
56 changes: 56 additions & 0 deletions docs/data/yolo_geometry_51/budapest_district5_pano.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,56 @@
Bundle: benchmark/budapest_district5 (125 scored panos) match radius 0.022 ground truth: reviewer-confirmed ramps + missed marks
Detection cache: /homes/gws/jonf/RampNet/.model_cache

[y11x_pano_h200] all 125 panos already cached; model load skipped
Operating point: predictions with confidence < 0.25 dropped (models without confidences are unaffected).
model P 95% CI R 95% CI F1 AP tp/fp/fn/ign
-------------------------------------------------------------------------------------------------------------------
y11x_pano 0.732 (0.581-0.843) 0.100 (0.071-0.139) 0.176 0.444 30/11/270/1
y11x_pano_h200 0.864 (0.733-0.936) 0.127 (0.094-0.169) 0.221 0.427 38/6/262/2
AP: all-point interpolated, over the recall-confirmed panos, from the full confidence range (--op-threshold does not truncate it); '-' = no calibrated per-box score.

[y11x_pano] threshold sweep (re-scored from cached detections)
thr P R F1 tp/fp/fn
0.05 0.467 0.630 0.536 189/216/111 <- best F1
0.10 0.668 0.437 0.528 131/65/169
0.15 0.720 0.317 0.440 95/37/205
0.20 0.726 0.177 0.284 53/20/247
0.25 0.732 0.100 0.176 30/11/270
0.30 0.750 0.070 0.128 21/7/279
0.40 0.733 0.037 0.070 11/4/289
0.50 0.750 0.030 0.058 9/3/291
0.60 0.778 0.023 0.045 7/2/293

[y11x_pano_h200] threshold sweep (re-scored from cached detections)
thr P R F1 tp/fp/fn
0.05 0.444 0.570 0.499 171/214/129
0.10 0.702 0.400 0.510 120/51/180 <- best F1
0.15 0.786 0.270 0.402 81/22/219
0.20 0.810 0.170 0.281 51/12/249
0.25 0.864 0.127 0.221 38/6/262
0.30 0.828 0.080 0.146 24/5/276
0.40 0.727 0.027 0.051 8/3/292
0.50 0.833 0.017 0.033 5/1/295
0.60 1.000 0.010 0.020 3/0/297

PR curves written to /homes/gws/jonf/RampNet/yolo_eval_results_geometry_51/pr_budapest_district5_pano (JSON + pr_curves.png)

--- RampNet verdict-based cross-check ---
Panos fully judged: 125 (of 125 seen)
Detections judged: 173 (correct 151, incorrect 22)
Detections duplicate: 7 (redundant hits on an already-counted ramp; folded into the numbers above per this run's duplicate scoring)
Detections unsure: 16 (abstained — not in precision or recall)
Missed ramps marked: 149 (+48 unsure, abstained)
Precision: 0.873 (95% CI 0.815-0.915)
Recall: 0.503 (95% CI 0.447-0.560) [vs ramps visible in the 125 recall-pool panos]

threshold kept precision recall
0.55 173 0.873 0.503
0.60 139 0.914 0.423
0.65 119 0.924 0.367
0.70 99 0.919 0.303
0.75 73 0.904 0.220
0.80 50 0.920 0.153
0.85 23 1.000 0.077
0.90 9 1.000 0.030
0.95 2 1.000 0.007
42 changes: 42 additions & 0 deletions docs/data/yolo_geometry_51/budapest_district5_tiles.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,42 @@
Bundle: benchmark/budapest_district5 (125 scored panos) match radius 0.022 ground truth: reviewer-confirmed ramps + missed marks
Detection cache: /homes/gws/jonf/RampNet/.model_cache

Operating point: predictions with confidence < 0.25 dropped (models without confidences are unaffected).
model P 95% CI R 95% CI F1 AP tp/fp/fn/ign
-------------------------------------------------------------------------------------------------------------------
y11x_tiles 0.843 (0.740-0.910) 0.197 (0.156-0.245) 0.319 0.458 59/11/241/4
AP: all-point interpolated, over the recall-confirmed panos, from the full confidence range (--op-threshold does not truncate it); '-' = no calibrated per-box score.

[y11x_tiles] threshold sweep (re-scored from cached detections)
thr P R F1 tp/fp/fn
0.05 0.645 0.570 0.605 171/94/129 <- best F1
0.10 0.701 0.453 0.551 136/58/164
0.15 0.757 0.343 0.472 103/33/197
0.20 0.784 0.267 0.398 80/22/220
0.25 0.843 0.197 0.319 59/11/241
0.30 0.881 0.123 0.216 37/5/263
0.40 0.958 0.077 0.142 23/1/277
0.50 0.900 0.030 0.058 9/1/291
0.60 1.000 0.013 0.026 4/0/296

PR curves written to /homes/gws/jonf/RampNet/yolo_eval_results_geometry_51/pr_budapest_district5_tiles (JSON + pr_curves.png)

--- RampNet verdict-based cross-check ---
Panos fully judged: 125 (of 125 seen)
Detections judged: 173 (correct 151, incorrect 22)
Detections duplicate: 7 (redundant hits on an already-counted ramp; folded into the numbers above per this run's duplicate scoring)
Detections unsure: 16 (abstained — not in precision or recall)
Missed ramps marked: 149 (+48 unsure, abstained)
Precision: 0.873 (95% CI 0.815-0.915)
Recall: 0.503 (95% CI 0.447-0.560) [vs ramps visible in the 125 recall-pool panos]

threshold kept precision recall
0.55 173 0.873 0.503
0.60 139 0.914 0.423
0.65 119 0.924 0.367
0.70 99 0.919 0.303
0.75 73 0.904 0.220
0.80 50 0.920 0.153
0.85 23 1.000 0.077
0.90 9 1.000 0.030
0.95 2 1.000 0.007
57 changes: 57 additions & 0 deletions docs/data/yolo_geometry_51/clovis_pano.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,57 @@
Bundle: benchmark/clovis (125 scored panos) match radius 0.022 ground truth: reviewer-confirmed ramps + missed marks
Detection cache: /homes/gws/jonf/RampNet/.model_cache

[y11x_pano_h200] all 125 panos already cached; model load skipped
Operating point: predictions with confidence < 0.25 dropped (models without confidences are unaffected).
model P 95% CI R 95% CI F1 AP tp/fp/fn/ign
-------------------------------------------------------------------------------------------------------------------
y11x_pano 0.918 (0.846-0.958) 0.456 (0.388-0.526) 0.610 0.720 89/8/106/1
y11x_pano_h200 0.938 (0.864-0.973) 0.390 (0.324-0.460) 0.551 0.711 76/5/119/0
AP: all-point interpolated, over the recall-confirmed panos, from the full confidence range (--op-threshold does not truncate it); '-' = no calibrated per-box score.

[y11x_pano] threshold sweep (re-scored from cached detections)
thr P R F1 tp/fp/fn
0.05 0.520 0.815 0.635 159/147/36
0.10 0.678 0.733 0.704 143/68/52
0.15 0.809 0.672 0.734 131/31/64 <- best F1
0.20 0.868 0.574 0.691 112/17/83
0.25 0.918 0.456 0.610 89/8/106
0.30 0.905 0.390 0.545 76/8/119
0.40 0.935 0.221 0.357 43/3/152
0.50 1.000 0.133 0.235 26/0/169
0.60 1.000 0.087 0.160 17/0/178
0.70 1.000 0.031 0.060 6/0/189

[y11x_pano_h200] threshold sweep (re-scored from cached detections)
thr P R F1 tp/fp/fn
0.05 0.573 0.785 0.662 153/114/42
0.10 0.788 0.687 0.734 134/36/61 <- best F1
0.15 0.866 0.595 0.705 116/18/79
0.20 0.917 0.508 0.653 99/9/96
0.25 0.938 0.390 0.551 76/5/119
0.30 0.948 0.282 0.435 55/3/140
0.40 1.000 0.164 0.282 32/0/163
0.50 1.000 0.103 0.186 20/0/175
0.60 1.000 0.046 0.088 9/0/186
0.70 1.000 0.005 0.010 1/0/194

PR curves written to /homes/gws/jonf/RampNet/yolo_eval_results_geometry_51/pr_clovis_pano (JSON + pr_curves.png)

--- RampNet verdict-based cross-check ---
Panos fully judged: 125 (of 125 seen)
Detections judged: 152 (correct 139, incorrect 13)
Detections unsure: 12 (abstained — not in precision or recall)
Missed ramps marked: 56 (+21 unsure, abstained)
Precision: 0.914 (95% CI 0.859-0.949)
Recall: 0.713 (95% CI 0.646-0.772) [vs ramps visible in the 125 recall-pool panos]

threshold kept precision recall
0.55 152 0.914 0.713
0.60 140 0.936 0.672
0.65 126 0.937 0.605
0.70 115 0.948 0.559
0.75 96 0.938 0.462
0.80 75 0.973 0.374
0.85 47 0.979 0.236
0.90 25 0.960 0.123
0.95 6 1.000 0.031
Loading
Loading