Skip to content

Recall collapses on flush / minimal-reveal curb ramps: Laurens R 0.390, and the experiment that isolates it #151

Description

@jonfroehlich

⚠️ AMENDED 2026-09-03 — the GSV arm has been reviewed, and it answers this issue partly against the text below

GSV recall is 0.509 (verdict scorer 0.505), between the two pre-registered outcomes, so the answer is the split verdict this issue listed as a real possible outcome.

  • Rig/season is real and worth ~12 recall points, confirmed on the paired subset of corners both rigs saw.
  • The deficit mostly survives anyway — 0.509 is still the worst US split by ~18 points (next lowest paterson 0.684) — so the escalation below stands.
  • The second bullet under "The finding" is withdrawn. The positive near-miss delta is arm-specific: laurens_gsv reads −0.0053, back to the normal far-field direction. "Near, well-resolved ramps are being missed" was substantially a GoPro-Max-in-November artifact, not a property of rural ramps. The flush-ramp hypothesis is not refuted, but its most specific evidence is gone.
  • Separately: the town is not the problem. Every zero-shot challenger is flat or worse on the arm RampNet prefers, while RampNet gains +0.115 F1 — so Laurens is RampNet meeting an out-of-domain rig, not a town that defeats detectors.

Full numbers, the paired-radius table and the caveats are in the comment below.

The finding

Laurens (#149), the first rural split, scores precision 0.898 / recall 0.390 — precision is mid-pack, recall is roughly half the next-worst split (clovis 0.650). The evidence says this is a distinct failure mode, not a weak sample:

  • Ramp-rich, not ramp-poor. 2.65 ramps/pano, 3rd of 10 splits. Misses/pano is 1.62 against 0.45–1.19 everywhere else.
  • Not the far-field falloff. [WITHDRAWN 2026-09-03 — arm-specific; see the amendment above.] In all nine other splits missed ramps sit nearer the horizon than detected ones (delta −0.013 to −0.030 in normalized y). Laurens is the only split with a positive delta (+0.004): near, well-resolved ramps are being missed. This is orthogonal to the annapolis far-field finding and to The recall-by-distance axis is stretched ~25-30%: recompute it from GSV depth #112.
  • Not shadow or leaf litter. The imagery is November Iowa and the reviewer flagged leaves on three panos, but over a 2%-width window missed and detected ramps sit in the same light — median luminance 94.6 vs 97.4, 47% vs 42% in shadow.

Inspecting the miss crops, the ramps being missed look flush or minimal-reveal: a street cross-section close to at-grade, little or no curb face, the ramp readable mostly from a subtle grade change and a joint line. That is the standard rural/small-town design, and it is absent from every other split — all nine are urban, suburban, or a college town.

Hypothesis: RampNet keys substantially on curb-face contrast, so recall degrades as curb reveal goes to zero — independent of range, resolution and lighting. If true it is a systematic blind spot over a large fraction of the rural US, and it is invisible in the current pooled numbers because no other split contains the geometry.

The experiment

Laurens is the only city in the benchmark with both imagery sources over one footprint (area hash d8dd392b…), which makes it a natural control:

arm panos capture rig reviewed
Mapillary 4,495 2025-11 (leaf litter, low sun) GoPro Max 5760×2880 ✅ R 0.390
GSV 2,137 2024-09 (leaf-free, high sun) GSV, 16384×8192 native

Reviewing the GSV arm — one ~94-pano bundle, an afternoon — discriminates cleanly:

  • GSV recall ≈ 0.35 → the cause is ramp geometry. Flush ramps are a real blind spot; the fix is training data, not imagery. Escalate.
  • GSV recall ≈ 0.75 → the cause is rig or season, and Laurens becomes evidence about consumer 360 rigs and autumn capture rather than about rural design.

Either outcome is worth having, and no other city can produce it.

Cost: python scripts/export_benchmark.py runs/laurens_gsv/results.jsonl --bundle benchmark/laurens_gsv in the auto-labeler, then the usual review. Detection is already done — 2,137 panos, 473 operational detections, 10.8% of panos.

Caveats

Related

🤖 Generated with Claude Code (claude-opus-5[1m], effort: high)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions