Point your phone at the world (or at a video), type what you're looking for, and SpotMe auto-snapshots every match. Open-vocabulary detection via Roboflow's hosted YOLO-World — no model training, no GPU, no Docker.
Two modes, one pipeline:
- Live camera — scans the rear camera ~2x/second while you walk around
- Upload video — samples a video file from your gallery every 0.5s (this is how you test "helmet-less riders" without any nearby)
phone browser ──frame (640px jpeg)──▶ /api/detect ──▶ Roboflow serverless
▲ (Next.js API /yolo_world/infer
└── boxes + auto-snapshots ◀──── route; key
stays here)
The Roboflow API key lives only in a server environment variable. The browser never sees it.
- Sign up at app.roboflow.com
- Settings → API Keys → copy your Private API Key
The free plan includes $60/month in credits — roughly 50–60k frames of YOLO-World, i.e. ~7-8 hours of continuous live scanning. The HUD's PROC readout shows per-frame processing time; per-frame cost ≈ (100ms + PROC) / 500,000 credits.
npm install
cp .env.example .env.local # then paste your key into .env.local
npm run devOpen http://localhost:3000. Note: on your laptop, localhost counts as a secure context so the camera works. But your phone can't use the camera against http://<laptop-ip>:3000 — camera access requires HTTPS. Two options for phone testing before deploying:
- Easiest: just deploy to Vercel first (2 minutes, below) and test against the deployed URL.
- Or tunnel:
npx ngrok http 3000and open the https URL on your phone.
git init && git add -A && git commit -m "SpotMe MVP"
gh repo create spotme --private --source=. --push # or push manually- vercel.com → Add New → Project → import the repo
- Framework Preset: Next.js (auto-detected — don't leave it on "Other"; you've hit that one before)
- Environment Variables → add
ROBOFLOW_API_KEY= your key - Deploy
Open the .vercel.app URL on your phone. Share → Add to Home Screen installs it as a full-screen app (manifest + icons are included).
- Download any traffic/street video from YouTube (or screen-record one) to your phone
- Open SpotMe → type
motorcycle helmet, person riding motorcycle→ Upload video → pick the clip → Scan - Matches appear in the gallery stamped with the video timestamp
- For live-mode testing: play the same video on your laptop screen, point the phone at it, and hit Start scanning — glare and moiré will cost you a little confidence, so drop sensitivity to ~0.15–0.2
Prompting tips (YOLO-World responds well to these):
- Use short noun phrases:
wrist watch,wall clock,red backpack - It's strong on common objects, weaker on niche ones — if a prompt misses, try synonyms (
motorbikevsmotorcycle) - Negative logic ("rider without helmet") isn't directly expressible; detect
motorcycle riderandhelmetseparately, then flag frames where riders appear but helmets don't (good v2 feature — the per-frame predictions already contain everything you need)
| What | Where | Default |
|---|---|---|
| Live scan rate | SCAN_INTERVAL_MS in components/LiveScanner.jsx |
450ms |
| Video sample step | STEP_SECONDS in components/VideoScanner.jsx |
0.5s |
| Frame size sent | CAPTURE_WIDTH in lib/detection.js |
640px |
| Snapshot cooldown | cooldownMs in app/page.js → useDetector |
2000ms |
| Detection model | app/api/detect/route.js (swap the endpoint path) |
YOLO-World |
- Supabase Storage for persistent snapshot history across sessions
- "Alert mode": vibrate (
navigator.vibrate) or beep on match — useful for the accessibility use case - Helmet-compliance logic: cross-reference
riderandhelmetdetections per frame - Web Share API on snapshots for one-tap sharing