The GPS filter that knows a tunnel from a teleport.
A consumer GPS receiver does not fail by going quiet. It fails by lying
confidently. It reports (0, 0) before it has a fix. It sits on a stale
almanac and puts you four hundred kilometres from where you are, for three
seconds, and then snaps back without comment. It jitters a metre back and forth
while you wait at a red light. Then you sum the distances between consecutive
fixes and a two-mile errand comes out at two hundred and fifty.
gpxkit removes that noise, and then answers questions about what is left.
Pure Python, standard library only, no dependencies, ever.
from gpxkit import filter_high_quality_points
clean = filter_high_quality_points(raw_track)There is no gpxkit.Point to convert to. Coordinates and times are read
through duck-typed accessors, so all of these are valid tracks — including all
of them in the same list:
{"time": 12.0, "lat": 45.0, "lon": -99.0} # a mapping, short keys
{"timestamp": t, "latitude": …, "longitude": …} # a mapping, GPX-schema keys
(12.0, 45.0, -99.0) # (t, lat, lon)
(45.0, -99.0) # (lat, lon), no time
Trackpoint(lat=45.0, lon=-99.0, timestamp=12.0) # anything with attributesWhatever your GPX parser, telemetry SDK or database row hands back, gpxkit
reads it. Fields it cannot find are None, and a test that needs a field it
cannot see is skipped rather than failed — a missing HDOP channel is not
evidence of a bad fix. See gpxkit.points.
from gpxkit import filter_high_quality_points, haversine_km
# Two minutes of real driving at 1 Hz and about 98 km/h, as (time, lat, lon).
drive = [(3.0 + i, 45.0 + i * 0.0002, -99.0 + i * 0.0002) for i in range(120)]
# The first three fixes came off a stale almanac: 400 km north, then a snap
# back to the real position one second later.
stale = [(t, 48.6, -99.0) for t in (0.0, 1.0, 2.0)]
raw = stale + drive
clean = filter_high_quality_points(raw)
def total_km(track):
return sum(haversine_km(a[1], a[2], b[1], b[2]) for a, b in zip(track, track[1:]))| fixes | total distance | |
|---|---|---|
raw |
123 | 403.5 km |
clean |
120 | 3.2 km |
Same two minutes of driving. One of those numbers is a trip summary; the other
is a bug report. clean == drive — the three stale fixes are gone and nothing
else moved.
| noise class | what it looks like | how it is caught |
|---|---|---|
| Null island | exactly (0.0, 0.0) |
it is the no-lock sentinel, not a spot in the Gulf of Guinea |
| High HDOP | a fix the receiver itself rates as poor | ceiling on the reported hdop field |
| Stationary jitter | a metre of drift at a red light | floor on the reported speed field |
| Sub-second duplicates | four fixes stamped the same second | minimum gap between consecutive fixes |
| Spatial outliers | one fix in another country | distance from the median of the surviving cluster |
| Teleport runs | a run of stale fixes joined to the track by an impossible seam | implied speed across the seam |
The last two are the interesting pair, because the obvious single test does neither job properly.
Fifty kilometres between consecutive fixes is either a catastrophe or a Tuesday, and the distance does not tell you which:
# Half an hour underground, and a 50 km jump on the far side of it.
# Nothing is wrong here.
before = [(float(i * 60), 45.0, -99.0) for i in range(10)]
after = [(2400.0 + i * 60, 45.35 + i * 0.001, -98.6 + i * 0.001) for i in range(10)]
len(filter_high_quality_points(before + after)) # 20 — every fix survivesFifty kilometres over half an hour is ninety-seven kilometres an hour. Fifty
kilometres in one second is a hundred and eighty thousand. gpxkit never
tests raw distance against a threshold; it tests implied speed, km / hours,
against a ceiling of 1000 km/h — far above any ground vehicle or ferry, far
below any instantaneous jump. Ferries, tunnels, underground car parks, a
receiver that simply stopped logging for a while: all preserved. Stale locks:
gone.
Two guards keep the speed test from firing on ordinary noise. A seam must also clear an absolute floor (25 km) and twenty times the track's own median step, so a dense 1 Hz trace can never trip it on jitter.
The intuitive filter is "drop any fix that is too far from its neighbours." It does not work, because a stale lock is not one bad fix. It is a run of consecutive bad fixes that all agree with each other.
real track ──────●──●──●──●──●
╲ ← the only anomalous step
●──●──●──●──● ← the stale run: internally
consistent, looks perfect
In a run of five bad fixes, only the two at the seam are far from a neighbour. The three in the middle sit a metre apart, exactly like good data. A per-point rule drops the seam fixes and declares victory — but the run is still there, now joined to the track by a new seam one fix along. Run the filter again and it eats one more fix. Plot the result and you get a sawtooth: the polyline shuttling between the real cluster and the stale one, adding hundreds of kilometres every time it crosses.
A fixed radius around the median has the mirror-image problem. Make it tight enough to catch a stale lock 50 km out and it will cut a legitimate long journey in half. Make it loose enough to keep the journey and it waves the stale lock through.
So gpxkit uses both, in order. The radius is the coarse net, deliberately
generous at 500 km, and it only has to catch the wild ones. Then
strip_teleport_runs looks at the seams — and when it finds an impossible
one near either end, it removes the whole run behind it, not just the fix at
the join.
Two limits on that, on purpose:
- Interior anomalies are never removed. A bad run in the middle of a track
has several defensible fixes — interpolate across it, split the track in two,
leave it and flag it — and which one is right depends on what you are doing.
gpxkitwill not guess on your behalf. - At most 5 % is stripped from either end, and at least two fixes always survive. A filter that can quietly eat an arbitrary amount of your data is worse than the noise it removes.
from gpxkit import (
trace_time_window,
locked_gps_coverage,
window_median_speed_mps,
window_stationary_fraction,
nearest_trace_distance_km,
)
w0, w1 = trace_time_window(clean) # (3.0, 122.0)
locked_gps_coverage(clean, w0, w1) # 1.0 — a fix every second
window_median_speed_mps(clean, w0, w1) # 27.24 — m/s, about 98 km/h
window_stationary_fraction(clean, w0, w1, stop_speed_mps=0.5) # (0.0, 119)
nearest_trace_distance_km(45.05, -98.95, clean, w0, w1) # 3.57 kmlocked_gps_coverage(track, start, end)— the share of a window that actually has fixes, bucketed so that an over-sampled trace saturates at 1.0 instead of exceeding it and a dropout lowers it in proportion to how long it lasted. This is the number to demote on when you are about to make a claim about a stretch of a journey and want to know whether you have the evidence.locked_gps_duration_coverage(track, duration)— the same measure for material that is not contiguous in any single clock: a recording with sections cut out, a track stitched from several files, a set of fixes timestamped in one clock while the length you care about is measured in another. A start-to-end window would count every removed section as dead GPS. This does not.window_median_speed_mps/window_stationary_fraction— were we moving? Both derive speed from position and time rather than from a reported speed channel, because a parsed GPX file usually has no speed field and where one exists it is the receiver's own smoothed guess. The median answers "was the middle interval stopped"; the fraction answers "what share of the time were we stopped", against any cut-off you choose, and returns the interval count so you can refuse to act on a number derived from three samples.nearest_trace_distance_km(lat, lon, track, start, end)— how close did we pass to this place? Distance to the nearest sampled fix, not to the path between them, so a sparse track reports a larger number than a dense one over the same route.trace_time_window(track)— the window a track implies about itself, when you have not got one to supply.
Every one of these returns None when the answer is genuinely unmeasurable,
rather than 0.0 or 1.0. A caller that reads "no data" as "passed" builds a
gate that opens widest exactly when it knows least.
If the recording a track accompanies has been cut — sections removed, a highlight reel assembled, a trip split into legs — the raw track no longer describes what is left, and statistics over it will include the roads you cut out.
from gpxkit import trim_gps_to_clip_timeline
survived = [(0.0, 25.0), (50.0, 65.0), (85.0, 100.0)] # ranges that made the cut
trimmed = trim_gps_to_clip_timeline(track, survived)None means no edit happened and returns everything; an empty list means
nothing survived and returns nothing. For datetime timestamps, which have no
single correct origin, pass a projection:
trim_gps_to_clip_timeline(
waypoints, survived,
time_key=lambda wp: (wp.timestamp - journey_start).total_seconds(),
)from gpxkit import select_significant_roads
select_significant_roads(segments, min_duration_sec=30, min_length_m=1000, max_count=3)A track matched against a road network becomes a long list of named segments, most of which are noise — the slip road, the roundabout counted as its own way, the forty metres of a side street you used to turn around in. This keeps the few that lasted long enough and ran far enough to be worth mentioning, ranks them by duration, and hands them back in the order they were travelled so you can write a sentence with them.
pip install gpxkitPython 3.11 or newer. No dependencies. From source:
git clone https://github.com/Back-Road-Creative/gpxkit.git
cd gpxkit
pip install -e '.[test]'
pytest- It does not parse GPX. It has no XML parser, no file I/O, and no opinion
about where your points came from. Bring your own parser;
gpxkitstarts at the list of points. That is also why it has no dependencies. - The defaults are tuned for ground vehicles at roughly 1 Hz. A 1000 km/h ceiling is nonsense for a hiker and useless for an aircraft; a 500 km radius is absurd for a park run. Every threshold is a keyword argument for exactly that reason — but you do have to set them.
- Clocks are your responsibility. The window functions cannot tell epoch
seconds from seconds-into-a-recording. Mix them and you get
0.0with no complaint. Points and window bounds must share a clock. - Distances are spherical, not ellipsoidal. Haversine on a mean-radius
sphere is up to ~0.5 % off from WGS-84. Every decision here is a threshold
comparison three orders of magnitude coarser than that, so it does not
matter — but do not use
haversine_kmfor survey work. - The filter is order-dependent and single-pass. Field tests run first, then the radius, then the seam test, each on the survivors of the last. It does not iterate to a fixed point.
filter_high_quality_pointsreturns your objects, not copies, and never mutates or interpolates. It only ever removes.
MIT — see LICENSE.