Conversation
Parameters are keyed by leaf identity, so a sample may now hold several point clouds of differing length (e.g. lidar and radar) and each is shuffled, sampled, and jittered independently.
RangeFilter3D still hand-rolls forward to reach the labels_getter, but every leaf now dispatches through the standard transform() hook and the no-input guard is shared with the base class.
The tests generated random box positions and sizes, so on some draws every jittered pose collided and nothing was pasted, silently skipping the assertions. Fixed unit cubes spaced wider than the offsets under test make the paste impossible to miss.
The object database used a lambda as its defaultdict factory, so the transform could not be pickled at all, which torch.save and DataLoader's spawn start method both require.
The default getter now resolves label-like keys the way torchvision does, so conventions like gt_labels work without a custom getter, and labels_getter=None resolves to a module-level function instead of a lambda so the transform stays pickleable.
Any plain tensor in the batch was treated as per-box labels, so a sample carrying something else such as a timestamp failed with a misleading count error instead of passing it through untouched.
The rotation matrix is built in the default dtype, so rotating a float64 point cloud or box set failed on the matmul rather than returning float64.
A batch with no boxes has nothing to copy or paste, but any point cloud or camera tensor in it was counted against zero box sets and rejected, so lidar-only, camera-only and multimodal batches all failed unless they were annotated.
Every transform exported by vision3d.transforms now has a registry entry, so the structure, pickling, pass-through and parameter-sharing audits run against each one automatically instead of being scattered per-transform and skipped for anything new.
Raise when the entry does not match the target sample during copy / paste.
yeetypete
force-pushed
the
fix/make-transforms-input-agnostic
branch
from
July 26, 2026 22:36
22659bd to
c0a0f63
Compare
1 task
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
vision3d's transforms are supposed to work on any sample structure (a dict, an
(inputs, targets)pair, a nested dict, aNamedTuple) deciding what to act on from each leaf's type rather than from a key name or position. In this PR we fix remaining issues in existing transforms and also enforce this contract going forwards.Structure-agnosticism fixes
PointShuffle,PointSampleandPointJitterdrew one set of parameters for the whole sample, so a sample holding both a lidar and a radar cloud broke as soon as the two differed in length. Parameters are now keyed by leaf identity and each cloud is shuffled, sampled and jittered independently.CopyPaste3Dtreated every plain tensor in a batch as per-box labels, so a batch carrying a timestamp failed with a misleading count error. It now locates labels throughlabels_getter.RangeFilter3Dstill hand-rollsforwardto reach thelabels_getter, but every leaf dispatches through the standardtransform()hook, and calling any transform with no inputs now raises.Correctness fixes found along the way
CopyPaste3Dcould not be pickled at all: its object database used a lambda as thedefaultdictfactory, whichtorch.saveandDataLoader'sspawnstart method both need.float64point cloud or box set failed on the matmul, because the rotation matrix is built in the default dtype.Test consolidation
Transform structure tests are now their own parametrized test suite driven by a registry. A coverage test fails if a transform is exported without an entry to ensure that all transforms get tested.
fyi: @simonschlaepfer