High-throughput, resume-safe vision and video embedding pipelines for VLM datasets—26 pinned encoders, streaming Hugging Face ingestion, and Safetensors shard commits.
-
Updated
Aug 9, 2026 - Python
High-throughput, resume-safe vision and video embedding pipelines for VLM datasets—26 pinned encoders, streaming Hugging Face ingestion, and Safetensors shard commits.
Open-source vision stack with stereo camera hardware, GPU processing, and AI agent for training video classifiers.
vjepa / vjepa2 / vjepa2.1 PCA visualization utility for dense features and world model inspection.
Can the V-JEPA2 model be used as a world model?
Assess Data Quality Before Annotation or Labelled Data Quality after Annotation (Txt files/Yolo Format). Visualise the patterns covered by each class/activity.
Masked Multi-Component Gated Decomposition Architecture
V-JEPA for Gray-Scott dynamics. Initial work produced during the 24-hour Hack the World(s) hackathon. 1st place 🏆
A physics-based video search engine using Meta's V-JEPA 2 world model to find videos with similar motion dynamics.
Patch-level predictive surprise from video foundation model embeddings. The embedding delta is the attention signal.
End-to-end engineering of a V-JEPA-style video world-model trainer: data curation at scale, distributed (FSDP) training, and inference optimization. From-scratch JEPA + SIGReg.
GranularVAR: Multi-scale video understanding with augmentation-graded contrastive learning and calibrated uncertainty. V-JEPA 2 encoder + granularity-conditioned decoder. GWU MS Data Science Capstone, Spring 2026.
SCOUT: frozen-encoder embedding-space prediction for sim-to-real text-based person retrieval. ECCV 2026 Workshop (AI City Challenge Track 4), accepted as a poster.
Locally-Hosted Media Gallery App with AI Similarity Search
Add a description, image, and links to the vjepa topic page so that developers can more easily learn about it.
To associate your repository with the vjepa topic, visit your repo's landing page and select "manage topics."