2025-06-11 05:50:50 -07:00
2025-06-11 05:50:50 -07:00
2025-06-11 05:50:50 -07:00
2025-06-13 13:54:27 +02:00
2025-06-13 13:54:27 +02:00
2025-06-11 05:50:50 -07:00
2025-06-11 05:50:50 -07:00
2025-06-11 05:50:50 -07:00
2025-06-11 05:50:50 -07:00
2025-06-11 05:50:50 -07:00
2025-06-11 05:50:50 -07:00
2025-06-11 05:50:50 -07:00
2025-06-11 05:50:50 -07:00
2025-06-11 05:50:50 -07:00
2025-06-11 05:50:50 -07:00
2026-02-26 19:41:19 -08:00
2025-06-11 05:50:50 -07:00
2025-06-11 05:50:50 -07:00
2025-06-11 05:50:50 -07:00

VJEPA2

Self-supervised visual representation learning from video. Part of the Zen LM ecosystem.

License

Overview

VJEPA2 implements Video Joint-Embedding Predictive Architecture for learning visual representations from unlabeled video data without relying on hand-crafted augmentations.

Features

  • Self-supervised learning from video
  • No hand-crafted augmentations required
  • Pre-trained visual encoder for downstream tasks
  • Efficient training with masking strategies
  • jin — Multimodal understanding framework
  • Zen LM — Full model family

License

See LICENSE file.

Part of the Zen LM ecosystem by Hanzo AI

S
Description
PyTorch code and models for VJEPA2 self-supervised learning from video.
Readme MIT
7.7 MiB
Languages
Python 95.5%
Jupyter Notebook 4.5%