I-JEPA: Yannn LeCun's First 'World Model' for Computer Vision

Tutorials

I-JEPA: Yannn LeCun's First 'World Model' for Computer Vision

Yann LeCun has been arguing for years that generative pretraining is the wrong bet for building machines that understand the world.

Yann LeCun has been arguing for years that generative pretraining is the wrong bet for building machines that understand the world. I-JEPA is the first concrete Meta AI release to put that thesis into code for computer vision. Instead of reconstructing pixels or aligning augmentations, it predicts abstract representations of masked target blocks from a single context block โ€” learning semantics without ever generating an image.

The walkthrough digs into why the masking strategy matters (target blocks have to be large enough to force semantic prediction, context blocks have to be spatially informative) and why the joint-embedding predictive architecture sidesteps the representation collapse that haunts contrastive methods. The efficiency numbers are the headline: a ViT-Huge/14 trained on ImageNet in under 72 hours on 16 A100s, hitting strong low-shot classification with just 12 labeled examples per class. For anyone tracking self-supervised learning beyond MAE and DINO โ€” or curious about what a LeCun-style world model actually looks like when it ships โ€” this is a useful primer on where representation learning may be heading next.