Draft:JEPA
Draft article not currently submitted for review.
This is a draft Articles for creation (AfC) submission. It is not currently pending review. While there are no deadlines, abandoned drafts may be deleted after six months. To edit or make changes to this draft, simply click on the "Edit" tab at the top of the window. To be accepted, a draft should:
It is strongly discouraged to write about either yourself or your business or employer. If you do so, you must declare it. Where to get help
How to improve a draft
You can also browse Wikipedia:Featured articles and Wikipedia:Good articles to find examples of Wikipedia's best writing on topics similar to your proposed article. Improving your odds of a speedy review To improve your odds of a faster review, tag your draft with relevant WikiProject tags using the button below. This will let reviewers know a new draft has been submitted in their area of interest. For instance, if you wrote about a female astronomer, you would want to add the Biography, Astronomy, and Women scientists tags. Editor resources
Last edited by Colorado001 (talk | contribs) 5 days ago. (Update) |
JEPA (Joint-Embedding Predictive Architecture) is a type of self-supervised, non-generative, non-contrastive AI architecture which predicts embeddings. The concept was popularized by Yann LeCun.
Background
The term "Joint-Embedding Predictive Architecture" was coined by Yann LeCun in a 2022 "position paper".[1][2] LeCun's paper is built on past work involving world models, or models with internal representations of the physical world;[1]: 2 [3][4] LeCun claimed that in addition to representing the dynamics of the world, an autonomous intelligence must also have the capacity to understand and predict future actions.[5] He also argued systems like large language models are lacking in common-sense understanding of the real world due to their lack of physical interaction.[1]: 45 [6]: 37 [7]
Self-supervised learning

JEPA is also defined by how it differs from other forms of self-supervised learning (SSL). While supervised learning takes in labeled inputs and attempts to predict the label from the input , self-supervised systems generate their own "labels" from their inputs, which are then predicted. SSL models can be categorized as generative/reconstruction-based or contrastive/invariance-based, with other categories sometimes included.[8][9][10]
Generative models attempt to reconstruct or synthesize a signal based on input data .[9]: 3 [10]: 2 For example, masked autoencoders (MAEs) for images remove patches of pixels from an image and attempt to reconstruct the missing patches.[11][12]: 9 LeCun suggests that generative models, in attempting to predict missing information as closely as possible (e.g. the next frame of a video), waste computing power on unhelpful noise which cannot be predicted (e.g. the movement of individual leaves on a tree in the wind).[1]: 27 [13]: 1 [14]
Contrastive models – or models using contrastive learning – encode and with the goal of producing similar outputs for related inputs and dissimilar (contrasting) outputs for unrelated inputs.[10]: 2 [15]: 3 [16] Typically, the paired inputs are considered as different "views" of the same underlying concept; for example, photos of the same object from different angles.[10] The goal is to make the model learn representation invariance: regardless of the specific view of an object, the model should produce the same output.[17] However, contrastive methods require exponentially more contrasting samples as the dimensionality of the data increases (i.e. as the inputs contain more information).[1]: 23 Non-contrastive invariance-based methods, such as Bootstrap Your Own Latent (BYOL),[18] instead maximize the agreement between views[clarification needed] while preventing collapse through statistical constraints, i.e. regularization.[19]: 4 [20]: 5

Joint Embedding Architectures (JEAs), as described by LeCun, are a form of SSL with two encoders which are given different representations of the same input, whose output encodings (called "embeddings") are then compared.[21] They are commonly used with contrastive learning,[22]: 2 for example in Siamese networks.[23][1]: 42
Theory
JEPA is a form of self-supervised learning for representations of objects,[13]: 1 [15]: 2 meaning that instead of being provided labeled data to predict, the model uses information from co-occurring data (e.g. inputs which commonly appear together) to create its own labeled pairs.[24]: 2–3 [25] The goal is to generate semantic embeddings, which encode abstract or high-level information about concepts (e.g. the idea of a cow).[15]: 1 This is done by creating neural networks which, given two semantically similar inputs (e.g. a video of a cow and the audio of a cow), produce outputs that are "predictive" of each other.[26]: 1 [clarification needed]
Unlike autoencoders, JEPAs operate entirely in latent space, avoiding pixel-level noise to focus on semantic structure. Rather than (just) learning invariance, JEPAs learn by predicting masked latent representations from visible context.[27][clarification needed] This is also in contrast to LLMs, which operate in token space.[28]: 18
Energy-based model
JEPA can also be described as an energy-based model.[8]: 5–6 Specifically, it defines an energy function which quantifies the "compatibility" between context and target ; a lower energy means a higher compatibility. The goal is then to construct the function so the energy is minimized with correct pairs and maximized with incorrect pairs.[13]: 3 [29]: 16
Given two inputs and , along with optional conditioning or auxiliary information , the general form of JEPA's energy function is
where is the prediction loss function, is the predictor, is the output of the conditioning encoder, and are the representations of the inputs.[13]: 3 [15]: 7 The general training objective or overall loss function can be written as
where is a regularization function and is a parameter which controls the importance of the regularization term.[13]: 3 [contradictory]
Implementations
The first implementation of JEPA, named I-JEPA (Image-JEPA), was published in 2023.[30][31] I-JEPA is trained by taking an image and cutting out ("masking") multiple rectangles, called "target blocks". The rest of the image, called the "context block", is fed into the "context encoder" (a vision transformer) which generates a representation of the context block in latent space. The "context decoder" (another vision transformer) takes the context block's latent representation, along with the pixel location of a target block, and outputs a prediction of the latent-space embedding of the target block.[32]: 38 This process allows learning of invariant features without hand-crafting data augmentation processes.[33]: 9 [clarification needed]
JEPA was extended to video data with V-JEPA, published 2024.[34][35] This architecture enabled hierarchical modelling of temporal dependencies across multiple time scales.[36]: 2 [clarification needed]
JEPA has also been applied to image analysis,[37] audio processing,[38] and motion in images and video,[39] among other domains.[40]
Criticism
Jürgen Schmidhuber says that LeCun's work uses but fails to properly credit his past research,[41] specifically noting the similarities with his work on predictability maximization.[42]
Related pages
Notes
References
- ^ a b c d e f LeCun, Yann (2022-06-27). "A Path Towards Autonomous Machine Intelligence". Open Review.
- ^ "Yann LeCun on a vision to make AI systems learn and reason like animals and humans". Meta AI Blog. 2022-02-23.
- ^ Chen, Caiwei (2026-01-22). "Yann LeCun's new venture is a contrarian bet against large language models". MIT Technology Review. Retrieved 2026-05-25.
- ^ Ha, David; Schmidhuber, Jürgen (2018-03-28). "Recurrent World Models Facilitate Policy Evolution" (PDF). Advances in Neural Information Processing Systems 31. doi:10.5281/zenodo.1207631.
- ^ Ding, Jingtao; Zhang, Yunke; Shang, Yu; Zhang, Yuheng; Zong, Zefang; Feng, Jie; Yuan, Yuan; Su, Hongyuan; Li, Nian; Sukiennik, Nicholas; Xu, Fengli; Li, Yong (2025-09-09). "Understanding World or Predicting Future? A Comprehensive Survey of World Models". ACM Comput. Surv. 58 (3): 57:1–57:38. doi:10.1145/3746449. ISSN 0360-0300.
- ^ Hadi, Muhammad Usman; Tashi, Qasem Al; Qureshi, Rizwan; Shah, Abbas; Muneer, Amgad; Irfan, Muhammad; Zafar, Anas; Shaikh, Muhammad Bilal; Akhtar, Naveed; Hassan, Syed Zohaib; Shoman, Maged; Wu, Jia; Mirjalili, Seyedali; Shah, Mubarak. "Large Language Models: A Comprehensive Survey of its Applications, Challenges, Limitations, and Future Prospects". TechRxiv. 2025 (210). doi:10.36227/techrxiv.23589741.v8.
- ^ Kuka, Valeriia (2024-06-12). "What Is JEPA? Joint Embedding Predictive Architecture". Turing Post. Retrieved 2026-05-25.
- ^ a b Brotee, Shamyo; Chhetri, Gaurab; Polock, Sazzad Bin Bashar; Bellamkonda, Venkata Surya; Rafe, Amir; Das, Subasish (2025). "A Survey on Joint Embedding Predictive Architectures and World Models". SSRN 5772122.
- ^ a b Liu, Xiao; Zhang, Fanjin; Hou, Zhenyu; Mian, Li; Wang, Zhaoyu; Zhang, Jing; Tang, Jie (January 2023). "Self-Supervised Learning: Generative or Contrastive". IEEE Transactions on Knowledge and Data Engineering. 35 (1): 857–876. arXiv:2006.08218. Bibcode:2023ITKDE..35..857L. doi:10.1109/TKDE.2021.3090866. ISSN 1558-2191.
- ^ a b c d Skenderi, Geri; Li, Hang; Tang, Jiliang; Cristani, Marco (2024). "Graph-level Representation Learning with Joint-Embedding Predictive Architectures". Open Review.
- ^ He, Kaiming; Chen, Xinlei; Xie, Saining; Li, Yanghao; Dollár, Piotr; Girshick, Ross (2021-12-19). "Masked Autoencoders Are Scalable Vision Learners". arXiv:2111.06377 [cs.CV].
- ^ Zhang, Jingyi; Huang, Jiaxing; Jin, Sheng; Lu, Shijian (2024-02-16). "Vision-Language Models for Vision Tasks: A Survey". IEEE Transactions on Pattern Analysis and Machine Intelligence. 46 (8): 5625–5644. arXiv:2304.00685. Bibcode:2024ITPAM..46.5625Z. doi:10.1109/TPAMI.2024.3369699. PMID 38408000.
- ^ a b c d e Terver, Basile; Balestriero, Randall; Dervishi, Megi; Fan, David; Garrido, Quentin; Nagarajan, Tushar; Sinha, Koustuv; Zhang, Wancong; Rabbat, Mike (2026-04-08). "A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures". arXiv:2602.03604 [cs.CV].
- ^ Bandaru, Rohit. "Rohit Bandaru | Deep Dive into Yann LeCun's JEPA". Rohit Bandaru. Retrieved 2026-05-25.
- ^ a b c d Monemi, Mehdi; Chinipardaz, Maryam; Rasti, Mehdi; Bennis, Mehdi; Latva-aho, Matti. "Tutorial on Joint Embedding Predictive Architectures (JEPA): Foundations, Applications, and Future Directions". TechRxiv. 2025 (1210). doi:10.36227/techrxiv.176469421.19270944/v2.
- ^ Xiao, Tete; Wang, Xiaolong; Efros, Alexei A.; Darrell, Trevor (2020). "What Should Not Be Contrastive in Contrastive Learning". arXiv:2008.05659 [cs.CV].
- ^ Cite error: The named reference
viewswas invoked but never defined (see the help page). - ^ Grill, Jean-Bastien; Strub, Florian; Altché, Florent; Tallec, Corentin; Richemond, Pierre H.; Buchatskaya, Elena; Doersch, Carl; Pires, Bernardo Avila; Guo, Zhaohan Daniel (2020-09-10). "Bootstrap your own latent: A new approach to self-supervised Learning". arXiv:2006.07733 [cs.LG].
- ^ Shwartz Ziv, Ravid; LeCun, Yann (12 March 2024). "To Compress or Not to Compress—Self-Supervised Learning and Information Theory: A Review". Entropy. 26 (3): 252. arXiv:2304.09355. Bibcode:2024Entrp..26..252S. doi:10.3390/e26030252. ISSN 1099-4300. PMC 10968883. PMID 38539763.
- ^ Gui, Jie; Chen, Tuo; Zhang, Jing; Cao, Qiong; Sun, Zhenan; Luo, Hao; Tao, Dacheng (2024-07-14). "A Survey on Self-supervised Learning: Algorithms, Applications, and Future Trends". IEEE Transactions on Pattern Analysis and Machine Intelligence. 46 (12): 9052–9071. arXiv:2301.05712. Bibcode:2024ITPAM..46.9052G. doi:10.1109/TPAMI.2024.3415112. PMID 38885108.
- ^ Bardes, Adrien; Ponce, Jean; LeCun, Yann (2021). "VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning". arXiv:2105.04906 [cs.CV].
- ^ Saito, Ayumu; Kudeshia, Prachi; Poovvancheri, Jiju (2024). "Point-JEPA: A Joint Embedding Predictive Architecture for Self-Supervised Learning on Point Cloud". arXiv:2404.16432 [cs.CV].
- ^ Assran, Mahmoud; Caron, Mathilde; Misra, Ishan; Bojanowski, Piotr; Bordes, Florian; Vincent, Pascal; Joulin, Armand; Rabbat, Michael; Ballas, Nicolas (2022). "Masked Siamese Networks for Label-Efficient Learning". arXiv:2204.07141 [cs.LG].
- ^ Gui, Jie; Chen, Tuo; Zhang, Jing; Cao, Qiong; Sun, Zhenan; Luo, Hao; Tao, Dacheng (2024-07-14). "A Survey on Self-supervised Learning: Algorithms, Applications, and Future Trends". IEEE Transactions on Pattern Analysis and Machine Intelligence. 46 (12): 9052–9071. arXiv:2301.05712. Bibcode:2024ITPAM..46.9052G. doi:10.1109/TPAMI.2024.3415112. PMID 38885108.
- ^ Rani, Veenu; Nabi, Syed Tufael; Kumar, Munish; Mittal, Ajay; Kumar, Krishan (2023-05-01). "Self-supervised Learning: A Succinct Review". Archives of Computational Methods in Engineering. 30 (4): 2761–2775. doi:10.1007/s11831-023-09884-2. ISSN 1886-1784. PMC 9857922. PMID 36713767.
- ^ Littwin, Etai; Saremi, Omid; Advani, Madhu; Thilak, Vimal; Nakkiran, Preetum; Huang, Chen; Susskind, Joshua (2024-07-03). "How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation Networks". arXiv:2407.03475 [cs.LG].
- ^ Assran, Mahmoud; Duval, Quentin; Misra, Ishan; Bojanowski, Piotr; Vincent, Pascal; Rabbat, Michael; LeCun, Yann; Ballas, Nicolas (2023-01-19). "Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture". arXiv:2301.08243 [cs.CV].
- ^ Ricciardi Celsi, Lorenzo; McCann, James (2026-02-26). "Beyond Next-Token Prediction: A Standards-Aligned Survey of Autoregressive LLM Failure Modes, Deployment Patterns, and the Potential Role of World Models". Electronics. 15 (5): 966. doi:10.3390/electronics15050966. ISSN 2079-9292.
- ^ Ricciardi Celsi, Lorenzo; McCann, James (2026-02-26). "Beyond Next-Token Prediction: A Standards-Aligned Survey of Autoregressive LLM Failure Modes, Deployment Patterns, and the Potential Role of World Models". Electronics. 15 (5): 966. doi:10.3390/electronics15050966. ISSN 2079-9292.
- ^ "The first AI model based on Yann LeCun's vision for more human-like AI". Meta AI Blog. 2023-06-13. Archived from the original on 2026-05-03. Retrieved 2026-05-26.
- ^ Assran, Mahmoud; Duval, Quentin; Misra, Ishan; Bojanowski, Piotr; Vincent, Pascal; Rabbat, Michael; LeCun, Yann; Ballas, Nicolas (2023-01-19). "Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture". Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR): 15619–15629. arXiv:2301.08243. Retrieved 2026-05-26.
- ^ Konstantakos, Sotirios; Cani, Jorgen; Mademlis, Ioannis; Chalkiadaki, Despina Ioanna; Asano, Yuki M.; Gavves, Efstratios; Papadopoulos, Georgios Th. (2025-03-01). "Self-supervised visual learning in the low-data regime: A comparative evaluation". Neurocomputing. 620 129199. doi:10.1016/j.neucom.2024.129199. ISSN 0925-2312.
- ^ Khan, Asifullah; Sohail, Anabia; Fiaz, Mustansar; Hassan, Mehdi; Afridi, Tariq Habib; Marwat, Sibghat Ullah; Munir, Farzeen; Ali, Safdar; Naseem, Hannan (2025-08-25). "A Survey of the Self Supervised Learning Mechanisms for Vision Transformers". arXiv:2408.17059 [cs.CV].
- ^ "V-JEPA: The next step toward Yann LeCun's vision of advanced machine intelligence (AMI)". Meta AI Blog. 2024-02-15.
- ^ Bardes, Adrien; Garrido, Quentin; Ponce, Jean; Chen, Xinlei; Rabbat, Michael; LeCun, Yann; Assran, Mahmoud; Ballas, Nicolas (2024-02-15), Revisiting Feature Prediction for Learning Visual Representations from Video, arXiv:2404.08471
- ^ Xie, Ningwei; Tian, Zizi; Yang, Lei; Zhang, Xiao-Ping; Guo, Meng; Li, Jie (2025-06-25). "From 2D to 3D Cognition: A Brief Survey of General World Models". arXiv:2506.20134 [cs.CV].
- ^ Garrido, Quentin; Assran, Mahmoud; Ballas, Nicolas; Bardes, Adrien; Najman, Laurent; LeCun, Yann (2024). "Learning and Leveraging World Models in Visual Representation Learning". arXiv:2403.00504 [cs.CV].
- ^ Fei, Zhengcong; Fan, Mingyuan; Huang, Junshi (2023). "A-JEPA: Joint-Embedding Predictive Architecture Can Listen". arXiv:2311.15830 [cs.SD].
- ^ Bardes, Adrien; Ponce, Jean; LeCun, Yann (2023). "MC-JEPA: A Joint-Embedding Predictive Architecture for Self-Supervised Learning of Motion and Content Features". arXiv:2307.12698 [cs.CV].
- ^ Zhang, Peng-Fei; Cheng, Ying; Sun, Xiaofan; Wang, Shijie; Li, Fengling; Zhu, Lei; Shen, Heng Tao (2025-10-31). "A Step Toward World Models: A Survey on Robotic Manipulation". arXiv:2511.02097 [cs.RO].
- ^ Bordoloi, Pritam (2022-08-06). "Why is LeCun's 2022 paper on Autonomous Machine Intelligence controversial?". Analytics India Magazine. Retrieved 2026-05-25.
- ^ Schmidhuber, Jürgen (2026-03-31). "Who invented "JEPA"?". Jürgen Schmidhuber's AI Blog.
Content Disclaimer
Informasi ini disarikan dari Wikipedia dan disajikan kembali untuk tujuan edukasi. Konten tersedia di bawah lisensi CC BY-SA 3.0. Kami tidak bertanggung jawab atas ketidakakuratan data yang bersumber dari kontribusi publik tersebut.
- The information displayed on this website is sourced in part or in whole from Wikipedia and has been adapted for the purpose of restating it. We strive to provide accurate and relevant information, however:
- There is no guarantee of absolute accuracy. Wikipedia is an open, collaborative project that can be edited by anyone, so information is subject to change.
- It is not intended to constitute professional advice. The content displayed is for informational and educational purposes only. For important decisions (e.g., medical, legal, or financial), please consult a professional.
- Content copyright. Wikipedia is licensed under the Creative Commons Attribution-ShareAlike License (CC BY-SA). This means that content may be reused with appropriate attribution and shared under a similar license.
- Responsible use. Any risk arising from the use of information from this website is entirely the responsibility of the user.
