| Mask R-CNN |
Mask R-CNN |
Vision |
2017 |
|
32,927 |
4355 |
| RoBERTa: A Robustly Optimized BERT Pretraining Approach |
RoBERTa |
Language & LLMs |
2019 |
arXiv.org |
31,147 |
6068 |
| LLaMA: Open and Efficient Foundation Language Models |
LLaMA |
Language & LLMs |
2023 |
arXiv.org |
21,436 |
2209 |
| The Llama 3 Herd of Models |
Llama 3 / 3.1 / 3.2 / 3.3 |
Language & LLMs |
2024 |
|
18,281 |
3374 |
| Llama 2: Open Foundation and Fine-Tuned Chat Models |
Llama 2 |
Language & LLMs |
2023 |
arXiv.org |
18,110 |
2260 |
| Segment Anything |
Segment Anything |
Vision |
2023 |
IEEE International Conference on C |
15,338 |
2017 |
| BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension |
BART |
Language & LLMs |
2019 |
Annual Meeting of the Association |
13,175 |
2394 |
| Enriching Word Vectors with Subword Information |
fastText |
Open Source Infra |
2016 |
Transactions of the Association fo |
10,941 |
1070 |
| DINOv2: Learning Robust Visual Features without Supervision |
DINOv2 |
Vision |
2023 |
Trans. Mach. Learn. Res. |
10,180 |
1470 |
| Emerging Properties in Self-Supervised Vision Transformers |
DINO |
Vision |
2021 |
IEEE International Conference on C |
10,104 |
1700 |
| wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations |
wav2vec 2.0 |
Speech & Sound |
2020 |
Neural Information Processing Syst |
9,159 |
3191 |
| Unsupervised Cross-lingual Representation Learning at Scale |
XLM-R |
Language & LLMs |
2019 |
Annual Meeting of the Association |
9,113 |
1824 |
| OPT: Open Pre-trained Transformer Language Models |
OPT-175B |
Language & LLMs |
2022 |
arXiv.org |
4,933 |
533 |
| SAM 2: Segment Anything in Images and Videos |
SAM 2 |
Vision |
2024 |
International Conference on Learni |
4,075 |
591 |
| Code Llama: Open Foundation Models for Code |
Code Llama |
Language & LLMs |
2023 |
arXiv.org |
3,531 |
394 |
| fairseq: A Fast, Extensible Toolkit for Sequence Modeling |
fairseq |
Language & LLMs |
2019 |
North American Chapter of the Asso |
3,419 |
286 |
| Make-A-Video: Text-to-Video Generation without Text-Video Data |
Make-A-Video |
Generative Media |
2022 |
International Conference on Learni |
2,155 |
110 |
| Ego4D: Around the World in 3,000 Hours of Egocentric Video |
Ego4D |
World Models & Embodied AI |
2021 |
Computer Vision and Pattern Recogn |
2,004 |
286 |
| No Language Left Behind: Scaling Human-Centered Machine Translation |
No Language Left Behind |
Speech & Sound |
2022 |
arXiv.org |
1,968 |
174 |
| ImageBind One Embedding Space to Bind Them All |
ImageBind |
Vision |
2023 |
Computer Vision and Pattern Recogn |
1,754 |
226 |
| VGGT: Visual Geometry Grounded Transformer |
VGGT |
Vision |
2025 |
Computer Vision and Pattern Recogn |
1,680 |
447 |
| DINOv3 |
DINOv3 |
Vision |
2025 |
|
1,375 |
240 |
| Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations |
Purple Llama |
Safety & Trust |
2023 |
arXiv.org |
1,286 |
205 |
| Accelerating 3D deep learning with PyTorch3D |
PyTorch3D |
Vision |
2019 |
SIGGRAPH Asia 2020 Courses |
1,189 |
90 |
| Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture |
I-JEPA |
World Models & Embodied AI |
2023 |
Computer Vision and Pattern Recogn |
1,175 |
143 |
| GAIA: a benchmark for General AI Assistants |
GAIA Benchmark |
Safety & Trust |
2023 |
arXiv.org |
1,164 |
199 |
| Galactica: A Large Language Model for Science |
Galactica |
Language & LLMs |
2022 |
arXiv.org |
1,088 |
97 |
| Chameleon: Mixed-Modal Early-Fusion Foundation Models |
Chameleon |
Vision |
2024 |
arXiv.org |
956 |
111 |
| The Open Catalyst 2020 (OC20) Dataset and Community Challenges |
Open Catalyst Project |
AI for Science |
2020 |
ACS Catalysis |
870 |
90 |
| The Faiss Library |
faiss |
Open Source Infra |
2024 |
IEEE Transactions on Big Data |
861 |
53 |
| SAM 3: Segment Anything with Concepts |
SAM 3 |
Vision |
2025 |
arXiv.org |
851 |
129 |
| Simple and Controllable Music Generation |
AudioCraft |
Speech & Sound |
2023 |
Neural Information Processing Syst |
787 |
107 |
| Scaling Speech Technology to 1, 000+ Languages |
Massively Multilingual Speech |
Speech & Sound |
2023 |
Journal of machine learning resear |
727 |
123 |
| Training Large Language Models to Reason in a Continuous Latent Space |
Coconut |
Language & LLMs |
2024 |
arXiv.org |
705 |
140 |
| V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning |
V-JEPA 2 |
World Models & Embodied AI |
2025 |
arXiv.org |
647 |
96 |
| Movie Gen: A Cast of Media Foundation Models |
Movie Gen |
Generative Media |
2024 |
|
632 |
50 |
| Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale |
Voicebox |
Speech & Sound |
2023 |
Neural Information Processing Syst |
577 |
69 |
| Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives |
Ego-Exo4D |
World Models & Embodied AI |
2023 |
Computer Vision and Pattern Recogn |
573 |
76 |
| ParlAI: A Dialog Research Software Platform |
ParlAI |
Open Source Infra |
2017 |
Conference on Empirical Methods in |
405 |
40 |
| Better & Faster Large Language Models via Multi-token Prediction |
Multi-token prediction |
Language & LLMs |
2024 |
International Conference on Machin |
380 |
44 |
| Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots |
Habitat |
World Models & Embodied AI |
2023 |
International Conference on Learni |
323 |
85 |
| Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack |
Emu |
Generative Media |
2023 |
arXiv.org |
315 |
12 |
| MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases |
MobileLLM |
On-Device & Silicon |
2024 |
International Conference on Machin |
306 |
20 |
| BlenderBot 3: a deployed conversational agent that continually learns to responsibly engage |
BlenderBot |
Language & LLMs |
2022 |
arXiv.org |
300 |
23 |
| Decoding speech perception from non-invasive brain recordings |
Brain & AI: Decoding Perception from MEG/EEG |
AI for Science |
2022 |
Nature Machine Intelligence |
298 |
38 |
| Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning |
Emu Video & Emu Edit |
Generative Media |
2023 |
European Conference on Computer Vi |
296 |
12 |
| Sapiens: Foundation for Human Vision Models |
Sapiens |
Vision |
2024 |
European Conference on Computer Vi |
284 |
44 |
| UMA: A Family of Universal Models for Atoms |
UMA — Universal Models for Atoms |
AI for Science |
2025 |
Neural Information Processing Syst |
234 |
37 |
| SAM 3D: 3Dfy Anything in Images |
SAM 3D (Objects + Body) |
Vision |
2025 |
arXiv.org |
221 |
50 |
| SeamlessM4T: Massively Multilingual&Multimodal Machine Translation |
SeamlessM4T & the Seamless family |
Speech & Sound |
2023 |
|
220 |
29 |
| Combining Deep Reinforcement Learning and Search for Imperfect-Information Games |
ReBeL |
Games & Strategy |
2020 |
Neural Information Processing Syst |
200 |
16 |
| Audiobox: Unified Audio Generation with Natural Language Prompts |
Audiobox |
Speech & Sound |
2023 |
arXiv.org |
184 |
11 |
| Byte Latent Transformer: Patches Scale Better Than Tokens |
Byte Latent Transformer |
Language & LLMs |
2024 |
arXiv.org |
160 |
21 |
| Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models |
Open Materials 2024 |
AI for Science |
2024 |
arXiv.org |
157 |
12 |
| SpiRit-LM: Interleaved Spoken and Written Language Model |
Spirit LM |
Speech & Sound |
2024 |
Transactions of the Association fo |
156 |
20 |
| ELF OpenGo: An Analysis and Open Reimplementation of AlphaZero |
ELF OpenGo |
Games & Strategy |
2019 |
International Conference on Machin |
122 |
13 |
| Large Concept Models: Language Modeling in a Sentence Representation Space |
Large Concept Models |
Language & LLMs |
2024 |
arXiv.org |
117 |
11 |
| The Open Molecules 2025 (OMol25) Dataset, Evaluations, and Models |
Open Molecules 2025 |
AI for Science |
2025 |
|
100 |
9 |
| From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations |
audio2photoreal |
Reality Labs Research |
2024 |
Computer Vision and Pattern Recogn |
99 |
14 |
| PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks |
PARTNR |
World Models & Embodied AI |
2024 |
arXiv.org |
83 |
14 |
| CWM: An Open-Weights LLM for Research on Code Generation with World Models |
Code World Model |
Language & LLMs |
2025 |
arXiv.org |
78 |
12 |
| Meta Large Language Model Compiler: Foundation Models of Compiler Optimization |
Meta LLM Compiler |
On-Device & Silicon |
2024 |
arXiv.org |
71 |
2 |
| VL-JEPA: Joint Embedding Predictive Architecture for Vision-language |
VL-JEPA |
World Models & Embodied AI |
2025 |
arXiv.org |
54 |
3 |
| Video Seal: Open and Efficient Video Watermarking |
Video Seal |
Vision |
2024 |
arXiv.org |
52 |
12 |
| emg2pose: A Large and Diverse Benchmark for Surface Electromyographic Hand Pose Estimation |
emg2pose |
Reality Labs Research |
2024 |
Neural Information Processing Syst |
43 |
4 |
| emg2qwerty: A Large Dataset with Baselines for Touch Typing using Surface Electromyography |
emg2qwerty |
Reality Labs Research |
2024 |
Neural Information Processing Syst |
39 |
7 |
| SAM Audio: Segment Anything in Audio |
SAM Audio |
Speech & Sound |
2025 |
arXiv.org |
38 |
3 |
| Memory Layers at Scale |
Memory Layers at Scale |
Language & LLMs |
2024 |
arXiv.org |
35 |
2 |
| Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages |
Omnilingual ASR |
Speech & Sound |
2025 |
arXiv.org |
35 |
5 |
| Brain-to-Text Decoding: A Non-invasive Approach via Typing |
Brain2Qwerty |
AI for Science |
2025 |
arXiv.org |
33 |
3 |
| MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes |
MobileLLM-R1 / R1.5 |
On-Device & Silicon |
2025 |
arXiv.org |
9 |
0 |
| MobileLLM-Pro Technical Report |
MobileLLM-Pro |
On-Device & Silicon |
2025 |
arXiv.org |
4 |
1 |
| Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining |
Codec Avatars |
Reality Labs Research |
2026 |
arXiv.org |
3 |
0 |
| MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale Deployment |
MobileLLM-Flash |
On-Device & Silicon |
2026 |
Proceedings of the 64th Annual Mee |
2 |
0 |