#DINOv2
1/ Can open-data models beat DINOv2? Today we release Franca, a fully open-sourced vision foundation model. Franca with ViT-G backbone matches (and often beats) proprietary models like SigLIPv2, CLIP, DINOv2 on various benchmarks setting a new standard for open-source research.
July 21, 2025 at 2:47 PM
Want strong SSL, but not the complexity of DINOv2?

CAPI: Cluster and Predict Latents Patches for Improved Masked Image Modeling.
February 14, 2025 at 6:05 PM
Are you using DINOv2 for tasks that require semantic features? DIY-SC might be the alternative!
It refines DINOv2 or SD+DINOv2 features and achieves a new SOTA on the semantic correspondence dataset SPair-71k when not relying on annotated keypoints! [1/6]
genintel.github.io/DIY-SC
June 26, 2025 at 12:56 PM
One weird trick for better diffusion models: concatenate some DINOv2 features to your latent channels!

Combining latents with PCA components extracted from DINOv2 features yields faster training and better samples. Also enables a new guidance strategy. Simple and effective!
1/n Introducing ReDi (Representation Diffusion): a new generative approach that leverages a diffusion model to jointly capture
– Low-level image details (via VAE latents)
– High-level semantic features (via DINOv2)🧵
April 25, 2025 at 1:03 PM
DINOv3.
Trained with hi-res refinement and GRAM-matrix supervision on gigantic dataset.

Time to retrain all dinov2 based models :)
Oriane Siméoni et 25 al.
ai.meta.com/research/pub...
August 14, 2025 at 6:09 PM
Loft🆙 Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models. We achieve SotA upsampling results for DINOv2. Paper and code:
andrehuang.github.io/loftup-site/
April 26, 2025 at 2:47 PM
(1/3) Happy to share LUDVIG: Learning-free Uplifting of 2D Visual features to Gaussian Splatting scenes, that uplifts visual features from models such as DINOv2 (left) & CLIP (mid) to 3DGS scenes. Joint work w. @dlarlus.bsky.social @jmairal.bsky.social
Webpage & code: juliettemarrie.github.io/ludvig
January 31, 2025 at 9:59 AM
Depth Anything fine-tunes DINOv2 for monocular depth and enforce a threshold on the cosine similarity between the frozen DINOv2 encoder and their fine-tuned encoder (initialized as DINOv2).
November 22, 2024 at 4:10 PM
🕳️🐇 𝙄𝙣𝙩𝙤 𝙩𝙝𝙚 𝙍𝙖𝙗𝙗𝙞𝙩 𝙃𝙪𝙡𝙡 – 𝙋𝙖𝙧𝙩 𝙄 (𝑃𝑎𝑟𝑡 𝐼𝐼 𝑡𝑜𝑚𝑜𝑟𝑟𝑜𝑤)

𝗔𝗻 𝗶𝗻𝘁𝗲𝗿𝗽𝗿𝗲𝘁𝗮𝗯𝗶𝗹𝗶𝘁𝘆 𝗱𝗲𝗲𝗽 𝗱𝗶𝘃𝗲 𝗶𝗻𝘁𝗼 𝗗𝗜𝗡𝗢𝘃𝟮, one of vision’s most important foundation models.

And today is Part I, buckle up, we're exploring some of its most charming features. :)
October 14, 2025 at 9:00 PM
🐇Into the Rabbit Hull — Part 1: A Deep Dive into DINOv2🧠
Our latest Deeper Learning blog post is an #interpretability deep dive into one of today’s leading vision foundation models: DINOv2.
📖Read now: bit.ly/4nNfq8D
Stay tuned — Part 2 coming soon.
#AI #VLMs #DINOv2
Into the Rabbit Hull – Part I - Kempner Institute
This blog post offers an interpretability deep dive, examining the most important concepts emerging in one of today’s central vision foundation models, DINOv2. This blogpost is the first of a […]
bit.ly
November 12, 2025 at 3:49 PM
🕳️🐇Into the Rabbit Hull – Part II

Continuing our interpretation of DINOv2, the second part of our study concerns the *geometry of concepts* and the synthesis of our findings toward a new representational *phenomenology*:

the Minkowski Representation Hypothesis
October 15, 2025 at 5:17 PM
Our Frechet Wavelet Distance paper just got accepted for ICLR … openreview.net/forum?id=Qin...

Let’s bring back the old knowledge and see what it can do today!
February 19, 2025 at 8:17 PM
Maybe this won’t come as a surprise but this tech is built, in part, on models made by OpenAI (CLIP) and Meta (DINOv2). So there’s that, too.
August 13, 2025 at 7:11 PM
Anyone any practical experience with DINOv2 derivatives or alternatives for 3D vision tasks? Monocular only :)
November 22, 2024 at 8:28 AM
Want stronger Vision Transformers? Use octic-equivariant layers (arxiv.org/abs/2505.15441).

TLDR; We extend @bokmangeorg.bsky.social's reflection-equivariant ViTs to the (octic) group of 90-degree rotations and reflections and... it just works... (DINOv2+DeiT)

Code: github.com/davnords/octic-vits
May 23, 2025 at 7:38 AM
1/n Introducing ReDi (Representation Diffusion): a new generative approach that leverages a diffusion model to jointly capture
– Low-level image details (via VAE latents)
– High-level semantic features (via DINOv2)🧵
April 25, 2025 at 7:23 AM
🚨New doctor in the house!🚨
Congrats to @timdarcet.bsky.social for his tremendous work (DINOv2, registers, CAPI) & successful PhD defense followed by ~2 hrs of questions -- he's got stamina!
Congrats to his incredible team of advisors from Inria & Meta: @jmairal.bsky.social, P. Bojanowski, M. Oquab
July 2, 2025 at 10:32 PM
Multimodal Autoregressive Pre-training of Large Vision Encoders
Enrico Fini et 15 al

tl;dr: in title. Scaling laws and ablations.
they claim to be better than SigLIP and DINOv2 for semantic tasks. I would be interested in monodepth performance though.

arxiv.org/abs/2411.14402
November 22, 2024 at 8:07 AM
DUNE 🏜️ - multi-teacher distillation extended to heterogeneous teachers: DiNOv2, Multi-HMR and MASt3R - #CVPR2025
Paper: arxiv.org/abs/2503.14405
Project page: europe.naverlabs.com/dune
June 8, 2025 at 7:21 PM
Happy to see FiT3D being used as an alternative feature extractor to DINOv2 for motion mask extraction🤓, which goes beyond the tasks we originally considered.
November 27, 2024 at 9:21 PM
1/n 🚀New paper out - accepted at #ICCV2025!

Introducing DIP: unsupervised post-training that enhances dense features in pretrained ViTs for dense in-context scene understanding

Below: Low-shot in-context semantic segmentation examples. DIP features outperform DINOv2!
June 25, 2025 at 7:21 PM
It would actually take closer to 550 years to reproduce the DINOv2 paper if your laptop has an A100 gpu.
February 17, 2025 at 1:34 PM
(3/3) LUDVIG uses a graph diffusion mechanism to refine 3D features, such as coarse segmentation masks, by leveraging 3D scene geometry and pairwise similarities induced by DINOv2.
January 31, 2025 at 9:59 AM
Today, we release Franca, a new vision Foundation Model that matches and often outperforms DINOv2.
The data, the training code and the model weights are open-source.

This is the result of a close and fun collaboration
@valeoai.bsky.social (in France) and @funailab.bsky.social (in Franconia)🚀
1/ Can open-data models beat DINOv2? Today we release Franca, a fully open-sourced vision foundation model. Franca with ViT-G backbone matches (and often beats) proprietary models like SigLIPv2, CLIP, DINOv2 on various benchmarks setting a new standard for open-source research.
July 21, 2025 at 2:58 PM
recent code releases from our group:

Sadra shared the code for our BMVC paper, StereoGS: github.com/sadrasafa/St...

Merve shared the code for robust BEV with DINOv2: github.com/mrabiabrn/ro...

more on the way!
December 7, 2024 at 3:52 PM