#SigLIP
needed a better way to traverse my tens of thousands of reference images so testing out SigLIP-based semantic image search
September 28, 2025 at 6:31 AM
pitted google's SigLIP with apple's MobileCLIP and the result are:
- if you prefer searching with danbooru-style tags, go with SigLIP
- if you prefer english sentences go with MobileCLIP

on my ryzen 7 7800x3d cpu, MobileCLIP is faster than SigLIP by a 3-4 seconds
needed a better way to traverse my tens of thousands of reference images so testing out SigLIP-based semantic image search
October 2, 2025 at 12:34 AM
Our big_vision codebase is really good! And it's *the* reference for ViT, SigLIP, PaliGemma, JetFormer, ... including fine-tuning them.

However, it's criminally undocumented. I tried using it outside Google to fine-tune PaliGemma and SigLIP on GPUs, and wrote a tutorial: lb.eyer.be/a/bv_tuto.html
December 3, 2024 at 12:18 AM
Enjoyed discussing how we used GPT and SigLIP to design search and access on digitaldocumerica.org @adho-org.bsky.social #DH2025
July 16, 2025 at 1:45 PM
Best multilingual SigLIP ever is now compatible with transformers huggingface.co/google/sigli...
google/siglip-so400m-patch16-256-i18n · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
November 18, 2024 at 8:37 AM
Nvidia's RADIOv2.5 = DFN_CLIP + DINOv2 + SAM + SigLIP + ToMe + multi-res training + teacher loss balancing + smart augmentations

RADIO is one encoder, one pass. Better features than DFN-CLIP, DINO, SAM, and SigLIP - all at once. Like a Swiss army knife for vision tasks.
April 5, 2025 at 5:21 AM
Siglip needs registers

Siglip has no global CLS token. It has no extra registers. It only has the spatial positions.
March 3, 2025 at 3:22 AM
Replace Variational Autoencoder (VAE) with pretrained representation encoders (e.g., DINO, SigLIP, MAE) paired with trained decoders, which they terms as Representation Autoencoders (RAE).
October 15, 2025 at 3:49 AM
Yay someone using SigLIP. Info on SigLIP huggingface.co/docs/transfo... #chr2024
SigLIP
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
December 5, 2024 at 3:53 PM
ProLip: A probabilistic trained SigLIP

ProLIP is the first from-scratch trained probabilistic vision-language model, which is comparable with CLIP or SigLIP

Paper: Probabilistic Language-Image Pre-Training ( arxiv.org/abs/2410.18857 )
Models: huggingface.co/collections/...
ProLIP - a SanghyukChun Collection
Official ProLIP weights
huggingface.co
January 24, 2025 at 7:39 AM
This CLIP doesn’t seem to work either…

Maybe SigLIP could help @giffmana.ai?
November 30, 2024 at 3:01 PM
Multimodal Autoregressive Pre-training of Large Vision Encoders
Enrico Fini et 15 al

tl;dr: in title. Scaling laws and ablations.
they claim to be better than SigLIP and DINOv2 for semantic tasks. I would be interested in monodepth performance though.

arxiv.org/abs/2411.14402
November 22, 2024 at 8:07 AM
GPU really fighting for its life over here. Really should have sprung for more VRAM....

Spending a couple hours embedding 25k Fluoddity screenshots with a SigLIP 400M model
August 13, 2026 at 7:01 PM
the BigVision repo is my current reference impl for gemma and ViT. such an underrated repo @giffmana.bsky.social and team are doing the lord's work

github.com/google-resea...

github.com/google-resea...
big_vision/big_vision/models/ppp/gemma.py at main · google-research/big_vision
Official codebase used to develop Vision Transformer, SigLIP, MLP-Mixer, LiT and more. - google-research/big_vision
github.com
November 24, 2024 at 5:25 PM
Yay, DINOv3 is out!

SigLIP (VLMs) and DINO are two competing paradigms for image encoders.

My intuition is that joint vision-language modeling works great for semantic problems but may be too coarse for geometry problems like SfM or SLAM.

Most animals navigate 3D space perfectly without language.
August 14, 2025 at 5:59 PM
ILIAS: Instance-Level Image retrieval At Scale

@gkordo.bsky.social, Vladan Stojnić @annetka.bsky.social Pavel Šuma, Nikolaos-Antonios Ypsilantis @nikos-efth.bsky.social Zakaria Laskar,Jiří Matas, Ondřej Chum, @gtolias.bsky.social

tl;dr: SigLIP rules. Lots of ablations
arxiv.org/abs/2502.11748
1/
February 24, 2025 at 10:03 AM
OmniVision-968M: a new local VLM for edge devices, fast & small but performant 👏

it's based on SigLIP-so-400M and Qwen-2.5-0.5B
💨 9x less image tokens, super efficient
📖 aligned with SFT and DPO for reducing hallucinations
🔥 Apache 2.0 license
Demo hf.co/spaces/NexaAIDev/omnivlm-dpo-demo
November 16, 2024 at 11:37 PM
OpenCity3D: What do Vision-Language Models know about Urban Environments?

Valentin Bieri, Marco Zamboni, Nicolas S. Blumer, Qingxuan Chen, Francis Engelmann
tl;dr: if you have aerial 3D reconstruction, use SigLIP to be happy.
arxiv.org/abs/2503.16776
March 24, 2025 at 10:21 AM
Leading computer vision researchers Lucas Beyer (@giffmana.ai), Alexander Kolesnikov (@kolesnikov.ch), Xiaohua Zhai have left Google DeepMind to join OpenAI!

They were behind recent SOTA vision approaches and open-source models like ViT, SigLIP, PaliGemma
December 4, 2024 at 8:07 AM
Excited to introduce LoftUp!

A strong (than ever) and lightweight feature upsampler for vision encoders that can boost performance on dense prediction tasks by 20%–100%!

Easy to plug into models like DINOv2, CLIP, SigLIP — simple design, big gains. Try it out!

github.com/andrehuang/l...
April 22, 2025 at 7:55 AM
VLMs go MoE ✨

DeepSeek AI dropped three new commercially permissive vision LMs based on SigLIP encoder and their DeepSeek-MoE decoder 🐳

the models come in 1.0B, 2.8B and 4.5B active params 🥹 models seem to catch up with state-of-the-art with less active parameters! huggingface.co/collections/...
December 13, 2024 at 1:33 PM
We maintain strong zero-shot transfer of CLIP / SigLIP across model size and data scale, while achieving up to 4x few-shot sample efficiency and up to +16% performance gains!

Fun project with @confusezius.bsky.social, @zeynepakata.bsky.social, @dimadamen.bsky.social and
@olivierhenaff.bsky.social.
🤔 Can you turn your vision-language model from a great zero-shot model into a great-at-any-shot generalist?

Turns out you can, and here is how: arxiv.org/abs/2411.15099

Really excited to this work on multimodal pretraining for my first bluesky entry!

🧵 A short and hopefully informative thread:
November 28, 2024 at 2:43 PM
Unfortunately, our submission to #NeurIPS didn’t go through with (5,4,4,3). But because I think it’s an excellent paper, I decided to share it anyway.

We show how to efficiently apply Bayesian learning in VLMs, improve calibration, and do active learning. Cool stuff!

📝 arxiv.org/abs/2412.06014
Post-hoc Probabilistic Vision-Language Models
Vision-language models (VLMs), such as CLIP and SigLIP, have found remarkable success in classification, retrieval, and generative tasks. For this, VLMs deterministically map images and text descripti...
arxiv.org
September 18, 2025 at 8:34 PM
SigLIP could be an interesting choice. However, note that in ILIAS we observe that DINOv2 mostly outperforms SigLIP on landmarks, which is probably our closest category to images in localization
February 26, 2025 at 6:06 PM
Google Deepmind's SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

They also released model checkpoints at four sizes (ViT-B (86M), L (303M), So400m (400M), and g-opt (1B)).

Repo: github.com/google-resea...
February 21, 2025 at 4:23 AM