#BLIP3
Salesforce Research's BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions

BLIP3-KALE, a dataset of 218 million image-text pairs that augments synthetic dense image captions with web-scale alt-text to generate factually grounded image captions.
November 14, 2024 at 4:55 AM
Salesforce AI open-sources (code, model, dataset, and paper) BLIP3-o, the unified multimodal models that excels at both image understanding and generation in a single autoregressive architecture!

It also includes:
▶️ Complete model implementations
▶️ Model weights
May 18, 2025 at 12:57 PM
▶️ 25M+ detailed caption pretrain dataset
▶️ 60K high-quality instruction tuning dataset

📊 Paper: www.arxiv.org/abs/2505.09568
🤗 Models: huggingface.co/BLIP3o/BLIP3...
🤗 Datasets: huggingface.co/BLIP3o
🧠 Code: github.com/JiuhaiChen/B...
📽️ Learn on the go (AI Generated): bit.ly/3EWDZQp
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
Unifying image understanding and generation has gained growing attention in recent research on multimodal models. Although design choices for image understanding have been extensively studied, the opt...
www.arxiv.org
May 18, 2025 at 12:57 PM
November 16, 2024 at 9:00 PM
🚨Clear your schedule for a recap of a tremendous week featuring DeepSeek-V3 along with BLIP3-o’s improvements in multimodal architecture 🚀

Check out the top 10 papers for the week👇

- Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
May 19, 2025 at 4:34 PM
Anas Awadalla, Le Xue, Manli Shu, An Yan, Jun Wang, Senthil Purushwalkam, Sheng Shen, Hannah Lee, Oscar Lo, Jae Sung Park, Etash Guha, Silvio Savarese, Ludwig Schmidt, Yejin Choi, Caiming Xiong, ...
BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions
https://arxiv.org/abs/2411.07461
November 13, 2024 at 5:31 AM
[23/30] 131 Likes, 15 Comments, 5 Posts
2505.09568, cs․CV | cs․AI, 14 May 2025

🆕BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Jiuhai Chen, Zhiyang Xu, Xichen Pan, Yushi Hu, Can Qin, Tom Goldstein, Lifu Huang, Tianyi Zhou, Saining Xie, Silvio Sav...
May 21, 2025 at 12:07 AM
(1/3) 76 Likes, 3 Comments, 15 May 2025, Hugging Face
Paper page - BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
Join the discussion on this paper page
huggingface.co
May 21, 2025 at 12:07 AM
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

BLIP3-o introduces a suite of open-source unified multimodal models that excel at both image understanding and image generation, combining autoregressive and diffusion architectures.
May 19, 2025 at 4:34 PM
- BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
- Learning Dynamics in Continual Pre-Training for Large Language Models
- Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs
May 19, 2025 at 4:34 PM
and datasets, we develop BLIP3-o, a suite of state-of-the-art unified multimodal models. BLIP3-o achieves superior performance across most of the popular benchmarks spanning both image understanding and generation tasks. To facilitate future research, [7/8 of https://arxiv.org/abs/2505.09568v1]
May 15, 2025 at 6:06 AM
Chen, Xu, Pan, Hu, Qin, Goldstein, Huang, Zhou, Xie, Savarese, Xue, Xiong, Xu: BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset https://arxiv.org/abs/2505.09568 https://arxiv.org/pdf/2505.09568 https://arxiv.org/html/2505.09568
May 15, 2025 at 6:06 AM
Introducing BLIP3-o: A Family of Fully Open Unified Multimodal Models Architecture, Training and ...

https://www.salesforce.com/blog/blip3/

Result Details
May 27, 2025 at 6:59 PM
Introducing BLIP3-o: A Family of Fully Open Unified Multimodal Models Architecture, Training and ...

https://www.salesforce.com/blog/blip3/

Result Details
May 23, 2025 at 6:47 PM
Introducing BLIP3-o: A Family of Fully Open Unified Multimodal Models Architecture, Training and ...

https://www.salesforce.com/blog/blip3/

Result Details
May 23, 2025 at 6:05 PM
Train Mipha, Aquila, BLIP3, Molmo, Qwen2VL, and InternVL2 variants which is awesome. Their variants beat specialists like VideoLlama2, Oryx. It seems to beat even the video capabilities of InternVL2 and Qwen2VL themselves!
December 10, 2024 at 12:12 AM