#VisionLanguage
#VisionLanguage models are increasingly used for a wide range of problems, but seem complex to build. I wrote some code and recorded a tutorial in my lab yesterday to help others demystify how to create these models. #keepbuilding
November 25, 2025 at 5:40 PM
DeepSeek-VL: Vision Meets Language

AI that can see and understand: DeepSeek-VL explains charts, diagrams, scanned docs, and natural images with high accuracy.
👉 deepseekai.free/models/deeps...

#AI #VisionLanguage #DeepSeekVL #ArtificialIntelligence #MachineLearning #ComputerVision #MultimodalAI
August 14, 2025 at 7:54 AM
I'm so excited to share that our paper "EmoGist: Efficient In-Context Learning for Visual Emotion Understanding", has been accepted to #emnlp2025 Findings! @emnlpmeeting.bsky.social

Read the pre-print from arXiv: tinyurl.com/emogist-paper

#emogist #multimodal #visionlanguage #icl
September 23, 2025 at 3:00 AM
July 18, 2025 at 3:01 PM
RP
これ私も悩んでます。
「生成AI」と呼ばれていたものを「LLM」とだけ呼ぶ人が増えてて、自分の認識が間違っていたのかとずっと疑問に感じてるんですよね。

私もそもそも大量の既存データ合成するツールの「生成AI」呼び自体にも疑問があるものの、海外産の技術でそれに対抗して裁判してるアーティストも「GenAI」と呼んでいたのでそれに倣ってきました。

前にLargeLanguageModel(大規模言語モデル)を調べたときに言語・テキストに特化したものという情報を見ていたのでVisionLanguage Model(視覚言語モデル)他もLLMに含めるのかどうかがわかってません。
June 28, 2026 at 1:53 AM
EMM1 evaluates how AI understands images and text together. It highlights where models excel and where they fall short, helping build more reliable multimodal systems.

#AI #Data #VisionLanguage
encord.com/multimodal-d...
E-MM1 Dataset: The World's Largest Multimodal AI Dataset
The E-MM1 dataset is the world's largest multimodal AI dataset, with more than 100 million groups of data in five modalities to foster the development of models that fuse multiple modalities.
encord.com
October 23, 2025 at 12:08 PM
Discover FastVLM, a breakthrough in Vision Language Models that boosts image resolution and speeds up processing, making text-rich image understanding more efficient! 🚀📸 #AI #VisionLanguage #TechInnovation https://rpst.cc/CojGYo
September 1, 2025 at 2:54 AM
A new study introduces cross-modal backward-compatible learning for vision-language models. Read more: https://getnews.me/cross-modal-backward-compatible-learning-for-vision-language-models/ #visionlanguage #crossmodal #machinelearning
October 8, 2025 at 12:32 PM
Hybrid pipeline merging Monte Carlo Tree Search with a strong vision‑language model makes reliable step‑level labels, boosting benchmarks like MMMU and MathVista. Read more: https://getnews.me/vision-language-process-reward-models-enhance-test-time-scaling/ #multimodal #visionlanguage
October 3, 2025 at 10:59 AM
The Multi‑Modal Explainable Learning (MMEL) framework boosts explanation precision on vision‑language benchmarks while preserving accuracy via a Relationship module. Read more: https://getnews.me/multi-modal-interpretability-boosts-vision-language-model-transparency/ #multimodal #visionlanguage
September 22, 2025 at 4:29 AM
❓Docs read by models—but can they prove it? // RAD² X ensures traceable, agency-first choices in VLM pipelines. https://glcnd.io/transforming-document-processing-how-vision-language-models-are-changing-the-game/ #GLCND #RAD2X #VisionLanguage #ExplainableAI
Transforming Document Processing: How Vision-Language Models are Changing the Game - GLCND.IO
glcnd.io
September 27, 2025 at 3:28 PM
The AGILE framework raised 2x2 jigsaw accuracy from 9.5% to 82.8% and added roughly 3% average gain across nine vision tasks, according to the authors. Read more: https://getnews.me/agile-boosts-visual-perception-and-reasoning-in-vision-language-models/ #visionlanguage #agile #multimodal
October 3, 2025 at 7:33 PM
Chimera, a new benchmark of 7,500 Wikipedia diagrams, shows many vision-language models rely on shortcuts, with 15 models evaluated across seven families. Read more: https://getnews.me/chimera-test-suite-reveals-shortcut-learning-in-vision-language-models/ #visionlanguage #diagrams #machinelearning
September 29, 2025 at 3:12 PM
Decoupled Proxy Alignment (DPA) uses a proxy LLM and visual‑relevance‑driven loss weighting to improve vision‑language alignment, with open‑source code on GitHub. Read more: https://getnews.me/decoupled-proxy-alignment-improves-vision-language-harmony-in-ai-models/ #dpa #visionlanguage #multimodal
September 19, 2025 at 10:18 PM
A new learned scoring model ranks synthetic remote-sensing image-text pairs, and fine-tuning on the top 30% of data boosts accuracy versus using the full set. Read more: https://getnews.me/learning-a-scoring-model-to-curate-remote-sensing-vision-language-data/ #remotesensing #visionlanguage
September 22, 2025 at 9:43 PM
"Exciting news from Liquid AI: LFM2.5-VL-3B is their most powerful vision-language model for on-device use, with improved screen/UI understanding, grounding, multi-image input, and function calling. 📸🔥 #AI #VisionLanguage #EdgeComputing"
LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge
A Blog post by Liquid AI on Hugging Face
huggingface.co
August 27, 2026 at 12:39 PM
Discover PaliGemma 2: Google's lightweight, multi-scale vision-language model, ideal for image-text tasks, content creation, and AI development projects.

#AI #PaliGemma #Google #LLM #visionlanguage
aidisruptionpub.com/p/paligemma-...
PaliGemma 2: Google's Multi-Scale Lightweight Vision-Language Model
Discover PaliGemma 2: Google's lightweight, multi-scale vision-language model, ideal for image-text tasks, content creation, and AI development projects.
aidisruptionpub.com
December 8, 2024 at 1:51 PM
Back from the break with Phillip Isola @phillipisola.bsky.social on
“On the Perceptual Distance Between Images and Text.”
A fascinating and interactive look at how models (and humans!) measure similarity 👏🏻

#HiCV2025 #ICCV2025 #VisionLanguage
October 20, 2025 at 9:09 PM
#Inaturalist has released a new #VisionLanguage #AI tool that let you search the huge Inaturalist foto pool in natural language like "Bird eating fruit" www.inaturalist.org/blog/95911
Search iNaturalist Photos With Text
We are excited to announce the launch of our Vision Language Demo, developed in collaboration with our long-time partners at the University of Massachusetts Amherst, the University of Edinburgh, the U...
www.inaturalist.org
June 28, 2024 at 10:33 AM
arXiv:2503.08144v1 Announce Type: new
Abstract: Recently, large language models (LLMs) and visionlanguage models (VLMs) have achieved significant success, demonstrating remarkable capabilities in understanding various images and videos, particularly [1/7 of https://arxiv.org/abs/2503.08144v1]
March 12, 2025 at 5:59 AM
assessment of such synthetically generated RS visionlanguage data is notably absent. To fill this gap, we propose a novel score model trained on large-scale RS visionlanguage preference data for automated quality assessment. [4/7 of https://arxiv.org/abs/2503.00743v1]
March 5, 2025 at 6:13 AM
assessment of such synthetically generated RS visionlanguage data is notably absent. To fill this gap, we propose a novel score model trained on large-scale RS visionlanguage preference data for automated quality assessment. [4/7 of https://arxiv.org/abs/2503.00743v1]
March 4, 2025 at 6:29 AM