#visionlanguage
"Exciting news from Liquid AI: LFM2.5-VL-3B is their most powerful vision-language model for on-device use, with improved screen/UI understanding, grounding, multi-image input, and function calling. 📸🔥 #AI #VisionLanguage #EdgeComputing"
LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge
A Blog post by Liquid AI on Hugging Face
huggingface.co
August 27, 2026 at 12:39 PM
RP
これ私も悩んでます。
「生成AI」と呼ばれていたものを「LLM」とだけ呼ぶ人が増えてて、自分の認識が間違っていたのかとずっと疑問に感じてるんですよね。

私もそもそも大量の既存データ合成するツールの「生成AI」呼び自体にも疑問があるものの、海外産の技術でそれに対抗して裁判してるアーティストも「GenAI」と呼んでいたのでそれに倣ってきました。

前にLargeLanguageModel(大規模言語モデル)を調べたときに言語・テキストに特化したものという情報を見ていたのでVisionLanguage Model(視覚言語モデル)他もLLMに含めるのかどうかがわかってません。
June 28, 2026 at 1:53 AM
🚀 The Qwen team just dropped Qwen3.6‑35B‑A3B, a 3B‑param vision‑language MoE model with linear & grouped query attention. Open‑source, agentic coding ready—see how this transformer pushes AI forward! #Qwen3_6_35B_A3B #MixtureOfExperts #VisionLanguage

🔗 aidailypost.com/news/qwen-te...
April 17, 2026 at 6:55 AM
Even the hottest multimodal models stumble—capped at 50% on simple visual entity tasks. What does this reveal about current vision‑language gaps? Dive into the benchmarks and see why AI still has a long way to go. #MultimodalLearning #VisionLanguage #AIPerformance

🔗 aidailypost.com/news/top-mul...
February 8, 2026 at 1:45 PM
New research shows how to fool CLIP‑style vision‑language models with fresh adversarial tricks. Could this expose hidden AI security gaps? Dive into the latest evasion techniques and what they mean for multimodal ML. #AdversarialAttacks #VisionLanguage #AIsecurity

🔗 aidailypost.com/news/researc...
January 28, 2026 at 5:09 PM
Beyond Memorization: A Multi-Modal Ordinal Regression Benchmark to Expose Popularity Bias in Vision-Language Models
Chia-Jui Chang, He Syu et al.
Paper
Details
#VisionLanguage #OrdinalRegression #BiasBenchmark
December 26, 2025 at 9:01 AM
Just saw an open‑source OCR model hit 82.4 on the olmOCR‑bench—handles equations, tables, multilingual docs, and scales like a champ with PaddleOCR VL and ERNIE‑4.5‑0.3B. Dive into the details! #OCR #olmOCRbench #VisionLanguage

🔗 aidailypost.com/news/open-so...
December 24, 2025 at 1:45 PM
Google’s PaLI learns across 100 languages—with images as context.
See the leap → https://glcnd.io/unlocking-multilingual-insights-googles-pali-for-language-image-learning-in-100-languages/
#multilingualAI #visionlanguage
Unlocking Multilingual Insights: Google’s PaLI for Language-Image Learning in 100 Languages - GLCND.IO
glcnd.io
December 8, 2025 at 1:37 PM
Black Forest Labs just dropped Flux 2, packing the new Mistral‑3 24B vision‑language model with a hybrid Rectified Flow Transformer + VAE encoder. The BFL API makes it super easy to experiment—check out the details! #Flux2 #Mistral324B #VisionLanguage

🔗 aidailypost.com/news/black-f...
November 25, 2025 at 7:40 PM
#VisionLanguage models are increasingly used for a wide range of problems, but seem complex to build. I wrote some code and recorded a tutorial in my lab yesterday to help others demystify how to create these models. #keepbuilding
November 25, 2025 at 5:40 PM
EMM1 evaluates how AI understands images and text together. It highlights where models excel and where they fall short, helping build more reliable multimodal systems.

#AI #Data #VisionLanguage
encord.com/multimodal-d...
E-MM1 Dataset: The World's Largest Multimodal AI Dataset
The E-MM1 dataset is the world's largest multimodal AI dataset, with more than 100 million groups of data in five modalities to foster the development of models that fuse multiple modalities.
encord.com
October 23, 2025 at 12:08 PM
Back from the break with Phillip Isola @phillipisola.bsky.social on
“On the Perceptual Distance Between Images and Text.”
A fascinating and interactive look at how models (and humans!) measure similarity 👏🏻

#HiCV2025 #ICCV2025 #VisionLanguage
October 20, 2025 at 9:09 PM
A training-free, explainable vision-language model for medical imaging has been announced. Read more: https://getnews.me/training-free-explainable-vision-language-model-for-medical-imaging/ #medimaging #visionlanguage #explainable
October 8, 2025 at 3:49 PM
A new probabilistic language-image pre-training approach is reported to boost performance of vision-language models. Read more: https://getnews.me/probabilistic-language-image-pre-training-boosts-vision-language-models/ #visionlanguage #pretraining #ai
October 8, 2025 at 2:28 PM
A new study introduces cross-modal backward-compatible learning for vision-language models. Read more: https://getnews.me/cross-modal-backward-compatible-learning-for-vision-language-models/ #visionlanguage #crossmodal #machinelearning
October 8, 2025 at 12:32 PM
Vision‑language models guide indoor robot navigation, selecting subgoals that reduce path length by about 10 % in simulation, working zero‑shot with the DYNUS planner. Read more: https://getnews.me/vision-language-models-boost-efficiency-of-indoor-robot-navigation/ #visionlanguage #robotics
October 8, 2025 at 6:23 AM
The study reframes zero‑shot classification as Q&A and adds an attention‑intervention, boosting top‑1 accuracy on bird, flower and vehicle benchmarks. Code on GitHub. Read more: https://getnews.me/zero-shot-fine-grained-classification-with-vision-language-models/ #visionlanguage #zeroshot
October 7, 2025 at 8:36 PM
Spatial‑ViLT adds depth maps, 3D coordinate grids and edge maps to vision‑language models, achieving top results on the Visual Spatial Reasoning benchmark. Read more: https://getnews.me/spatial-vilt-improves-3d-spatial-reasoning-with-multi-task-learning/ #spatialvilt #visionlanguage
October 7, 2025 at 2:21 PM
Fine‑tuned LLaVa‑NeXT‑Vicuna with LoRA boosted specificity and balanced accuracy in carotid plaque stroke‑risk prediction, especially when paired with patient data. 3 Oct 2025. https://getnews.me/large-vision-language-models-boost-carotid-plaque-risk-prediction/ #visionlanguage #carotid #stroke
October 6, 2025 at 8:58 AM
MaskCD, a new contrastive decoding method that masks the image head, cuts hallucination rates in LVLMs like LLaVA‑1.5‑7B and Qwen‑VL‑7B without hurting overall performance. Read more: https://getnews.me/maskcd-cuts-hallucinations-in-vision-language-models/ #maskcd #lvlm #visionlanguage
October 6, 2025 at 7:54 AM
A study of 221 rebus puzzles shows vision‑language models excel at visual composition but falter on missing elements and cultural symbols. The paper was submitted on 3 Oct 2025. https://getnews.me/explainability-shows-limits-of-vision-language-models-on-rebus-puzzles/ #visionlanguage #rebuspuzzles
October 6, 2025 at 7:50 AM
AdaRD‑Key selects query‑relevant, diverse keyframes in real time on a single GPU, achieving state‑of‑the‑art results on LongVideoBench and Video‑MME. https://getnews.me/adard-key-boosts-query-driven-frame-selection-for-long-form-video-ai/ #adardkey #visionlanguage
October 6, 2025 at 7:48 AM
The AGILE framework raised 2x2 jigsaw accuracy from 9.5% to 82.8% and added roughly 3% average gain across nine vision tasks, according to the authors. Read more: https://getnews.me/agile-boosts-visual-perception-and-reasoning-in-vision-language-models/ #visionlanguage #agile #multimodal
October 3, 2025 at 7:33 PM
AgenticIQA uses a planner‑executor‑summarizer workflow and released AgenticIQA‑200K with 200,000 examples. It beats strong baselines on Pearson and Spearman correlation. https://getnews.me/agenticiqa-adaptive-interpretable-image-quality-assessment-framework/ #agenticiqa #imagequality #visionlanguage
October 3, 2025 at 12:22 PM