#MLLMs
i swear to god
April 7, 2025 at 6:17 PM
Back to self-supervised basics (rotation prediction, colorization) in the era of MLLMs.

We discovered that simply interleaving SSL tasks with instruction tuning ones improves visual grounding of MLLMs, especially on vision-centric tasks.

Check it out 👇
1/n New paper - V-GIFT 🎁

Self-supervised tasks like rotation prediction or colorization were big in 2018.
Do they still matter?

Yes.
We turn them into visual instruction tuning data for MLLMs.

Result: models rely more on the image and perform better on vision tasks 👀
April 17, 2026 at 5:44 PM
Xuechen Li
One Domain, Many Tongues: Composing Domain and Language LoRAs for Cross-Lingual Remote-Sensing MLLMs without Paired Data
https://arxiv.org/abs/2609.26097
September 24, 2026 at 10:01 AM
Die Studie zeigt eine gravierende Schwäche aktueller KI-Systeme. Selbst multimodale Sprachmodelle versagen bei einfachen visuellen Aufgaben, die sogar Kleinkinder mühelos bewältigen. Aber das absolute Schlusslicht ist Grok. Na ja, Grok ist wohl mit Deepfakes ausgelastet.

unipat.ai/blog/BabyVis...
January 18, 2026 at 7:54 PM
I buy all of my AI products from MLLMs.
February 9, 2026 at 2:51 PM
A study shows that multimodal large language models (MLLMs) risk suboptimal performance due to their reliance on text. Researchers suggest combining curriculum learning with KL-based self-distillation to enhance visual reasoning and bridge this gap. https://arxiv.org/abs/2510.22836
Rethinking the Text-Vision Reasoning Imbalance in MLLMs through the Lens of Training Recipes
ArXiv link for Rethinking the Text-Vision Reasoning Imbalance in MLLMs through the Lens of Training Recipes
arxiv.org
January 10, 2026 at 6:21 AM
Neat idea and great insight into the inner workings of MLLMs. MLLMs are good at high level vision, but fail at low- and mid-level visual tasks.
If GPT-4o walked into a neuro-opthalmology clinic, what would it be diagnosed with?

Here we administered 51 tests from 6 clinical and experimental batteries to assess vision in commercial AI models.

Very proud to share this first work from @genetang.bsky.social's PhD!

arxiv.org/abs/2504.10786
Visual Language Models show widespread visual deficits on neuropsychological tests
Visual Language Models (VLMs) show remarkable performance in visual reasoning tasks, successfully tackling college-level challenges that require high-level understanding of images. However, some recen...
arxiv.org
April 18, 2025 at 7:33 AM
Why Do Multimodal LLMs (MLLM) Struggle with Spatial Understanding?

Is it just because of not enough data? Nope!

Researchers finds that current MLLMs have fundamental limitations, and they suggest that real progress in spatial reasoning will require targeted reasoning injection
September 11, 2025 at 11:54 PM
A survey paper on what is probably the hottest area in AI right now: memory.

"The AI Hippocampus: How Far are We From Human Memory?"

This survey presents a comprehensive and structured synthesis of memory in LLMs and MLLMs, organizing the literature into a cohesive taxonomy
January 25, 2026 at 1:57 AM
Microsoft's FlorenceVL

A new family of multimodal large language models (MLLMs) with enriched visual representations produced by Florence-2 (huggingface.co/microsoft/Fl...).

See links provided by‪ @gm8xx8.bsky.social
December 6, 2024 at 5:57 PM
1/n New paper - V-GIFT 🎁

Self-supervised tasks like rotation prediction or colorization were big in 2018.
Do they still matter?

Yes.
We turn them into visual instruction tuning data for MLLMs.

Result: models rely more on the image and perform better on vision tasks 👀
April 17, 2026 at 2:52 PM
New paper in Nature Human Behaviour.

I use a conjoint experiment to test multimodal large language models (MLLMs) for context-sensitive content moderation and compare with human subjects. Methodologically, this demonstrates how social science techniques can enhance AI auditing. 💻🤖💬
December 15, 2025 at 3:04 PM
🍄 OpenEMMA: Un paso adelante en conducción autónoma con MLLMs 🍄
January 8, 2025 at 8:09 PM
🛟 Reliable & reliability folks @eccv.bsky.social!
Join our workshop on Uncertainty Quantification for Computer Vision #eccv2026

We have a super lineup of speakers (from computer vision to MLLMs & LLMs) and cool posters.

🗓️ Date: Tue Sep 8
📌Room: Quality View Hotel - Helsingborg
September 6, 2026 at 9:19 PM
📄👀: cross-modal information flow in multimodal large language models

neat interpretability work on how visual and linguistic information is integrated in MLLMs: "the model first transfers the more general visual features of the whole image into the representations of (linguistic) question tokens."
Cross-modal Information Flow in Multimodal Large Language Models
The recent advancements in auto-regressive multimodal large language models (MLLMs) have demonstrated promising progress for vision-language tasks. While there exists a variety of studies investigatin...
arxiv.org
December 1, 2024 at 7:24 PM
Very impressive visual reasoning benchmark. They found that most models still fall below the capacities of a 3-year-old

unipat.ai/blog/BabyVis...
BabyVision: Visual Reasoning Beyond Language
State-of-the-art MLLMs achieve PhD-level language reasoning but struggle with visual tasks that 3-year-olds solve effortlessly. We introduce BabyVision, a benchmark revealing the infancy of AI vision.
unipat.ai
February 15, 2026 at 7:43 PM
Yijun Hu, Heng Fan, Libo Zhang: Exploiting Target Knowledge from MLLMs for Robust Few-Shot Segmentation https://arxiv.org/abs/2609.28949 https://arxiv.org/pdf/2609.28949 https://arxiv.org/html/2609.28949
September 25, 2026 at 6:40 AM
Sam is 100% correct on this. Indeed, human babies have essential cognitive priors such as permanence, continuity, and boundary of objects, 3D Euclidean understanding of space, etc.

We spent 2 years to systematically to examine and show the lack of such in MLLMs: arxiv.org/abs/2410.10855
May 24, 2025 at 5:55 AM
Multimodal large language models (MLLMs) often struggle with small visual details, but do we need to retrain them to fix this?

In our #ICLR'25 paper, we found that MLLMs already know where to look—even when their final answers are wrong!
April 26, 2025 at 7:09 AM
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

abs: arxiv.org/abs/2412.05271
model: huggingface.co/OpenGVLab/In...

Introduces new InternVL-2.5 model, the first open-source MLLMs to surpass 70% on the MMMU benchmark
December 9, 2024 at 4:57 AM
Join me for AI and Image Analysis at ESU DH 2025 in Besançon, France in July! The one week course will introduce AI methods such as object detection and MLLMs for analyzing images such as art, manuscripts, photography, and film. More details here: esudh.github.io
March 25, 2025 at 5:05 PM
🚀🌍The rapid advancement of multilingual large language models (mLLMs) is exciting, but are we evaluating them effectively?

Our new paper explores how we can improve generative evaluations for mLLMs by learning from machine translation (MT) evaluation practices. 🔎
April 17, 2025 at 6:09 PM
🧵1/10 Excited to share our #SIGGRAPH paper "MonetGPT: Solving Puzzles Enhances MLLMs' Image Retouching Skills" 🌟
We explore how to make MLLMs operation-aware by solving visual puzzles and propose a procedural framework for image retouching
#MLLM
May 27, 2025 at 3:13 PM