#VLLMs
VLLMs don't "see", they remember (and they are pretty stubborn about their memories too!). I had the pleasure to supervise two master students at @itu.dk that decided to investigate this fascinating problem of VLLMs. Ayat Khudoir and Rabia Sufian just presented their work at the FAILED workshop
September 10, 2026 at 7:13 AM
I wonder who's working on VLLMs.

It's the next obvious thing, right? LLMs are working out like gangbusters. Throw some zeros on the end of some hyperparameters. Start talking VLLMs. That could be super interesting.

<cat_reading_newspaper>
I should spend more time on arXiv.
</cat_reading_newspaper>
July 10, 2026 at 10:23 PM
Honestly i do think robots and VLLMs could make things like electronics recycling maybe less hacky?
April 12, 2026 at 4:08 PM
vision is too weak in vLLMs
December 5, 2024 at 12:01 AM
I'm pretty happy to present our work on using VLLMs to achieve semantic clustering at #ica25. (Centennial A at 9.00 if you are curious ;-))
June 13, 2025 at 2:39 PM
Are Vision Language Models ready for scientific research?
🧑‍🔬🧪

We compared leading VLLMs on the three pillars of chemical and material science discovery: data extraction, lab experimentation and data interpretation.
arxiv.org/abs/2411.16955
November 27, 2024 at 4:46 PM
really important to keep in mind: VLLMs don't see.
September 10, 2026 at 7:13 AM
It is kind of a shame that Claude Code has made making data entry portals so easy, given that LLMs and VLLMs, etc are killing the need for human RAs to punch in data for me in the first place
March 10, 2026 at 7:00 PM
Clear illustration of the limitations of VLLMs. We still have a long way to go (most likely going beyond current training and model paradigms) before "solving vision".
thinking of calling this "The Illusion Illusion"

(more examples below)
December 1, 2024 at 6:10 PM
New research shows how Vision-Language Models (VLLMs) represent image concepts within hidden layers, discovering distinct features that enhance multimodal learning. Sparse Autoencoders reveal a shared representation of images and text evolving deeper in the model. https://arxiv.org/abs/2506.04706
Line of Sight: On Linear Representations in VLLMs
ArXiv link for Line of Sight: On Linear Representations in VLLMs
arxiv.org
June 7, 2025 at 2:30 AM
Two PolarVis presentations this week: (1) on the (mis)uses of VLLMs for analysing climate change visuals, today at the CCVision Network Meeting, and (2) on methods for longitudinal coordination detection applied to PolarVis data, tomorrow at SunBelt. @lrossi.bsky.social @matmagnani.bsky.social
June 24, 2025 at 11:42 AM
Fine-Tuning vLLMs for Document Understanding | Towards Data Science in @towardsdatascience.com

towardsdatascience.com/vllm-fine-tu...
Fine-Tuning vLLMs for Document Understanding | Towards Data Science
Learn how you can fine-tune visual language models for specific tasks
towardsdatascience.com
May 12, 2025 at 11:01 AM
#computationalsocialscience Our work on using VLLMs for semantic visual clustering is out on SSCR: journals.sagepub.com/doi/10.1177/...

There is a lot happening in the space of computational analysis of visual content as we acknowledge more and more the visual nature of contemporary communication.
Sage Journals: Discover world-class research
Subscription and open access journals from Sage, the world's leading independent academic publisher.
journals.sagepub.com
September 19, 2025 at 9:48 AM
deep learning methods can totally read images but not generalist vLLMs and they almost certainly never will, just because there's no real reason to train them to
September 18, 2026 at 1:35 AM
This iteration of AI tech (LLMs and VLLMs) has as much chance of becoming sentient as I have of actually becoming a border collie. Any suggestion otherwise is idiocy.

A future tech might but the roadblocks aren't learning or memory, they're the inability to experience reality.

That's decades away.
July 30, 2026 at 1:18 PM