Yingxuan Zhuang et al.
#arXiv #cs.AI #cs.LG
We discovered that simply interleaving SSL tasks with instruction tuning ones improves visual grounding of MLLMs, especially on vision-centric tasks.
Check it out 👇
Self-supervised tasks like rotation prediction or colorization were big in 2018.
Do they still matter?
Yes.
We turn them into visual instruction tuning data for MLLMs.
Result: models rely more on the image and perform better on vision tasks 👀
We discovered that simply interleaving SSL tasks with instruction tuning ones improves visual grounding of MLLMs, especially on vision-centric tasks.
Check it out 👇
One Domain, Many Tongues: Composing Domain and Language LoRAs for Cross-Lingual Remote-Sensing MLLMs without Paired Data
https://arxiv.org/abs/2609.26097
One Domain, Many Tongues: Composing Domain and Language LoRAs for Cross-Lingual Remote-Sensing MLLMs without Paired Data
https://arxiv.org/abs/2609.26097
unipat.ai/blog/BabyVis...
unipat.ai/blog/BabyVis...
Here we administered 51 tests from 6 clinical and experimental batteries to assess vision in commercial AI models.
Very proud to share this first work from @genetang.bsky.social's PhD!
arxiv.org/abs/2504.10786
Is it just because of not enough data? Nope!
Researchers finds that current MLLMs have fundamental limitations, and they suggest that real progress in spatial reasoning will require targeted reasoning injection
Is it just because of not enough data? Nope!
Researchers finds that current MLLMs have fundamental limitations, and they suggest that real progress in spatial reasoning will require targeted reasoning injection
"The AI Hippocampus: How Far are We From Human Memory?"
This survey presents a comprehensive and structured synthesis of memory in LLMs and MLLMs, organizing the literature into a cohesive taxonomy
"The AI Hippocampus: How Far are We From Human Memory?"
This survey presents a comprehensive and structured synthesis of memory in LLMs and MLLMs, organizing the literature into a cohesive taxonomy
A new family of multimodal large language models (MLLMs) with enriched visual representations produced by Florence-2 (huggingface.co/microsoft/Fl...).
See links provided by @gm8xx8.bsky.social
A new family of multimodal large language models (MLLMs) with enriched visual representations produced by Florence-2 (huggingface.co/microsoft/Fl...).
See links provided by @gm8xx8.bsky.social
Self-supervised tasks like rotation prediction or colorization were big in 2018.
Do they still matter?
Yes.
We turn them into visual instruction tuning data for MLLMs.
Result: models rely more on the image and perform better on vision tasks 👀
Self-supervised tasks like rotation prediction or colorization were big in 2018.
Do they still matter?
Yes.
We turn them into visual instruction tuning data for MLLMs.
Result: models rely more on the image and perform better on vision tasks 👀
I use a conjoint experiment to test multimodal large language models (MLLMs) for context-sensitive content moderation and compare with human subjects. Methodologically, this demonstrates how social science techniques can enhance AI auditing. 💻🤖💬
I use a conjoint experiment to test multimodal large language models (MLLMs) for context-sensitive content moderation and compare with human subjects. Methodologically, this demonstrates how social science techniques can enhance AI auditing. 💻🤖💬
Join our workshop on Uncertainty Quantification for Computer Vision #eccv2026
We have a super lineup of speakers (from computer vision to MLLMs & LLMs) and cool posters.
🗓️ Date: Tue Sep 8
📌Room: Quality View Hotel - Helsingborg
Join our workshop on Uncertainty Quantification for Computer Vision #eccv2026
We have a super lineup of speakers (from computer vision to MLLMs & LLMs) and cool posters.
🗓️ Date: Tue Sep 8
📌Room: Quality View Hotel - Helsingborg
neat interpretability work on how visual and linguistic information is integrated in MLLMs: "the model first transfers the more general visual features of the whole image into the representations of (linguistic) question tokens."
neat interpretability work on how visual and linguistic information is integrated in MLLMs: "the model first transfers the more general visual features of the whole image into the representations of (linguistic) question tokens."
unipat.ai/blog/BabyVis...
unipat.ai/blog/BabyVis...
We spent 2 years to systematically to examine and show the lack of such in MLLMs: arxiv.org/abs/2410.10855
We spent 2 years to systematically to examine and show the lack of such in MLLMs: arxiv.org/abs/2410.10855
In our #ICLR'25 paper, we found that MLLMs already know where to look—even when their final answers are wrong!
In our #ICLR'25 paper, we found that MLLMs already know where to look—even when their final answers are wrong!
Paper: www.arxiv.org/abs/2509.023...
Paper: www.arxiv.org/abs/2509.023...
abs: arxiv.org/abs/2412.05271
model: huggingface.co/OpenGVLab/In...
Introduces new InternVL-2.5 model, the first open-source MLLMs to surpass 70% on the MMMU benchmark
abs: arxiv.org/abs/2412.05271
model: huggingface.co/OpenGVLab/In...
Introduces new InternVL-2.5 model, the first open-source MLLMs to surpass 70% on the MMMU benchmark
Our new paper explores how we can improve generative evaluations for mLLMs by learning from machine translation (MT) evaluation practices. 🔎
Our new paper explores how we can improve generative evaluations for mLLMs by learning from machine translation (MT) evaluation practices. 🔎