Read more here! wjbmattingly.com/blog/where-t...
Read more here! wjbmattingly.com/blog/where-t...
We show how to efficiently apply Bayesian learning in VLMs, improve calibration, and do active learning. Cool stuff!
📝 arxiv.org/abs/2412.06014
We show how to efficiently apply Bayesian learning in VLMs, improve calibration, and do active learning. Cool stuff!
📝 arxiv.org/abs/2412.06014
it's up on youtube by popular demand www.youtube.com/embed/_TlhKH...
it's up on youtube by popular demand www.youtube.com/embed/_TlhKH...
with @chloesu07.bsky.social, @thomasfel.bsky.social, @shamkakade.bsky.social and Stephanie Gil
arxiv.org/abs/2504.11695
with @chloesu07.bsky.social, @thomasfel.bsky.social, @shamkakade.bsky.social and Stephanie Gil
arxiv.org/abs/2504.11695
Do multimodal LLMs (VLMs) reason about high-level visual perception like humans do? We asked over 2000 human observers and 18 VLMs to describe scenes using 15 different tasks, ranging from general knowledge, affordances, affect, sensory experiences, and future prediction. 1/
Do multimodal LLMs (VLMs) reason about high-level visual perception like humans do? We asked over 2000 human observers and 18 VLMs to describe scenes using 15 different tasks, ranging from general knowledge, affordances, affect, sensory experiences, and future prediction. 1/
Non-paywalled version:
arxiv.org/abs/2504.10786
Tweet thread below from first author @genetang.bsky.social...
Non-paywalled version:
arxiv.org/abs/2504.10786
Tweet thread below from first author @genetang.bsky.social...
- arxiv.org/abs/2410.06468 (does spatial cognition ...)
- arxiv.org/abs/2307.06281 (MMBench)
- arxiv.org/abs/2411.13543 (BALROG) - additional points for the LOTR ref.
- arxiv.org/abs/2410.06468 (does spatial cognition ...)
- arxiv.org/abs/2307.06281 (MMBench)
- arxiv.org/abs/2411.13543 (BALROG) - additional points for the LOTR ref.
Most of us know that LLMs, VLMs, foundation models etc. are just statistical language models trained in big data and that they neither have emerged capabilities nor are they able to really reason.
You absolutely can lie via data & math & everything about how you present data is an argument for how to interpret/understand it
Everything we know is founded on best guesses & analogies
All technoscience is enmeshed with human values
Most of us know that LLMs, VLMs, foundation models etc. are just statistical language models trained in big data and that they neither have emerged capabilities nor are they able to really reason.
huggingface.co/bartowski/Le...
github.com/apple-aiml-r...
arxiv.org/abs/2605.07019
huggingface.co/bartowski/Le...
github.com/apple-aiml-r...
arxiv.org/abs/2605.07019
arxiv.org/abs/2410.06468 does spatial cognition
arxiv.org/abs/2307.06281 MMBench
arxiv.org/abs/2411.13543 BALROG
arxiv.org/abs/2410.07765 GameTraversalBenchmark
3dsrbench.github.io 3DSRBenchmark
open-eqa.github.io Open-EQA
arxiv.org/abs/2410.06468 does spatial cognition
arxiv.org/abs/2307.06281 MMBench
arxiv.org/abs/2411.13543 BALROG
arxiv.org/abs/2410.07765 GameTraversalBenchmark
3dsrbench.github.io 3DSRBenchmark
open-eqa.github.io Open-EQA
LLMs/VLMs are the currently the state of the art for every content moderation use case.
LLMs/VLMs are the currently the state of the art for every content moderation use case.
get started with sota VLMs (gemma 3, Qwen2.5VL, InternVL3 & more) and serve them wherever you want 🤩
learn more github.com/ggml-org/lla... 📖
get started with sota VLMs (gemma 3, Qwen2.5VL, InternVL3 & more) and serve them wherever you want 🤩
learn more github.com/ggml-org/lla... 📖
🤨Interested? Check out our latest work at #AAAI25:
💻Code and 📝Paper at: github.com/CompVis/DisCLIP
🧵👇
🤨Interested? Check out our latest work at #AAAI25:
💻Code and 📝Paper at: github.com/CompVis/DisCLIP
🧵👇
TL;DR: many modern VLMs are unsafe across various types of queries and languages.
arxiv.org/abs/2501.10057
huggingface.co/datasets/fel...
MSTS is exciting because it tests for safety risks *created by multimodality*. Each prompt consists of a text + image that *only in combination* reveal their full unsafe meaning.
🧵
TL;DR: many modern VLMs are unsafe across various types of queries and languages.
arxiv.org/abs/2501.10057
huggingface.co/datasets/fel...
🔗 to extended versions:
1. 🙋 "How can we make predictions in BDL efficiently?" 👉 arxiv.org/abs/2411.18425
2. 🙋 "How can we do prob. active learning in VLMs" 👉 arxiv.org/abs/2412.06014
🔗 to extended versions:
1. 🙋 "How can we make predictions in BDL efficiently?" 👉 arxiv.org/abs/2411.18425
2. 🙋 "How can we do prob. active learning in VLMs" 👉 arxiv.org/abs/2412.06014
github.com/davanstrien/...
github.com/davanstrien/...
SmolVLM (256M & 500M) runs on <1GB GPU memory.
Fine-tune it on your laptop and run it on your toaster. 🚀
Even the 256M model outperforms our Idefics 80B (Aug '23).
How small can we go? 👀
SmolVLM (256M & 500M) runs on <1GB GPU memory.
Fine-tune it on your laptop and run it on your toaster. 🚀
Even the 256M model outperforms our Idefics 80B (Aug '23).
How small can we go? 👀
🖼️ We studied how unified VLMs, trained to generate both text and images (e.g., Meta's Chameleon), exchange information between modalities, comparing them to standard VLMs.
📄 Paper: arxiv.org/abs/2412.06646
Deep dive: 👇
🖼️ We studied how unified VLMs, trained to generate both text and images (e.g., Meta's Chameleon), exchange information between modalities, comparing them to standard VLMs.
📄 Paper: arxiv.org/abs/2412.06646
Deep dive: 👇