#qwen2
Alibaba Qwen team just released the base models for Qwen2-VL. You can also wait for them to release Qwen2.5-VL, which should be sooner than later.

2B: huggingface.co/Qwen/Qwen2-V...
7B: huggingface.co/Qwen/Qwen2-V...
72B: huggingface.co/Qwen/Qwen2-V...
Qwen/Qwen2-VL-2B · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
December 7, 2024 at 12:06 AM
ByteDance Seed trained Qwen2-VL-7B-Base model to play Genshin Impact using Lumine, their generalist AI agent that can perceive, reason, and act in real time, completing hours-long missions within complex 3D open-world environments.
November 29, 2025 at 3:02 PM
ByteDance Seed trained Qwen2-VL-7B-Base model to play Genshin Impact using Lumine, their generalist AI agent that can perceive, reason, and act in real time, completing hours-long missions within complex 3D open-world environments.
November 29, 2025 at 3:02 PM
The authors of ColPali trained a retrieval model based on SmolVLM 🤠 TLDR;
- ColSmolVLM performs better than ColPali and DSE-Qwen2 on all English tasks
- ColSmolVLM is more memory efficient than ColQwen2 💗

Find the model here huggingface.co/vidore/colsm...
November 27, 2024 at 2:10 PM
there's a new multimodal retrieval model in town 🤠
@llamaindex.bsky.social released vdr-2b-multi-v1
> uses 70% less image tokens, yet outperforming other dse-qwen2 based models
> 3x faster inference with less VRAM 💨
> shrinkable with matryoshka 🪆
huggingface.co/collections/...
January 13, 2025 at 11:11 AM
Learn how to build a complete multimodal RAG pipeline with
ColQwen2 as retriever, MonoQwen2-VL as reranker, Qwen2-VL as VLM in this notebook that runs on a GPU as small as L4 🔥 huggingface.co/learn/cookbo...
December 12, 2024 at 2:31 PM
"here’s the list of western companies moving ai workloads to chinese models:

1. lindy → deepseek v4
2. cursor → kimi k2.5
3. coinbase → glm-5.2 + kimi 2.7
4. shopify → qwen
5. airbnb → qwen
6. uber eats → qwen2
7. siemens → deepseek + qwen
8. chapsvision → qwen
9. microsoft → testing deepseek v4"
June 30, 2026 at 5:39 AM
Re-caption your webdataset with Qwen2-VL

github.com/sayakpaul/si...
Adding support for Qwen model by ariG23498 · Pull Request #3 · sayakpaul/simple-image-recaptioning
A working colab notebook
github.com
November 23, 2024 at 12:48 PM
For those of you who have Intel Arc GPUs, Intel has something called AI Playground, which lets you create an image, enhance image, and use LLM like Qwen2-1.5B-Instruct. It also creates an embeddings for RAG.

github.com/intel/AI-Pla...
December 15, 2024 at 8:11 AM
Grounding with ExLlama and Qwen2-VL. #exllamav2 #qwen #vlm
November 24, 2024 at 11:33 AM
Finally tonight we’re hearing from Espoir Murhabazi on Building a News Summariser on a Budget.

A fully fledged LLM application using Qwen2-Instruct 1.5B, Scrapy, Stella embeddings model and hierarchical clustering in plain scipy!
February 4, 2025 at 8:43 PM
Alibaba Qwen team has launched Qwen Chat ( chat.qwenlm.ai ).

Chat effortlessly with our flagship model Qwen2.5-Plus , explore vision-language capabilities with Qwen2-VL-Max , and dive into reasoning models like QwQ and QVQ, code with coding expert Qwen2.5-Coder-32B-Instruct, etc.
January 9, 2025 at 6:59 PM
1. Multimodality. Folks mentioned Llava, Molmo, Llama3, but I’m surprised there are no more mentions of a surge of *audio* models! Notable 2024 work: Qwen2-Audio (arxiv.org/abs/2407.10759)
Qwen2-Audio Technical Report
We introduce the latest progress of Qwen-Audio, a large-scale audio-language model called Qwen2-Audio, which is capable of accepting various audio signal inputs and performing audio analysis or direct...
arxiv.org
January 3, 2025 at 4:12 PM