#VideoMME
Videomme Pegged The Radio Sub
May 21, 2025 at 8:26 PM
Mulla on ollut nyt 5 kk portugalinkielinen YouTube -kanava, jossa esittelen vaimoni kanssa Suomea ja muita maita joihin satumme matkustamaan. Suosituimmat videot ovat kuitenkin ne jotka käsittelevät politiikkaa, eikä nähtävyyksiä. Esim suosituin videomme on vappumarssista 🤔
July 12, 2024 at 7:51 AM
VideoLLaMA3, latest MLLMs for image and video understanding.

🖐️ 7B models: DocVQA: 94.9, MathVision: 26.2, VideoMME: 66.2/70.3, MLVU: 73.0
🤏 2B models for edge devices: MMMU: 45.3, VideoMME: 59.6/63.4
👊 Frontier-class video model with ONLY 3M video-text pairs
February 7, 2025 at 3:34 PM
- Achieves state-of-the-art scores like 87.8% on MMLU-Pro and 87.5% on VideoMME, outperforming models in multimodal tasks.
- Supports 201 languages, expanded visual and STEM data, and capabilities in GUI automation and video-to-code generation.
February 17, 2026 at 3:29 AM
Our approach also outperforms #GPT-4o by 21 points in accuracy on EPIC-KITCHENS-100-MQA and demonstrates improvements across other action-related video benchmarks, including #VideoMME, #PerceptionTest, and #MVBench.
March 25, 2025 at 8:46 AM
Google has launched an updated Gemini 2.5 Pro (I/O edition) AI model. The company reports "significantly" improved coding capabilities, leading the WebDev Arena Leaderboard benchmark, and 84.8% performance on the VideoMME benchmark for video understanding.
May 7, 2025 at 3:48 PM
relationships and long-range temporal dynamics. Experiments on standard video benchmarks show significant relative performance gains of 3.05% on VideoMME, 1.97% on NeXTQA, and 1.31% on LongVideoBench, over the baseline Qwen2.5-VL model. These results [4/6 of https://arxiv.org/abs/2504.14096v1]
April 22, 2025 at 5:56 AM
Uskallanko luottaa? -videomme ovat saaneet pelkästään YouTubessa yli 30 000 katselukertaa! 🎉 Jos et ole vielä nähnyt videoita, löydät ne YouTube-kanavaltamme. 👇 #vakehyva #uskallankoluottaa youtube.com/playlist
Uskallanko luottaa?
Tarinoita viranomaispelosta ja sen voittamisesta darin, albanian, venäjän ja suomen kielillä. Julkaisemme uudet videot viikoittain 23.9.–14.10.2025.
youtube.com
October 21, 2025 at 10:08 AM
Näitkö jo kaikki uudenvuoden videomme? Viime viikolla mukana olivat Haruka, THE SOUND BEE HD, THE MICRO HEAD 4N'S, IRON ATTACK!, DEXCORE, I love you Orchestra, Shinya ja DARRELL. <a href="http://www.jame-world.com/fi/themes-1170-vuodenvaihteen-videoviestit-2018.html" class="hover:underline text-blue-600 dark:text-sky-400 no-card-link" target="_blank" rel="noopener" data-link="bsky">http://www.jame-world.com/fi/themes-1170-vuodenvaihteen-videoviestit-2018.html
November 16, 2024 at 4:52 PM
Yhdysvaltalainen APA-järjestö julkaisi vuonna 2024 päivitetyn ohjeistuksen aikuisiän monimuotoisen trauman kanssa työskentelylle. Tällä videolla 79 avataan APA:n seitsemää keskeistä teema-aluetta tarkemmin. Jos aihe kiinnostaa katso myös videomme 78 Youtubessa.

www.youtube.com/watch?v=kjNm...
Video 79. Traumainformoitu lähestymistapa ja APA:n uudet ohjeet, Osa 2.
YouTube video by Yhteinen kieli – Traumainformoitu kulttuuri
www.youtube.com
November 10, 2025 at 6:15 AM
Scores 84.8% on the VideoMME benchmark for video understanding, and is available through the Gemini API via Google AI Studio, Vertex AI, and in the Gemini app

blog.google/products/ge...
Build rich, interactive web apps with an updated Gemini 2.5 Pro
Our updated version of Gemini 2.5 Pro has improved capabilities for coding.
blog.google
May 6, 2025 at 9:09 PM
CoS: Chain-of-Shot Prompting for Long Video Understanding

arxiv.org/abs/2502.06428

Datasets used: VideoMME, MLVU, LongVideoBench, MVBench, NEXT-QA
CoS: Chain-of-Shot Prompting for Long Video Understanding
Multi-modal Large Language Models (MLLMs) struggle with long videos due to the need for excessive visual tokens. These tokens exceed massively the context length of MLLMs, resulting in filled by redun...
arxiv.org
February 19, 2025 at 9:59 AM
2507.07966
強化学習を活用し、視覚言語モデル(VLM)の推論を長時間の動画にスケールアップするフルスタックフレームワークを紹介する。我々は3つの重要なコンポーネントを統合することで、ロングビデオ推論のユニークな課題に対処する:(1) LongVideo-Reasonという大規模なデータセット。スポーツ、ゲーム、ブログなど...
July 14, 2025 at 12:07 AM
Google dropped Gemini 2.5 Pro.

🔥 Tops WebDev Arena
🎯 84.8% on VideoMME
💻 Way better at coding, building web apps, editing, and agents

Was supposed to launch at I/O… but the buzz made it come early. 😀
May 9, 2025 at 11:15 AM
Gemini 2.5 Pro update: Coding, web apps with Gemini

blog.google/products/gem...

- 구글은 코딩 기능이 크게 향상된 제미나이 2.5 Pro 프리뷰(I/O 에디션)의 조기 액세스를 출시

- 특히 인터랙티브 웹 앱 구축에 강점을 보임

- 이 업데이트는 WebDev Arena 리더보드에서 선두(+147 Elo 포인트)를 차지

- VideoMME에서 84.8% 점수를 기록

(계속)
Build rich, interactive web apps with an updated Gemini 2.5 Pro
Our updated version of Gemini 2.5 Pro has improved capabilities for coding.
blog.google
May 6, 2025 at 3:41 PM
What’s new with Gemini 2.5 Pro (I/O Preview):
– #1 on WebDev Arena, beating Claude 3.7 in UI + frontend tasks
– Turns YouTube videos into working apps (84.8% VideoMME score)
– Helps devs polish UX with better CSS, layout, and animation
– Auto-upgrades in API and Vertex AI—zero setup, same cost
May 7, 2025 at 1:46 PM
performance on the mainstream long video QA benchmarks, e.g., it achieves 77.0 on VideoMME and 70.1 on EgoSchema, outperforming its strong baselines (e.g., Intern2.5VL-8B and InternVideo2.5-8B), by up to 10.8\% and 6.2\%. Compared to leading [7/8 of https://arxiv.org/abs/2506.06097v1]
June 9, 2025 at 6:12 AM
granularity approach enables processing of hour-long videos while maintaining semantic fidelity. Experimental validation on LongVideoBench and VideoMME demonstrates significant performance improvements, establishing state-of-the-art results for not [4/5 of https://arxiv.org/abs/2506.04953v1]
June 6, 2025 at 6:07 AM
FlexSelect delivers strong gains across multiple long-video benchmarks including VideoMME, MLVU, LongVB, and LVBench. Moreover, it achieves significant speed-ups (for example, up to 9 times on a LLaVA-Video-7B model), highlighting FlexSelect's promise [5/6 of https://arxiv.org/abs/2506.00993v1]
June 3, 2025 at 6:11 AM
achieves 62.0% on VideoMME, 69.8% on MLVU, and 67.4% on TempCompass, all with fewer than 6,000 visual tokens per video. The code will be publicly available on the homepage. [6/6 of https://arxiv.org/abs/2505.15529v1]
May 22, 2025 at 6:10 AM
including VideoMME, LVBench, and MLVU, demonstrate that ViaRL consistently delivers superior temporal grounding performance and robust generalization across diverse video understanding tasks, highlighting its effectiveness and scalability. Notably, [6/7 of https://arxiv.org/abs/2505.15447v1]
May 22, 2025 at 6:08 AM
Comprehensive evaluations on public benchmarks, LVBench and VideoMME-Long, demonstrate that AVA achieves state-of-the-art performance, attaining 62.3% and 64.1% accuracy, respectively, significantly surpassing existing VLM and video [5/7 of https://arxiv.org/abs/2505.00254v1]
May 2, 2025 at 5:56 AM
a real-time mode. Meanwhile, it achieves state-of-the-art results at the 7B/8B scale on popular video QA benchmarks such as VideoMME and OVOBench, demonstrating the broad generalizability of our approach. All resources of this paper have been released [7/8 of https://arxiv.org/abs/2504.16030v1]
April 23, 2025 at 6:05 AM
Compared to Qwen2.5-VL-7B, VideoChat-R1 boosts performance several-fold in tasks like temporal grounding (+31.8) and object tracking (+31.2). Additionally, it significantly improves on general QA benchmarks such as VideoMME (+0.9), MVBench (+1.0), and [5/6 of https://arxiv.org/abs/2504.06958v1]
April 10, 2025 at 6:07 AM
EPIC-KITCHENS-100-MQA. Lastly, we show improvements on other action-related video benchmarks such as EgoSchema, PerceptionTest, LongVideoBench, VideoMME and MVBench, suggesting that MLLMs are a promising path forward for complex action tasks. Code and [5/6 of https://arxiv.org/abs/2503.18712v1]
March 25, 2025 at 6:20 AM