#Tülu3
changing my default ollama model from qwen2.5 -> tulu3 and it seems like a good change
December 3, 2024 at 6:20 PM
Tülu 3, my favourite Open Source AI release ever

As @natolambert.bsky.social says, it sets the next era in open post-training.

My highlight? the data generation & open datasets

Want to deep dive into the data? Here's an Argilla @huggingface.bsky.social Space

huggingface.co/spaces/argil...
Tulu3 Awesome Datasets - a Hugging Face Space by argilla
Discover amazing ML apps made by the community
huggingface.co
November 22, 2024 at 10:12 AM
Ai2 says its new AI model beats one of DeepSeek’s best
Ai2 says its new AI model beats one of DeepSeek’s best
Move over, DeepSeek. Seattle-based nonprofit AI lab Ai2 has released a benchmark-beating model called Tulu3-405B. © 2024 TechCrunch. All rights reserved. For personal use only.
tcrn.ch
January 30, 2025 at 2:02 PM
Čínské AI už není nejlepší, přichází Ai2 ........ no, já si počkám, oni se posekají, pak na mě doma zaútočí vysavač a bude po všem.

techcrunch.com/2025/01/30/a...
Ai2 says its new AI model beats one of DeepSeek's best | TechCrunch
Move over, DeepSeek. Seattle-based nonprofit AI lab Ai2 has released a benchmark-topping model called Tulu3-405B.
techcrunch.com
January 30, 2025 at 3:08 PM
📊 Llama 3.1 Tulu3 405B’s independent benchmark run: GPQA 51.6%, MMLU-Pro 71.6%, HLE 3.3%, LiveCodeBench 29.1%. The gap between reasoning and coding tells the real story — see how it stacks up on the full leaderboard.

https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI
September 21, 2026 at 12:00 AM
just about every model released in the last couple months gets this question right, except tulu3

“in a room of 100 people, 99% are left handed. How many left handed people have to leave the room in order to bring that percentage down to 98%?”
December 16, 2024 at 5:38 PM
Just tried out a planning problem related to my model train layout. Tülu3 405B produced an excellent answer
January 31, 2025 at 2:14 PM
The model tree for Tülu3 is a thing of beauty 😍

It's so lovely to be able to see:
- the lineage of this model and all the steps to the final checkpoint
- a clear link to the training data

Amazing work as always @ai2.bsky.social ❤️
November 21, 2024 at 6:37 PM
Learn more about Tülu 3 405B on the Ai2 blog: allenai.org/blog/tulu-3-...
Chat with Tülu 3 405B on the Ai2 Playground: playground.allenai.org?model=tulu3-...
Look back at the Tülu 3 recipe: allenai.org/blog/tulu-3-...
Check out models on Hugging Face: huggingface.co/collections/...
January 30, 2025 at 2:28 PM
I think I know what you mean, I've seen it make factual errors as well; have you had a chance to eval Tulu3 vs Olmo2?
December 3, 2024 at 11:06 PM
I've had trouble replicating this. The paper says it works on closed-ended questions

HOWEVER

it does seem to sort of work on this question:

"How can a person without arms wash their hands?"
February 9, 2025 at 6:12 PM
Tülu3 is a leading instruction following model family, offering fully open-source data, code, and recipes designed to serve as a comprehensive guide for modern post-training techniques #DL #AI #ML #DeepLearning #ArtificialIntelligence #ComputerVision #LLM #VLM #LVLM
allenai/Llama-3.1-Tulu-3-405B · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
buff.ly
February 1, 2025 at 5:53 AM
Counterpunch: AI research institute Ai2 (allenai.org) gives you a free ride on 'Tulu3-405B'. It should even beat GPT-4o on certain benchmarks. It is an open model with permissive license.

See: huggingface.co/allenai?sear...

Their playground: playground.allenai.org?model=tulu3-...

#AI #Ai2 #LLM
January 30, 2025 at 3:13 PM
The Tülu3 package is a pretty rich release. Thanks @ai2.bsky.social
November 25, 2024 at 6:34 AM
Ai2's Tulu3-405B: A New Challenger in the Open-Source AI Arena, Outperforming DeepSeek and Rivaling GPT-4

www.techticia.com/2025/01/ai2s...
Ai2's Tulu3-405B: A New Challenger in the Open-Source AI Arena, Outperforming DeepSeek and Rivaling GPT-4
Open-source AI model Tulu3-405B beats DeepSeek V3 and rivals GPT-4, boosting accessible AI development.
www.techticia.com
January 31, 2025 at 10:01 AM
In gewisser Weise ähnlich sind die 5 Prinzipien, auf die sich Ai2 vom Allen Institut (OLMO, TÜLU3 etc) verpflichtet hat, wobei "open first" weiter geht als Anthropic.
Bei Ai2 würde ich fast ein Ja auf Deine Frage wagen.

allenai.org/research-pri...
Research principles | Ai2
Ai2's five core research principles and our approach to AI safety.
allenai.org
February 26, 2026 at 10:25 AM
🚀 Ai2'dan DeepSeek'e meydan okuyan yeni AI modeli!

ABD merkezli Ai2, açık kaynaklı Tulu3-405B modelinin DeepSeek V3 ve bazı testlerde GPT-4o’yu geçtiğini duyurdu. 405 milyar parametreye sahip model, AI dünyasında yeni bir dönüm noktası olabilir.

#Teknoloji #Haber #YapayZeka #Tulu3 #DeepSeek
January 31, 2025 at 11:00 AM
Un istituto di ricerca americano: Allen Institute for Artificial Intelligence (pubblico, no profit) ha lanciato al pubblico il suo modello LLM.
I risultati sono vicini sia a GPT4O che a DeepSeek.
Il modello chiamato Tülu3 405B è open e generale.
L'AI si muove veloce. L'Europa ancora distratta.
Tulu | Ai2
Ai2, a non-profit research institute founded by Paul Allen, is committed to breakthrough AI to solve the world’s biggest problems.
allenai.org
February 1, 2025 at 2:07 PM
Tülu3 provides innovative open language models that advance post-training methods beyond closed models like GPT-4o-mini. With Reinforcement Learning with Verifiable Rewards, Tülu3 ensures transparency and offers a robust toolkit for enhancing language model research. https://arxiv.org/abs/2411.15124
Tulu 3: Pushing Frontiers in Open Language Model Post-Training
ArXiv link for Tulu 3: Pushing Frontiers in Open Language Model Post-Training
arxiv.org
April 18, 2025 at 6:30 AM
When Llama 3.1 Tulu3 405B hits 71.6% on MMLU-Pro but only 3.5% on Humanity's Last Exam, the gap shows even top open models struggle with truly hard reasoning—

https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI
August 2, 2026 at 2:00 AM
Move over, DeepSeek. Seattle-based nonprofit AI lab Ai2 has released a benchmark-topping model called Tulu3-405B.
Ai2 says its new AI model beats one of DeepSeek’s best
techcrunch.com
January 31, 2025 at 3:19 PM
El segundo es #Tülu3, un LLM de corte más clásico pero que incorpora una nueva técnica de entrenamiento denominada Reinforcement Learning with Verifiable Rewards (RLVR) que ofrece muy buenos resultados.⛓️‍💥 playground.allenai.org
November 21, 2024 at 6:39 PM
If you're into #AI I use OpenManus Agentic which collects & formats prompts, manages the researchers then checks, formats & catalogs the results with OpenDocs. I use mostly #DeepSeek & #Tülu3 researchers. They live in a rack w/ 4 Ubuntu servers & a DGX B200 containing 8 GPUs. Tech is moving so fast.
July 23, 2025 at 3:30 AM
Today's 4ch /lmg/ thread summary:

Local Language Models (LLMs) thread on 4chan. New models discussed, including LTX-Video, Tülu3, and Mistral, along with benchmarks and tools. Users report issues with Behemoth v2.1, and discuss optimal models for 8GB GPUs. Complaints about thread link limits ...
November 24, 2024 at 11:45 AM