A new blog from Rafal Wojdyla, MTS, explains how a change in sharding strategy, Expert Parallelism, made room for a ~50% larger model at the same ~23B active parameters, and what it costs.
🔗 openathena.ai/blog/expert-...
A new blog from Rafal Wojdyla, MTS, explains how a change in sharding strategy, Expert Parallelism, made room for a ~50% larger model at the same ~23B active parameters, and what it costs.
🔗 openathena.ai/blog/expert-...
bit.ly/535b
bit.ly/535b
🔗 openathena.ai/blog/cluster...
🔗 openathena.ai/blog/cluster...
Delphi is Marin’s first step: we pretrained many small models with one recipe, then extrapolated 300× to predict a 25B-param / 600B-token run with just 0.2% error.
Getting there took some work 🧵
Delphi is Marin’s first step: we pretrained many small models with one recipe, then extrapolated 300× to predict a 25B-param / 600B-token run with just 0.2% error.
Getting there took some work 🧵
And about the Models we are releasing in @dlwh.bsky.social's training retro: marin.readthedocs.io/en/latest/re...
And about the Models we are releasing in @dlwh.bsky.social's training retro: marin.readthedocs.io/en/latest/re...
x.com/percyliang/s...
x.com/percyliang/s...