⚡100x Training Throughput
🎯Fast Convergence
🔢Pure Int8 Pretraining of RNN LLMs
⚡100x Training Throughput
🎯Fast Convergence
🔢Pure Int8 Pretraining of RNN LLMs
今日KPCでいくINT8の男
今日KPCでいくINT8の男
Sol-AttnやらINT8 ConvRot VAEやら3StepLoRAやら入れて
960 × 544p
1秒1分から1秒20秒に((( ;゚Д゚)))
3べぇ界王拳だ❕
これならガチャが許容範囲かな?🤔
きゃぴ💕
Sol-AttnやらINT8 ConvRot VAEやら3StepLoRAやら入れて
960 × 544p
1秒1分から1秒20秒に((( ;゚Д゚)))
3べぇ界王拳だ❕
これならガチャが許容範囲かな?🤔
きゃぴ💕
Blog: fengyao.notion.site/flash-rl
Repo: github.com/yaof20/Flash...
Blog: fengyao.notion.site/flash-rl
Repo: github.com/yaof20/Flash...
The trick: Binary search with int8 rescoring.
I'll show you a demo & how it works in the 🧵:
The trick: Binary search with int8 rescoring.
I'll show you a demo & how it works in the 🧵:
int8 + int8 = ?
int8 * int8 = ?
int8 + int8 = ?
int8 * int8 = ?
pplx-embed-v1 and pplx-embed-context-v1
Specifically trained for int8 and binary embeddings, they'll be viable for massive search problems.
Details in 🧵
pplx-embed-v1 and pplx-embed-context-v1
Specifically trained for int8 and binary embeddings, they'll be viable for massive search problems.
Details in 🧵
"... reducing from float32 (32 bits) to float16 (16 bits) or int8 (8 bits) should yield ideal 2× or 4× gains, respectively. However, such improvements are not observed in practice."
"... reducing from float32 (32 bits) to float16 (16 bits) or int8 (8 bits) should yield ideal 2× or 4× gains, respectively. However, such improvements are not observed in practice."
Codestral Embed can output embeddings with different dimensions and precisions. Codestral Embed with dimension 256 and int8 precision still performs better than any model from our competitors.
Codestral Embed can output embeddings with different dimensions and precisions. Codestral Embed with dimension 256 and int8 precision still performs better than any model from our competitors.