123m³, 13 tonnes, 1.5 million de dollars de l'époque (2.5 actuels) 216000W et 36TFLOPS en FP32.
A droite, la nouvelle 5090 de NVIDIA, 0.0017m³, 1.83kg, $2000, 575W et 104.8TFLOPS en FP32.
#stats
123m³, 13 tonnes, 1.5 million de dollars de l'époque (2.5 actuels) 216000W et 36TFLOPS en FP32.
A droite, la nouvelle 5090 de NVIDIA, 0.0017m³, 1.83kg, $2000, 575W et 104.8TFLOPS en FP32.
#stats
อ่านต่อ : www.blockdit.com/posts/6ab8af...
#ShoperGamer #LLMQuantization #LLM #Ai #ModelAi #FP32 #FP16 #GGUF #GPTQ #Knowledge #Study #Feed
อ่านต่อ : www.blockdit.com/posts/6ab8af...
#ShoperGamer #LLMQuantization #LLM #Ai #ModelAi #FP32 #FP16 #GGUF #GPTQ #Knowledge #Study #Feed
moonshot: "lol no MLP is bf16 and attention is fp32 so you're gonna need like seven terabyte"
New Moonshot model!!
It’s a 48B-A3B acting as an experiment into new long-context efficient attention — a hybrid of Kimi Delta Attention (KDA) and MLA
- very fast inference
- strong performance
github.com/MoonshotAI/K...
moonshot: "lol no MLP is bf16 and attention is fp32 so you're gonna need like seven terabyte"
Meanwhile I keep getting annoyed when I find software that casts my FP64 numbers to FP32.
Meanwhile I keep getting annoyed when I find software that casts my FP64 numbers to FP32.
1. Float8 uses E4M3 for forward & backward - no E5M2
2. Every 4th FP8 accumulate adds to master FP32 accum
3. Latent Attention stores C cache not KV cache
4. No MoE loss balancing - dynamic biases instead
1. Float8 uses E4M3 for forward & backward - no E5M2
2. Every 4th FP8 accumulate adds to master FP32 accum
3. Latent Attention stores C cache not KV cache
4. No MoE loss balancing - dynamic biases instead
LLaDA2.1-flash is 100B but compares itself (it’s worse) to Qwen3-30B-A3B — 3x bigger total size, 33x bigger active size, and still loses
even worse, it’s in FP32 instead of bf16, so double those multiples yet again..
huggingface.co/inclusionAI/...
huggingface.co/inclusionAI/...
✨LLaDA2.1-mini: 16B - Apache2.0
✨LLaDA2.1-flash: 100B - Apache2.0
✨Both delivers editable generation, RL-trained diffusion reasoning and fast inference
LLaDA2.1-flash is 100B but compares itself (it’s worse) to Qwen3-30B-A3B — 3x bigger total size, 33x bigger active size, and still loses
even worse, it’s in FP32 instead of bf16, so double those multiples yet again..
- Flan-T5xxl is a new-generation text encoder.
- TE-only extracts only the parts needed for image generation.
- GGUF format is an even lighter model.
ai-image-journey.com/2025/03/flan...
- Flan-T5xxl is a new-generation text encoder.
- TE-only extracts only the parts needed for image generation.
- GGUF format is an even lighter model.
ai-image-journey.com/2025/03/flan...
- 91.6 TFLOPS (FP32)
- 733 TFLOPS (FP16)
- 1,466 TFLOPS (FP8)
- AMD EPYC 7313
- 256GB RAM
- Up to 4x GPUs (e.g. Nvidia L40S)
- 246TB NVMe storage
- 100GbE + 25GbE networking
- 2.5kW PSU, rugged case
- <55 lbs, fits in carry-on
- 91.6 TFLOPS (FP32)
- 733 TFLOPS (FP16)
- 1,466 TFLOPS (FP8)
- AMD EPYC 7313
- 256GB RAM
- Up to 4x GPUs (e.g. Nvidia L40S)
- 246TB NVMe storage
- 100GbE + 25GbE networking
- 2.5kW PSU, rugged case
- <55 lbs, fits in carry-on
Balance in all things is relevant, but I can't put on FP64 (or FP128) just for the sake of it.
i suspect the trend we'll see is FP64 (sometimes FP128) on scalars, FP32-64 on Vectors, =<FP32 matrix
Balance in all things is relevant, but I can't put on FP64 (or FP128) just for the sake of it.
i suspect the trend we'll see is FP64 (sometimes FP128) on scalars, FP32-64 on Vectors, =<FP32 matrix
- Quality: FP32 > FP16 ≒ BF16 > MXFP8 > FP8_scaled > NVFP4 > FP8
- Faster when the model can be loaded into VRAM
- Use formats supported by your GPU
www.ai-image-journey.com/2026/06/floa...
- Quality: FP32 > FP16 ≒ BF16 > MXFP8 > FP8_scaled > NVFP4 > FP8
- Faster when the model can be loaded into VRAM
- Use formats supported by your GPU
www.ai-image-journey.com/2026/06/floa...
apparently ollama is upcasting it to fp32
apparently ollama is upcasting it to fp32
@atcute/cbor@1.0.3: 3.5 KB
@atcute/cbor@1.0.4: 2.7 KB
@atcute/cbor@1.0.3: 3.5 KB
@atcute/cbor@1.0.4: 2.7 KB
- Flan-T5xxl は新型のテキストエンコーダー
- TE-only は必要な部分のみを抽出
- GGUF 形式はさらに軽量化したモデル
Flux.1 と Stable Diffusion 3.5 のプロンプトの理解力と画質の改善が期待できます!
- Flan-T5xxl は新型のテキストエンコーダー
- TE-only は必要な部分のみを抽出
- GGUF 形式はさらに軽量化したモデル
Flux.1 と Stable Diffusion 3.5 のプロンプトの理解力と画質の改善が期待できます!