123m³, 13 tonnes, 1.5 million de dollars de l'époque (2.5 actuels) 216000W et 36TFLOPS en FP32.
A droite, la nouvelle 5090 de NVIDIA, 0.0017m³, 1.83kg, $2000, 575W et 104.8TFLOPS en FP32.
#stats
123m³, 13 tonnes, 1.5 million de dollars de l'époque (2.5 actuels) 216000W et 36TFLOPS en FP32.
A droite, la nouvelle 5090 de NVIDIA, 0.0017m³, 1.83kg, $2000, 575W et 104.8TFLOPS en FP32.
#stats
moonshot: "lol no MLP is bf16 and attention is fp32 so you're gonna need like seven terabyte"
New Moonshot model!!
It’s a 48B-A3B acting as an experiment into new long-context efficient attention — a hybrid of Kimi Delta Attention (KDA) and MLA
- very fast inference
- strong performance
github.com/MoonshotAI/K...
moonshot: "lol no MLP is bf16 and attention is fp32 so you're gonna need like seven terabyte"
Meanwhile I keep getting annoyed when I find software that casts my FP64 numbers to FP32.
Meanwhile I keep getting annoyed when I find software that casts my FP64 numbers to FP32.
1. Float8 uses E4M3 for forward & backward - no E5M2
2. Every 4th FP8 accumulate adds to master FP32 accum
3. Latent Attention stores C cache not KV cache
4. No MoE loss balancing - dynamic biases instead
1. Float8 uses E4M3 for forward & backward - no E5M2
2. Every 4th FP8 accumulate adds to master FP32 accum
3. Latent Attention stores C cache not KV cache
4. No MoE loss balancing - dynamic biases instead
LLaDA2.1-flash is 100B but compares itself (it’s worse) to Qwen3-30B-A3B — 3x bigger total size, 33x bigger active size, and still loses
even worse, it’s in FP32 instead of bf16, so double those multiples yet again..
huggingface.co/inclusionAI/...
huggingface.co/inclusionAI/...
✨LLaDA2.1-mini: 16B - Apache2.0
✨LLaDA2.1-flash: 100B - Apache2.0
✨Both delivers editable generation, RL-trained diffusion reasoning and fast inference
LLaDA2.1-flash is 100B but compares itself (it’s worse) to Qwen3-30B-A3B — 3x bigger total size, 33x bigger active size, and still loses
even worse, it’s in FP32 instead of bf16, so double those multiples yet again..
- Flan-T5xxl is a new-generation text encoder.
- TE-only extracts only the parts needed for image generation.
- GGUF format is an even lighter model.
ai-image-journey.com/2025/03/flan...
- Flan-T5xxl is a new-generation text encoder.
- TE-only extracts only the parts needed for image generation.
- GGUF format is an even lighter model.
ai-image-journey.com/2025/03/flan...
The model's activation range exceeds FP16's dynamic range. The card warns of NaNs or silently degraded embeddings.
BF16 is recommended where natively supported. Use FP32 elsewhere, including most CPUs.
The model's activation range exceeds FP16's dynamic range. The card warns of NaNs or silently degraded embeddings.
BF16 is recommended where natively supported. Use FP32 elsewhere, including most CPUs.
- 91.6 TFLOPS (FP32)
- 733 TFLOPS (FP16)
- 1,466 TFLOPS (FP8)
- AMD EPYC 7313
- 256GB RAM
- Up to 4x GPUs (e.g. Nvidia L40S)
- 246TB NVMe storage
- 100GbE + 25GbE networking
- 2.5kW PSU, rugged case
- <55 lbs, fits in carry-on
- 91.6 TFLOPS (FP32)
- 733 TFLOPS (FP16)
- 1,466 TFLOPS (FP8)
- AMD EPYC 7313
- 256GB RAM
- Up to 4x GPUs (e.g. Nvidia L40S)
- 246TB NVMe storage
- 100GbE + 25GbE networking
- 2.5kW PSU, rugged case
- <55 lbs, fits in carry-on
Balance in all things is relevant, but I can't put on FP64 (or FP128) just for the sake of it.
i suspect the trend we'll see is FP64 (sometimes FP128) on scalars, FP32-64 on Vectors, =<FP32 matrix
Balance in all things is relevant, but I can't put on FP64 (or FP128) just for the sake of it.
i suspect the trend we'll see is FP64 (sometimes FP128) on scalars, FP32-64 on Vectors, =<FP32 matrix
- Quality: FP32 > FP16 ≒ BF16 > MXFP8 > FP8_scaled > NVFP4 > FP8
- Faster when the model can be loaded into VRAM
- Use formats supported by your GPU
www.ai-image-journey.com/2026/06/floa...
- Quality: FP32 > FP16 ≒ BF16 > MXFP8 > FP8_scaled > NVFP4 > FP8
- Faster when the model can be loaded into VRAM
- Use formats supported by your GPU
www.ai-image-journey.com/2026/06/floa...
apparently ollama is upcasting it to fp32
apparently ollama is upcasting it to fp32
@atcute/cbor@1.0.3: 3.5 KB
@atcute/cbor@1.0.4: 2.7 KB
@atcute/cbor@1.0.3: 3.5 KB
@atcute/cbor@1.0.4: 2.7 KB
อ่านต่อ : www.blockdit.com/posts/6ab8af...
#ShoperGamer #LLMQuantization #LLM #Ai #ModelAi #FP32 #FP16 #GGUF #GPTQ #Knowledge #Study #Feed
อ่านต่อ : www.blockdit.com/posts/6ab8af...
#ShoperGamer #LLMQuantization #LLM #Ai #ModelAi #FP32 #FP16 #GGUF #GPTQ #Knowledge #Study #Feed
- Flan-T5xxl は新型のテキストエンコーダー
- TE-only は必要な部分のみを抽出
- GGUF 形式はさらに軽量化したモデル
Flux.1 と Stable Diffusion 3.5 のプロンプトの理解力と画質の改善が期待できます!
- Flan-T5xxl は新型のテキストエンコーダー
- TE-only は必要な部分のみを抽出
- GGUF 形式はさらに軽量化したモデル
Flux.1 と Stable Diffusion 3.5 のプロンプトの理解力と画質の改善が期待できます!