#FP32
A gauche, le Blue Gene d'IBM, meilleur supercalculateur du monde en 2004.

123m³, 13 tonnes, 1.5 million de dollars de l'époque (2.5 actuels) 216000W et 36TFLOPS en FP32.

A droite, la nouvelle 5090 de NVIDIA, 0.0017m³, 1.83kg, $2000, 575W et 104.8TFLOPS en FP32.

#stats
January 23, 2025 at 9:57 PM
today's cocaine is only getting cheaper because it's being served at such a low quant. it's not that fp32 80's coke
September 26, 2026 at 11:00 PM
schnellだと上手くいかないけどDev fp32だと上手くいく。 #メカ子
January 23, 2025 at 10:07 AM
September 27, 2026 at 6:32 AM
update: model gets 6.4% bettter at fp32 just by directly pointing h_ones directly at the logit identity matrix.
error because h_ones isn't pointed directly at the ones vector at logits. okay. how about i just do a rank-one correction to point you straight at the ones vector. see how you like that.
March 8, 2026 at 9:29 PM
me: "oh, i bet i can get this to run in 48GB of VRAM"

moonshot: "lol no MLP is bf16 and attention is fp32 so you're gonna need like seven terabyte"
Kimi-Linear: more efficient attention

New Moonshot model!!

It’s a 48B-A3B acting as an experiment into new long-context efficient attention — a hybrid of Kimi Delta Attention (KDA) and MLA

- very fast inference
- strong performance

github.com/MoonshotAI/K...
October 30, 2025 at 3:50 PM
It falls out of supporting fp32 multiplication (fp32 has 24-bit mantissa so to multiply two numbers you need a 24x24=>48 product, and you can reuse it for a 24-bit integer multiplication if you don't have a dedicated full speed int32 multiply unit). Earlier NV HW had this too, deprecated now.
May 27, 2026 at 11:10 PM
So Nvidia are talking about one Petaflop in a desktop box, with the caveat being this is FP4, i.e. 4 bit 'floating point' numbers! Which apparently works for LLMs just fine.
Meanwhile I keep getting annoyed when I find software that casts my FP64 numbers to FP32.
January 7, 2025 at 7:04 AM
Fun fact: half of all the fp32 values are between -1 and 1.
November 15, 2023 at 12:14 PM
A well known important feature to stabilize RL training is implementing the LM head in fp32 precision to help with gradients. Reproduced the plot from the MiniMax M1 paper entirely on my dgx spark and in Ai2's post-training research codebase.
January 23, 2026 at 3:04 PM
Daniel Han ( of @unsloth.bsky.social )'s diagram of DeepSeek v3 Architecture.

1. Float8 uses E4M3 for forward & backward - no E5M2
2. Every 4th FP8 accumulate adds to master FP32 accum
3. Latent Attention stores C cache not KV cache
4. No MoE loss balancing - dynamic biases instead
December 29, 2024 at 9:36 PM
this is embarrassing

LLaDA2.1-flash is 100B but compares itself (it’s worse) to Qwen3-30B-A3B — 3x bigger total size, 33x bigger active size, and still loses

even worse, it’s in FP32 instead of bf16, so double those multiples yet again..
LLaDA 2.1 is out 🔥 MoE diffusion language models released by AntGroup

huggingface.co/inclusionAI/...
huggingface.co/inclusionAI/...

✨LLaDA2.1-mini: 16B - Apache2.0
✨LLaDA2.1-flash: 100B - Apache2.0
✨Both delivers editable generation, RL-trained diffusion reasoning and fast inference
February 9, 2026 at 10:46 PM
Releasing Flan-T5xxl_TE-only in FP32, FP16, and GGUF Formats!
- Flan-T5xxl is a new-generation text encoder.
- TE-only extracts only the parts needed for image generation.
- GGUF format is an even lighter model.
ai-image-journey.com/2025/03/flan...
March 8, 2025 at 9:22 AM
Preview builds of v5.2 are coming soon! Big performance improvements, an upgraded fully FP32 rendering pipeline, and a newly modernized FileType plugin system
March 8, 2026 at 6:57 PM
GigaIO Gryf - portable AI Supercomputer:
- 91.6 TFLOPS (FP32)
- 733 TFLOPS (FP16)
- 1,466 TFLOPS (FP8)
- AMD EPYC 7313
- 256GB RAM
- Up to 4x GPUs (e.g. Nvidia L40S)
- 246TB NVMe storage
- 100GbE + 25GbE networking
- 2.5kW PSU, rugged case
- <55 lbs, fits in carry-on
May 18, 2025 at 6:47 PM
If I had to guess the ray reconstruction format. They are using FP8 with FP32 Accumulate, which for the 3090 has to default to FP16 with FP32 instead. The 4090 is 8.25x faster than a 3090 Ti at this type of workload...
January 25, 2025 at 4:04 PM
But what you net out is you're still loosing out on performance/area vs other designs that don't.

Balance in all things is relevant, but I can't put on FP64 (or FP128) just for the sake of it.

i suspect the trend we'll see is FP64 (sometimes FP128) on scalars, FP32-64 on Vectors, =<FP32 matrix
June 24, 2025 at 2:51 PM
Which is Better: FP8_scaled or MXFP8? A Thorough Comparison of Image Generation AI Model
- Quality: FP32 > FP16 ≒ BF16 > MXFP8 > FP8_scaled > NVFP4 > FP8
- Faster when the model can be loaded into VRAM
- Use formats supported by your GPU
www.ai-image-journey.com/2026/06/floa...
June 13, 2026 at 7:16 AM
(zoomed in a little) quarter-res (2048^2), fp64 vs fp32, final renders & absolute differences of converged vs diverged raw data; the loss of precision doesn't impact the convergent region all that badly, just a faint layer of noise. rate of divergence in divergent region gets messed up badly tho
February 23, 2026 at 12:22 AM
i grabbed the biblically correct bf16 version of R1-0529-8b off unsloth and... oh my dog this shit is slow

apparently ollama is upcasting it to fp32
May 30, 2025 at 1:48 AM
okay, got rid of fp16 and fp32 support, not sure why i kept it around
@atcute/cbor@1.0.3: 3.5 KB
@atcute/cbor@1.0.4: 2.7 KB
October 19, 2024 at 2:41 AM
Flan-T5xxl_TE-only の FP32、FP16、GGUF 形式 を公開!
- Flan-T5xxl は新型のテキストエンコーダー
- TE-only は必要な部分のみを抽出
- GGUF 形式はさらに軽量化したモデル
Flux.1 と Stable Diffusion 3.5 のプロンプトの理解力と画質の改善が期待できます!
March 8, 2025 at 9:15 AM
?? fp32 training is fast enough to be fine with this big ass threadripper ?? why wasn't i doing that before
August 3, 2026 at 7:33 PM