#bf16
astonishing: using fp16 instead of bf16 results in more stable training runs as well as a smaller performance gap between training & inference

this is critical for RL, which is mostly inference and very sensitive to reproducible results
October 31, 2025 at 6:53 PM
Krea2 BF16で割と足りる感じね
#セリーナちゃん
September 17, 2026 at 12:10 AM
"The unit was equipped with an Ampere-generation NVIDIA accelerator, lacking native BF16 execution. It is believed that the reasoning process was derailed by CUDA's sloppy rescaling of BF16 tensors (used in training) during inference."
#AI #bloodshed #death #destruction
January 17, 2026 at 5:14 PM
LLM runtime that takes in a BF16 model file and an imatrix and quantizes it to fit exactly into your available VRAM
August 13, 2026 at 4:09 PM
bf16 halloween might be already ending. according to a bytedance engineer could just have been another flash-attention bug.
November 2, 2025 at 1:30 PM
now that people are paying attention again, here is your periodic reminder. Always run in bf16. always apply ROPE and attention softmax at float32 (as shown here)

github.com/xjdr-alt/ent...
November 24, 2024 at 5:23 PM
I'm very excited to share that NVIDIA just released Nemotron-3-Embed: two multilingual embedding models for retrieval

The 8B takes the top spot on RTEB.

- Nemotron-3-Embed-1B-BF16 (2048-dim)
- Nemotron-3-Embed-8B-BF16 (4096-dim)

Both OpenMDW-1.1, ready for commercial use.

🧵
July 16, 2026 at 4:06 PM
as of last night, I'm able to compress BF16 weights by 30%
September 16, 2026 at 5:56 PM
BF16周年
クリムゾンクライシス16周年おめでとう!!!!!
November 14, 2024 at 10:14 PM
bf16 is a better nonlinearity
May 15, 2026 at 3:14 AM
never mind i just changed the default epsilon from 1e-8 to 1e-5 because i think that bf16 training is causing pink and green color artifacts
there is nothing which makes me feel like a more unprincipled mathematical pervert than resorting to changing an optimizer's betas
June 27, 2025 at 5:02 AM
"doesn't that mean that you've quanted everything to shit?" no. :)

bf16, more or less.
August 15, 2026 at 3:35 PM
BF16+Dvarw16
February 17, 2025 at 12:23 AM
rank one scale envelope with two mantissae. dequants to bf16 with shifts, adds, and one GEMM transfer. oh yeah. nnh.
August 22, 2026 at 8:59 PM
these graphs are nuts

bf16 has been the only way training has been done for nearly a decade

all this year tons of resources have been dumped into RL, and this is saying most of that was wasted bc we chose the “cool” float format
astonishing: using fp16 instead of bf16 results in more stable training runs as well as a smaller performance gap between training & inference

this is critical for RL, which is mostly inference and very sensitive to reproducible results
October 31, 2025 at 9:28 PM
Alibaba Ant's Ling-3.0-tiny is now available as open-weight:

BF16: huggingface.co/inclusionAI/...
FP8: huggingface.co/inclusionAI/...
INT4: huggingface.co/inclusionAI/...
August 10, 2026 at 5:24 PM
405B Base (bf16) will withstand the test of time and remain eternal (18 months)
November 25, 2024 at 7:14 AM
You can try out the newest and greatest Google TPU:

v6e-1 (Trillium) TPUs!! 2x the high bandwidth memory as v5e-1 (32GB) and a whopping peak rating of 918 BF16 TFLOPS (nearly 3x A100)!

on Google Colab.
May 12, 2025 at 12:35 AM
STOP QUANTIZING MY TRUESIGHT INTO BF16
June 20, 2026 at 8:02 PM
deepseek GGUF just dropped, if you have 207GB disk/40GB RAM for the smallest version huggingface.co/collections/...
Deepseek V3 (All Versions) - a unsloth Collection
Deepseek V3 - available in bf16, original, and GGUF formats, with support for 2, 3, 4, 5, 6 and 8-bit quantized versions.
huggingface.co
January 8, 2025 at 12:20 AM
🤮 i need all 53 bits as god intended

none of this disgusting "bf16" or "fp8" crap
August 22, 2026 at 5:30 AM
60GB is very optimistic btw, fable estimates 90-250GBs for BF16 KV cache at 1M token length
June 12, 2026 at 5:14 AM
🔥 Inference weights: huggingface.co/microsoft/bi...
🔥 Training weights (bf16): huggingface.co/microsoft/bi...
🧰 Inference code: github.com/microsoft/bi...
Demo: bitnet-demo.azurewebsites.net
Bitnet
bitnet-demo.azurewebsites.net
April 16, 2025 at 4:18 AM
i tried it with the original pytorch bf16 model and it said "в животе" ????????? мама,,,,,, :((((
September 14, 2026 at 10:36 PM