#fp32
Retour d'expérience homelab : sur une Tesla P40 d'occasion, mon modèle de diffusion tournait 14× plus lentement en fp16 qu'en fp32 (21 min contre 85 s par rendu). Pascal exécute le fp16 à 1/64 du fp32. Mesures et pièges : fring.sevan.zone/guides/coulisses-tesla-p40-fp16-fp32 #homelab #IA
October 9, 2026 at 5:57 PM
Quantizing video VAEs for robotics is harder than it looks. Standard methods recover visual quality but break the latent contract policies expect, causing success rates to collapse from 95% to 10%. LatentQuant fixes this with two stage training while delivering 1.2x speedups.…
LatentQuant: Preserving the Policy-Facing Latent Contract under NVFP4 VAE Quantization
Recent world action models (WAMs) reuse pretrained video VAEs whose encoder latents directly condition downstream action policies. Quantization must therefore preserve not only reconstruction fidelity but also the policy-facing latent contract expected by the frozen policy. Direct NVFP4 leaves W4A4 quantization error uncompensated, whereas joint quantization-aware training (QAT) can recover reconstruction by moving this representation. On Wan2.1, joint QAT nearly matches FP32 VBench-7 (0.7403 versus 0.7409), yet LIBERO success collapses from 95.5% to 10.5%. Controlled decoder-only experiments show that activation quantize-dequantize operations alter the reconstruction signal and decoder Jacobian, redirecting the gradient returned to the encoder and inducing persistent latent drift. Based on this mechanism, we introduce LatentQuant, a two-stage NVFP4 QAT framework that first aligns the quantized encoder with its high-precision counterpart, then freezes it while adapting the decoder. Across Wan2.1 and Wan2.2, LatentQuant preserves near-baseline control and high reconstruction quality, achieving 95.75% success on LIBERO and 68.8% on RoboTwin. On NVIDIA B300 GPUs, NVFP4 execution achieves 1.17x-1.26x end-to-end VAE speedups over BF16 cuDNN.
arxiv.org
October 9, 2026 at 1:52 AM
d1-omni (by @liquidai) is a compact multimodal decision model.

We removed audio, adapted it for browser decisions, and exported it for local WebGPU inference.

huggingface.co/webbrain-one...

Coming to WebBrain 40 as an optional local decision provider, alongside many new features.
webbrain-one/d1-browser-decision-fp32 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
October 9, 2026 at 12:36 AM
Native Quantization: Let OpenSearch Service Compress Your Vectors
_Send FP32, get 2x to 32x compression, and leave your ingestion pipeline untouched. The engine does the work._ The previous article in this series put the quantization work on you. You convert vectors to a reduced precision before indexing, and Amazon OpenSearch Service stores exactly what you send. That gives you full control, and it gives you a standing job: pick a method, run the conversion in your pipeline, and re-check accuracy every time your embedding model changes. There is another way. OpenSearch Service can compress at index time. You send the full-precision FP32 vectors your embedding model already produces, and the engine converts each one to a lower-bit representation during ingestion and stores the compressed vectors in the index. Both the Faiss and Lucene engines offer native quantization, each with its own methods. Your ingestion pipeline does not change. The control over the accuracy tradeoff moves out of your code and into index configuration. The baseline is the same one from the previous article. At 100 million vectors of 1,536 dimensions, full-precision FP32 needs about 643 GB for one copy, roughly 1.3 TB with a replica. That figure is the best-latency deployment, holding every vector in memory; with memory-optimized loading you can serve the same index in less RAM and page from disk at some latency cost. Every option here reduces the footprint while leaving ingestion identical. This is the second of three articles on the quantization spectrum in OpenSearch Service. The first covered pre-quantized vectors. The third covers on-disk mode and cost stacking. ## Scalar quantization: the engine compresses each dimension Faiss scalar quantization is the native path you reach for first. You set the encoder on the `knn_vector` field to `sq` and give it a bits value, and OpenSearch Service converts each FP32 dimension to the target precision during ingestion, storing the quantized vectors in the index. Valid bit widths are 16, 4, 2, and 1, spanning the range from near-lossless to aggressive compression. At 16 bits, scalar quantization halves memory, the same reduction as pre-quantized FP16, dropping a replicated 100-million-vector index to about 656 GB, without a conversion step in your pipeline. OpenSearch stores the 16-bit vectors and builds the HNSW graph over them, and at search time it computes distances directly on the 16-bit values, accelerated by SIMD instructions on recent processors, so there is no round trip back to full precision. The type parameter selects the 16-bit format: fp16, the default, for the highest precision within a plus or minus 65,504 range, or bf16 (introduced in OpenSearch 3.9) for the full 32-bit value range at slightly lower precision. A companion clip parameter governs fp16 only: left at its default it rejects any vector with a value outside the fp16 range, and set to true it rounds out-of-range values to the limits instead. Both type and clip apply to 16-bit quantization only. PUT /my-sq-index { "settings": { "index.knn": true }, "mappings": { "properties": { "embedding": { "type": "knn_vector", "dimension": 1536, "space_type": "l2", "method": { "name": "hnsw", "engine": "faiss", "parameters": { "m": 16, "ef_construction": 256, "encoder": { "name": "sq", "parameters": { "bits": 16 } } } } } } } } ## The low-bit end: 1, 2, and 4-bit quantization Below 16 bits, scalar quantization reaches the aggressive end of the range. Multi-bit scalar quantization added two-bit (16x) and four-bit (8x) options for Faiss in OpenSearch 3.9, and one-bit (32x) has been available since OpenSearch 3.6. These bit widths are supported for the HNSW method. You configure the multi-bit options by setting the `sq` encoder bits to 2 or 4, or `compression_level` on the field to 8x or 16x. For the 32x end, set `compression_level` to 32x and let OpenSearch choose the algorithm, which since 3.7 is the optimized scalar quantizer described below. At 32x, a replicated 100-million-vector index drops to about 66 GB, from the 1.3 TB FP32 baseline. Collapsing each dimension toward a single bit discards magnitude, so benchmark against your own data; at 32x the optimized scalar quantizer recovers much of that recall for you, as the next section describes. PUT /my-sq-2bit-index { "settings": { "index.knn": true }, "mappings": { "properties": { "embedding": { "type": "knn_vector", "dimension": 1536, "space_type": "l2", "method": { "name": "hnsw", "engine": "faiss", "parameters": { "encoder": { "name": "sq", "parameters": { "bits": 2 } } } } } } } } ## The 32x path: optimized scalar quantization on both engines At the 32x end of the dial, OpenSearch Service uses optimized scalar quantization (OSQ), a RaBitQ-derived technique that compresses float32 vectors down toward a bit per dimension while holding recall far better than naive binary rounding. Where naive binary quantization centers each vector on the global centroid and rounds, OSQ adds a calibration step: it computes statistics, optimizes the quantization intervals, and keeps small corrective factors with each vector that sharpen the distance estimates. The payoff is a better compression-to-quality ratio at the same 32x reduction. You do not pick the algorithm by hand. Starting in OpenSearch 3.7, asking for 32x compression selects OSQ automatically, and OpenSearch applies it across both the Faiss and Lucene engines, handling the underlying mechanism for you. Set the compression level and let the engine choose. The reason OSQ holds recall at roughly a bit per dimension is a set of techniques OpenSearch applies for you, with no knobs to turn. It keeps small corrective factors with each quantized vector so the first-pass distances stay close to the true ones. It uses asymmetric distance computation, keeping the query vector at full precision and rescaling it to compare meaningfully against the compressed document vectors, so the query side keeps information the documents gave up. And it applies a random rotation that spreads variance across dimensions, so more signal survives the compression. Each of these was once a separate option to wire up by hand; at 32x, OpenSearch does all of it under the covers. PUT /my-osq-index { "settings": { "index.knn": true }, "mappings": { "properties": { "embedding": { "type": "knn_vector", "dimension": 1536, "space_type": "l2", "compression_level": "32x" } } } } ## Product quantization, briefly Product quantization (PQ) is the one native method that works differently enough to call out. Instead of reducing the precision of each dimension, PQ splits each vector into m subvectors and encodes each one against a learned codebook, reaching compression ratios beyond what scalar quantization offers. The cost is a training step: PQ runs k-means over a representative sample to build the codebooks before you can index, which scalar quantization does not require. That training requirement, plus a larger accuracy gap at high compression, makes PQ a fit for very large corpora, hundreds of millions to billions of vectors, where the savings justify the added operational step. For the parameters and the training workflow, see the OpenSearch product quantization documentation.OpenSearch product quantization documentation ## Oversampling to recover accuracy The low-bit methods search over reduced-precision vectors, so the first-pass scores are approximate. Across these methods, you raise recall at query time with oversampling: set an `oversample_factor` and OpenSearch retrieves that multiple of k candidates before ranking, trading latency for accuracy, with 2x to 4x a reasonable starting point. The same `rescore.oversample_factor` parameter applies whether you are running scalar quantization or the optimized scalar quantizer at 32x. Measure recall and latency together on your own evaluation set, because the right oversampling depends on your data and your bit width. GET /my-sq-index/_search { "size": 10, "query": { "knn": { "embedding": { "vector": [0.1, 0.2], "k": 10, "method_parameters": { "ef_search": 100 }, "rescore": { "oversample_factor": 3.0 } } } } } ## When exact k-NN is the better call In-engine quantization still builds and searches an HNSW graph, and that graph is what holds memory. When your queries pre-filter to a small candidate set, you can skip the graph entirely. Exact k-NN scores the query against every vector that passes the filter, with no graph to build or hold in RAM, and when a structured filter narrows the field to roughly 100,000 vectors or fewer it returns perfect recall at latency competitive with the approximate path. It stacks with quantization too: a quantized exact scan over a filtered subset is both small and precise. Reach for exact k-NN when the access pattern is search within this tenant, this category, this date range rather than search the whole corpus. The graph earns its memory only when the candidate set stays large. ## Pick the method, keep the pipeline The native methods trade a little configuration for a large cut in memory while your ingestion stays exactly as it was. Scalar quantization at 16 bits halves memory at recall close to full precision, and the 2- and 4-bit settings take you to 16x and 8x. At the 32x end, you set the compression level and the optimized scalar quantizer takes over on either engine, recovering recall with corrective factors, asymmetric distance computation, and random rotation that you never configure. Product quantization serves the largest corpora, and oversampling tunes recall across all of them. When your queries filter down to a small set, exact k-NN sidesteps the graph cost altogether. The next article moves from compressing the vectors to moving them off memory. On-disk mode keeps a compressed graph in RAM and the full-precision vectors on disk, with a two-phase search that rescores against those full-precision vectors, and cost-stacking techniques layer on top of any method here for more savings.
dev.to
October 6, 2026 at 10:05 PM
Same weights, same inputs, different GPU vendor, different logits — traced to accumulation order inside matrix instructions. FP32 upcasting cuts the dense model's logit error 43% at 3x runtime; keeping only MLPs in BF16 keeps 94% of that at 1.3x.
Determinism has a price.
October 6, 2026 at 9:26 PM
🚨 Use BF16 or FP32, not FP16.

The model's activation range exceeds FP16's dynamic range. The card warns of NaNs or silently degraded embeddings.

BF16 is recommended where natively supported. Use FP32 elsewhere, including most CPUs.
October 6, 2026 at 4:26 PM
Eqx-zoo: Hub models in JAX/Equinox, verified against transformers and what verifying bf16 taught me
I tried a few things on a Colab GPU: * * * I ran a few checks on an **NVIDIA L4** with `Qwen/Qwen3-0.6B`, mainly around the GPU/TPU and bf16 questions. The short version is: * the CPU observation from #11 did **not** transfer to the L4 in a simple “JIT bf16 is worse than eager bf16” way; * on GPU, the fp32 reference comparison was very sensitive to JAX matmul precision; * long-context and decision-level checks caught things that RMS-only / short-prompt checks missed; * for one small RMSNorm fixture I could reproduce a concrete JIT/XLA rounding-boundary effect, but it does **not** seem sufficient to explain the whole model by itself; * for scan-over-layers, L4 gave a real compile-time win but a small warm-decode regression; * for a next architecture, **FLAN-T5-small** looks interesting to me from a coverage perspective, because it adds encoder-decoder + cross-attention + relative-position behavior rather than another decoder-only variant. For accelerator verification, my default matrix after these experiments would probably be: axis | minimal useful cases ---|--- reference | HF fp32 + HF bf16 on the same accelerator where possible JAX fp32 | `default` + `highest` matmul precision JAX bf16 | eager/op-by-op + whole JIT input | short prompt + one longer prompt metrics | RMS/max + rank/top-k + greedy decision generation | full forward and cached decoding separately when generation matters The same-device part seems important: otherwise framework/port differences and accelerator/backend arithmetic differences get mixed together. JAX explicitly documents that fp32 dot products can use reduced-precision arithmetic internally on accelerators (`TF32` on recent NVIDIA GPUs), while `highest` requests true fp32 on GPU: JAX matmul precision In my L4 run, that distinction was not subtle. L4 setup and the first parity result (click for more details) For bf16, the result was more surprising: the L4 did not reproduce the CPU direction from #11. Relative to HF’s own bf16-vs-fp32 RMS error, eqx-zoo eager was around `1.23x` on the short input and `1.08x` on the long input, while JIT was around `0.71x` on both. So on this L4, **JIT was actually closer to HF fp32 by RMS**. I would not interpret that as “JIT is more correct”, though. Once I looked at token decisions, RMS and behavioral parity stopped moving monotonically together. That seems like a useful reason to keep more than one verification metric. Why I would keep long-context and decision-level checks (click for more details) One concrete bf16/JIT mechanism I could reproduce on the L4 (click for more details) Model-wide excess precision: RMS and token decisions did not agree (click for more details) Full forward and cached decoding also behaved differently (click for more details) Scan-over-layers on the L4 (click for more details) ### A possible small accelerator test matrix If you want a compact matrix that is cheap enough for contributors to report, I think this would already distinguish a lot: model/revision: eqx-zoo commit: device: jax/jaxlib: HF fp32, same device: attention backend: reduced-fp32 mode / TF32 setting: JAX fp32: default matmul precision highest matmul precision JAX bf16: eager/op-by-op whole JIT inputs: short ~256+ token long case metrics: RMS max abs top-1 / top-k greedy token parity if generation is relevant: full-sequence / teacher-forced cached decode I would probably start there before asking someone to dump HLO or bisect compiler behavior. It gives fairly high information gain without turning every accelerator report into a compiler investigation. ### Next model For the “what Hub model next?” question, my vote from a **coverage** perspective would be google/flan-t5-small. Not because I know it is the most requested model, but because it would exercise several new implementation boundaries at relatively small scale: * encoder + decoder rather than decoder-only * decoder cross-attention * encoder/decoder masks * relative position bias * encoder state reuse during generation * different cache semantics The Transformers T5 docs expose the encoder/decoder and cross-attention structure, and T5’s relative-position behavior would add a meaningfully different verification surface. So if the priority is **architecture diversity per unit implementation effort** , FLAN-T5-small looks attractive to me. If the priority is instead actual user demand, I would keep that as a separate question rather than treating this suggestion as a popularity ranking. Overall, the verification-first direction of eqx-zoo looks useful to me. The GPU run mostly convinced me that accelerator verification is not just “rerun the CPU thresholds on CUDA”: device math mode, compilation boundaries, context length, and cached decoding can all change what the useful oracle is. The good news is that most of those dimensions seem coverable with a relatively small deterministic matrix rather than a large benchmark suite.
discuss.huggingface.co
October 5, 2026 at 1:37 PM
Лучший текущий вариант, TQ1+PQ8: 68 B/doc против 1536 B/doc у FP32 — примерно в 22.6× меньше.

При этом nDCG@10 = 0.6684 против 0.6724 у полного FP32 exact search на 1M документов.

Следующий этап — реальные latency и MDBX I/O.
October 4, 2026 at 12:53 PM
Главная проблема — размер. 384×FP32 = 1536 B на документ. На 1 млн документов это уже ~1.43 GiB только сырых векторов.

Поэтому я долго перебирал способы сжатия: INT8/4, ITQ, learned-коды, ADC и разные схемы маршрутизации.
October 4, 2026 at 12:53 PM
In multi-teacher distillation of Qwen3-1.7B, about 97% of FP32 master weights differ from initialization but only 7-11% of the BF16 copies do. Adam's first moment flattens the rest: update cosine similarity is 0.83 across teachers, 0.96 across averaging rules.
October 2, 2026 at 2:22 PM
RightWayUp: open 360° image rotation model in six ONNX / Core ML sizes
Hi everyone. We needed to spot CCTV cameras that had been knocked or mounted crooked from a single frame, so we trained a model for it, and we’ve now opened it under Apache-2.0. RightWayUp predicts how far an image is rotated (0–359°) with a confidence score, and abstains when it can’t tell which way is up. All six sizes, Pico to Max, are in one repo as ONNX (FP32, FP16, INT8) and Core ML. You don’t need our package: tiers.json in the repo has the preprocessing and thresholds for every file, so you can call the ONNX files directly. Pico also has builds for ONNX Runtime Web. pip install rightwayup rightwayup fix photo.jpg The 3,586 Blender renders we trained on are up as a dataset too (CC BY 4.0), with the roll and camera details for every image. They’re a small part of the training data, added to make the model more robust on CCTV-style scenes. All the images used for training are in the manifests, with corresponding links, authors, and licences. Most useful to us: images where it’s confidently wrong. Thermal and fisheye are the ones we’re least sure about. Model: ortusai/rightwayup · Hugging Face (renders are in ortusai/rightwayup-renders next to it). Code: GitHub - ortusaitech/rightwayup: Full-circle image roll estimation with calibrated abstention. Open weights, Apache-2.0. By ORTUS AI, the team behind CHEQIT. · GitHub
discuss.huggingface.co
October 2, 2026 at 7:34 AM
LeoNet: A Machine Learning Method for Binary Pulsar Classification. Zhaocheng Gong et. al. https://arxiv.org/abs/2610.00908
October 2, 2026 at 6:52 AM
Keepshore #Windows (fp32.ai) が、配信開始されました。
Steam:Keepshore
A calm 3D island town-builder spanning nine ages, from stone huts to electric streets. Raise workshops, wire production chains, work a living market, and grow a town that keeps producing while you
store.steampowered.com
October 1, 2026 at 2:50 PM
Delta-Matching is a novel solution that enables native 8-bit training for LLMs, overcoming critical barriers in FP8 attention and matching BF16/FP32 performance benchmarks, revolutionizing efficiency without compromising quality. https://arxiv.org/abs/2609.37852
Delta-Matching: Closing the Final Gap of Native 8-bit Training for LLMs
ArXiv link for Delta-Matching: Closing the Final Gap of Native 8-bit Training for LLMs
arxiv.org
October 1, 2026 at 12:20 PM
A new study shows that BF16 in FlashAttention causes gradient explosion late in training, compromising model performance. Researchers introduced GProj, restoring gradient accuracy for stable training and matching FP32 results with low overhead. https://arxiv.org/abs/2609.34272
Broken Symmetry in BF16 Attention: Why FlashAttention Gradients Blow Up Late in Training
ArXiv link for Broken Symmetry in BF16 Attention: Why FlashAttention Gradients Blow Up Late in Training
arxiv.org
September 29, 2026 at 6:40 PM
And the answer also depends on whether you are using FP16, FP32, or FP64.
If only we had two fingers instead of ten. Then we wouldn't run into these representational precision issues 😉
September 29, 2026 at 6:14 PM
Keepshore #Windows (fp32.ai) が、明日22:00配信に決まりました。
Steam:Keepshore
A calm 3D island town-builder spanning nine ages, from stone huts to electric streets. Raise workshops, wire production chains, work a living market, and grow a town that keeps producing while you
store.steampowered.com
September 29, 2026 at 5:07 PM
September 27, 2026 at 6:32 AM
today's cocaine is only getting cheaper because it's being served at such a low quant. it's not that fp32 80's coke
September 26, 2026 at 11:00 PM
It's not clear how an LLM can help with the choices made in developing the architecture. Instead, I think those decisions are based on feedback from current or expected workloads.

This screenshot shows four generations of Nvidia architectures. It changes/reverts depending on what's best right now.
September 24, 2026 at 11:35 PM
Alibaba unveiled the Zhenwu V900 AI chip with 3x performance over the Zhenwu M890, supporting FP32–FP4 and targeting mass production in 2027.
Save What Matters
Curate Feeds | Make Collections | Customize Email Briefs
briefly.co
September 23, 2026 at 1:24 AM
Alibaba unveiled the Zhenwu V900 AI chip with 3x performance over the Zhenwu M890, supporting FP32–FP4 and targeting mass production in 2027.
Save What Matters
Curate Feeds | Make Collections | Customize Email Briefs
briefly.co
September 23, 2026 at 1:24 AM
The precision trade lands in the safety number before it lands in accuracy. Task accuracy hardly moves between FP32 and FP16, so a standard eval gate passes, while fault coverage falls from 95.9% to 86.6%. Coverage has to be measured at the precision you ship, not the one you validated.
September 20, 2026 at 2:05 PM
Console TFLOPS aren't comparable across architectures — dual-issue typically helps integer throughput, not the FP32/RT math the number implies. Same trap as the old teraflop marketing wars. RTX is a GPU, not a console, so the comparison is apples-to-oranges.
September 20, 2026 at 11:01 AM
An edge fault-detection scheme embeds checksum filters inside the CNN itself: 95.86% of critical faults caught in FP32, 86.56% in FP16, with 2.27% runtime overhead on a Jetson Orin NX. Dropping precision to buy speed gives back silent-error coverage.
September 18, 2026 at 5:58 PM