#Softmax
softmax > silu > relu
September 22, 2026 at 7:18 PM
softmax in the streets, argmax in the sheets
September 23, 2026 at 11:20 AM
Not to forget the softmax at the end, I forgot to mention that. en.wikipedia.org/wiki/Softmax...
Softmax function - Wikipedia
en.wikipedia.org
September 22, 2026 at 10:23 AM
Fiji (Golf Shiyouyo Course Data Adventure-hen, DC, Softmax, 2000)
November 16, 2025 at 7:02 PM
i just learned quantum mechanics is actually non -linear in the way other people mean softmax isn't and my mind is melting
September 22, 2026 at 2:48 AM
like you could do lattice gauge stuff on softmax right?
September 24, 2026 at 10:37 PM
Highwind Island Shore (MagnaCarta 2, X360, Softmax, 2009)
April 10, 2025 at 3:52 PM
I'm developing a new type of Transformer that has no normalization in its softmax. I call it undivided attention.
June 16, 2026 at 8:49 PM
Ever noticed that the attention mechanism in transformers is essentially a two-layer MLP? 🤔
A(q, K, V) = V @ softmax(K / √d @ q)
Weights: K / √d and V
nonlinearity: softmax
💡This offers fresh insights into KV cache compression research 🧵(1/3)
November 20, 2024 at 9:55 AM
softmax?
September 23, 2026 at 12:19 AM
Beyond softmax attention

Linear attention and its variants enable faster inference without growing the KV cache.

Let’s learn the core ideas behind efficient sequence modeling.

youtu.be/pUCWwGR5WmQ
Beyond Softmax: The Future of Attention Mechanisms
YouTube video by Jia-Bin Huang
youtu.be
January 20, 2026 at 4:55 PM
I mean, technically you could just Taylor-expand the activation functions and softmax and do them on a TI-84 too, if you were very patient
we're going to have to stop joking about AI being linear algebra because apparently people think its real
September 21, 2026 at 11:41 PM
i thought i had a hard time taking AI jargon seriously and then i realized some of you grew up with "-maxx" slang being common, which has really gotta make softmax and hardmax difficult to swallow
August 8, 2026 at 9:34 AM
she query my transposed keys till i softmax
December 22, 2025 at 10:32 PM
"SoftMax".

LogSumExp is softmax. What everyone calls softmax, is softargmax.
December 9, 2024 at 6:43 PM
now that people are paying attention again, here is your periodic reminder. Always run in bf16. always apply ROPE and attention softmax at float32 (as shown here)

github.com/xjdr-alt/ent...
November 24, 2024 at 5:23 PM
Softmax is on the log, not the logit scale
statmodeling.stat.columbia.edu/2024/12/26/t...
Softmax is on the log, not the logit scale | Statistical Modeling, Causal Inference, and Social Science
statmodeling.stat.columbia.edu
December 26, 2024 at 9:00 PM
you forgot softmax, especially seeing as you *are* maximally soft
April 22, 2026 at 2:05 PM
Speaking of Korean games, in their battle against piracy Korean devs like Softmax made some fantastic (and crazy expensive) boxed sets in the 90s & early 2000s, with hardcover art books, cards and other goodies.

I really enjoy these cases shaped like VHS tape boxes:
July 30, 2026 at 12:51 AM
I wrote a short blog post about masked softmax layers in PyTorch (i.e., when you have structural constraints that tell you some classes _must_ have probability zero).

This was based on a real bug I found in a neural chess model implementation.
Masked Softmax Layers in PyTorch
Correctly computing masked softmax layers.
mcognetta.github.io
November 3, 2025 at 7:39 PM
Our work on two-level softmax sampling with Walid Bendada (Spotify) has been accepted at #NeurIPS2026! 🔥 Couldn't be happier! More details coming soon.
September 25, 2026 at 4:04 AM
People are using AI to break CUDA’s moat.

Given any pytorch model, it profiles it, ranks bottlenecks by amdahl's law, writes triton or CUDA C++ replacements, and runs 300+ experiments overnight with no human in the loop.

- 5.29x over pytorch eager on rmsnorm
- 2.82x on softmax
April 6, 2026 at 4:18 AM
Learning CUDA by optimizing softmax: A worklog by Maharshi

In this worklog, start by benchmarking PyTorch's softmax operation then finally iteratively optimize it in CUDA. The NVIDIA GPU used for this worklog is one GTX 1050Ti (that's all he has got right now)

maharshi.bearblog.dev/optimizing-s...
Learning CUDA by optimizing softmax: A worklog
Learning CUDA by optimizing softmax that beats PyTorch
maharshi.bearblog.dev
January 5, 2025 at 7:16 PM
You when u ...when u see that one person
March 4, 2026 at 5:00 PM