#MatMul
matmul users are so cringe. like, learn a real algorithm??? just because you *can* reduce any problem to matmul doesn't mean you should
August 5, 2025 at 12:14 AM
"Inside NVIDIA GPUs: Anatomy of high performance matmul kernels", includes a great intro to GPU architecture and PTX/SASS: www.aleksagordic.com/blog/matmul
Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordić
From GPU architecture and PTX/SASS to warp-tiling and deep asynchronous tensor core pipelines.
www.aleksagordic.com
October 1, 2025 at 7:36 PM
i think i'm just biased in favor of matmul
yeah orwell predicted this. the expansion of state surveillance? no what are you talking about. i mean the kalecki socialists are much smarter and cooler than the bukharin socialists
January 20, 2025 at 6:54 AM
basically anything you can imagine doing with bulk compute over the next fifteen years is more or less matmul
January 27, 2025 at 8:33 PM
Exa’s vector DB implementation is super interesting

unlike HNSW (standard algo), they can partition very well

they aggressively quantize (1-bit) and do matmul on CPU via lookup tables loaded into CPU registers

finally do real matmul on small data. Absolutely brilliant
April 28, 2026 at 10:52 PM
software companies attempting to build a moat around matmul while hardware companies are investing heavily in commoditizing matmul.

someone who is good at the economy please help me budget this my business venture is dying!
February 24, 2025 at 10:56 AM
What the fuck do you mean that loss is insensitive to width-vs-depth as long as you do the same amount of matmul. fuck you. fuck you
September 18, 2026 at 5:10 PM
i don't think there's any world in which big tech says "i think we bought too much matmul" even if LLMs go tits up
January 27, 2025 at 8:34 PM
things we know about LLMs and large DL models in general:

- how they are trained (gradient descent)
- the structure into which they are placed (architecture)
- the base arithmetic (matmul, norm, batch norm, and so on)
as a girl with a PhD in natural language processing and machine learning it's actually offensive to me when you say "we don't know how LLMs work so they might be conscious"

I didn't spend 10 years in mines of academia to be told ignorance is morally equal knowledge.

We know exactly how LLMs work.
October 5, 2025 at 1:22 AM
everything is matmul
October 14, 2025 at 10:08 PM
separately there is a brief accusation that rats intend to do Aum Shinrikyo 2 Matmul Edition. just so no one misses that
September 24, 2026 at 5:30 PM
Inside NVIDIA GPUs: Anatomy of high performance matmul kernels

To deeply understand how one writes state of the art NVIDIA GPU matrix-multiplication (matmul) kernels in CUDA

www.aleksagordic.com/blog/matmul
September 30, 2025 at 1:24 PM
AGI is hidden inside the commutative matmul 😔
November 8, 2025 at 7:13 AM
honestly trying to fuck matmul is probably the most reasonable response to *motions at everything*
ed3d.net Ed @ed3d.net · 11d
??? the rust programmers are the most normal people i know, at least the ones not trying to have carnal relations with the matmul machine
September 15, 2026 at 3:04 AM
??? the rust programmers are the most normal people i know, at least the ones not trying to have carnal relations with the matmul machine
September 15, 2026 at 3:00 AM
so now there's more hardware shipping with beefy matmul accelerators (for "AI"), is there a way to use that for cryptography?
September 8, 2024 at 1:01 PM
everyone is using the same CLIP from a million years ago because no one wants to do a century of matmul at CPU speed
If CLIP is so great why haven't they made a CLIP 2
March 4, 2025 at 2:33 AM
"carnal relations with the matmul machine" wheres the bleach
September 15, 2026 at 3:02 AM
"It's just matmul" you can literally play doom on human neurons
August 23, 2026 at 5:04 PM
Llamafile doesn’t require GPUs and works on most architectures justine.lol/matmul/
LLaMA Now Goes Faster on CPUs
I wrote 84 new matmul kernels to improve llamafile CPU performance.
justine.lol
December 26, 2024 at 8:55 PM
FlashAttention-4

I hope it is not pain to work with. It changes the algorithm & pipeline so that softmax & SMEM bandwidth no longer dictate speed. Attn reaches ~1600 TFLOPs, pretty much at matmul speed!
March 5, 2026 at 6:47 PM
you are not obligated to complete the matmul but neither are you free to abandon it
May 21, 2026 at 8:05 PM
#picotron 0.2.1c is up on lexaloffle & humble: www.picotron.net

- tracker fx: fade, oscillate, fast arps
- batch matmul() // demo: treegen.p64
- fast tline3d / spr / sspr / map
- external .p64 change reloading
- trinkets wallpaper

full changelog: www.lexaloffle.com/dl/docs/pico...
November 1, 2025 at 2:22 AM
So, there was this guy who took Karpathy's code for llama2.c and made it work on his powerbook G4, just for fun (2 tok/s). With the help of Gemini2, he transformed the pure C matmul function into Altivec code (SIMD instruction for the powerpc G4, 30% faster) and I decided to play along as well ../
April 6, 2025 at 9:39 PM
Also this is all matmul, which even if AI completely explodes, you can never have enough matmul lol
February 10, 2025 at 1:41 AM