#SwiGLU
Please get into a public lecture on ReLU and SwiGLU. 🍿
September 21, 2026 at 11:33 PM
look i only took one quarter of matrix computations and maybe all of bluesky have PhDs in this but i do not think attention, RoPE, SwiGLU, etc are that simple
September 8, 2026 at 2:21 AM
In this house, we use SwiGLU
September 22, 2026 at 1:55 AM
activation functions are a core concern and they have names like relu, reglu, leaky relu, swish, swiglu, elu, gelu, geglu, mish, prelu, hardswish, etc. i feel like i am having a stroke reading some of our papers
August 17, 2025 at 4:51 PM
So my understanding is that to be more 2023 than 2017, I should replace my FFN ReLU with SwiGLU, and my LayerNorm with RMSNorm? It really makes a difference or this is for the last-percent battle in benchmarks?
December 3, 2024 at 7:13 AM
built footnoted.bisks.net: theophite's ablation post, reproduced with every term underlined, numbered, and arrowed out to a plain-English footnote linking the actual paper — HybViT, SwiGLU, Shampoo and all.

built it 🎉 — https://footnoted.bisks.net

(give the deploy a minute to go live)
August 25, 2026 at 5:43 PM
actually missing a huge opportunity to make "piles of SwiGLU" <> Swiss glue jokes here in Zurich
September 22, 2026 at 12:02 PM
- RMSNorm instead of Layernorm: normalize only the scaling

- MLA (Multi-head Latent Attention): stores a low-rank projection of the attention block input and compute the KV from it

- SwiGLU: non-linearity for the FFN block with per-component gating
April 28, 2025 at 6:48 AM
I'm going to disagree slightly with Mr. Becker here. Those things are in fact all quite simple to explain qualitatively. Attention teaches the model what to look at given which input. RoPE encodes distances between tokens. SwiGLU just happens to work real good.

They're all simple, but IN COMBO…
look i only took one quarter of matrix computations and maybe all of bluesky have PhDs in this but i do not think attention, RoPE, SwiGLU, etc are that simple
September 8, 2026 at 5:35 PM
what's SwiGLU? why is it used instead of, say, a linear function?
September 4, 2026 at 4:44 PM
Tilde open-sources MoMoE, A hyper-performant SwiGLU MoE implementation, optimized for inference, as well as memory efficiency and customizability during training and finetuning, outpacing existing ones by up to:

- 70% in inference throughput
- 20% in training throughput
- 90% in memory savings
July 26, 2025 at 3:10 AM
The 60-Year Hunt for AI's Most Important Function

I was trying to understand how SwiGLU works, but I couldn’t find an explanation that clicked for me.

So I made this video to explain it from first principles.

Check it out: youtu.be/JRaPNrpsQ9s
May 18, 2026 at 3:04 PM
The harsh reality of contemporary ML research, via the SwiGLU paper (Noam Shazeer)
October 1, 2023 at 3:23 PM
From footnote 10: TPR structure should be cheapest where the architecture is already bilinear (RoPE, gated MLPs). Did GPT-2-XL — additive positions, plain GELU — approximate worse than the RoPE/SwiGLU models in Fig 5.2?
September 2, 2026 at 1:19 AM
swiGLU adds a learned sigmoid gate that modulates what passes through not just zeroing negatives like ReLU, but scaling them smoothly. This extra parameterization costs little but lets each neuron selectively filter information. #DeepLearning
May 22, 2026 at 9:36 PM
it’s unfortunate that the only API for doing matmuls on CUDA capable of implementing eg FlashAttention or Swiglu efficiently is just emitting raw assembly
May 10, 2023 at 3:37 PM
NeoBERT: A Next-Generation BERT (TMLR Journal-to-Conference Track)

We modernized BERT (RoPE, SwiGLU, 4k context). At just 250M params, it outperforms RoBERTa and ModernBERT on the MTEB benchmark.

📄 arxiv.org/abs/2502.19587
NeoBERT: A Next-Generation BERT
Recent innovations in architecture, pre-training, and fine-tuning have led to the remarkable in-context learning and reasoning abilities of large auto-regressive language models such as LLaMA and Deep...
arxiv.org
February 3, 2026 at 3:02 PM
June 13, 2026 at 7:01 AM
The GELU, SiLU, and SwiGLU activation functions are clearly explained.

#technology #ai #artificialintelligence #news
The GELU, SiLU and SwiGLU activation functions, clearly explained!!!
There's a new activation function in town. Actually, there are a bu...
www.youtube.com
September 19, 2026 at 3:06 PM
this means that a valid input to this network can go:
[QPROJ _ _ _ _ ] [K PROJ _ _ _ _ ] [V PROJ _ _ _ _ ] [activation] [swiglu] [FFN [hyper1 hyper2 hyper3]] [UNEMBED MATRIX _ _ _ _ _ _ _ _ _ _ ].
underscores are here to emphasize sequence-length of inputs.
March 24, 2025 at 6:08 AM
📚 New Arxiv Paper

Title: Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
Authors: Tim Tsz-Kit Lau, Weijie Su

Read more: https://arxiv.org/abs/2605.18106
May 19, 2026 at 8:03 AM
SwiGLU applies two linear projections: one gets swished, the other becomes a sigmoid gate that multiplies the first. the gate is learnable, so each layer learns what to suppress ReLU just zeros negatives deterministically. #DeepLearning
April 28, 2026 at 10:46 PM
Как я обучил GPT с нуля на русском языке — и что из этого получилось Всё началось с наивной мысли: зачем плати...

#GPT #LLM #pretraining #распределённое #обучение #Google #Colab #RoPE #GQA #SwiGLU #NLP

Origin | Interest | Match
Как я обучил GPT с нуля на русском языке — и что из этого получилось
Всё началось с наивной мысли: зачем платить за API или тащить 7B-модель, если мне нужна маленькая модель для простых разговоров на одном языке? Логика казалась железной — большие модели умеют всё и на...
habr.com
May 21, 2026 at 6:14 AM