built it 🎉 — https://footnoted.bisks.net
(give the deploy a minute to go live)
built it 🎉 — https://footnoted.bisks.net
(give the deploy a minute to go live)
- MLA (Multi-head Latent Attention): stores a low-rank projection of the attention block input and compute the KV from it
- SwiGLU: non-linearity for the FFN block with per-component gating
- MLA (Multi-head Latent Attention): stores a low-rank projection of the attention block input and compute the KV from it
- SwiGLU: non-linearity for the FFN block with per-component gating
They're all simple, but IN COMBO…
They're all simple, but IN COMBO…
- 70% in inference throughput
- 20% in training throughput
- 90% in memory savings
- 70% in inference throughput
- 20% in training throughput
- 90% in memory savings
I was trying to understand how SwiGLU works, but I couldn’t find an explanation that clicked for me.
So I made this video to explain it from first principles.
Check it out: youtu.be/JRaPNrpsQ9s
I was trying to understand how SwiGLU works, but I couldn’t find an explanation that clicked for me.
So I made this video to explain it from first principles.
Check it out: youtu.be/JRaPNrpsQ9s
We modernized BERT (RoPE, SwiGLU, 4k context). At just 250M params, it outperforms RoBERTa and ModernBERT on the MTEB benchmark.
📄 arxiv.org/abs/2502.19587
We modernized BERT (RoPE, SwiGLU, 4k context). At just 250M params, it outperforms RoBERTa and ModernBERT on the MTEB benchmark.
📄 arxiv.org/abs/2502.19587
#technology #ai #artificialintelligence #news
#technology #ai #artificialintelligence #news
[QPROJ _ _ _ _ ] [K PROJ _ _ _ _ ] [V PROJ _ _ _ _ ] [activation] [swiglu] [FFN [hyper1 hyper2 hyper3]] [UNEMBED MATRIX _ _ _ _ _ _ _ _ _ _ ].
underscores are here to emphasize sequence-length of inputs.
[QPROJ _ _ _ _ ] [K PROJ _ _ _ _ ] [V PROJ _ _ _ _ ] [activation] [swiglu] [FFN [hyper1 hyper2 hyper3]] [UNEMBED MATRIX _ _ _ _ _ _ _ _ _ _ ].
underscores are here to emphasize sequence-length of inputs.
Title: Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
Authors: Tim Tsz-Kit Lau, Weijie Su
Read more: https://arxiv.org/abs/2605.18106
Title: Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
Authors: Tim Tsz-Kit Lau, Weijie Su
Read more: https://arxiv.org/abs/2605.18106
#GPT #LLM #pretraining #распределённое #обучение #Google #Colab #RoPE #GQA #SwiGLU #NLP
Origin | Interest | Match