not sure what to do with this discovery but there you go
not sure what to do with this discovery but there you go
embedding-space.github.io/sparse-netwo...
The subject is WHY neural networks work, and I think the answer I offer is kind of interesting. Maybe even a little correct, possibly.
embedding-space.github.io/sparse-netwo...
The subject is WHY neural networks work, and I think the answer I offer is kind of interesting. Maybe even a little correct, possibly.
emilhvitfeldt.com/talk/2025-08...
emilhvitfeldt.com/talk/2025-08...
absolutely incredible, and also someone needs to stop these people
arxiv.org/abs/2507.23186
absolutely incredible, and also someone needs to stop these people
arxiv.org/abs/2507.23186
emilhvitfeldt.github.io/talk-slc-spa...
emilhvitfeldt.github.io/talk-slc-spa...
LLMs are neurosymbolic models. the structure of the representation space (specifically, sparsity + orthogonalization) creates an explicit albeit opaque learned ontology. leakages from that learned ontology are processed by things which look like PCA.
LLMs are neurosymbolic models. the structure of the representation space (specifically, sparsity + orthogonalization) creates an explicit albeit opaque learned ontology. leakages from that learned ontology are processed by things which look like PCA.
Our latest work with NVIDIA introduces new CUDA kernels & data formats for faster inference and training of sparse transformer language models:
Blog: pub.sakana.ai/sparser-fast...
This work introduces new open-source GPU kernels and data formats for faster inference and training of sparse transformer LLMs:
🧵 Thread 👇
Our latest work with NVIDIA introduces new CUDA kernels & data formats for faster inference and training of sparse transformer language models:
Blog: pub.sakana.ai/sparser-fast...
We now support sparse data during the whole process, generate data sparsely in recipes steps when your model supports spare data structures. and you don't have to change anything!
www.tidyverse.org/blog/2025/03...
We now support sparse data during the whole process, generate data sparsely in recipes steps when your model supports spare data structures. and you don't have to change anything!
www.tidyverse.org/blog/2025/03...
Piotr Nawrot!
A repo & notebook on sparse attention for efficient LLM inference: github.com/PiotrNawrot/...
This will also feature in my #NeurIPS 2024 tutorial "Dynamic Sparsity in ML" with André Martins: dynamic-sparsity.github.io Stay tuned!
Piotr Nawrot!
A repo & notebook on sparse attention for efficient LLM inference: github.com/PiotrNawrot/...
This will also feature in my #NeurIPS 2024 tutorial "Dynamic Sparsity in ML" with André Martins: dynamic-sparsity.github.io Stay tuned!
this “blog” is such an amazing way to announce your work. i imagine it’s all LLM generated, but wow, so incredibly clear
reminds me of one of those Financial Times infographic stories
pumpkin-co.github.io/SparsityAndC...
this “blog” is such an amazing way to announce your work. i imagine it’s all LLM generated, but wow, so incredibly clear
reminds me of one of those Financial Times infographic stories
pumpkin-co.github.io/SparsityAndC...
Language Models: A Guide for the Perplexed. (arXiv:2311.17301v1 [cs.CL])
http://arxiv.org/abs/2311.17301