#MLSys
Very deep dive on Blackwell architecture and performance tuning. Seriously go read it. Even if you exclusively do non-AI graphics, go read Part 1. mlc.ai/modern-gpu-p...
Modern GPU Programming For MLSys — Modern GPU Programming For MLSys
mlc.ai
June 26, 2026 at 12:52 PM
General Matrix-Matrix Multiplication (GEMM) and FlashAttention. Along the way, we will also study the core ingredients behind GPU optimization: data layout, asynchronous data movement, and asynchronous coordination."

mlc.ai/modern-gpu-p...
Modern GPU Programming For MLSys — Modern GPU Programming For MLSys
mlc.ai
July 3, 2026 at 4:15 PM
this is a nice higher-level talk on the subject: www.youtube.com/watch?v=dO4T...
Compression for AGI - Jack Rae | Stanford MLSys #76
YouTube video by Stanford MLSys Seminars
www.youtube.com
May 12, 2025 at 7:23 PM
Sponsor registration is open for #MLSys 2025. We have the most submissions ever to MLSys so it promises to be a great conference! mlsys.org/Sponsors/spo...
2025 Sponsor / Exhibitor Information
mlsys.org
January 20, 2025 at 2:20 AM
Modern GPU Programming for MLSys | Discussion
Modern GPU Programming For MLSys — Modern GPU Programming For MLSys
mlc.ai
June 26, 2026 at 6:40 PM
Modern GPU Programming For MLSys

"First understand the GPU hardware, then learn the programming model we will use, and finally build state-of-the-art kernels step by step. Our main target is the Blackwell generation, and our main running examples are
July 3, 2026 at 4:15 PM
great to see more specialized ML conferences! Mega conferences are fun, but at least in my experience with MLSys, I've had much better scientific conversations at smaller ones.
🦕The 19th conference on Neurosymbolic AI will be in beautiful Santa Cruz (CA, USA), September 8-10, 2025!

CFP is now out: 2025.nesyconf.org/call-for-pap...
🚨 Paper deadline: Feb 28 (abstract), March 7 (full)

#neurosymbolic #NeSy2025
Call for papers
19th International Conference on Neurosymbolic Learning (NeSy 2025, 8-10 September 2025, Santa Cruz, CA, USA)
2025.nesyconf.org
December 11, 2024 at 7:51 PM
This is a good starting point (he says, giving a link that will take several days to fully process)

github.com/AIoT-MLSys-L...
GitHub - AIoT-MLSys-Lab/Efficient-LLMs-Survey: [TMLR 2024] Efficient Large Language Models: A Survey
[TMLR 2024] Efficient Large Language Models: A Survey - AIoT-MLSys-Lab/Efficient-LLMs-Survey
github.com
December 31, 2024 at 9:29 PM
#MLSys 2025 is next week! You can still register at mlsys.org.
May 5, 2025 at 4:37 PM
ICLRとNeurIPSよりもう少し小さめで専門性が高い,AISTATS,MLSys,CoLLAs,UAIあたりのproになりたい
August 25, 2026 at 4:51 AM
Uber open-sourced ADR: enterprise security for AI agents. Includes Sensor, ADR-Bench (300+ tasks, 133 MCP servers), and two-tier detector. Accepted to MLSys 2026. Evaluate open-source Sensor and Detector to monitor agent behavior.
[GitHub Trending] uber/ADR
Four Signals — The Wire
www.foursignals.dev
August 4, 2026 at 1:00 PM
REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving

Mines an LLM's past attention into reusable document scores, rendering compressed evidence with far less overhead than per-query compressors.

📝 arxiv.org/abs/2609.11209
👨🏽‍💻 github.com/UIUC-MLSys/R...
REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving
Retrieval-augmented generation (RAG) improves knowledge-intensive large language model (LLM) applications by conditioning generation on retrieved documents, but longer contexts increase latency, key-v...
arxiv.org
September 11, 2026 at 3:21 AM
New LAPS System Cuts LLM Prefill Latency Over 30% with Load-Aware Deflection

30% faster LLM prefill—no model changes needed. LAPS just routes prompts smarter.

#AI #AIResearch #MachineLearning

https://autonainews.com/new-laps-system-cuts-llm-prefill-latency-over-30-with-load-aware-deflection/
New LAPS System Cuts LLM Prefill Latency Over 30% with Load-Aware Deflection
Most LLM serving research focuses on splitting prefill and decode across separate GPUs. LAPS, presented at MLSys 2026, goes a level deeper: it splits the prefill stage itself, routing long and short …
autonainews.com
July 3, 2026 at 2:21 PM
📌 Supports ALL MoE parallel modes: TP/EP/EP+TP
📌 MLSys'25 top scores (5/5/5/4) - battle-tested at scale

📄 Paper: arxiv.org/abs/2502.19811
📦 Code: github.com/bytedance/fl...
March 5, 2025 at 5:11 AM
Research led by UC Davis Ph.D. student Cameron Shinn shows that AI models can maintain response quality while using far less energy on existing hardware. The work received the Best Paper Award at MLSys 2026.

Read more: https://ow.ly/pPUi50Zunxi

#UCDavisEngineering
July 29, 2026 at 4:11 PM
Building Custom Modules with Rust and PyTorch. Wrote this one a while ago, hope you enjoy it. Keep experimenting. #ai #ml #mlengineering #pytorch #rust #mlsys
boredmle.blogspot.com/2026/08/adva...
Advanced Rust ML: Custom Modules with Tch-rs | ML Engineering
Custom Modules with Tch-rs
boredmle.blogspot.com
September 10, 2026 at 8:17 AM
June 26, 2026 at 6:16 PM
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism Article URL: https://mlsys.wuklab.io/posts/nitsum/ Comments URL: https://news.ycombinator.com/item?id=48188152 Points: 1 # Comme...

Origin | Interest | Match
MLSys @ WukLab - Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
mlsys.wuklab.io
May 19, 2026 at 2:04 AM
Understand DeepSeek V3.2: Pushing the Frontier of Open LLMs Recently, I joined the MLSys 2026 NVIDIA competition track! So I’m trying to understand DeepSeek V3.2, sparse attention, and learn GPU...

#gpu #sparse-attention #llm #machine-learning #deepseek

Origin | Interest | Match
Awakari App
awakari.com
February 25, 2026 at 12:20 AM
UC San Diego Packs a Punch of AI Research Power with Gift from NVIDIA May 27, 2025 — The MLSys ...

https://www.hpcwire.com/off-the-wire/uc-san-diego-packs-a-punch-of-ai-research-power-with-a-gift-from-nvidia/

Result Details
May 27, 2025 at 6:35 PM
A paper published by @utexasece.bsky.social students and faculty in collaboration with Meta has received an Outstanding Paper Honorable Mention at MLSys 2025 www.ece.utexas.edu/news/researc...
Researchers Receive Outstanding Paper Award at MLSys 2025
A paper published by students and faculty from the Chandra Family Department of Electrical Engineering in collaboration with Meta has received an Outstanding Paper Honorable Mention at MLSys 2025, the...
www.ece.utexas.edu
June 5, 2025 at 4:24 PM
I dreamed last night that I agreed to be PC chair for *another* conference (MLSys). What is wrong with me
September 5, 2024 at 2:53 PM