#executorch
ExecuTorch Hackathon comes to San Francisco October 17-18. Build and deploy PyTorch models with ExecuTorch across Compute, Mobile + XR, and IoT tracks using Snapdragon-powered PCs, Samsung Galaxy devices, and Arduino hardware.

Submit a proposal: https://luma.com/executorch-hackathon
ExecuTorch Hackathon · Luma
NOTE: By registering for the Luma, you are submitting an application to participate. We will reach out on a rolling basis to let you know whether you've been…
luma.com
September 24, 2026 at 8:52 PM
Build AI on the edge with ExecuTorch at a Meta + Arm hackathon at SFSU. Get hands-on experience on Arm-powered devices with talks, demos, workshops, and mentorship.

Learn more: https://arm-executorch.devpost.com/?_gl=1*lvaup2*_gcl_au*MjA3MDMyNDUxNS4xNzg5NzQ0MDk3*_ga*MTA4MTg0NjcyMy4xNzg5NTQ4NDE3
Arm + ExecuTorch: Edge AI Challenge
Build intelligent experiences at the edge. Sponsored by Arm, Meta and SFSU
arm-executorch.devpost.com
September 22, 2026 at 10:34 PM
This Week In React 297

⚛️
- DevTools
- Motion
- Apollo
- React Router
- React Hook Form
- shadcn lint
📱
- Shopify
- Expo 58 beta
- ExecuTorch
- Screenmap
- Lynx
- Stim
- Voltra

🍿 Read: thisweekinreact.com/newsletter/297

✍️ @jwr.ski & I
September 17, 2026 at 8:16 AM
Join Thomas Cottenier from Arm this October 20-21. Thomas will present how AI agents can synthesize customized, target-specific runtimes using PyTorch components like torch.export, ExecuTorch, and torchao.

Register now for PyTorch Conference North America: https://hubs.la/Q04v4SL60

#PyTorchCon
September 10, 2026 at 10:00 PM
ExecuTorch에서 Muse Glimmer로 구현하는 빠른 온디바이스 에이전틱 AI | 파이토치 한국 사용자 모임

오늘 Meta가 Muse Glimmer를 공개했습니다. 온디바이스(on-device) 에이전틱(agentic) 워크플로우를 위해 Meta의 Muse Spark에서 증류한, 매개변수 300억 개 규모의 오픈 웨이트(open-weight) 모델입니다. 이와 함께 ExecuTorch는 NVIDIA GPU와 Apple 실리콘 기반 Mac에서 Muse Glimmer를 실행할 수 있도록 엔드투엔드(end-to-end) 지원…
ExecuTorch에서 Muse Glimmer로 구현하는 빠른 온디바이스 에이전틱 AI | 파이토치 한국 사용자 모임
오늘 Meta가 Muse Glimmer를 공개했습니다. 온디바이스(on-device) 에이전틱(agentic) 워크플로우를 위해 Meta의 Muse Spark에서 증류한, 매개변수 300억 개 규모의 오픈 웨이트(open-weight) 모델입니다. 이와 함께 ExecuTorch는 NVIDIA GPU와 Apple 실리콘 기반 Mac에서 Muse Glimmer를 실행할 수 있도록 엔드투엔드(end-to-end) 지원을 추가합니다. Today, Meta introduced Muse Glimmer, an open-weight, 30-billion-parameter model di...
discuss.pytorch.kr
August 16, 2026 at 1:00 AM
Fast, On Device Agentic AI with Muse Glimmer on ExecuTorch

This article heavily emphasizes the technical novelty of ExecuTorch for backend portability, but it downplays the significant engineering challenge of ensuring performance parity and stability across disparate, highly optimized kernels …
arc-codex.com
August 15, 2026 at 5:13 AM
ExecuTorch (0.1.0) for zephyr by Meta Platforms

➡️ https://github.com/meta-pytorch/executorch-arduino

Run PyTorch models on Arduino microcontrollers with ExecuTorch.

#ArduinoLibs #ArduinoLibraries #zephyrLibraries
August 13, 2026 at 12:10 AM
Inside Meta's Muse Glimmer 30B and ExecuTorch: How to compile, quantize, and run complex agentic loops locally on edge hardware without cloud latency or costs. Read the full technical breakdown:

https://ixuvo.com/blog/on-device-agentic-ai-meta-muse-glimmer-executorch
On-Device Agentic AI: Inside Meta's Muse Glimmer 30B and ExecuTorch Core Architectures
An in-depth technical analysis of Meta's Muse Glimmer 30B and the ExecuTorch runtime. Learn how to compile, quantize, and optimize large agentic models for low-latency, offline execution on consumer-grade and edge hardware.
ixuvo.com
August 12, 2026 at 4:00 PM
Meta 推出 300 亿参数的开放权重模型 Muse Glimmer

Meta 在 Apache 2.0 许可证下推出了 300 亿参数的开放权重模型 Muse Glimmer。Muse Glimmer 为本地运行的智能体工作流程进行了优化,能在配备了单个消费级 GPU 的 Mac 或 PC 上运行,支持从本地智能体和函数调用到本地编程和 LLM 自动化评估等多种应用场景。Meta 将在未来几天推出针对 llama.cpp、MLX 和 ExecuTorch 的集成。
August 10, 2026 at 3:30 PM
Distilled from Meta's Muse Spark model, Muse Glimmer is supported end-to-end on NVIDIA GPUs and Apple silicon via ExecuTorch, with Mark Zuckerberg urging Washington to remove regulatory barriers to open-source AI.
August 10, 2026 at 2:44 PM
You can download and start running quickly the ExecuTorch weights on Hugging Face here: huggingface.co/meta-models/...
meta-models/Muse-Glimmer-30B-ExecuTorch-PTE · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
August 10, 2026 at 2:13 PM
Today Meta introduced Muse Glimmer, an open-weight, 30-billion-parameter model distilled from Meta’s Muse Spark for on-device agentic workflows & ExecuTorch is adding end-to-end support for running Muse Glimmer on NVIDIA GPUs and Macs with Apple silicon. Learn more: pytorch.org/blog/fast-on...
August 10, 2026 at 2:11 PM
💡 Summary:

MetaのMuse Glimmerは、30Bパラメータのローカルエージェント向けモデルをオープンソース化(Apache 2.0)したもので、Mac/PC上の単一GPUで動作可能な高性能を狙います。長期的な推論・ツール呼び出し・マルチモーダル対応などエージェント機能を強化し、量子化とドラフターによる高速化を実現、オフライン運用も視野に入れています。現在HuggingFaceで weights を公開、今後llama.cppやExecuTorchなどとの統合も提供予定です。
August 10, 2026 at 1:44 PM
Running Qwen3 Through the ExecuTorch MLX Delegate: Up to 4.52x Faster on M1 Max
Hello, everyone. There are now many ways to run an LLM on a Mac, but exporting a PyTorch model for Apple Silicon and executing it in a lightweight runtime is still an evolving path. How much faster is it, and does 4-bit quantization change the output? Today, I am looking at ExecuTorch's experimental MLX delegate, released in May 2026. It enables PyTorch models to run on Apple Silicon GPUs. I use ExecuTorch 1.3.1 to run Qwen3-0.6B and compare it with PyTorch MPS. The short result is that decode throughput was **41.8 tokens/s** with PyTorch MPS BF16, **134.8 tokens/s** with MLX BF16, and **188.9 tokens/s** with MLX INT4. MLX INT4 was 4.52x faster, and its file was 71.8% smaller than BF16. However, INT4 changed the generated output in two of three simple prompts. ## What Is the ExecuTorch MLX Delegate? ExecuTorch is a runtime for running trained PyTorch models on desktops, phones, and embedded devices. Inference means feeding input into a trained model to obtain an output. The MLX delegate was released on **May 18, 2026**. It sends a PyTorch computation graph to Apple's MLX framework for execution on an Apple Silicon GPU. It supports BF16, FP16, FP32, and 2/4/8-bit quantization, among other formats. It is currently experimental, so its APIs and supported scope may change. Quantization stores weights with fewer bits to reduce size and computation. I compared ordinary BF16 with INT4, which quantizes the linear layers and embeddings to four bits. ## The Qwen3-0.6B Model Used Here Qwen3 is an LLM family announced by the Qwen Team on **April 29, 2025**. A single model can switch between a thinking mode for step-by-step reasoning and a non-thinking mode for shorter answers. The family supports more than 100 languages and dialects. The Qwen3-0.6B used here is one of the smaller dense models in the family. “Dense” means that it generally uses the full model for each input, unlike a mixture-of-experts model that activates only selected parts. It has a nominal 0.6 billion parameters, 28 layers, and a context length of 32,768 tokens. A token is a small unit of text processed by the model. I disabled thinking mode for this test and pinned the model revision to `c1899de289a04d12100db370d81485cdf75e47ca`. ### Licenses Component | License ---|--- Qwen3-0.6B weights | Apache License 2.0 ExecuTorch | BSD 3-Clause License MLX | MIT License PyTorch | BSD 3-Clause License These permissive open-source licenses allow broad use, including commercial use, but their conditions—such as retaining copyright notices when redistributing—still apply. Check the linked license text for your use case. ## The Three Paths and Their Data Flow Only one trained model, Qwen3-0.6B, was used. The same weights were tested through three execution paths. same prompt -> tokenizer: convert text to token IDs -> Qwen3-0.6B ├─ ExecuTorch MLX BF16 -> .pte -> MLX / Metal GPU ├─ ExecuTorch MLX INT4 -> 4-bit .pte -> MLX / Metal GPU └─ PyTorch MPS BF16 ----------------> MPS / Metal GPU -> greedy generation: select the most likely next token each time -> tokenizer: convert token IDs back to text A `.pte` file is a model program exported for ExecuTorch. The complete computation graph could be lowered into a single MLX subgraph. PyTorch MPS BF16 ran the same BF16 weights through the ordinary Transformers/PyTorch path and served as the reference. ## What I Tested * Whether the complete Qwen3-0.6B graph could be lowered to MLX and executed * Whether MLX BF16 and PyTorch MPS BF16 generated the same tokens * How INT4 changed model size, speed, and process memory * Whether the experiment could be reproduced using published wheels only The complete code and JSON report are available in the executorch-mlx-qwen3 lab in kiarina/labs. ## Reproducing the Lab You need an Apple Silicon Mac, Xcode Command Line Tools, `mise`, `uv`, and an internet connection. The first run downloads the Qwen3 weights and creates about 1.53 GB of PTE files in total. git clone --depth 1 --filter=blob:none --sparse \ https://github.com/kiarina/labs.git cd labs git sparse-checkout set .gitignore .mise/tasks Makefile mise.toml \ 2026/07/22/executorch-mlx-qwen3 mise -C 2026/07/22/executorch-mlx-qwen3 run The export and benchmark steps can also be run separately. mise -C 2026/07/22/executorch-mlx-qwen3 run export mise -C 2026/07/22/executorch-mlx-qwen3 run benchmark ## Test Conditions machine: MacBook Pro (Apple M1 Max, 32 GPU cores, 64 GB) OS: macOS 26.5.2 Python: 3.13.7 ExecuTorch: 1.3.1 PyTorch: 2.12.1 Transformers: 4.56.1 model: Qwen/Qwen3-0.6B generation: greedy, batch 1, up to 16 tokens PTE: custom MLX SDPA / KV cache, requested maximum sequence 128 For performance, each backend was given the same Japanese prompt and forced to generate 16 tokens. The reported values are medians from five trials after warm-up. Prefill reads the input and produces the first token; decode generates the remaining tokens one at a time. Each backend ran in a separate process. The MLX path loaded a fresh forward method and initialized its KV cache for each trial, while the PyTorch path created a fresh cache for every trial. ## Results ### PTE Size PTE | Size | Relative to BF16 | Export time | SHA-256 ---|---|---|---|--- MLX BF16 | 1,192,264,196 bytes | 100.0% | 41.61 s | `83da47c2…bfb8c0` MLX INT4 | 335,662,976 bytes | 28.2% | 54.55 s | `0e30a054…71267` INT4 was 856,601,220 bytes, or **71.8% smaller** , than BF16. Quantization itself takes work, so exporting INT4 took about 13 seconds longer. ### Generation Speed The input was the Japanese prompt “Briefly explain local inference on Apple Silicon.” Backend | Load | Median prefill | Median decode | 16-token total | Peak RSS increase ---|---|---|---|---|--- ExecuTorch MLX BF16 | 0.002 s | **0.020 s** | 134.8 tokens/s | 0.131 s | 1.27 GiB ExecuTorch MLX INT4 | 0.003 s | 0.028 s | **188.9 tokens/s** | **0.108 s** | **0.47 GiB** PyTorch MPS BF16 | 0.726 s | 0.038 s | 41.8 tokens/s | 0.396 s | 0.14 GiB MLX BF16 decoded **3.22x** as fast as PyTorch MPS BF16, while MLX INT4 was **4.52x** as fast. INT4 was also 1.40x faster than MLX BF16 and reduced the RSS increase by about 63%. There are important qualifications. The MLX load figure measures only opening the PTE program, not all work needed to materialize weights for GPU use. RSS measures the process's main memory, not GPU memory itself. PyTorch reported 1.20 GiB of MPS driver-allocated memory at the end of the run. Because GPU memory was not measured on the same basis, the table does not prove that PyTorch used the least memory. On the first invocation only, BF16 prefill took 0.434 seconds and INT4 took 1.014 seconds. Cold starts that include Metal setup and initial compilation were much slower than the warmed-up figures. ### Generated Output I compared the generated token sequences on three short prompts. Prompt | PyTorch MPS BF16 | MLX BF16 | MLX INT4 ---|---|---|--- Answer Japan's capital in one word | `日本の首都は、**大阪**です。` (Japan's capital is **Osaka**.) | Exact token match | `日本の首都は、**东京**です。` (Japan's capital is **Tokyo**.) Answer 1+1 with one digit | `1+1=2` | Exact token match | `1` Copy `MLX` unchanged | `MLX` | Exact token match | Exact token match MLX BF16 matched PyTorch MPS BF16 token for token in all three cases. Within this narrow test, changing the execution path to MLX did not introduce a difference. INT4 matched in only one case. Quantization represents numbers more coarsely, so when candidate next tokens have similar scores, their ranking can change. The capital answer was wrong in both BF16 and INT4. The small 0.6B model, prompt, and greedy decoding could all contribute, but three questions are not enough to identify the cause. This probe checks differences between backends; it does not certify the model's knowledge or answer quality. ## A Plain-Language Reading of the Results 1. **The same small LLM ran substantially faster through MLX.** Even BF16 decoded 3.22x as fast as PyTorch MPS. Exporting a PyTorch model to a lightweight Mac runtime looks promising. 1. **Four-bit weights are smaller and faster, but answers can change.** INT4 reduced the file from about 1.19 GB to 336 MB and produced the highest decode rate. Yet it changed two of only three token sequences. It should be evaluated on a task-specific quality set before adoption. 1. **The gap cannot be attributed to MLX kernels alone.** This comparison covers an ExecuTorch MLX pipeline versus a Transformers/PyTorch MPS pipeline. Their runtimes, cache handling, and execution paths differ, so the numbers describe the complete pipelines. ## Reproducibility Findings The dependency metadata for `executorch==1.3.1` allowed PyTorch 2.13.0, but importing the published ExecuTorch extension failed because the `materialize_cow_storage` symbol was missing. Pinning PyTorch to 2.12.1 made the same wheel work. The failure can be reproduced with this optional task: mise -C 2026/07/22/executorch-mlx-qwen3 run probe-torch-2-13 The bundled PTE inspector also failed because its included `flatc` did not recognize the `--json` option. I verified full-graph delegation from the partitioner log emitted during export instead. ## Limitations * Only one M1 Max, Qwen3-0.6B, batch 1, and short Japanese prompts were tested. * Five short trials do not control thermal state, power consumption, or other GPU workloads. * Only BF16 and INT4 were compared; FP16, 2/8-bit, Core ML, and other paths were not tested. * The quality probe had only three questions; no standard benchmark or perplexity was measured. * The requested maximum sequence was 128, and long text was not tested. * GPU memory could not be measured on the same basis for MLX and MPS. * Cold-start latency was observed only once. * The MLX delegate is experimental. ## My Takeaway Lowering the entire Qwen3-0.6B graph to MLX and increasing decode throughput by more than 3x without changing the BF16 tokens was a good result. INT4 reduced the file to less than one-third of its BF16 size and reached about 189 tokens/s, which is attractive when embedding a small model on a Mac. The output changes from 4-bit quantization appeared immediately, even in this tiny probe. Quantization is not a free speedup. Paired with a quality suite for the real task, this path could work well for local helper features or small, responsive agents.
dev.to
July 22, 2026 at 4:24 AM
ExecuTorch es el framework oficial de PyTorch para llevar modelos de inteligencia artificial a dispositivos edge, desde teléfonos hasta microcontroladores.
ExecuTorch de PyTorch/Meta: el framework con 4,800 estrellas que corre Llama 3 y Qwen 3 en tu celular sin internet - Sinaptica
ExecuTorch es el framework oficial de PyTorch para llevar modelos de inteligencia artificial a dispositivos edge, desde teléfonos hasta microcontroladores. Lo d
sinapti.ca
July 15, 2026 at 8:21 AM
ExecuTorch is the official PyTorch framework for bringing artificial intelligence models to edge devices, from phones to microcontrollers.
ExecuTorch by PyTorch/Meta: the framework with 4,800 stars that runs Llama 3 and Qwen 3 on your phone without internet - Sinaptica
ExecuTorch is the official PyTorch framework for bringing artificial intelligence models to edge devices, from phones to microcontrollers. It was developed by M
sinapti.ca
July 15, 2026 at 8:21 AM
We are proud to share the results of our security audit of PyTorch ExecuTorch. Thanks to @trailofbits.bsky.social and Alpha-Omega, this project underwent a custom engagement of security review and testing. Read more on our blog- ostif.org/pytorch-exec...

#OSTIF #PyTorch #TrailofBits #AlphaOmega
July 13, 2026 at 3:24 PM
ExecuTorch 해커톤에서 온디바이스 AI의 미래를 만들어가다 | 파이토치 한국 사용자 모임

지난 주말, 샌프란시스코에서는 빌더와 연구자, 모바일 개발자, AI 실무자들이 한자리에 모여 ExecuTorch 해커톤을 열었습니다. 강력한 AI를 손안의 기기에서 로컬로 실행하면 무엇을 만들 수 있을까라는 실용적이면서도 갈수록 중요해지는 질문에 초점을 맞춘 이틀간의 현장 행사였습니다. 2026년 6월 27~28일에 열린 이번 해커톤에서 참가팀들은 ExecuTorch를 사용해 Snapdragon 기반 모바일 기기에서 직접 동작하는 실…
ExecuTorch 해커톤에서 온디바이스 AI의 미래를 만들어가다 | 파이토치 한국 사용자 모임
지난 주말, 샌프란시스코에서는 빌더와 연구자, 모바일 개발자, AI 실무자들이 한자리에 모여 ExecuTorch 해커톤을 열었습니다. 강력한 AI를 손안의 기기에서 로컬로 실행하면 무엇을 만들 수 있을까라는 실용적이면서도 갈수록 중요해지는 질문에 초점을 맞춘 이틀간의 현장 행사였습니다. 2026년 6월 27~28일에 열린 이번 해커톤에서 참가팀들은 ExecuTorch를 사용해 Snapdragon 기반 모바일 기기에서 직접 동작하는 실시간 AI 애플리케이션을 만들고 최적화하는 과제에 도전했습니다. 참가자들은 Snapdragon으로 구동되는 삼성 갤럭시 S25 Ultra 기기에서 개발했으며, Qualcomm과 Meta 전문가들의 워크숍과 멘토링, 실습 지원을 받았습니다. This past weekend in San Francisco, builders, researchers, mobile developers, and AI practitioners...
discuss.pytorch.kr
July 11, 2026 at 10:41 AM
PyTorch Foundation supported the ExecuTorch Hackathon, where teams built on-device AI apps with PyTorch and ExecuTorch.

Congratulations to SafeScreen AI, SixthSense and Toddle AI, the winners.

Read recap: https://pytorch.org/blog/building-the-future-of-on-device-ai-at-the-executorch-hackathon/
July 2, 2026 at 9:00 PM
If you're exploring Edge AI with PyTorch, these new Jupyter labs are a practical place to start.

Learn how to deploy and optimise models with ExecuTorch on Arm CPUs and NPUs, with hands-on examples for platforms including Raspberry Pi. pytorch.org/blog/efficie...
June 29, 2026 at 1:55 PM