#autoregressive
oh look, an excuse to share a compvis paper (demonstrates segmentation learning emerging implicitly from masked prediction)
[MASK] is All You Need
In generative models, two paradigms have gained attraction in various applications: next-set prediction-based Masked Generative Models and next-noise prediction-based Non-Autoregressive Models, e.g., ...
arxiv.org
October 1, 2026 at 2:23 PM
Jian Chen, You Zhang, Mark Vinton: Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning https://arxiv.org/abs/2609.38658 https://arxiv.org/pdf/2609.38658 https://arxiv.org/html/2609.38658
October 1, 2026 at 6:45 AM
Yilun Liu, Yi Zhang, Ganyu Wu, Sikuan Yan, Mengyue Wang, Alois Knoll, Volker Tresp, Yunpu Ma: Learning Chaos Without Seeing Chaos: Extrapolation of Global Dynamics in Autoregressive Transformers https://arxiv.org/abs/2609.38814 https://arxiv.org/pdf/2609.38814 https://arxiv.org/html/2609.38814
October 1, 2026 at 6:43 AM
Jeon, Kim, Kakade, Du, Bedi, Chithanar, Lee, Kim, Chen: Blackboard Intelligence Can Surpass Autoregressive on Globally Constrained Problems https://arxiv.org/abs/2609.38806 https://arxiv.org/pdf/2609.38806 https://arxiv.org/html/2609.38806
October 1, 2026 at 6:43 AM
Umer Gupta, Saku Peltonen, Martin Ritzert: Autoregressive Frontier Expansion: Growing Trees with Graph Machine Learning https://arxiv.org/abs/2609.38506 https://arxiv.org/pdf/2609.38506 https://arxiv.org/html/2609.38506
October 1, 2026 at 6:42 AM
Lin, Ge, Zhu, Zhang, Liu, Wang, Li, Zhang, Song, Liu, Zhang: Enhancing Autoregressive Video Generation via Representation Adversarial Distillation https://arxiv.org/abs/2609.40037 https://arxiv.org/pdf/2609.40037 https://arxiv.org/html/2609.40037
October 1, 2026 at 6:42 AM
Qin Yan, Ruixiao Dong, Yutao Xie, Li Li, Ying Chen, Kai Li, Daowen Li, Houqiang Li: ResARC: Residual-Aware AutoRegressive Coding for Ultra-Low Bitrate Image Compression https://arxiv.org/abs/2609.39451 https://arxiv.org/pdf/2609.39451 https://arxiv.org/html/2609.39451
October 1, 2026 at 6:41 AM
Zeqi Xiao, Qingle Liu, Kaiwen Zhang, Yifan Zhou, Zihan Ding, Xingang Pan: DeCoPrune: Efficient KV-Cache Pruning for Autoregressive Video Diffusion via Denoising Consistency https://arxiv.org/abs/2609.39096 https://arxiv.org/pdf/2609.39096 https://arxiv.org/html/2609.39096
October 1, 2026 at 6:41 AM
Residual Trajectory Distillation for Generative Retrieval

Feeds residual information discarded when building Semantic IDs back into retrieval training, leaving the index and inference unchanged.

📝 arxiv.org/abs/2609.39319
👨🏽‍💻 github.com/Nevaeh7/iclr...
Residual Trajectory Distillation for Generative Retrieval
Generative retrieval has emerged as a general retrieval paradigm, representing items with discrete Semantic IDs (SIDs) and retrieving them through autoregressive identifier generation. When SIDs are c...
arxiv.org
October 1, 2026 at 4:34 AM
unless they have found a way to make attention non-autoregressive which... no they did not lol
October 1, 2026 at 2:17 AM
HyperZip: Efficient Data Compression through Personalized Diffusion LLMs with Hypernetworks

Thai Nguyen, Khang Tran, NhatHai Phan

#arXiv #cs.AI
HyperZip: Efficient Data Compression through Personalized Diffusion LLMs with Hypernetworks
Large language models (LLMs) have shown strong potential for lossless data compression, but existing approaches are constrained by the high computational cost and low throughput of autoregressive decoding. We propose HyperZip, an efficient and scalable LLM-based compression framework that leverages…
arxiv.org
October 1, 2026 at 2:05 AM
HyperZip: Efficient Data Compression through Personalized Diffusion LLMs with Hypernetworks

Thai Nguyen, Khang Tran, NhatHai Phan

#arXiv #cs.AI
HyperZip: Efficient Data Compression through Personalized Diffusion LLMs with Hypernetworks
Large language models (LLMs) have shown strong potential for lossless data compression, but existing approaches are constrained by the high computational cost and low throughput of autoregressive decoding. We propose HyperZip, an efficient and scalable LLM-based compression framework that leverages…
arxiv.org
October 1, 2026 at 12:39 AM
Can you use autoregressive diffusion to generate market data?
tech_blogs_jane_street

#MachineLearning
Can you use autoregressive diffusion to generate market data?
The following is part of a series of posts about 2026 summer intern projects – for more, see “What the interns have wrought, special jumbo 2026 edition”
blog.janestreet.com
September 30, 2026 at 11:06 PM
Calibrate the Decisions That Change the Future: On-Policy Post-Training Quantization for Multimodal Large Language Models

Wenxiao Fan et al.

#arXiv #cs.AI
Calibrate the Decisions That Change the Future: On-Policy Post-Training Quantization for Multimodal Large Language Models
Post-training quantization (PTQ) lowers deployment cost for multimodal large language models, but calibration typically reconstructs fixed sequences with local objectives. This overlooks autoregressive feedback: a quantization-induced token change redirects the prefix and changes future states. Yet…
arxiv.org
September 30, 2026 at 8:10 PM
Calibrate the Decisions That Change the Future: On-Policy Post-Training Quantization for Multimodal Large Language Models

Wenxiao Fan et al.

#arXiv #cs.AI
Calibrate the Decisions That Change the Future: On-Policy Post-Training Quantization for Multimodal Large Language Models
Post-training quantization (PTQ) lowers deployment cost for multimodal large language models, but calibration typically reconstructs fixed sequences with local objectives. This overlooks autoregressive feedback: a quantization-induced token change redirects the prefix and changes future states. Yet…
arxiv.org
September 30, 2026 at 6:36 PM
A single autoregressive neural model now learns molecular ground states from sparse anchor geometries, achieving chemical accuracy at untrained configurations with 25.8× GPU-cost reduction compared to independent optimization.

#NeuralQuantumStates #MolecularChemistry #Research
Geometry-Conditioned Neural-Network Quantum States for Molecular Potential Energy Surfaces
arxiv.org
September 30, 2026 at 6:04 PM
New arXiv paper proposes map-conditioned autoregressive generation of human mobility, using a road raster to condition a decoder emitting 31.25 m mesh-cell tokens, tested on trajectories from Ishikawa Prefecture.

#Semiconductors #AIInfrastructure #GPUComputing
https://arxiv.org/abs/2609.32360
September 30, 2026 at 6:01 PM
Open source Python library. Runs via SGLang with Qwen3, MiniCPM5 and more. Full docs on the listing. www.everydev.ai/tools/typellm
TypeLLM - Type Safe LLM Output Library | EveryDev.ai
TypeLLM is an open-source Python library, licensed under Apache 2.0, that brings type-safe structured output to existing autoregressive LLMs without…
www.everydev.ai
September 30, 2026 at 1:12 PM
arXiv📈🤖
Beyond the Coast: an Empirical Assessment of the Kaldor-Verdoorn Law in Chinese Provinces
By G\'oes, Barabuffi
September 30, 2026 at 9:38 AM
Using an Autoregressive Model to Predict the Price-to-Earnings Ratio and Develop an Investment Strategy
#finance #trading #investing

In a previous post, we highlighted an article that showed how useful accounting numbers are. In this post, we will prese…
Using an Autoregressive Model to Predict the Price-to-Earnings Ratio and Develop an Investment Strategy
In a previous post, we highlighted an article that showed how useful accounting numbers are.
ift.tt
September 30, 2026 at 9:20 AM
"This has never sat right with me: a bounded yes/no or routing decision was often passed through the same autoregressive decoder used to generate a paragraph."

Lambert Leong digs into the architectural constraints that are steering the field towards "decision" models like Jev.
When All You Have Are Decoders, Every Decision Looks Like Generation
Not every decision needs a decoder, generation Is not always a decision
towardsdatascience.com
September 30, 2026 at 9:02 AM
Chunyi Hou, Xiangfei Qiu, Hanyin Cheng, Yutong Li, Bin Yang: HALO: Enhancing Time Series Generation via Hyperspherical Latents and Masked AutoregRessive Modeling https://arxiv.org/abs/2609.34511 https://arxiv.org/pdf/2609.34511 https://arxiv.org/html/2609.34511
September 30, 2026 at 6:50 AM
Julian Boesch, Andrew Wee, Alexander Stranzl: DreamingGoose: Staged Distillation from Autoregressive Transformers to Bidirectional Recurrent Diffusion Language Models https://arxiv.org/abs/2609.34253 https://arxiv.org/pdf/2609.34253 https://arxiv.org/html/2609.34253
September 30, 2026 at 6:49 AM
Suqin Yuan, Runqi Lin, Kevin Qinghong Lin, Junchi Yu, Lei Feng, Chris Russell, Tongliang Liu: Decoupling Token Roles in Autoregressive Pretraining https://arxiv.org/abs/2609.33405 https://arxiv.org/pdf/2609.33405 https://arxiv.org/html/2609.33405
September 30, 2026 at 6:47 AM
Thomas S. Robinson: Shared Autoregressive Context Can Distort Relationships in Synthetic Data https://arxiv.org/abs/2609.32546 https://arxiv.org/pdf/2609.32546 https://arxiv.org/html/2609.32546
September 30, 2026 at 6:44 AM