Fast and Expressive Multi-Byte Prediction with Probabilistic Circuits
By Andreas Grivas, Lorenzo Loconte, Emile van Krieken, Piotr Nawrot, Yu Zhao, Euan Wielewski, Pasquale Minervini, Edoardo Ponti, and Antonio Vergari, July 6, 2026
In ICML 2026
Multi-token prediction can accelerate language-model generation, especially for tokeniser-free byte-level models, but common approaches either assume that future outputs are independent or generate them sequentially within each prediction window. The former limits expressiveness, while the latter adds latency.
MTPC represents the joint distribution over future bytes with probabilistic circuits. Choosing different circuit structures recovers or generalises mixture models, hidden Markov models, and tensor networks, making the trade-off between dependency modelling and inference latency explicit.
The method retrofits byte-level and byte-converted language models and combines with speculative decoding. Experiments show faster generation than independent multi-token prediction while preserving the verifier model’s output quality, alongside an analysis of circuit architecture and layer sharing.
Paper: https://arxiv.org/abs/2511.11346
Stay ahead with research-backed solutions
From papers to production, we translate cutting-edge AI research into practical systems that give your business a competitive edge.
Book a Consultation