Low-Rank Compression of Language Models via Differentiable Rank Selection
By Sidhant Sundrani, Francesco Tudisco, and Pasquale Minervini, December 14, 2025
In arXiv 2025
Low-rank decomposition can make language models smaller, but choosing a rank independently for every layer creates a difficult trade-off between compression and downstream accuracy. Existing approaches often rely on limited heuristic searches or require fine-tuning after compression.
Learning to Low-Rank Compress instead learns differentiable masks over singular values. Using a calibration dataset, it optimises only the mask weights to reduce rank while keeping intermediate activations close to those of the original model.
Across common-sense reasoning and open-domain question-answering tasks, the method improves on competing rank-selection approaches that do not use post-compression fine-tuning.
Paper: https://arxiv.org/abs/2512.13733
Stay ahead with research-backed solutions
From papers to production, we translate cutting-edge AI research into practical systems that give your business a competitive edge.
Book a Consultation