Low-Rank Compression of Language Models via Differentiable Rank Selection

By Sidhant Sundrani, Francesco Tudisco, and Pasquale Minervini, December 14, 2025

In arXiv 2025

Low-rank decomposition can make language models smaller, but choosing a rank independently for every layer creates a difficult trade-off between compression and downstream accuracy. Existing approaches often rely on limited heuristic searches or require fine-tuning after compression.

Learning to Low-Rank Compress instead learns differentiable masks over singular values. Using a calibration dataset, it optimises only the mask weights to reduce rank while keeping intermediate activations close to those of the original model.

Across common-sense reasoning and open-domain question-answering tasks, the method improves on competing rank-selection approaches that do not use post-compression fine-tuning.

Paper: https://arxiv.org/abs/2512.13733

Stay ahead with research-backed solutions

From papers to production, we translate cutting-edge AI research into practical systems that give your business a competitive edge.

Book a Consultation