Per-Matrix Optimality Is Not Enough: Three-Level Optimization for Low-Rank LLM Compression

Three-level optimization for low-rank language model compression.

Abstract

Per-matrix singular value decomposition (SVD) truncation is Eckart-Young optimal in the whitened Frobenius norm, but errors from independently compressed matrices compound through the block’s nonlinear forward pass. Inspired in part by hierarchical variational optimization in quantum many-body methods, we introduce a three-level chain that widens optimization scope from individual matrices to Transformer blocks to the full model: whitened SVD (L1), block-level joint optimization (L2), and end-to-end language-modeling loss refinement (L3), all from 256 calibration sequences, with no instruction or recovery data. On LLaMA-7B at 60% compression, the chain reduces WikiText-2 perplexity from 42.1 to 19.1 to 11.4. The block-level stage acts as a regularizer: skipping it worsens Penn Treebank (PTB) perplexity by 24 points, a gap that additional end-to-end training did not close in our experiments. Perplexity gains hold across 20-80% compression, five architectures up to 13B parameters, and both in-distribution and out-of-distribution benchmarks, though the cross-architecture rows use architecture-specific configurations and the ratio sweep was not run under one common protocol. With more calibration data, skipping the block-level stage becomes competitive, revealing an offline compute-data trade-off. We therefore claim improvements only in perplexity and compression fidelity; downstream accuracy remains well below the dense model.

Publication
arXiv:2609.15838

To appear in Findings of EMNLP 2026.

Huicheng Zhang
Huicheng Zhang
PhD Student (2025)

I obtained my BS degree from Nankai University. My research interests include quantum information theory, quantum computation, large language models, and artificial intelligence.

Xiyao Feng
Xiyao Feng
MPhil (2024)

I obtained my Bachelor of Science in Physical Oceanography from Ocean University of China under the supervision of Prof. Jie Su. My research interests include quantum computing and artificial intelligence for quantum information science.

Chengkai Zhu
Chengkai Zhu
PhD Graduate
Xiao Shi
Xiao Shi
Research Associate

I obtained my BS in Software Engineering from Beijing information science and technology university under the supervision of Prof. Xiulei Liu. I obtained my PhD degree in Computer Software and Theory from University of Chinese Academy of Sciences under the supervision of Prof. Yun Shang. My research interests include tensor network and variational quantum algorithms.

Xin Wang
Xin Wang
Associate Professor

Prof. Xin Wang founded the QuAIR Lab at HKUST (Guangzhou) in June 2023. His research aims to advance our understanding of the limits of information processing with quantum systems and the potential of quantum artificial intelligence. His current interests include quantum algorithms, quantum resource theory, quantum machine learning, quantum computer architecture, and quantum error processing. Prior to establishing the QuAIR Lab, Prof. Wang was a Staff Researcher at the Institute for Quantum Computing at Baidu Research, where he focused on quantum computing research and the development of the Baidu Quantum Platform. Notably, he led the development of Paddle Quantum, a Python library for quantum machine learning. From 2018 to 2019, he was a Hartree Postdoctoral Fellow at the Joint Center for Quantum Information and Computer Science (QuICS) at the University of Maryland, College Park. Prof. Wang received his Ph.D. in quantum information from the University of Technology Sydney in 2018, under the supervision of Prof. Runyao Duan and Prof. Andreas Winter. He obtained his B.S. in mathematics (Wu Yuzhang Honors) from Sichuan University in 2014.