Papers Explained 187e: Quantized Llama 3.2, Llama 3.3
Unknown
Meta's quantized Llama 3.2 and Llama 3.3 models reduce size and memory via QAT with LoRA and SpinQuant, enabling deployment on mobile devices.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
Meta's quantized Llama 3.2 and Llama 3.3 models reduce size and memory via QAT with LoRA and SpinQuant, enabling deployment on mobile devices.
D. Du, Gu Gong, Xiaowen Chu
A comprehensive survey of model quantization and hardware acceleration techniques for vision transformers.
Ruikang Liu, Haoli Bai, Haokun Lin, et al.
Intactkv improves large language model quantization by preserving pivot tokens, achieving state-of-the-art results.
Jinuk Kim, Marwa El Halabi, Wonpyo Park, et al.
GuidedQuant improves post-training quantization of large language models by using end-loss guidance to better preserve model accuracy.
Unknown
This paper introduces Llmc, a versatile toolkit for benchmarking LLM quantization methods, enabling standardized evaluation and comparison.
Tianyi Zhang, Anshumali Shrivastava
Leanquant introduces a loss-error-aware grid quantization method for LLMs that achieves accurate and scalable compression by minimizing quantization error with respect to the model's loss function.
Yuhang Li, Mingzhu Shen, Yan Ren, et al.
Mqbench provides a reproducible and deployable benchmark for model quantization to accelerate deep learning inference.
Xing Hu, Yuan Cheng, Dawei Yang, et al.
Ostquant improves LLM quantization by applying orthogonal and scaling transformations to better fit weight distributions, reducing accuracy loss.
Unknown
Proposes a retraining-free model quantization method using one-shot weight-coupling learning to achieve efficient compression without fine-tuning.
Kai Liu, Qian Zheng, Kaiwen Tao, et al.
This survey provides a comprehensive overview of low-bit model quantization techniques for deep neural networks, including a curated list of resources.
Unknown
This survey comprehensively reviews model quantization techniques for deep neural networks in image classification, covering methods, challenges, and future directions.
Panagiotis Papantonakis, Georgios Kopanas, Bernhard Kerbl, et al.
This paper reduces the memory footprint of 3D Gaussian splatting by 27x via resolution-aware pruning, adaptive spherical harmonic coefficients, and codebook quantization.