JournalACM Transactions on Reconfigurable Technology and Systems2024
Understanding the Potential of FPGA-based Spatial Acceleration for Large Language Model Inference
Hongzheng Chen, Jiahao Zhang, Yixiao Du, et al.
This paper investigates FPGA-based spatial acceleration for LLM inference, achieving up to 13.4x speedup over prior FPGA accelerators and 5.7x energy efficiency vs. A100 GPU.
96Apr 4, 2024Large Language ModelsTransformers
arXiv