Dynamic sparse attention for scalable transformer acceleration
Unknown
Proposes Dynamic Sparse Attention (DSA) to efficiently exploit dynamic sparse patterns in attention for scalable transformer acceleration.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
Proposes Dynamic Sparse Attention (DSA) to efficiently exploit dynamic sparse patterns in attention for scalable transformer acceleration.
Unknown
This paper introduces a novel sparse attention mechanism that improves efficiency and performance for long-range transformers.
Unknown
This paper introduces a content-based sparse attention mechanism that combines the benefits of content-based and local temporal sparse attention for improved efficiency in Transformer models.
Unknown
This paper systematically evaluates six training-free sparse attention methods for transformer LLMs, proposing a taxonomy along four design axes and deriving actionable insights.
Unknown
Introduces speculative decoding, an algorithm for faster sampling from autoregressive models by computing multiple tokens in parallel without altering outputs.
D. Du, Gu Gong, Xiaowen Chu
A comprehensive survey of model quantization and hardware acceleration techniques for vision transformers.
Unknown
FuseMoE introduces a mixture-of-experts framework with a novel gating function for integrating diverse numbers of modalities.
Unknown
This paper investigates how transformers implement in-context learning by showing that linear models can perform gradient descent through their forward pass.
Yunpeng Huang, Jingwei Xu, Zixu Jiang, et al.
A comprehensive survey of transformer architecture advancements for extending context length in large language models.
Unknown
Planning Transformer introduces planning tokens for dual time-scale prediction to enable long-horizon offline reinforcement learning with implicit planning.
Oludare Isaac Abiodun, Muhammad Ubale Kiru, Aman Jantan, et al.
A comprehensive review of ANN applications to pattern recognition, highlighting current models like GAN, CNN, and Transformer, and identifying unresolved challenges such as whimsical orientation and neuron behavior analysis.
Mohaimenul Azam Khan Raiaan, Md. Saddam Hossain Mukta, Kaniz Fatema, et al.
A comprehensive review of LLMs covering history, architectures, transformers, training methods, applications, and open challenges.