Sparser is faster and less is more: Efficient sparse attention for long-range transformers
Unknown
This paper introduces a novel sparse attention mechanism that improves efficiency and performance for long-range transformers.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
This paper introduces a novel sparse attention mechanism that improves efficiency and performance for long-range transformers.
Unknown
This paper introduces a content-based sparse attention mechanism that combines the benefits of content-based and local temporal sparse attention for improved efficiency in Transformer models.
Unknown
Introduces speculative decoding, an algorithm for faster sampling from autoregressive models by computing multiple tokens in parallel without altering outputs.
D. Du, Gu Gong, Xiaowen Chu
A comprehensive survey of model quantization and hardware acceleration techniques for vision transformers.
Unknown
FuseMoE introduces a mixture-of-experts framework with a novel gating function for integrating diverse numbers of modalities.
Unknown
This paper investigates how transformers implement in-context learning by showing that linear models can perform gradient descent through their forward pass.
Mohaimenul Azam Khan Raiaan, Md. Saddam Hossain Mukta, Kaniz Fatema, et al.
A comprehensive review of LLMs covering history, architectures, transformers, training methods, applications, and open challenges.