Dynamic sparse attention for scalable transformer acceleration
Unknown
Proposes Dynamic Sparse Attention (DSA) to efficiently exploit dynamic sparse patterns in attention for scalable transformer acceleration.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
Proposes Dynamic Sparse Attention (DSA) to efficiently exploit dynamic sparse patterns in attention for scalable transformer acceleration.
Unknown
This paper introduces a novel sparse attention mechanism that improves efficiency and performance for long-range transformers.
Unknown
This paper introduces a content-based sparse attention mechanism that combines the benefits of content-based and local temporal sparse attention for improved efficiency in Transformer models.
Unknown
This paper systematically evaluates six training-free sparse attention methods for transformer LLMs, proposing a taxonomy along four design axes and deriving actionable insights.
Unknown
Introduces speculative decoding, an algorithm for faster sampling from autoregressive models by computing multiple tokens in parallel without altering outputs.
D. Du, Gu Gong, Xiaowen Chu
A comprehensive survey of model quantization and hardware acceleration techniques for vision transformers.
Junsong Chen, Jincheng Yu, Yitong Li, et al.
SANA-Video 2.0 introduces hybrid linear-softmax attention and attention residuals to achieve softmax-level video quality with linear-complexity scaling, enabling 720p generation on a single GPU.
Danna Lesley Cruz Reyes, Juan Camilo Camargo Prieto, Andres David Leon Hernandez, et al.
This paper compares BERT and LSTM deep learning models for spam email classification on the Enron dataset, achieving 97% accuracy with BERT slightly outperforming LSTM but at higher computational cost.
Korota Arsène Coulibaly, Mohamed Hamlich, Khalid Hmali, et al.
A synthetic data generation framework for rotogravure printing defects enables training object detectors to achieve 80.9% mAP on real samples, eliminating manual data collection.
Unknown
FuseMoE introduces a mixture-of-experts framework with a novel gating function for integrating diverse numbers of modalities.
Unknown
This paper uses synthetic data to study how pretrained large language models form representations during in-context learning.
Unknown
This paper investigates how transformers implement in-context learning by showing that linear models can perform gradient descent through their forward pass.