Attention Is All You Need
Ashish Vaswani, Noam Shazeer et al.
976
Citations
112
Influential Citations
Annual Meeting of the Association for Computational Linguistics
Venue
2019
Year
… This chapter focuses on delineating the mixture of experts modelling framework and demonstrates the utility and flexibility of mixture of experts models as an analytic tool. …
This paper provides a foundational overview of the mixture of experts (MoE) framework, which has become increasingly important in modern machine learning. MoE models allow for the combination of multiple specialized sub-models (experts) controlled by a gating network, enabling efficient scaling and task-specific performance. The paper's emphasis on the framework's utility and flexibility highlights its relevance for both research and practical applications, such as large-scale neural networks and multi-task learning.
The key innovation is the systematic delineation of the MoE framework, covering:
The abstract does not present specific experimental results, metrics, or comparisons. As a framework-level paper, its contribution is conceptual rather than empirical. The value lies in formalizing MoE as a general analytic tool rather than reporting state-of-the-art numbers.
This paper solidifies the mixture of experts as a core technique in machine learning, influencing subsequent work in large language models (e.g., Mixtral 8x7B), multi-task learning, and conditional computation. By clarifying the framework, it enables researchers and practitioners to apply MoE more effectively, leading to models that are both more powerful and computationally efficient.
Ashish Vaswani, Noam Shazeer et al.
Jakubův, Jan, Chvalovský, Karel et al.
Pauli Virtanen, Ralf Gommers et al.
Tom B. Brown, Benjamin Mann et al.