Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
FreeScaling transformers to trillions of parameters with sparse mixture-of-experts
FreeFree tier
About Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Switch Transformers is a research paper that introduces a sparse mixture-of-experts (MoE) architecture for scaling transformer models to trillions of parameters with simple and efficient sparsity. The paper is available as an open-access preprint on arXiv.