Gemini 1.5 Pro
Unknown
Gemini 1.5 Pro is a compute-efficient multimodal mixture-of-experts model excelling in long-context retrieval and understanding across text, video, and audio.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
Gemini 1.5 Pro is a compute-efficient multimodal mixture-of-experts model excelling in long-context retrieval and understanding across text, video, and audio.
Unknown
MoE-LLaVA introduces a sparse mixture-of-experts framework for vision-language models, activating only top-k experts per token to match larger model performance with fewer parameters.
Unknown
Introduces Mixtral 8x7B, a sparse mixture-of-experts language model trained on multilingual data with a 32k-token context window.
Unknown
FuseMoE introduces a mixture-of-experts framework with a novel gating function for integrating diverse numbers of modalities.
Unknown
This paper empirically studies the design of vision Mixture-of-Experts (MoE) models, providing insights into routing strategies and expert architectures for improved scalability.
Zixiang Chen, Yihe Deng, Yue Wu, et al.
This paper formally proves that Mixture-of-Experts layers provably improve test accuracy over a single expert in two-layer CNNs.
Siyuan Mu, Sen-Fon Lin
A comprehensive survey of Mixture-of-Experts architectures covering algorithms, theory, and applications.