ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2025
Year
… In this paper, we conduct a comprehensive survey of the most recent work on multimodal RAG (MM-RAG) in the sense that it has a full coverage of almost all the combinations of …
Multimodal retrieval-augmented generation (MM-RAG) is an emerging paradigm that extends traditional RAG to handle multiple modalities (text, image, audio, video, etc.) in both input and output. This survey is significant because it provides a structured overview of a rapidly growing field, which is crucial for researchers and practitioners to navigate the fragmented literature. By covering all combinations of input and output modalities, the paper fills a gap in existing surveys that often focus on specific modality pairs (e.g., text-to-image) or single-modality RAG.
The paper's comprehensive scope is timely, as multimodal AI systems are becoming increasingly prevalent in applications like visual question answering, multimodal chatbots, and content generation. Understanding the design space of MM-RAG is essential for building robust systems that can retrieve and generate across modalities. This survey serves as a roadmap, helping the community identify what has been done and what remains unexplored.
The paper's main contribution is a taxonomy that categorizes MM-RAG approaches based on the modality of the query (input) and the modality of the generated response (output). This includes combinations such as text-to-text, text-to-image, image-to-text, image-to-image, and more complex multi-modal inputs and outputs. The survey also discusses key components of MM-RAG systems, including retrieval strategies, fusion mechanisms, and generation models.
As a survey, the paper does not present experimental results but rather synthesizes findings from existing literature. It provides a structured analysis of the state of the art, identifying trends such as the dominance of text-image combinations and the growing interest in video and audio modalities. The survey also notes the lack of standardized benchmarks for MM-RAG, which is a critical gap for progress.
The broader impact of this survey is to accelerate research in multimodal RAG by providing a clear map of the field. It enables researchers to quickly identify relevant work and gaps, fostering innovation. For practitioners, it offers a guide to selecting appropriate MM-RAG architectures for their use cases. As multimodal AI continues to evolve, this survey will serve as a foundational reference, shaping future research directions and contributing to the development of more capable and versatile AI systems.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba