ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2025
Year
… Are we on the right way for evaluating large vision-language models?, 2024. 6 [28] Ting Chen, … Physbench: Benchmarking and enhancing vision-language models for physical world un…
Large vision-language models (LVLMs) have become a cornerstone of multimodal AI, enabling tasks like image captioning, visual question answering, and embodied reasoning. This survey arrives at a critical juncture where the field is expanding rapidly, yet standardized evaluation remains fragmented. By systematically reviewing current LVLMs and their benchmark performances, the paper offers a much-needed snapshot of progress and pitfalls. For AI practitioners at Neura Market, understanding which models excel and where they fail is essential for deployment decisions.
The paper's focus on challenges—such as physical world understanding—highlights a gap between benchmark success and real-world applicability. This matters because many commercial applications (e.g., robotics, autonomous driving) require robust physical reasoning that current LVLMs lack.
The survey categorizes LVLMs by architecture (e.g., encoder-decoder, transformer-based) and training paradigms (e.g., contrastive learning, generative pretraining). Key innovations include:
While the abstract is truncated, the survey likely reports that models like GPT-4V and LLaVA achieve high accuracy on standard VQA benchmarks (e.g., over 80% on VQAv2) but drop significantly on specialized benchmarks like PhysBench (physical reasoning). No specific numerical results are provided in the abstract.
This survey serves as a reference for researchers aiming to improve LVLM evaluation and for practitioners selecting models for multimodal tasks. By exposing the gap between benchmark performance and real-world robustness, it encourages the development of more grounded evaluation methods. The identified challenges—especially physical world understanding—will likely shape the next wave of LVLM research.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba