Sharp: Steering hallucination in lvlms via representation engineering
Unknown
SHARP is a representation-level intervention framework that modulates hallucination in LVLMs by steering internal representations.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
SHARP is a representation-level intervention framework that modulates hallucination in LVLMs by steering internal representations.
Unknown
Proposes MRRE, a training-free inference-time method using representation engineering to enhance multilingual reasoning in LLMs and LVLMs without additional training data.
Unknown
This paper explores how generative world models can provide priors to vision-language models (VLMs) for improved understanding of world dynamics in common vision tasks.
Unknown
FastVLM introduces an efficient vision encoding method that reduces computational cost while maintaining high performance in text-rich image understanding tasks.
Unknown
This paper introduces LVLM-eHub, a comprehensive evaluation benchmark for Large Vision-Language Models, systematically assessing their capabilities across diverse multimodal tasks.
Unknown
This paper assesses the effectiveness of recent large vision-language models (LVLMs) in achieving artificial general intelligence, focusing on their performance across various tasks.
Unknown
This paper proposes an efficient and unbalanced multi-agent collaboration framework for long-horizon planning, using a final-outcome-based reward and a VLM evaluator to provide stable feedback.
Unknown
This paper evaluates closed- and open-source VLMs as planners, showing that spatially grounded long-horizon planning remains a major unsolved challenge.
Yang Zhou, Zixuan Huang, Sunzhu Li, et al.
SpatialCLI teaches VLMs to reason with spatial tools and progressively internalize specialist perceptual capabilities, raising Qwen3-VL-8B-Instruct from 29.3% to 84.6% on MindCube.
Ioannis Maniadis Metaxas, Adrian Bulat, Alberto Baldrati, et al.
UltraViT is a latency-optimized vision encoder for LVLMs using a pyramidal architecture with heterogeneous spatial mixers and a two-stage generative pre-training strategy, achieving 1.7x speedup on-device.
James Y. Huang, Sheng Zhang, Qianchu Liu, et al.
Proposes BeMyEyes, a multi-agent framework that uses a small VLM as a perceiver and a text-only LLM as a reasoner to achieve multimodal reasoning without training large-scale models.
Yifei Zhao, Xiangxin Zhou, Wenhao Yang, et al.
SceneActBench benchmarks VLM agents on acting in multi-object 3D scenes via a unified agent-environment loop, revealing no model performs consistently across five tasks.