Llava-Mini
Unknown
Llava-Mini reduces vision tokens by pre-fusing CLIP visual information into text tokens and using query-based compression for efficient multimodal processing.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
Llava-Mini reduces vision tokens by pre-fusing CLIP visual information into text tokens and using query-based compression for efficient multimodal processing.
Unknown
LLaVA 1.6 enhances multimodal AI with dynamic high-resolution processing, improved OCR, and scalable LLM integration.
Unknown
MoE-LLaVA introduces a sparse mixture-of-experts framework for vision-language models, activating only top-k experts per token to match larger model performance with fewer parameters.
Unknown
LLaVA 1.5 enhances multimodal AI by integrating CLIP-ViT-L-336px with MLP projection and academic VQA data, achieving state-of-the-art results.
Unknown
LLaVA connects CLIP and Vicuna via end-to-end training on GPT-4-generated instruction-following data from image-text pairs.