L-MAGIC by Intel logo

L-MAGIC by Intel

Paid

Language Model Assisted Generation of Images with Coherence

4.8
Inputs: image, text
Type
Saas
Company
Intel Labs

About L-MAGIC by Intel

L-MAGIC is a novel method presented at CVPR 2024 by Intel Labs for generating coherent 360-degree panoramic scenes from a single input image or text. It leverages large language models for guidance while diffusing multiple coherent views, using pre-trained diffusion and language models without fine-tuning (zero-shot). The output quality is enhanced by super-resolution and multi-view fusion. It accepts various input modalities including text, depth maps, sketches, and colored scripts, and can enable 3D point cloud generation and dynamic scene exploration.

Key Features

Leverages large language models for guidance
Zero-shot performance without fine-tuning
Accepts multiple input modalities (text, depth maps, sketches, colored scripts)
Super-resolution and multi-view fusion for enhanced output
Enables 3D point cloud generation and dynamic scene exploration

Pros & Cons

Pros
  • Better scene layouts and perspective view rendering quality compared to related works
  • Over 70% preference in human evaluations
  • Zero-shot: no fine-tuning required
  • Accepts diverse input modalities

Best For

Generate panoramic scenes from a single input imageText-to-panorama generation for various scenes like libraries, rooms, landscapes3D point cloud generation and dynamic camera motion exploration

Alternatives to L-MAGIC by Intel