Thinking with Generated Images
FreeEnabling LMMs to think visually by generating intermediate image-based reasoning steps.
About Thinking with Generated Images
Thinking with Generated Images is a novel research paradigm that enables large multimodal models (LMMs) to natively reason across text and vision by spontaneously generating intermediate visual thinking steps. It introduces two complementary mechanisms: (1) vision generation with intermediate visual subgoals, where complex visual tasks are decomposed into manageable components generated and integrated progressively; and (2) vision generation with self-critique, where models generate an initial visual hypothesis, analyze its shortcomings through textual reasoning, and produce refined outputs. The approach achieves up to 50% relative improvement on complex multi-object scenarios in vision generation benchmarks. An open-source suite is released to support further research and applications in fields such as biochemistry, architecture, forensics, and sports strategy.
Key Features
Pros & Cons
- Significant improvement on complex multi-object scenarios (up to 50% relative gain)
- Enables models to engage in visual imagination and iterative refinement
- Combines text and vision reasoning natively without external tools
- Open-source suite available for community use and extension
- Currently a research paradigm; may not be production-ready
- Requires substantial computational resources for image generation and iterative refinement
- Effectiveness may vary across different types of visual reasoning tasks
- Dependence on underlying LMM capabilities for generation and critique