Thinking with Generated Images logo

Thinking with Generated Images

Free

Enabling LMMs to think visually by generating intermediate image-based reasoning steps.

FreeFree tier
Inputs: text, imageOutputs: image, text
Type
Open Source

About Thinking with Generated Images

Thinking with Generated Images is a novel research paradigm that enables large multimodal models (LMMs) to natively reason across text and vision by spontaneously generating intermediate visual thinking steps. It introduces two complementary mechanisms: (1) vision generation with intermediate visual subgoals, where complex visual tasks are decomposed into manageable components generated and integrated progressively; and (2) vision generation with self-critique, where models generate an initial visual hypothesis, analyze its shortcomings through textual reasoning, and produce refined outputs. The approach achieves up to 50% relative improvement on complex multi-object scenarios in vision generation benchmarks. An open-source suite is released to support further research and applications in fields such as biochemistry, architecture, forensics, and sports strategy.

Key Features

Native reasoning across text and vision through generated intermediate visual thoughts
Vision generation with intermediate visual subgoals for decomposing complex tasks
Vision generation with self-critique for iterative refinement of visual hypotheses
Up to 50% relative improvement on complex multi-object visual reasoning benchmarks
Open-source suite for further research and application development

Pros & Cons

Pros
  • Significant improvement on complex multi-object scenarios (up to 50% relative gain)
  • Enables models to engage in visual imagination and iterative refinement
  • Combines text and vision reasoning natively without external tools
  • Open-source suite available for community use and extension
Cons
  • Currently a research paradigm; may not be production-ready
  • Requires substantial computational resources for image generation and iterative refinement
  • Effectiveness may vary across different types of visual reasoning tasks
  • Dependence on underlying LMM capabilities for generation and critique

Best For

Biochemists exploring novel protein structuresArchitects iterating on spatial designsForensic analysts reconstructing crime scenesBasketball players envisioning strategic playsGeneral visual reasoning tasks requiring intermediate visual steps

FAQ

What is Thinking with Generated Images?
It is a paradigm that allows large multimodal models to reason across text and vision by spontaneously generating intermediate visual thinking steps as part of their reasoning process.
How does Thinking with Generated Images work?
It works through two complementary mechanisms: vision generation with intermediate visual subgoals, and vision generation with self-critique, where models generate, critique, and refine visual hypotheses.
What improvements does it offer?
The approach achieves up to 50% relative improvement (from 38% to 57%) in handling complex multi-object scenarios on vision generation benchmarks.
Is the code or suite available?
Yes, the authors release an open-source suite for the paradigm as mentioned in the paper.