Sora logo

Sora

Free

Presentation of Sora, a large video generation model. OpenAI, February 15, 2024.

FreeFree tier
Inputs: textOutputs: video, image
Type
Open Source
Company
OpenAI

About Sora

Sora is a large-scale text-conditional diffusion model developed by OpenAI for video generation. It uses a transformer architecture that operates on spacetime patches of video and image latent codes, enabling generation of up to one minute of high-fidelity video with variable durations, resolutions, and aspect ratios. The model is trained on a compressed latent space via a video compression network and can also generate images as single-frame videos. The technical report discusses the method and provides a qualitative evaluation of Sora's capabilities and limitations, positioning it as a step toward building general-purpose world simulators.

Key Features

Text-conditional diffusion model trained on videos and images
Transformer architecture operating on spacetime patches
Generates up to one minute of high-fidelity video
Supports variable durations, resolutions, and aspect ratios
Trained on compressed latent space via a video compression network
Can generate images as single-frame videos
Scalable diffusion transformer design
Generalist model for diverse visual data

Pros & Cons

Pros
  • High-fidelity video generation up to 60 seconds
  • Flexible across different durations, resolutions, and aspect ratios
  • Scalable architecture based on diffusion transformers
  • Generalist model not limited to narrow visual categories
Cons
  • Technical report lacks model and implementation details
  • Limited to one minute of video generation
  • Still in research stage, not a commercial product
  • May have limitations in accurately simulating physical dynamics

Best For

Video generation from text promptsImage generation as single-frame outputsResearch in generative video models and world simulationExploring scalable training for physical world simulators

FAQ

What is Sora?
Sora is a large-scale text-conditional diffusion model developed by OpenAI for video generation. It can generate up to one minute of high-fidelity video from text prompts, using a transformer architecture on spacetime patches.
How does Sora work?
Sora compresses videos into a latent space, extracts spacetime patches as tokens, and uses a diffusion transformer trained to predict clean patches from noisy ones, conditioned on text prompts.
What are Sora's capabilities?
Sora can generate videos and images of variable durations, resolutions, and aspect ratios, up to 60 seconds. It is designed as a generalist model for diverse visual data.
Is Sora publicly available?
As of the technical report (February 2024), Sora is presented as a research model. The website does not indicate public availability or a product release.