MusicLM logo

MusicLM

Free

A model by Google Research for generating high-fidelity music from text descriptions.

FreeFree tier
Inputs: textOutputs: audio
Type
Open Source
Company
Google Research

About MusicLM

MusicLM by Google Research generates high-fidelity music at 24 kHz from text descriptions such as 'a calming violin melody backed by a distorted guitar riff'. It uses a hierarchical sequence-to-sequence modeling approach to produce music that remains consistent over several minutes. The model outperforms previous systems in both audio quality and adherence to text descriptions. MusicLM can be conditioned on both text and a melody, allowing it to transform whistled and hummed melodies according to a text caption. It also supports 'story mode' where sequential text prompts influence the continuation of semantic tokens. To support future research, the authors publicly released MusicCaps, a dataset of 5.5k music-text pairs with rich descriptions by human experts.

Key Features

Generates high-fidelity music at 24 kHz from text descriptions
Hierarchical sequence-to-sequence modeling for coherent long-form music
Consistent music generation over several minutes
Conditioning on both text and melody (whistled/hummed)
Story mode: sequential text prompts influence continuation of music
Outperforms previous systems in audio quality and text adherence
Publicly released MusicCaps dataset of 5.5k music-text pairs

Pros & Cons

Pros
  • Produces high-fidelity (24 kHz) audio with musical coherence
  • Maintains consistency over several minutes of music
  • Flexible conditioning: text alone or text with melody input
  • Outperforms prior state-of-the-art systems in quality and text adherence
  • Includes publicly available benchmark dataset (MusicCaps)

Best For

Generating music from rich text captions describing instruments, genres, mood, etc.Long music generation from a single text promptStory mode for creating narrative-driven music with sequential promptsTransforming whistled or hummed melodies according to a text style descriptionGenerating music conditioned on painting titles and descriptionsExploring generation diversity with same text prompt or same semantic tokens

FAQ

What is MusicLM?
MusicLM is a model by Google Research that generates high-fidelity music from text descriptions.
How does MusicLM work?
It casts conditional music generation as a hierarchical sequence-to-sequence modeling task, generating audio at 24 kHz.
Can MusicLM be conditioned on a melody?
Yes, by adding melody embeddings to the conditioning, the model can generate music that respects the text prompt while following a provided whistled or hummed melody.
What is MusicCaps?
MusicCaps is a dataset of 5.5k music-text pairs with rich descriptions provided by human experts, released alongside MusicLM to support future research.
Does MusicLM support long music generation?
Yes, the model generates music that remains consistent over several minutes, and story mode allows sequential prompts to guide longer compositions.