About MoMask
MoMask is a novel masked modeling framework for text-driven 3D human motion generation, developed by researchers at the University of Alberta and Google Research and published at CVPR 2024. It employs a hierarchical quantization scheme to represent human motion as multi-layer discrete motion tokens with high-fidelity details. A Masked Transformer predicts randomly masked base-layer motion tokens conditioned on text input during training and iteratively fills in missing tokens during inference. A Residual Transformer progressively predicts next-layer tokens based on current layer results. MoMask achieves state-of-the-art performance on text-to-motion generation, with an FID of 0.045 on HumanML3D (vs 0.141 for T2M-GPT) and 0.228 on KIT-ML (vs 0.514). It also supports text-guided temporal inpainting (e.g., inbetweening, prefix, suffix) without requiring additional fine-tuning.
Key Features
Pros & Cons
- Outperforms existing methods on text-to-motion benchmarks
- Supports temporal inpainting without additional model fine-tuning
- Hierarchical quantization enables high-fidelity motion detail
Best For
Alternatives to MoMask
Reallusion Character Creator
Create realistic 3D characters, animate for movies and games, with a library of thousands of drag-and-drop items.
AI Shopify Product Reviews
Boost Sales Instantly With Automated Social Proof
Vector Magic
Convert JPG, PNG images to SVG, EPS, AI vectors
Plask AI
Plask.ai uses AI motion capture to animate characters from video, simplifying 3D animation workflows for creators and studios.
PlugSugar
Automate conversations, answer questions with Web Search plugin, and customize ChatGPT experience using powerful AI plugins.
Apploi
Simplify recruitment, improve job seeker visibility, and help employers find the perfect candidates efficiently and affordably.