MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
FreeAbout MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
MiniMax-M1 is an open-weight, large-scale hybrid-attention reasoning model developed by MiniMax. It is built on a hybrid Mixture-of-Experts (MoE) architecture with a lightning attention mechanism, based on the MiniMax-Text-01 model (456B total parameters, 45.9B activated per token). The model natively supports a context length of 1 million tokens, 8 times that of DeepSeek-R1, and enables efficient scaling of test-time compute. It is trained using large-scale reinforcement learning (RL) on diverse problems including sandbox-based and real-world software engineering environments. The training process is accelerated by the proposed CISPO RL algorithm and completed on 512 H800 GPUs in three weeks with a rental cost of $534,700. Two versions are released with 40K and 80K thinking budgets, achieving comparable or superior performance to models like DeepSeek-R1 and Qwen3-235B on standard benchmarks, particularly excelling in complex software engineering tasks.
Key Features
Pros & Cons
- Open-weight model with state-of-the-art performance
- Massive 1 million token context length
- Efficient training with reduced cost and time
- Strong performance on complex software engineering benchmarks