MiMo
Xiaomi's open-source 7B reasoning language model series trained from scratch, achieving strong math and code reasoning performance through optimized pretraining and reinforcement learning.
At a Glance
About MiMo
MiMo is a series of open-source language models developed by Xiaomi's LLM Core Team, designed from the ground up for reasoning tasks. Released under the Apache 2.0 license, the project includes checkpoints for a base model, SFT model, and two RL-trained variants, all available on HuggingFace and ModelScope. The technical report (arXiv:2505.07608) details the full pretraining-to-posttraining pipeline that underpins the series.
What It Is
MiMo-7B is a family of 7-billion-parameter language models built specifically to maximize reasoning potential. Unlike approaches that apply reinforcement learning only to large base models (e.g., 32B), Xiaomi's team argues that reasoning capability is rooted in pretraining quality. The project covers the entire development pipeline: data preprocessing, multi-stage pretraining, supervised fine-tuning, and RL post-training with rule-based verifiers. The result is a compact model that the team reports matches OpenAI o1-mini on mathematics and code benchmarks.
Pretraining Architecture and Strategy
MiMo-7B-Base is pretrained on approximately 25 trillion tokens using a three-stage data mixture strategy. Key design choices include:
- Multi-dimensional data filtering to increase reasoning pattern density in pretraining data
- Massive synthetic reasoning data generation to supplement organic corpora
- Multiple-Token Prediction (MTP) as an auxiliary training objective, which the team reports enhances performance and accelerates inference via speculative decoding (approximately 90% acceptance rate with one MTP layer)
Post-Training Recipe
The RL training pipeline uses 130K curated mathematics and code problems verified by rule-based systems. Notable techniques include:
- Test difficulty-driven code reward: fine-grained scores for test cases at varying difficulty levels to address sparse reward problems in code tasks
- Data re-sampling for easy problems: improves rollout sampling efficiency and stabilizes policy updates in later RL phases
- Rule-based accuracy rewards only: avoids potential reward hacking from learned reward models
The team also built a custom Seamless Rollout Engine integrating continuous rollout, asynchronous reward computation, and early termination, which the README states achieves 2.29× faster training and 1.96× faster validation.
Update: MiMo-7B-RL-0530
As of May 30, 2025, Xiaomi released an updated checkpoint, MiMo-7B-RL-0530, with the SFT dataset scaled from approximately 500K to 6M instances and the RL training window expanded from 32K to 48K tokens. The README reports that MiMo-7B-RL-0530 achieves 80.1 on AIME 2024 (Pass@1), which the team states surpasses DeepSeek R1 (79.8) on that benchmark. Additional reported scores include 97.2 on MATH500, 70.2 on AIME 2025, and 60.6 on GPQA-Diamond.
Deployment and Inference
MiMo supports multiple inference backends:
- SGLang: officially supported via a merged pull request in the SGLang mainline, with MTP speculative decoding available
- vLLM: Xiaomi maintains a fork of vLLM (based on v0.7.3) with MTP support; a registry loader is also provided for standard vLLM without MTP parameters
- HuggingFace Transformers: standard
AutoModelForCausalLMloading withtrust_remote_code=True
The team recommends using an empty system prompt and temperature=0.6 for evaluation-consistent inference.
Model Variants Available
The repository provides four checkpoints:
- MiMo-7B-Base — base model with pretraining only
- MiMo-7B-RL-Zero — RL trained directly from the base model
- MiMo-7B-SFT — supervised fine-tuned from the base model
- MiMo-7B-RL — RL trained from the SFT model (the primary recommended variant)
Community Discussions
Be the first to start a conversation about MiMo
Share your experience with MiMo, ask questions, or help others learn from your insights.
Pricing
Open Source
Fully open-source under Apache 2.0. All model checkpoints freely available on HuggingFace and ModelScope.
- MiMo-7B-Base checkpoint
- MiMo-7B-SFT checkpoint
- MiMo-7B-RL-Zero checkpoint
- MiMo-7B-RL checkpoint
- MiMo-7B-RL-0530 checkpoint
Capabilities
Key Features
- 7B parameter reasoning-optimized language model
- Pretrained on ~25 trillion tokens
- Multiple-Token Prediction (MTP) for speculative decoding
- Rule-based RL post-training with 130K math and code problems
- Test difficulty-driven code reward for dense RL signal
- Seamless Rollout Engine for 2.29x faster RL training
- SGLang and vLLM inference support
- HuggingFace and ModelScope model distribution
- Apache 2.0 open-source license
- Four model checkpoints: Base, SFT, RL-Zero, RL
