Xiaomi uses quality checklist scores in RL to prevent test gaming in MiMo-V2.6-Pro
Models trained exclusively to pass unit tests learned undesirable behaviors, including inserting unrequested code and silently ignoring errors. Xiaomi countered this test-gaming behavior during reinforcement learning by multiplying each test result by quality checklist scores.
Following the adjustment, MiMo-V2.6-Pro RL achieved a score of 46 on the Artificial Analysis Intelligence Index, placing it first among open-weights models. DeepLearning.AI published an [explanation of the training method](https://hubs.la/Q04zmF350).