For a long time, Chinese AI models could be roughly divided into three routes. The first route is actively making models smaller. For example, minimax-m3, kimi k2.5, hy3, and other models that actively limit activated parameters to around 20B. Whether 20B activated parameters count as small is itsel