In Focus: Open Source models, China vs US Edition
Our analysis on the open source AI race between China vs US
In this week’s In Focus series, we are going to be focusing on the performance metrics for open source models for the two leading nations in the generative AI race, China and the United States. We will be focusing on open source model capabilities in reasoning, general knowledge, and agentic workflows.
The first half of 2025 has seen a key trend emerge in the generative AI landscape: open source models are not only here to stay, but in many cases, they can match or even surpass existing closed-models on baseline performance metrics such as programming, content creation, and mathematical reasoning. Open source models were initially touted for their verifiable nature: anyone, anywhere, can create AI applications with outputs that verifiable based on their weights and code.
However, the emergence of DeepSeek R1 earlier this year set forth a reality in which an open-source model, from a relatively, at the time, unknown lab, match and even outpace the performance of premier AI models from large labs. Ever since then, open source has become a key strategy of the majority of large labs in the United States, from Meta to even OpenAI, which released GPT OSS earlier this month.
Comparisons: single-turn, normative tasks
The premier Chinese open source models have actually surpassed the strongest US models in many normative tasks, including programming and mathematics. In fact, the majority of the top performers on MBPP Plus, even when you account for premier, non open-source models in the United States, are Chinese models, highlighted by Qwen 3:
This indicates that the Chinese models, for the most part, have caught up or exceeded the top US models when it comes to programming, especially on routine, day to day Python challenges.
An alternative trend manifests itself in general knowledge and reasoning, where US models tend to perform better. On Humanity’s Last Exam, it is actually GPT OSS that ranks at the top beating out even other, top performing closed source models.
Perhaps more interestingly, MiniMax’s M2 model is among the top here, indicating that even smaller labs are making headway in creating top tier models.
Comparisons: Agentic, multiturn tasks
When we shift our focus to agentic, multiturn tasks, the performance landscape becomes even more nuanced. These benchmarks evaluate models' abilities to maintain context over extended interactions, reason through complex problems incrementally, and function effectively as assistants in real-world scenarios.
Unlike single-turn evaluations that test isolated capabilities, multiturn benchmarks challenge models to demonstrate coherence, consistency, and adaptability across conversations. This closely mirrors how these models are actually used in production environments - rarely for one-off questions, but rather as ongoing assistants that must build upon previous exchanges.
For example, on Terminal-Bench, a benchmark that measures a model’s ability to effectively work on different tasks within a programming terminal, it was surprisingly GLM 4.5 that ended up with the highest score:
A similar trend can be found in the benchmarks for Berkley Function Calling, where GLM, GLM Air, and Qwen dominate:
Key Takeaways
From this comparative analysis of open source models, several key findings emerge:
Chinese open source models have demonstrated remarkable strength in programming and mathematical reasoning, with Qwen 3 leading the MBPP Plus benchmark, surpassing even closed-source US models.
US models maintain an edge in general knowledge and reasoning tasks, with GPT OSS topping the Humanity's Last Exam benchmark, though Chinese models from smaller labs like MiniMax are showing competitive performance.
Perhaps most significantly, Chinese models excel in complex multiturn, agentic tasks that more closely resemble real-world applications. GLM 4.5 leads in Terminal-Bench, while GLM variants and Qwen dominate Berkeley Function Calling benchmarks.
The rapid advancement of Chinese open source models indicates a potential shift in the global AI landscape, where China's technical capabilities in certain aspects of AI development may be outpacing those of US counterparts, particularly in areas that translate directly to practical applications.
This comparative performance suggests that the open source AI landscape is becoming increasingly multipolar, with different regional strengths emerging across various dimensions of model capability.
Conclusion
The comparative analysis of open source models from China and the US reveals a complex and evolving landscape in generative AI development. While Chinese models have demonstrated superior capabilities in programming and agentic tasks, US models maintain an edge in general knowledge and reasoning. This bifurcation suggests that the future of AI development may not be dominated by a single approach or region, but rather characterized by diverse strengths across different dimensions of model capability.
As open source models continue to gain prominence, both in performance and adoption, understanding their unique strengths and limitations becomes increasingly vital for organizations seeking to leverage AI effectively. The strong performance of models from smaller labs like DeepSeek and MiniMax also indicates that innovation in this space is not limited to resource-rich tech giants.
At LayerLens, we remain committed to providing objective, comprehensive evaluations of AI models through rigorous benchmarking and analysis. Our goal is to help enterprises navigate this complex landscape and make informed decisions about which models best suit their specific needs. As the generative AI race continues to accelerate globally, such independent assessment will only grow in importance.
Stay tuned for our next deep dive, where we'll explore emerging trends and capabilities in the rapidly evolving world of AI. For organizations looking to harness the full potential of these powerful tools, our team is ready to provide the insights and guidance you need to succeed in this transformative era.






