Abstract
China’s open-weight LLM ecosystem extends well beyond DeepSeek to more than a dozen organisations releasing near-state-of-the-art, compute-efficient models under increasingly permissive licences, and their global downstream adoption now outpaces US open models — a shift with consequences for technology dependencies, AI governance, safety, and geopolitical competition that policymakers should ground in granular deployment data.
Key takeaways (as published)
- After years of lagging behind, Chinese AI models — especially open-weight LLMs — seem to have caught up or even pulled ahead of their global counterparts in advanced AI model capabilities and adoption.
- The brief profiles and compares four notable Chinese open-weight language model families, highlighting that China’s ecosystem of open-weight LLMs is driven by a wide range of actors prioritising computationally efficient models optimised for flexible downstream deployment.
- Diverse commercial strategies for translating open-weight model adoption into business success are emerging, yet their long-term viability remains uncertain.
- The Chinese government’s support of open-weight model development — while not the sole determinant of its success — has played a substantial role, though there is no guarantee it will continue.
- The widespread global adoption of Chinese open-weight models may reshape global technology access and reliance patterns, and impact AI governance, safety, and competition. Policymakers should ground their policy actions in a granular understanding of real-world deployment.
Background
- In late 2022 Chinese developers faced dual shocks: US export controls on semiconductor manufacturing equipment and top training chips (October), and the launch of ChatGPT (November).
- Chinese labs initially built on Meta’s Llama and its architecture, but did not stay reliant on US-trained models for long.
- DeepSeek’s R1 reasoning model, released just over two years after ChatGPT, demonstrated capability and efficiency that shocked investors and contributed to Nvidia suffering the largest single-day loss in US stock market history.
- Baidu, previously committed to closed models, has since opened the weights of some flagship models.
- The brief focuses on language models while noting strong Chinese performance in speech, visual, and video models.
Defining openness
- Model release sits on a gradient from fully closed to publicly downloadable under various licences. Sharing training data and pre-/post-training code is often argued necessary for true “open source”, but that level of openness is uncommon.
- “Open-weight” here means weights are free to download, use, and modify — enabling autonomous operation outside the developer’s app or API, and adaptation to new use cases while adopters retain ownership of their data environments.
- Openness offers practical advantages such as reproducibility and efficiency in low-resource environments that can outweigh marginal raw-performance gains. DeepSeek R1 and V3 appealed not as best-in-class but as free, cheap to deploy, and good enough.
Evidence of diffusion
- In September 2025 Alibaba’s Qwen family surpassed Llama as the most downloaded LLM family on Hugging Face.
- Between August 2024 and August 2025, Chinese open-model developers accounted for 17.1% of all Hugging Face downloads, slightly ahead of US developers at 15.8%.
- Since January 2025, derivative model uploads based on Alibaba and DeepSeek models have outpaced those based on major US and European models; derivatives of Alibaba models now exceed those from Google, Meta, Microsoft, and OpenAI combined.
- In September 2025, Chinese fine-tuned or derivative models made up 63% of all new fine-tuned or derivative models released on Hugging Face.
- Measurement is inherently hard: open models run on any sufficiently resourced infrastructure, so no central source captures full diffusion. Available proxies are download counts, derivative-model counts, and cloud inference usage (e.g. OpenRouter).
Technical realities: four conclusions
1. A wide range of actors
- More than a dozen Chinese organisations develop and release powerful models openly; the four profiled families are examples, and the leading set could change within months.
- Beyond DeepSeek and Alibaba: “tiger” AI unicorns (Z.ai, formerly Zhipu AI; Moonshot AI; MiniMax; Baichuan AI; StepFun; 01.AI); non-profit and university labs (e.g. Beijing Academy of Artificial Intelligence); and large tech companies (Tencent, Baidu, Huawei, ByteDance).
2. Near-state-of-the-art performance
- On Chatbot Arena as of 4 December 2025, top closed models from Google DeepMind, xAI, and OpenAI edged out the best from Z.ai, Moonshot AI, Alibaba, and Baidu — but by margins small enough that all were listed tied for first, alongside Anthropic’s top offering.
- Among open models only, 22 releases from five Chinese labs beat the top-ranked US open model, OpenAI’s gpt-oss-120b; Mistral was the only other non-Chinese developer in the top 25.
- Distinct strengths: Qwen3 in multimodal and multilingual work; DeepSeek R1 in step-by-step reasoning; Kimi K2 in coding and tool use; GLM-4.5 in balanced generalist capability via multi-expert training.
- Caveats: leaderboards are susceptible to gaming and hidden dynamics; benchmark results often rely on self-reported figures; CAISI found many leading models perform differently under independent verification; benchmarks lack standardisation and often fail to measure what they claim.
- Composite indices such as the Epoch Capabilities Index (39 benchmarks) and the Artificial Analysis Intelligence Index both show Chinese models catching up.
3. Prioritising computational efficiency
- Many Chinese developers release Mixture of Experts (MoE) models, whose efficiency squeezes better performance from limited compute — valuable under US export controls on advanced chips.
- MoE predates Chinese use (e.g. Mistral’s Mixtral), but DeepSeek’s December 2024 and January 2025 releases introduced a particularly effective fine-grained, shared-experts architecture; other Chinese developers optimise for chips available domestically.
- Release strategies track organisational scale: smaller developers (DeepSeek, Moonshot AI, Z.ai) concentrate on optimising a single flagship model for deployment stability, while Alibaba, backed by in-house cloud, sustains diversified families across sizes and modalities.
4. More permissive licences and diverse variants
- Qwen 2.5’s smallest 3B and largest 72B variants (2024) were research-only; DeepSeek V3 (December 2024) limited redistribution and large-scale commercial use. In 2025, Qwen3 and DeepSeek R1 are both more capable and released under Apache 2.0 and MIT respectively.
- Motivations: credibility in the global AI community, sustaining a vibrant open-source community, and competitive pressure — Baidu, whose CEO had championed proprietary models, opened Ernie 4.5 in June 2025.
- Variants now span smartphone-scale models (~1B parameters) to flagship MoE models (235B or even 1T). Qwen3, Kimi K2, and GLM-4.5 all introduced dual-reasoning modes toggling between step-by-step reasoning and direct response.
- Alibaba and Z.ai have added multimodal and vision-reasoning branches alongside code- and math-specialised text variants.
The four profiled families
| Qwen3 | DeepSeek-R1 | Kimi-K2 | GLM-4.5 | |
|---|---|---|---|---|
| Developer | Alibaba Cloud (Big Tech, group market cap ~$380B) | DeepSeek (startup, spun out of hedge fund High-Flyer; ~$15B speculative valuation) | Moonshot AI (startup, ~$4B, funded partly by Alibaba/Tencent) | Z.ai (startup from Tsinghua, ~$5.6B, big-tech and state backing) |
| Collection | Eight MoE and dense models, 0.6–235B; largest has 3B and 22B active-parameter variants; task-specialised versions; ~128K context | 671B MoE reasoning models (37B active), R1-Zero and R1; distilled Qwen/Qwen3/Llama models 1.5–70B; ~128K context | 1T MoE (32B active), ~256K context, Base and Instruct | Air (106B, 12B active) and Base (355B, 32B active); multimodal GLM-4.5V |
| First release | 29 April 2025 | 20 January 2025 | 11 July 2025 | 28 July 2025 |
| Licence | Apache 2.0 | MIT | Modified MIT (attribution for large-scale commercial deployment) | MIT |
| Emphasis | Multilingual (119 languages), multimodal, cost-efficient MoE, dual reasoning modes | Math and complex reasoning; exposed chains-of-thought; distillation scalability | Coding and agentic tasks; stabilised 1T training; fast low-cost inference | Generalist across reasoning, coding, agentic, visual; multistage expert training unified by self-distillation |
- All four release open weights, architecture, and code, with technical reports published between the day of launch and 2.5 weeks after. DeepSeek later published a peer-reviewed report including a 12-page safety report; Kimi K2’s report has a dedicated safety-evaluations section.
- Moonshot’s modified MIT licence defines large-scale commercial deployment as more than 100M monthly active users or more than $20M monthly revenue — an attribution requirement some argue counters the spirit of open source.
Commercial strategies
- Adoption does not track benchmarks alone: usage costs, deployability on existing infrastructure, and performance in specific workflows also drive decisions.
- Alibaba, China’s leading cloud provider, markets Qwen as an “AI operating system” underpinning enterprise solutions, citing HP and AstraZeneca as clients. Singapore’s national AI programme building its flagship LLM on Qwen3 may drive Southeast Asian traffic to Alibaba’s cloud.
- DeepSeek and Z.ai lack their own large-scale compute and pursue collaborative deployment across providers. DeepSeek has supplied on-premises deployments to Chinese public-sector clients; one March 2025 report counted at least 72 local government agencies running localised DeepSeek models, fine-tuned with local data administrations while maintaining a feedback loop with DeepSeek’s model team.
- Like Western counterparts, Chinese open-weight developers still rely on indirect monetisation — building large user bases that can be directed to paid products and services built on the models. Profitability remains unproven.
The Chinese government’s role
- Support for open technology traces at least to the 2017 New Generation Artificial Intelligence Development Plan, which championed “open source” and “open” collaboration to aggregate domestic strengths.
- A recent policy document encouraged state-directed academic organisations to weigh open-source contributions alongside publications in performance metrics.
- Internationally, the Global AI Governance Initiative (October 2023) and Global AI Governance Action Plan (July 2025) promote open-source AI and “equal rights to develop and use AI”, framed in implicit contrast to US export controls and closed models. Open-model sharing has featured in the Forum on China–Africa Cooperation.
- DeepSeek appears to have succeeded with limited — if any — direct state support; political acknowledgment (Premier Li Qiang inviting founder Liang Wenfeng to a high-level symposium) came only after the V3 release drew attention.
- Support has largely taken the form of an enabling environment rather than direct subsidy, though some local governments now offer tailored support to projects engaging open-source communities. Longstanding education policy and public research funding built a deep, increasingly domestically educated talent pool; the state has also supported computing infrastructure build-out.
- Continuation is not guaranteed: DeepSeek executives have reportedly faced travel restrictions, and the tech-platform crackdown from late 2020 illustrates the balance between enabling innovation and controlling the private tech sector. A major adverse event or national security threat could prompt rapid constraint.
Global policy implications
Global access and dependencies
- Wide availability opens pathways for less computationally resourced organisations and countries to access advanced AI, shaping diffusion and cross-border reliance.
- With frontier performance converging, adopters with limited resources may prioritise affordable, dependable access; “good enough” Chinese models under permissive licences with competitive usage costs meet that demand. US companies, from large tech firms to hyped startups, are also adopting them — potentially decreasing global reliance on US API providers.
- Open weights do not eliminate dependency: Huawei and ZTE’s history of supplying connectivity and data-centre projects in Global South markets carries over, and Huawei markets DeepSeek integration in its cloud offerings in South Africa and elsewhere in sub-Saharan Africa. For “sovereign AI” efforts, an open model delivered through a foreign firm’s cloud still raises long-term dependency concerns, depending on how much of the stack is bundled.
AI governance
- Concerns include spreading Chinese political censorship or propaganda, data security, and cybersecurity risk.
- Chinese content restrictions and domestic regulation (e.g. model registration) shape work from data collection through pre-training and fine-tuning. Investigations have found censorship guardrails removable fairly easily, and many enterprise adopters are undeterred, but the potential for Chinese state preferences to shape user experience worldwide is real.
- Although open models can be run outside creators’ infrastructure, many users will rely on the developers’ apps, APIs, and integrated solutions — placing data under those companies’ control and potentially moving it to China. Running locally, on a trusted cloud, or via a trusted inference provider such as Hugging Face substantially mitigates this.
- There is no verified evidence of deliberate backdoors in advanced AI systems, but past cases involving Chinese technology such as Hikvision will continue to fuel speculation.
AI safety
- Some evidence suggests Chinese developers are less focused on AI safety: a CAISI evaluation found DeepSeek models on average 12 times more susceptible to jailbreaking than comparable US models, and independent evaluations show DeepSeek guardrails are easily bypassed.
- Given far-reaching diffusion, what happens on AI safety in China will not stay in China; policymakers seeking to understand or reduce risks likely need to engage with Chinese labs and regulators.
Geopolitical competition and the open-closed debate
- Until recently major US developers prioritised closed releases while the US policy community debated open-release risks.
- After R1, President Trump called it “a wake-up call”; AI czar David Sacks cited it to justify a deregulatory federal approach.
- July 2025’s America’s AI Action Plan elevates open-weight models as a strategic asset while emphasising stronger export controls on adversaries. The following month OpenAI released two open-weight models under Apache 2.0 — its first open release since GPT-2 nearly six years earlier. Several newer US startups and academic labs have since committed to open-weight development.
- Some of these shifts were already underway, but the DeepSeek moment likely expanded willingness to accept open-weight risks in order to compete, and fostered recognition that AI leadership depends on the reach, adoption, and normative influence of open-weight models, not only proprietary systems.
Conclusion
- Every implication is marked by uncertainty, but questions about who uses which technologies, how, and under what architectures of control or reliance are both answerable and urgent.
- Policymakers and researchers need to move beyond basic Hugging Face and GitHub data: outcomes depend on architecture and control of the hardware and software systems that turn models into services.
- China’s edge in open-weight adoption is only about a year old and may not last; continued collection of quantitative and qualitative deployment data is needed, requiring technical, industry, and regional domain knowledge.
- Policies aimed at competing with China or spreading US alternatives should be grounded in granular understanding of actual deployment and local needs, and should stay flexible. Researchers should look beyond language models to vision and code-assist models.
- Selective engagement with Chinese labs, academics, and policymakers should not be avoided or unnecessarily constrained; engagement on nontechnical topics such as incident reporting and risk management frameworks is possible without leaking sensitive technical insight. The nature of open-model releases itself enables better scrutiny, leaving space for academic collaboration on open-weight risks and guardrail efficacy.