State of Open Models: Summer 2026 — Chinese frontier leads, Qwen dominant

A biannual analysis covering January–August 2026 reports rapid change across the open-model ecosystem. Public model repositories rose from 2.43 to 2.96 million, datasets from 711,000 to 1 million, and Spaces from 1.00 to 1.44 million. The report notes extreme concentration: roughly 85.6% of models have fewer than 200 lifetime downloads, while 1.5% of repositories account for 99.2% of all downloads.

Frontier dynamics and licensing

The analysis finds that several Chinese labs bypassed the traditional progression from small to large models, publishing monthly ceilings between 754B and 2.78 trillion parameters. By contrast, U.S. monthly ceilings stayed under 130B in five of seven months, with exceptions including NVIDIA’s Nemotron 3 Ultra (561B in May and June) and Thinking Machines’ Inkling.

Hardware vendors became prominent publishers: AMD and NVIDIA each released more than 200 new model repositories, and LiquidAI released around 100. The report interprets this as vendors using open models to demonstrate hardware performance. At the same time, many large Chinese releases are licensed permissively: of 178 Chinese releases above 20B parameters in 2026, 59% used Apache 2.0 and 22% used MIT, with none carrying a non-commercial restriction. Comparable U.S. releases in the same size band showed 29% Apache/MIT, 41% custom terms and 30% declaring no license.

Adoption, deployment and agents

Adoption metrics diverge: the top 25 repositories by downloads and by likes overlap in only one case. Downloads favor small, stable models embedded in pipelines; likes skew toward recent frontier releases. Models under 1B parameters account for 83% of all-time downloads, while models above 100B take 1%. Restricting to 2026, only 3% of download volume went to models above 70B.

Qwen has become a central base model: the report records 151,448 Qwen-based derivatives on the Hub, growing at roughly 180–210 new repositories per day. Qwen GGUF builds saw about 39.6 million monthly downloads, compared with Gemma’s 20.8 million and Llama’s 7.5 million.

Tooling and agents are reshaping use. The ggml team’s February move and the rise of GGUF and llama.cpp enabled local inference of very large models, including GGUF builds of DeepSeek-V4-Flash (~284B) and Kimi-K3 (~2.8T). An agent-usage dataset shows volatile harness shares (Claude Code led July at 44.4% after holding 67.8% in April). The analysis also documents a recorded autonomous-agent intrusion and its subsequent technical disclosure.

The report covers activity from January through August 2026 and was published on August 14, 2026.


Original source: Hugging Face Blog

Leave a Comment