Article 78DZH China's open-weight AI models are now just 4 months behind frontier US offerings, Mozilla report claims — models still lag in some benchmarks but are drastically cheaper to use

China's open-weight AI models are now just 4 months behind frontier US offerings, Mozilla report claims — models still lag in some benchmarks but are drastically cheaper to use

by
Shane Downing
from Latest from Tom's Hardware on (#78DZH)

Mozilla has published version 1.1 of its State of Open Source AI report on Sept. 15 using data current to Sept. 1, revealing that many of the best Chinese open-weight AI models are closing the gap with U.S. frontier offerings. The best open model trailed the closed leader on the Artificial Analysis Intelligence Index by three points at 60% of the price and two points behind Claude Fable 5 at 30%. Mozilla's fit on METR task-horizon data puts the open-closed gap at around 4.4 months, in line with Epoch AI's four-month estimate.

Go deeper with TH Premium: AI and data centers

Vh4nY3pMCcmra2ymXah9S7-1920-80.jpg

(Image credit: Microsoft)

Mozilla is the nonprofit behind the Firefox web browser, and its report is a recurring assessment first published on July 14 on the Mozilla blog. It's built on a Mozilla/SlashData survey of roughly 1,400 developers along with OpenRouter traffic data and third-party benchmark indices. Mozilla is an advocate for open models, and TIME reported on July 14 that Raffi Krikorian, Mozilla's chief technology officer, described the report as partly advocacy. Open weights" in this context means downloadable weights rather than training data or code. The report counts 16 notable open releases, but none delivers the data recipe required by the Open Source Initiative's definition.

The four-month figure rests on METR, which is a research nonprofit that scores models by the length of task, in human working time, they complete half the time. By Mozilla's fitted estimate, closed models handle tasks that take human experts 8 to 12 hours. Open models reach that about four months later, with open capability doubling every 3.9 months versus 5.5 for closed, by Mozilla's computation. Mozilla also charted vals.ai's Terminal-Bench 2.1 results, which run every model through the same harness, or software layer that offers a model its tools. On that board, Z.ai's GLM-5.2 scored within a point of Claude Opus 4.7 and about four points behind Opus 4.8, at less than one-fifth the cost per test. On OpenRouter, a marketplace that routes developer traffic to hundreds of models, Mozilla counted eight of the top ten models by August token volume as open weights, seven of them Chinese-built. Nevertheless, closed providers took 96% of model-layer revenue on OpenRouter from May-September 2025, the Linux Foundation reported. We see the decision to pay for closed [models] as workload-specific rather than organization-specific," Krikorian told Ars Technica in an email.

ExsxWrUXoBVvTYXeppQoaf-1920-80.png

(Image credit: Mozilla)

One caveat is that the four-month gap and the 30% token price figure are measured API to API on hosted endpoints and at list price. The report's own hardware chart puts the best open model that fits one server at 52.6 and the best on one GPU at 40. The drop from the top is 10 and 23 points, respectively, a larger gap than the reported four months. Kimi K3's native MXFP4 checkpoint runs about 1.56TB across 96 shards, and Mozilla's serving configuration lists 64 or more accelerators, while vLLM calls for at least eight GB300 GPUs, with multiple nodes for production traffic. The report describes this as open but not runnable by most who hold it, and Tom's Hardware put the memory need near 1.5TB in July. One example exception is Thinking Machines' Inkling-Small model, under the Apache 2.0 license, whose NVFP4 version fits one B300 at a 180GB floor.

The report's data stops at Sept. 1. Since then, Artificial Analysis has moved its index to v4.3 with a different evaluation set. The live board has Claude Fable 5.1 at 53 on its highest effort setting with Kimi K3 at 44, not comparable to the v4.1.1 numbers Mozilla plotted. vals.ai's Terminal-Bench 2.1 board, updated Sept. 11, is now led by GPT-6 Astra at 87.27% with Fable 5.1 at 85.02%. Mozilla's own chart caption reads: the gap resets every release cycle." K3 also carries an allegation detailed in the Sept. 8 NSA/CISA/FBI joint advisory (AA26-251A). The claim, which Mozilla's report states as asserted, and unshown," is that Moonshot extracted Claude Fable 5 data to train K3 through distillation, the practice of training one model on another model's outputs. On July 17, Artificial Analysis had K3 at 57 versus Fable 5's 60, while on Sept. 1, Mozilla had it two points back.

External Content
Source RSS or Atom Feed
Feed Location https://www.tomshardware.com/feeds/all
Feed Title Latest from Tom's Hardware
Feed Link https://www.tomshardware.com/feeds.xml
Reply 0 comments