← Back to the blog

China and Meta just made frontier AI free to download

Every few weeks a Chinese AI lab drops a new model and the same thing happens. The markets wobble, the headlines call it a crisis for America, and by Thursday everyone's moved on.

Toby Remond14 August 20268 min read
China and Meta just made frontier AI free to download

China and Meta just made frontier AI free to download

Every few weeks a Chinese AI lab drops a new model and the same thing happens. The markets wobble, the headlines call it a crisis for America, and by Thursday everyone's moved on.

The very top American and Chinese labs are close now, and the gap keeps shrinking. That isn't the bit I'd watch though. I'd watch the open weight models, the ones you can download and run for free on hardware you control.

Handing models out for free isn't new. Zhipu open-sourced ChatGLM-6B in March 2023, Meta's first Llama a month before that, Alibaba shipping Qwen from August 2023 and DeepSeek from November 2023. What's changed is the tier they arrive at. Free downloads used to trail well behind the paid frontier. Now they turn up within weeks of it, each claiming to rival what the big American labs charge serious money for, while costing a lot less to run.

That's genuinely remarkable. You can download a model its makers say is close to the best available, pay nothing for it, run it on machines you own, and never let a client file leave your building. Two years ago that was a fantasy.

The catch is the machines. Doing it properly costs proper money in hardware, and most businesses underestimate how much. I'll come back to the cost, because it never makes the headline.

What's been released, oldest first

It starts in late June, when Zhipu, which also trades as Z.ai, released GLM-5.2 as a genuine open weight model, weights and all. Researchers found it goes toe to toe with a leading commercial model at hunting software bugs and on some security work, though it trails Anthropic and OpenAI on the broader stuff. That pattern holds for most of them, excellent at one narrow job and fairly ordinary at the rest.

Late July, and Moonshot AI released Kimi K3. The tech news site The Verge reported it was said to outperform some of the top American systems at a much lower running cost. "Said to" is doing a lot of work there, and it's their claim, not mine. MIT Technology Review, writing about the political fallout, called Kimi free and open source and said it reportedly rivals paid models from OpenAI and Anthropic.

Then on 3 August, Alibaba released Qwen3.8-Max, calling it their largest and most capable model to date, with the same claim that it rivals the US frontier labs and their own domestic rival, Kimi K3.

A week later, on 10 August, Meta released Muse Glimmer, and this is the one I'd want a UK business owner thinking about. It handles both text and images, it's built for reading documents and running assistants on your own machines, and it works from day one with the standard free tools. It's under an Apache 2.0 licence, which means any business can use it for free, no permission needed. Meta says it runs on your own laptop. They mean a seriously well specced one though, not the average office machine.

Which brings us to 14 August and GLM-5.3, Zhipu again. This is where things stand now, and why GLM-5.2 is already history. Zhipu call 5.3 "the strongest open-weights coding model" and claim roughly a 50% jump on 5.2 for coding. You can't download it yet, though. It went out through their coding service and API first, weights to follow about two weeks later once a security review is done. So today it's something you rent, and check back when the files turn up.

Five releases in about seven weeks. The pace is the story, not any single one of them.

Don't trust the test scores too much

A figure marks their own exam paper with a gold star, a faint mirror image of themselves sitting opposite

Almost every number published in the last two months has one thing in common. It's marked by the same people selling the model. Muse Glimmer's spec sheet is a decent example, and to Meta's credit they publish the table. On the standard tests labs use to score models, called benchmarks, Glimmer scores 75.5 on MCP Atlas against 54.2 for Gemma4-31B and 62.5 for Qwen3.6-27B, 51.2 on SWE-Bench Pro and 94.7 on the AIME 2026 maths test. All public, go and check them. They're also numbers Meta picked, against comparison models Meta picked.

So treat a test score the way you'd treat a competitor's case study on their own website. It tells you the model can do a task that looks like the test. It tells you nothing about whether it can handle your invoice reconciliation, your case file summary or your quoting process, on your own messy paperwork. The only test that actually decides anything is one you write yourself, on your own examples, with a right answer agreed in advance.

Why give the models away for free

Part of it's about who gets used the most. If your model becomes the free default developers build on, you set the standard even while you're behind on raw ability. Giving it away costs less than out-marketing OpenAI.

On the day Glimmer went live, Mark Zuckerberg said the quiet part out loud. In a video on his own Instagram he opens with this: "I think that the key to building a positive future for everyone is to make sure that everyone has access to personal superintelligence." He calls Glimmer the best performing model of its size, 30 billion of the settings that make an AI model work, small enough for your own laptop rather than a data centre, and says Muse Spark 1.2 goes free in the coming weeks. Then he gets to the app, where Meta's latest models are already live and where it's building a personal superintelligence agent to run around the clock on your health, relationships, career and finances. Free download and paid always-on service, same breath. Take the free model, it's a good thing to have. Just be clear which of those two is the business.

The other reason is computing power. US rules still limit what chips Chinese labs can buy, and Ars Technica reported in July that DeepSeek plans its own chips to cut reliance on Nvidia and Huawei. The US side isn't settled either. MIT Technology Review reported current and former Trump advisers publicly having a go at the country's own leading labs, a White House check on model safety before release, loosened chip export rules, and an April crackdown on labs copying a big model's abilities into a smaller, cheaper one. That's a lot of infighting for something people keep calling a coordinated race.

You can already run frontier AI in your own building

If your answer to AI has always been "our client data can't leave the building", this run of releases has made that objection much weaker. These models cost nothing to download, you own the hardware, and no data goes out to anyone. Muse Glimmer was built for exactly that. The software side used to be the hard bit and isn't any more. So if you're a regulated firm, or holding client files you've promised never to hand to a third party, the conversation that ended with "we can't, because of the data" is worth reopening.

And here's what that costs you

The model is free. The machine is not. For the smaller models you're looking at a high-spec machine with plenty of memory, real money up front but not frightening. For the big ones, the ones that get you frontier quality, you're into proper server grade kit. That's a capital purchase, landing on the balance sheet rather than disappearing into next month's software spend.

Then there's everything after. The models keep moving, so somebody has to keep yours current and own it once the novelty wears off. That's the part that gets waved through in the planning meeting and quietly falls on whoever is nearest to IT.

So be honest about which of these you are. Running a handful of document jobs a week, and renting access stays cheaper and far less hassle. Processing serious volume, or holding data that can't leave the building, and owning the hardware starts to make sense, and the sums are worth doing properly.

None of that makes it a bad idea, it just puts a price on the decision, and you want that price in front of you before you commit. What's changed is that you get to make the decision at all. That's the actual news.

What this means for your business

A hand swaps a dull grey lightbulb for a bright gold one while the desk lamp stays lit

Almost nothing here is about who's winning. Three things follow for whatever you're building.

First, the price will keep falling. Whatever you're quoted today for pay-as-you-go AI, where you pay per bit of text the model reads and writes, will look expensive in a year. Build on that assumption and don't sign a long deal on today's rates.

Second, swapping the model underneath your workflow should be as easy as changing a setting. If it means a rebuild, you bought the wrong thing. What decides whether the system works is everything around the model. Your own tests to check it's getting things right, the way it finds and reads your documents, the human check before anything goes to a customer. The model itself is close to a commodity now. This is why we build systems that don't depend on any one AI model, the model changes far more often than the process it serves.

Third, rent or own is a live decision again. It used to answer itself. Now run the numbers properly, once, on your own figures.

The scoreboard will keep moving whether you watch or not. Whether Kimi or Qwen or a US lab is ahead this month won't change a single decision in your business. Whether the thing you've built can swap models without falling over will, and so will knowing what it would cost to bring one in-house. If you're not sure where you stand on either, that's exactly what an AI audit is for.

Written with AI assistance, reviewed and edited by Toby Remond.

Explore an AI audit · Talk through your process