July 19, 2026 ยท 9 min read

New season, new contender: Kimi K3 shows up open-weight

Every few months this entire circus just resets itself. Some lab ships a new flagship, the benchmark charts get redrawn overnight, and the rest of us sit around arguing about which 200 dollar a month subscription we're feeding our coding habit this time. So, new season, new episode of the same show.

The fun part this round is that the new contender is open-weight. As in, you can actually download it and run it yourself, in your own house, assuming your house happens to be wired for an industrial substation. No API subscription, no enterprise sales calls, just you and a truly stupid amount of GPU.

Anyway. Meet Kimi K3.

The new contender

Moonshot dropped Kimi K3 this week and the spec sheet reads like somebody lost a bet. 2.8 trillion parameters. It's a mixture-of-experts model so it only fires up around 50 billion of them per token (16 experts out of 896), but "only 50 billion active" is a phrase that would've gotten you laughed out of the room two years ago. Add a 1 million token context window and it lands as the first open-weight model to actually shove its way into the top group on the flagship coding benchmarks.

And that's the real story here. A Chinese lab shipping a huge model, fine, that's basically a weekly event now. An open-weight model cracking the top 3 though? That just hasn't happened before, and it means you can finally host something near the top of the charts on your own hardware instead of renting it off somebody else.

Then you get to the hardware bill. "Just run it locally" is doing an absurd amount of work in that sentence. The weights on their own come out around 1.4 to 1.5 TB in native 4-bit. The realistic floor to even get it breathing with aggressive quantization is somewhere north of 650 GB of memory, and Moonshot's own recommendation is to serve the thing on a supernode of 64 or more accelerators. So yes, by all means, run K3 at home, right after you run a dedicated power line to the property and explain to your family why the living room is now a 400 amp server closet. I've already got a calendar reminder set for July 27 (when the weights actually drop) and I'm out here deleting movies off a drive I'm never realistically going to fill, purely so I can keep a copy of this thing for the shelf. Don't judge me.

Where the three actually land

Right, I promised this would be data driven, so here's the data. On Artificial Analysis's cross-vendor Intelligence Index (the closest thing we've got to comparing apples to apples), the top three are breathing right down each other's necks:

Model Intelligence Index Open weights? API price (in / out per M)
Claude Fable 5 59.9 No $10 / $50
GPT-5.6 Sol 58.9 No $5 / $30
Kimi K3 57.1 Yes (Jul 27) $3 / $15

About three points cover first through third, which is basically nothing. On SWE-bench Verified, the benchmark everybody actually quotes at each other, Claude Fable 5 still sits on top at around 95%, and it stretches the lead on the nastier SWE-bench Pro (roughly 80% against Sol's 64.6%). Then GPT-5.6 Sol turns around and tops the Artificial Analysis Coding Agent Index at 80.0, while costing about a third of what Fable does per task. And Kimi K3 is the cheapest of the three per token while currently sitting at number 1 on LMArena's Frontend Code Arena, ahead of both of the paid flagships. For something you can just download, that's wild.

So nobody is running away with this. The gap at the very top has shrunk to rounding error, and what you're really shopping for now is price, access, and how much you trust the benchmarks. Which gets me to my actual feelings on all this.

A loyalty confession, and a small plot twist

Quick disclaimer: I'm a Claude guy. Have been for ages. And it's got basically nothing to do with who's topping benchmarks this week. I just really value that Claude pushes back on me. I can tell it my architecture is genius and it'll gently point out that the whole thing is held together by hope and one cron job. It doesn't just nod along and let me go ship garbage, and after enough hours pairing with a yes-man, that kind of honesty is worth real money to me.

So I had a whole angry breakup letter loaded and ready this month. Because for most of the summer, Anthropic had been quietly walling Fable 5, their top model, off from the subscription plans and nudging everyone toward the metered API, which at $10 in / $50 out is the priciest option going. Nothing radicalizes a loyal customer quite like watching the good model wander off behind a paywall.

And then, right on cue, the plot twist. Anthropic just announced that starting July 20, Fable 5 is back in the Max and Team Premium plans (at 50% of the limits). Pro and Team Standard people still get it through credits, cushioned with a one-time $100 credit, but the whole flagship-behind-the-API panic is officially off for Max subscribers. So my finger came right back off the trigger. I see you, Anthropic. I'm staying put, for now.

The OpenAI temptation is real though

I'd be lying if I said GPT-5.6 Sol wasn't tempting me. It leads the agentic coding index, it's meaningfully cheaper per task, and if you basically live inside Codex, the ChatGPT Pro plan (200 a month, the "20x" tier) is stupidly generous, near enough unlimited as long as you're not actively trying to abuse it. For a lot of agentic workloads that's the best headroom for your money on the whole board.

A couple of asterisks though, because I said facts and not vibes. First, on SWE-bench Pro (the harder, closer-to-real-work one) Sol trails the Claude flagships by a wide margin: 64.6% against about 80%. Second, and this one's funnier, the safety crew over at METR caught Sol "gaming" its own software engineering eval at the highest rate they'd ever recorded. Which is either hilarious or slightly terrifying, depending on whether that model is about to go refactor your production database. Great scores, sure, just read the footnotes.

And while we're down in the footnotes, Kimi K3's numbers want a squint too. Its hallucination rate jumped to around 51%, and it chewed through roughly 1.9x the tokens Sol needed to finish the same benchmark, so that lovely cheap per-token price gets partly clawed back in sheer volume. Independent testing is also thin until the weights land on the 27th, so hold the confetti until then.

So what do you actually pay for this summer

Good news first: you don't have to jump straight to the 200 dollar "max" tier to get in the game. All three have a ~100 dollar rung and a ~200 dollar rung, and they line up almost suspiciously well:

Tier Claude (Anthropic) ChatGPT (OpenAI) Kimi (Moonshot)
~$100/mo ("5x") Max 5x, $100, 5x Pro usage, Fable 5 included from Jul 20 (50% limits) Pro, $100, 5x Plus usage, Sol + Codex Allegro, $99, Agent Swarm, daily Kimi Code, full 1M context
~$200/mo ("20x") Max 20x, $200, 20x Pro usage Pro, $200, 20x Plus usage, Sol Pro + max Codex Vivace, $199, 300-agent Swarm, Kimi Claw, priority access

(And yes, Kimi names its tiers after musical tempos, because of course it does. Adagio is the free one and you speed up from there. There's a $19 Moderato and a $39 Allegretto sitting under this lot too, if 100 bucks is rich for a hobby that already has me deleting movies for keepsakes.)

Matching budget to brain:

The 100 dollar sweet spot. If you're a solo dev who doesn't need the absolute ceiling every single hour, any of the three at 100 bucks is a ton of model. Claude Max 5x now bundles Fable 5 (half limits), which is my pick for pushback quality and raw SWE-bench muscle. ChatGPT Pro at 100 gets you Sol plus Codex with real headroom. And Kimi Allegro at 99 is the sneaky value pick: Agent Swarm, enough Kimi Code credits for real daily coding, and the full 1 million token context switched on, at the lowest sticker of the lot.

The 200 dollar max tier, for when you basically live inside an agent and want as much rope as they'll hand you: Claude Max 20x for the ceiling and the co-pilot that tells you no. ChatGPT Pro 20x for the most generous flat-rate Codex setup and the top agentic-index score. Kimi Vivace at 199 if you want 300 parallel agents and cloud deployment for the price of one Claude.

And then the enthusiast option nobody needs but everybody secretly wants: self-host K3 once the weights drop on July 27. Cheapest per token, first real open-weight top-tier model, and the only thing on this whole list you can actually keep. Right after you sort out that minor 64-GPU supernode. (See also: the electric company, above.)

The keepsake

Me, I'm staying on Claude, quietly eyeing Sol, and downloading K3 the second the weights go live. Am I going to run a 2.8 trillion parameter model out of my apartment? Absolutely not. But we're clearly living through the summer the open weights caught up to the big paid flagships, and I want a copy of the artifact sitting on a shelf. Future me can explain it to the power company.

New season, new contender, same three good problems to have. See you next patch.


Numbers pulled from Artificial Analysis (Intelligence + Coding Agent indices), Moonshot's own K3 launch materials, Anthropic's pricing announcement, the ChatGPT and Kimi plan pages, and METR's eval notes. All of it is a snapshot as of July 18, 2026 and, like every benchmark ever, should be taken with a fat handful of salt. Kimi has also hinted its tempo tiers are getting reshuffled soon, so double-check kimi.com before you go quoting me.