SignalSpawn
Kimi Agent demo — Gargantua black-hole visualization built with Kimi K3
TechnologyReview

Review: Kimi K3

The open frontier finally feels close — 2.8T weights, 1M context, and a real fight with closed labs. Not the crown. The clearest path to it.

8.7/10

Moonshot AI did not whisper Kimi K3 into existence. On July 16, 2026 the company dropped a 2.8-trillion-parameter Mixture-of-Experts flagship into Kimi.com, Kimi Work, Kimi Code, and the API — then, eleven days later, put the full weights on Hugging Face. Call it Kimi 3 if you want the consumer nickname; the badge on the box is K3, and the claim is blunt: first open model in the 3T class.

We lived with it for two weeks after the weight drop — coding agents, long-context research dumps, multimodal briefs, and enough API burns to feel the $3 / $15 per-million pricing in the wallet. The question that matters for SignalSpawn readers is not “is it clever?” Every frontier lab is clever. The question is whether open weights this large change who gets to ship.

“Kimi K3 is a 2.8T-parameter model built on our Kimi Delta Attention and Attention Residuals, with native vision capabilities and a 1-million-token context window.”

— Moonshot AI — Kimi K3 launch blog, July 16, 2026

Kimi Agent demo — Gargantua black-hole visualization built with Kimi K3

What actually shipped

Architecture is the story Moonshot wants you to remember: Kimi Delta Attention, Attention Residuals, Stable LatentMoE that lights up 16 of 896 experts per token (~104B active). Native text, image, and video in one stack. A 1,048,576-token context window with no long-context surcharge on the official API. Reasoning effort defaults to max at launch, with lower tiers promised as follow-ups.

Independent boards back the swagger. Artificial Analysis puts K3 near the top of the Intelligence Index among open-weight models — trailing only a short list of closed systems (Claude Fable 5, GPT-5.6 Sol, Claude Opus 5 on some harnesses depending on the week). Arena-style frontend coding leaderboards had K3 punching above its “open” label. Cyber-offense evals from joint AISI-style assessments still put it behind the most restricted US frontier stacks. That gap is not a footnote if your threat model cares.

Pricing is the other lever: $0.30 / MTok cached input, $3 cache-miss, $15 output on Moonshot’s API. Cached coding loops with a 90%+ hit rate (Moonshot’s claim for Mooncake-backed serving) turn “frontier-adjacent” into something a mid-size studio can actually leave on overnight.

Kimi K3 game-dev case — cyberpunk web-swing scene from Moonshot’s K3 blog

Hands-on: where it sings

Long-horizon coding is the headline skill. Point Kimi Code at a multi-file refactor with a fat repo map and K3 keeps the plot longer than most open peers we have used this year. Frontend generation — landing pages, dashboards, motion-adjacent UI — is where community arenas lit up; our own prompts matched that: tasteful defaults, fewer “AI purple gradient” reflexes than last year’s open cohort.

The million-token window is not a stunt if you actually fill it. Research packets, design docs, and transcript piles that used to require rag theater now fit in one shot. Multimodal input is usable, not decorative — screenshot-to-patch workflows land more often than they used to on K2-era stacks. Agent frameworks that already speak OpenAI-compatible endpoints drop in with less glue than expected.

Where it scrapes

Closed frontier models still win the last three points on the hardest reasoning boards, and you feel it on brittle tool-use chains that need perfect discipline. The custom Kimi K3 License is not MIT — fine for many products, a paperwork tax for others. Serving 1.5TB+ of shards is not a laptop hobby; “open” here means open to labs and clouds with serious iron, not every indie fine-tune night.

Max-effort defaults can be chatty and over-eager — Moonshot has flagged instability and over-proactivity in its own materials. Dialing effort down helps cost and temperament; leaving it on max for every Slack-like reply is how invoices and tangents happen together.

Who it’s for

Teams that want frontier-near coding without renting a closed lab’s policy surface. Researchers who need weights they can inspect, distill, and host. Product orgs already on Kimi Work / Kimi Code who want one model across chat, agents, and API. Pause if you need the absolute peak closed-model score, or if your counsel refuses anything but Apache/MIT.

The verdict

Kimi K3 is the open-weight moment the industry has been advertising since Llama 3 — not because it dethrones every proprietary flagship, but because the gap is finally small enough to matter for shipping. It is the best open model we have put through real agent loops this summer. Score it like a near-miss at the summit: the view is already worth the climb.

Score: 8.7 / 10 — open frontier intelligence you can actually download, with a price card that keeps the lights on.

Start at kimi.com / Kimi K3.

Author

Cristiano Lima

Published

Keep reading

View all