
Kimi K3: Moonshot's 2.8T Open Frontier Model Drops
Moonshot just shipped Kimi K3 — a 2.8-trillion-parameter open model that claims frontier-level performance on long-horizon coding, agentic knowledge work, and reasoning. It is the first open model to cross the 3T threshold and will release full weights on July 27, 2026.
Architecture
Kimi K3 uses Kimi Delta Attention (KDA) + Attention Residuals (AttnRes) and a heavily sparsified Stable LatentMoE (16 of 896 experts active). The company claims 2.5× scaling efficiency over K2. Native 1M context with automatic caching, vision, and tool calling built in.
Coding Performance
Strong on long-running engineering tasks, kernel optimization, GPU compiler construction, game dev with vision-in-the-loop, and even autonomous chip design on Nangate 45nm.
Knowledge Work & Agentic
Internal evaluations show consistent gains in real production workflows. The model produces interactive dashboards, research reports, and video edits from natural language.
Pricing (API)
- Cache-hit input: $0.30 / M tokens
- Cache-miss input: $3.00 / M tokens
- Output: $15.00 / M tokens
Flat pricing regardless of context length. 90%+ cache hit rate reported on coding workloads.
Availability
Live now on kimi.com, Kimi Work, Kimi Code, and the Kimi API (platform.kimi.ai). Full weights July 27.
Two Frontier Models in One Week
Kimi K3 arrived in the same week as Grok 4.5 — xAI's Cursor-native frontier model — making July 2026 one of the most eventful stretches in frontier AI this year. Where Kimi K3 is open-weight, 3T-class, and built for agentic knowledge work, Grok 4.5 is a closed, IDE-native coding model optimized for cost and speed. Two radically different approaches, same week. See our full Grok 4.5 coverage.
Sources: Kimi K3 announcement · API docs · Pricing
Comments
Leave a comment
Related articles

Ox Alpha: The Anonymous 1M-Context Coding Model That Has Everyone Guessing Z.ai
Ox Alpha is a free, anonymous 1M-context reasoning model on OpenRouter. Community fingerprinting points to Zhipu AI's GLM-5.x. We break down the benchmarks, the mystery, and the adoption.

DeepSeek V4 Pro Is Now an Ollama Cloud Model: The 1.6T Flagship on a $20 Flat Plan
DeepSeek's 1.6T flagship V4 Pro is now a flat-rate Ollama Cloud model — $20 a month, one line of shell, drop-in for Claude Code and OpenCode. Here's what the 0813 refresh changed and when the plan beats the API.
DeepSeek V4 Flash 0731: The 13B-Active Agent Workhorse That Outranks Its Own Flagship
DeepSeek's July 31 official release of V4 Flash 0731 is the small-model moment: a 284B-total / 13B-active re-post-trained MoE that outscores the 1.6T flagship across nine agent benchmarks at roughly a third of the price.
