Overnight, K3's capabilities dropped to the free tier

On September 11, Moonshot AI threw a new model called K2.8 Preview into Kimi Code and Kimi Work. The biggest change is not in model capability but in tier access: every membership tier — from Adagio (free) to Allegro (¥699/month) — now gets a 1M token context window, while K3's 1M context remains locked at Allegretto (¥199/month) and above.

Moonshot did not publish benchmark numbers for K2.8 Preview, but stated "overall performance close to K3", with significantly improved thinking efficiency over K2.7 Code, and full upgrades to coding and agent capability. The Model ID is unchanged — still kimi-for-coding — and clients and third-party tools upgrade with zero config changes. The reasoning effort levels (low/high/max, default max) align directly with K3.

This is a "bread-and-butter" model

K3 is the flagship: 2.8 trillion parameters, KDA linear attention, 1679 Elo on the Frontend Code Arena pushing Claude Fable 5 (1631) and GPT-5.6 Sol aside, jumping from K2.6's #18 straight to #1. But K3 is genuinely expensive — $3 per million input tokens, $15 output, one-third of competitors' prices — and still resource-hungry to run. Moonshot itself acknowledged in the K3 technical blog that K3 is sensitive to historical thinking content; if the model is switched mid-session or thinking records are missing, output quality can degrade noticeably, and on fuzzy tasks it sometimes takes actions the user did not ask for.

K2.8 Preview is clearly aimed at a different job: catching the routine code completion and conventional development tasks that don't need K3's full firepower, with a cheaper version. Moonshot also announced that once thinking is disabled, both K3 series and K2.8 Preview requests route to K2.8 Preview (no-thinking) — meaning users writing code daily are already running K2.8, they just don't know it.

ARR jumped from $100M to $1B in five months

The commercial backdrop for K2.8 is equally important. According to Bloomberg on September 11, Moonshot's ARR went from roughly $100M in March, to $200M in April, $300M+ in June, and broke through $1B in August. The internal target is $2B by year-end. K3 series currently generates roughly 300 billion tokens per day.

Hitting the doubling requires more than one expensive flagship. K3 fits long-horizon programming, kernel optimization, chip design — the "must use the strongest" scenarios. But routine code completion, file edits, single-step debugging — the high-frequency tasks users run every day — are exactly the battleground for low-cost-per-call workhorse models like K2.8. Pairing K3's capabilities being dropped to all tiers with the late-August Pre-IPO round at a $50B pre-money valuation, the rhythm is clear: stack up user volume first, push per-call cost down, then push the ARR number toward $2B.

The scissors between valuation and revenue

In 8 months, Moonshot's valuation climbed from $4.3B to $50B — roughly 8x. But at the $50B valuation against the publicly disclosed $300M ARR at the time, the P/S ratio is around 167x. For comparison, Anthropic runs roughly 20x, OpenAI roughly 40x, and Zhipu at its trillion-yuan market cap was around 94x — even the highest-valued Zhipu is barely half of Moonshot's multiple.

Moonshot clearly knows this number needs explaining. Earlier this month Bloomberg also disclosed that the company had confidentially submitted A1 listing application documents to HKEX, starting the Hong Kong IPO process. At the current $1B ARR, the P/S ratio drops to around 50x — still notably above OpenAI, but considerably more presentable than 167x.

Why this timing for K2.8

Four timestamps read together:

  • July 16: K3 official release, 1679 Elo debut
  • Late August: Pre-IPO round at $50B pre-money valuation
  • Early September: confidential filing to HKEX
  • September 11: Bloomberg story on ARR doubling, same day as K2.8 Preview launch

One last capability drop before listing, telling the primary market, secondary market and developer community simultaneously that "K3 is expensive" and "K2.8 is also strong." K2.8 not publishing a standalone benchmark is a smart move: avoids stealing K3's spotlight while K3 is still climbing leaderboards, and leaves pricing room for future commercialization.

One judgment

I'd frame K2.8 Preview as "K3's commercial recycler": converting the influence K3 earned on LMArena into frequency coverage by K2.8 in Kimi Code's daily requests. In other words: Moonshot has already proven the ceiling of capability. From now until year-end, the key metric is no longer "can the strongest model get even stronger" but "how many people are using the cheap model, and how many tokens does it eat per day." K2.8 Preview is the first bullet in that fight.

One open question for the comments: K3's thinking chains are unstable, and K2.8 must take over all no-thinking requests — will this traffic-switching make long-session users feel like the experience has regressed?