←── back to feed
/topics/kimi-k3-chinese-model-capabilities-and-limitations

Kimi K3 Chinese model capabilities and limitations

22 items3 sourcesupdated 16d agotrend 0

Moonshot AI released Kimi K3, a Chinese trillion-scale MoE model that closes the frontier gap with leading Western models on standard benchmarks. The release has sparked debate about Chinese AI labs' capital efficiency and the model's actual capabilities on hard problems, with mixed real-world testing results showing both strengths and notable limitations.

  • Kimi K3 is a trillion-parameter mixture-of-experts model released by Chinese company Moonshot AI in July 2026.
  • Chain-of-thought reasoning reaches 32 pages on complex tasks; 95.5% of CoT characters in Chinese prompts are in English.
  • Performs comparably to DeepSeek V4 Pro and GLM-5.2 on standard benchmarks, but struggles with creative tasks like writing murder mysteries.
  • Failed on complex statistical auditing tasks, misapplying statistics and making errors that GPT-5.6 Pro identified.
  • Chinese labs demonstrate higher capital efficiency than Western competitors in converting compute and data into model capability.
[BSKY]bluesky14
Where Kimi K3 puts the balance of power — and bends the trajectory — of the AI ecosystem.
@natolambert · @natolambert.bsky.social · ▲28 · 17d
If recent events with Kimi K3 have finally convinced you that you need to try and understand how the Chinese labs approach AI - and how it differs than the SF center of power - you should read my post from a few months ago:
@natolambert · @natolambert.bsky.social · ▲88 · 18d
When I asked Kimi K3 "I want you to suggest two poems that you think apply to the current state of GenAI models like you. Don’t just pick popular poems. Think hard" the CoT was 32 pages long (& interesting): docs.google.com/document/d/1...
@emollick · @emollick.bsky.social · ▲34 · 18d
Interestingly, when I made a request in Chinese for Kimi K3 to pick two non-cliched poems that apply to LLMs, 95.5% of the characters (88% of the words) in the chain-of-thought were in English, even when it was explicitly considering Chine…
@emollick · @emollick.bsky.social · ▲47 · 18d
I think what is pretty clear is that the Chinese labs are far more capital efficient. In a world where scaling labs are intelligence is proportional to effective capital (buys compute, data, & talent) that may be the greatest strength your…
@natolambert · @natolambert.bsky.social · ▲37 · 19d
Kimi K3, like Claude, loves drowned cities, ancient apocalypses, and vast dying gods.
@emollick · @emollick.bsky.social · ▲70 · 19d
A lot of swift conclusions are being drawn about Kimi K3 based on fairly saturated benchmarks and ELOs, rather than actually testing it on very hard problems. The AI frontier has already moved so far that a good model that is a still month…
@emollick · @emollick.bsky.social · ▲52 · 20d
A GPT-4 powered (& thus quite obsolete today) assistant for Pakistani judges increased the amount of cases they saw by 6% with no impact on quality. elliottash.com/papers/Mehmo...
@emollick · @emollick.bsky.social · ▲17 · 20d
Kimi K3 cannot write a good murder mystery (though neither can any other model). That remains the jaggedest of frontiers.
@emollick · @emollick.bsky.social · ▲25 · 20d
Let's talk about China's next top model: Moonshot AI's Kimi K3 www.platformer.news/kimi-k3-laun...
@caseynewton · @caseynewton.bsky.social · ▲14 · 20d
A note of caution: I will say that when doing some complex statistical auditing of some of my prior academic work, Kimi K3 Max messed up in a bunch of ways, including misapplying statistics and applying some stuff badly.
@emollick · @emollick.bsky.social · ▲62 · 21d
My notes on Kimi K3, plus some thoughts on what we can still learn from the pelican benchmark even while it becomes further detached from how good the models are at the things that matter (like agentic tool calling across longer conversati…
@simonw · @simonwillison.net · ▲113 · 21d
Post-Kimi K3 and open weights models getting closer to the frontier again, I wonder if Anthropic and OpenAI will be allowed to increase their release cadence by the government. Mythos came out in April (before Opus 4.7) which means Fable 5…
@emollick · @emollick.bsky.social · ▲73 · 21d
My benchmark where I have AIs create one file procedurally-generated harbor towns through history in one shot now has GPT-5.6 Pro, Fable, Kimi K3, and Inkling. You can play with all the simulations: ai-harbor-town-gallery.netlify.app#kimi…
@emollick · @emollick.bsky.social · ▲49 · 21d