Kimi K3 (Image © Moonshot)
Model Specifications and Release of the Open Weights
Kimi K3 features a massive architecture with 2.8 trillion parameters. The company plans to release the model’s weights on July 27 to make it available for broader use. While Moonshot acknowledges that Kimi K3’s overall performance currently lags behind proprietary high-end systems such as OpenAI’s GPT-5.6 Sol and Anthropic’s Claude Fable 5, internal tests suggest that the performance gap is narrowing. Our experience with the K3 was an eye-opener.
Kimi K3 vs. Claude Opus 4.8: A Hands-On Review
We frequently experiment with Google’s models, such as Gemini 3.1 Pro, but also ChatGPT and, above all, Claude. Whereas we used to pay for a ChatGPT subscription, we’ve since switched to Claude, and Kimi K3 is bringing about massive changes. For more complex programming tasks, we’d currently recommend Claude Sonnet 5 or Fable 5, but since the release of Kimi K3, the balance has shifted. This is partly because Anthropic has scaled back Fable 5’s performance and even excluded it from the budget-friendly package; Sonnet 5 also crashed during our tests, and we were often forced to fall back on Opus 4.8. We didn’t even have enough tokens left in our budget, yet the model is clearly “worse/dumber” compared to Kimi K3.
The extent to which the Kimi K3 is better is evident in the complexity it can handle. For tasks such as Wireshark PCAP packet analysis and static code analysis—tasks at which Opus 4.8 failed—the Kimi K3 spends more time working on them and then actually delivers usable results that Opus couldn’t achieve. What became clearly apparent during the tests is that Kimi K3 simply works on the problem for much longer throughout the process, often breaking it down into smaller subtasks, then validating these multiple times, resulting in a lower error rate. Where Anthropic excels is in tool usage: Opus and Sonnet were able to read PCAP files directly, analyze and evaluate data from them immediately, and ask follow-up questions. In our experiments, Kimi K3 runs in an Ubuntu VM with Chrome and can install many packages later as needed, but has trouble accessing GitHub.
Kimi K3’s token efficiency is king!
With the most affordable Moderato subscription at 19 USD per month, we’ve come an incredibly long way. With Anthropic, we would have hit a wall quickly with long contexts and would easily have had to pay several times more. With Kimi K3, you pay $3/M tokens for input, $15/M tokens for output, and $0.30/M tokens for cached input. In contrast, Opus 4.8 costs $5/M tokens for input,
$25 per million tokens for output, and $0.50 per million tokens for cached input. We’re very pleasantly surprised by the K3 model—even though it’s clearly tailored for China and tends to use Chinese as a second language instead of German—it handles German just fine.
The K3 model didn’t shy away from the work either. With Opus and Sonnet, we encountered a problem in some chats where the model tried to convince us to “call it a day” or not to carry out the project because it was too time-consuming.
Front-End Development Benchmarks
Beyond our own experiences, there are, of course, performance benchmarks, and Kimi K3 excels particularly in programming tasks. In the Arena.ai rankings for front-end development, the model secured first place, outperforming leading U.S. models. This represents a significant leap in the rankings compared to its predecessor, Kimi K2.6.
These results challenge the prevailing industry assumption that Chinese AI labs lag significantly behind their American counterparts. The release comes shortly after the launch of Claude Fable 5 and GPT-5.6, indicating a rapid acceleration of development cycles in Chinese labs.
Geopolitical Implications and Market Consequences
The emergence of powerful Chinese models has already caused volatility in global technology markets in the past. In January 2025, the release of the Deepseek-R1 model led to significant valuation losses across major technology indices. The launch of Kimi K3 is expected to continue this trend and further disrupt established market dynamics.
For Europe, Kimi K3 is a very good alternative, as the Americans have taken Anthropic’s most powerful model Fable 5 offline for over a week—and it apparently came back “dumber,” which was likely intentional. Since it’s an open-weights model, Kimi K3 can also be run on other, self-hosted infrastructure, giving users an incredibly powerful tool that, in our view, currently outshines Anthropic’s models.



