Press Esc to close

New coding test reveals Kimi K3 outperforms Fable and GPT-5.6 Sol, but takes three times longer per task

Moonshot AI's new Kimi K3 has been out less than a week, and a fresh independent coding test just put real numbers behind the launch-week hype. The short version: K3 edges out Anthropic's Claude Fable 5 on task-solving, costs a fraction of the price, but takes roughly three times as long to finish each job.

The test ran Kimi K3 max against Fable 5 xhigh on identical coding tasks and tracked how many attempts each model needed to land the correct answer.

On the first try, Fable 5 actually comes out slightly ahead, solving 69.9% of tasks against K3's 68.5%. That's close enough to call a tie. Give both models a second or fourth shot though, and K3 takes the lead, hitting 82% at pass@2 and 89.4% at pass@4 against Fable's 80.2% and 88.5%.

Money is where the gap really opens up. Each solved task runs about $4.65 on K3 versus $13.41 on Fable 5, meaning K3 delivers close to three times more solved work per $100 spent. The trade-off is speed. A typical K3 task takes about 66 minutes to wrap up, while Fable 5 finishes in roughly 21.

Split the results by programming language and the story gets messier. K3 pulls ahead noticeably in Go, scoring 79% against Fable's 71%. Fable claws its lead back everywhere else though, beating K3 by 10 points in Rust and staying ahead in Python, Typescript, and Javascript too.

The two models also fail in nearly identical ways. Both come up short on close calls, near misses rather than total wipeouts, about 65% of the time each. And out of the full task set, both models solved 96 of the same problems, with only 5 unique wins for K3, 4 for Fable, and 8 that stumped both entirely. That kind of overlap is unusually high for two models built by completely different labs.

None of this happens in isolation. K3 launched July 16 at $3 per million input tokens and $15 per million output tokens, a steep discount against Fable 5's $10 and $50. AI commentator Simon Willison still called it "the most expensive model released by a Chinese AI lab", which says a lot about how far token prices have climbed even on the cheaper side of this race.

This particular test skipped GPT-5.6 Sol, but broader numbers from Artificial Analysis, reported by The Decoder, put it and Fable 5 slightly ahead of K3 on general intelligence, even as K3 keeps winning on coding specifically.

For developers choosing between them, it really comes down to a budget versus speed question. Pick K3 if cost per task matters more. Pick Fable 5 if you need the answer fast.

Comments