Kimi K3, and what we can still learn from the pelican benchmark
Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.
Analysis of Kimi K3 model and pelican benchmark by Simon Willison, hot AI topic with strong community engagement
Moonshot AI's Kimi K3 (2.8T parameters, open weights by July 27) surpasses DeepSeek's 1.6T v4 Pro and beats Claude Opus 4.8 and GPT-5.5 on self-reported benchmarks, but trails Claude Fable 5 and GPT-5.6 Sol. Priced at $3/$15 per million tokens (double Kimi K2.6), it leads Arena.ai's Frontend Code arena. The author's 'pelican' SVG benchmark, once correlated with model quality, now shows weaker correlation and fails to test agentic tool calling—yet running it via OpenRouter and LLM CLI reveals model characteristics like reasoning token costs (25 cents for 16,658 output tokens).
Simon Willison