2 min read

Kimi K3 Ranks No. 2 on AA-Briefcase

Moonshot AI’s Kimi K3 posts the second-highest AA-Briefcase score, but it is costly and slow, averaging $10.57 and 56.4 minutes per task.

Image: Hacker News

Moonshot AI’s Kimi K3 has posted the second-highest score on AA-Briefcase, Artificial Analysis' benchmark for agentic knowledge work, trailing only Claude Fable 5. The 2.8T-parameter model scored 57 on the Artificial Analysis Intelligence Index, which the firm says puts it alongside models such as Opus 4.8 and GPT-5.5.

On AA-Briefcase, Kimi K3 reached an Elo of 1543—a +727 jump from Kimi K2.6 at 816—and finished behind only Claude Fable 5 at 1574.

AA-Briefcase benchmark chart
AA-Briefcase benchmark chart

Artificial Analysis describes AA-Briefcase as a proprietary benchmark built around a private dataset of realistic tasks spread across thousands of complex input files. The tasks require outputs such as spreadsheets, presentations, and UI mock-ups, with results rolled into a single Elo based on correctness, analytical quality, and presentation quality.

Kimi K3's standing came with strong underlying scores:

Recommended reading

Musk says Grok Imagine could make an Odyssey film in 2026

  • Rubric pass rate: 51%, second only to Claude Fable 5 at 56%
  • Analytical quality Elo: 1754, compared with Claude Fable 5 at 1744
  • Presentation Elo: 1471, behind GPT-5.6 Sol (max) at 1660 and Claude Opus 4.8 (max) at 1492
AA-Briefcase leaderboard

1 / 2

The tradeoff is cost and speed. Artificial Analysis says Kimi K3 averages $10.57 per task, making it one of the most expensive models on AA-Briefcase and roughly a 10x increase in cost per task. It averages 83 turns per task, versus 67 for Claude Fable 5 and 50 for GPT-5.6 Sol (max). Kimi K3 is priced at $3/$15 per 1M input/output tokens, with a 90% discount for cached tokens.

Cost per task chart
Cost per task chart

It is also slow. Kimi K3 averages 56.4 minutes per AA-Briefcase task, one of the highest times recorded on the benchmark. According to Artificial Analysis, that is driven by the model’s 83 turns per task, 120k output tokens per task, and slower speeds on the first-party Kimi API. The firm says that is about 2.5x the time of Claude Fable 5 and about 3.8x higher than Grok 4.5 (high).

Ava Chen

AI Editor

Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.

via Hacker News

// Keep reading