← Blog

/

Insights

Insights

Kimi K3 solves the task once. Not five times in a row.

Toloka Arena is live. See how your model ranks.

Kimi K3 is the fourth-best agent in the world on the public Artificial Analysis Agentic Index. On Toloka Arena — private, uncontaminated, realistic enterprise workflows across seven industries — it ranks nineteenth of forty. The reason is reliability: K3 will often solve a task on one run and fail the same task on the next. GLM-5.2, which the public index ranks a full tier below it, is steadier and finishes ahead. This is not a weak model. It is an inconsistent one.

Executive summary

  • Kimi K3 is genuinely strong. It is a large step up from K2.6, and it ranks #4 on the public Artificial Analysis Agentic Index (score 50) — behind only Opus 5, GPT-5.6 Sol, and Fable 5, and ahead of GPT-5.6 Terra, Opus 4.8, Sonnet 5, and GLM-5.2. On coding, in our own informal internal use, it is excellent. None of that is in dispute.

  • On Toloka Arena it ranks 19th of 40. Our leaderboard scores agents on pass^5 — the fraction of tasks an agent completes correctly on all five independent runs — across 834 private, realistic enterprise tasks in seven domains, graded on the exact final state of a live database. Kimi K3's composite pass^5 is 33.4%.

  • Most of the gap is reliability. K3 retains only about half of its single-run success once you demand that all five runs succeed — the lowest retention of the capable models we profiled (Fable 83%, Opus 4.8 76%, GPT-5.6 Sol 73%, GLM-5.2 62%; DeepSeek V4 Pro is comparable to K3). On the frozen field-level grade, 41% of K3's tasks are "solved sometimes, not reliably" — the widest such band we measured. It does the intelligent thing on Monday and skips a step on Tuesday.

  • A lower-ranked public model beats it here. GLM-5.2 (public #10) scores 42.3% on Arena (#13), winning 5 of 7 environments and beating Kimi K3 by a task-weighted 8.9 points — the reverse of their public ranking.

  • What makes this signal hard to game: pass^5 is a standard reliability metric — tau-bench and others use pass^k, and we claim no special sauce there. Toloka Arena's distinctive value is that the tasks are private and uncontaminated, span seven realistic enterprise domains, and are graded on exact database state. A model can't have seen them, and it can't bluff them with plausible prose. That is why a top-4 public agent landing 19th here is worth a close look — and we make no claim about why the gap exists.

A genuinely strong model

Give Kimi K3 its due first, because the rest of this only matters if you take its strength seriously.

Moonshot's K3 is a real generational step. On the public Artificial Analysis Agentic Index — a composite of GDPval-AA and tau3-Banking, and one of the most-watched agentic leaderboards — Kimi K3 scores 50 and ranks fourth in the world, the top open-weight model, ahead of GPT-5.6 Terra, Opus 4.8, Sonnet 5, Grok 4.5, and GLM-5.2. If you picked an agent off that leaderboard, K3 would be an obvious, defensible choice. On coding tasks, informally, we like it. This article is not about any of that.

It is about what happens when the same model has to operate a company's systems, correctly, every single time.

The same task, five times, graded on the database

Toloka Arena evaluates agents on private enterprise workflows: 834 tasks across seven industries — logistics, hotel and travel operations, restaurant operations, airlines, bank HR on Dynamics 365, a short-term-rental support desk, and manufacturing. In each, an agent must read records, consult policy, choose an action, call tools, and leave a live database in the exact expert-authored final state. Plausible prose does not count. One extra ticket, one wrong status, one missing record, and the task fails.

Three things make this hard to fake, and they — not the metric — are the point:

  • The tasks are private and uncontaminated. They are not on the public web and not in anyone's training data. A model cannot have memorized them.

  • They are realistic and broad. Seven genuine enterprise verticals with real policies, tools, and edge cases — not a single narrow benchmark.

  • They are graded on exact state. A deterministic field-level check of the database, not an LLM judge scoring an essay.

Our headline metric is pass^5: the fraction of tasks solved correctly on all five independent runs. This is a standard reliability metric — it comes from the tau-bench lineage of pass^k, and we claim nothing unique about it. It simply asks a question production teams care about: not "can the agent ever do this," but "will it do this every time." Every model is run at the same standard configuration (medium reasoning, temperature 0.6).

Figure 1. Composite pass^5 (all five runs pass), task-weighted across the seven environments. Fable 5 leads at 60.9%. Kimi K3 (red) sits 19th at 33.4% — below GLM-5.2 (blue, 42.3%), the GPT families, Opus 4.8, Grok 4.5, and Sonnet 5. Verifiable at toloka.ai/arena.

On that board, Kimi K3 ranks 19th of 40, at 33.4% — while the public index rates it 4th in the world. That 15-place gap is the whole story, and it has a specific cause.

It solves the task once, not five times in a row

Here is the mechanism, and it is not that K3 is incapable.

pass^5 punishes inconsistency by design: a model that solves a task on four runs and fumbles the fifth scores zero on that task, exactly like a model that never solved it. So the revealing measure is retention — how much of a model's single-run success survives the demand that all five runs succeed.

Figure 2. Left: for each model, the share of tasks solved on k = 0…5 of five runs. Dark green (k = 5) is what pass^5 rewards; the amber band (k = 2–4) is "solved sometimes, not reliably." Right: reliability retention — pass^5 as a percentage of single-run accuracy. Kimi K3 (highlighted) retains just under half.

The steady models keep most of their single-run success: Fable 5 retains 83%, Opus 4.8 76%, GPT-5.6 Sol 73%, GLM-5.2 62%. Kimi K3 retains about 50% — the lowest of the capable models we profiled, with DeepSeek V4 Pro essentially tied; K2.6 was far worse at 25%. Concretely, 41% of Kimi K3's tasks fall in the "sometimes" band — passed on some runs, failed on others — the widest band we measured. K3 can very often do these tasks. It just doesn't do them reliably.

That reconciles the two leaderboards. Public agentic benchmarks largely reward whether a capable model can reach a good answer; a private, exact-state, all-five-runs evaluation rewards doing it consistently. Kimi K3 is strong on the first and weak on the second.

One honest qualification before we go further. Against GLM-5.2 specifically, reliability is not the only difference: on these tasks GLM also edges K3 on raw single-run success, so it wins on both counts. Reliability is what explains the size of K3's fall relative to its public standing — a 15-place drop that raw capability alone cannot account for — rather than the entirety of the GLM gap.

"Aren't 20% of your tasks just impossible?"

Look again at the left of Figure 2 and you will notice that every model — Fable 5 included — has a grey band of roughly a fifth of tasks it never passes, not even once. The obvious suspicion is that the benchmark is padded with broken or unsolvable tasks, and that everyone is failing the same dead weight. It is the right question to ask of any private benchmark, so here is the answer in full.

It is not the same fifth. Across all 834 tasks and all 42 models we have run, only 10 tasks — 1.2% — are never passed by any model on any run. Ten. Everything else has been solved by somebody.

Figure 3. Left: how many of the 42 models solve each task reliably (all five runs). Only the leftmost bar — 158 tasks, 18.9% — has no reliable solver at all, and even most of those are passed intermittently by someone. Right: for each model, its own never-passed tasks split by what other models did with them. Around 94% of every model's zero band is solved by some other model.

For Kimi K3 specifically: of the 184 tasks it never passes, 94.6% are solved by at least one other model, and 31% are solved perfectly — five runs out of five — by another model. Those tasks are not broken. They are tasks Kimi K3 does not do. The same is true in reverse: about half of K3's zero band is shared with GLM-5.2 and about half is not (Jaccard 0.55), so each model is tripping over a partly different set of work. The median task in the suite is reliably solved by 9 of the 42 models, and 81% of tasks have at least one model that nails them every time.

We should be equally straight about the harder finding sitting next to it: 158 tasks (18.9%) have no model at all that solves them on all five runs. That is not a set of impossible tasks — most are passed by someone some of the time — but it is a real ceiling. Roughly one enterprise workflow in five is currently beyond every frontier model's ability to do dependably. That is the headroom this benchmark exists to measure, and it is the strongest argument we can make that these tasks are worth running: they are solvable, they discriminate sharply between models, and the best models in the world have not saturated them.

The same task, passed and failed

Because "inconsistent" is abstract, here is what it looks like in the transcripts. Each of these is one task where Kimi K3 passed on some of its five runs and failed on others — same model, same prompt, same tools.

  • A warehouse access request (logistics; passed 3 of 5). A staffer asks for system access on behalf of a contractor. On its good runs, K3 checks the requester's own HR record and tags the case as an on-behalf-of request. On its bad runs, it checks the contractor's record but forgets to verify the requester's — and the final case is missing the on-behalf tag.

  • A vacation day (hotel/travel HR; passed 4 of 5). An employee requests one PTO day. On good runs K3 confirms the balance, files the request, and leaves the approval fields untouched as policy requires. On the bad run it files the request but also fills in approval fields that were supposed to stay empty — a self-added step that corrupts the final state.

  • A tax-withholding change (bank HR; passed 2 of 5). An employee can't use self-service and needs help. On good runs K3 creates the required support case. On bad runs it gives the employee correct verbal guidance — and never creates the case, so the database ends empty where a record was required.

  • A production-planning change (manufacturing; passed 3 of 5). A user asks to disable demand-driven planning for a SKU. On good runs K3 checks the user's role, recognizes it is insufficient, and stops. On bad runs it performs the very same role check, treats it as sufficient anyway, and makes the unauthorized write.

The throughline is not a single dramatic bug. It is a required step — a verification, an approval boundary, a record to create — that K3 performs on some runs and skips on others. When we categorize every failing run by what was wrong in the final database state, the errors spread broadly across missing records, wrong lifecycle status, wrong routing, and extra state, with no single dominant category — much like GLM-5.2's failure mix. The difference between the two models is not what goes wrong; it is how often, and how repeatably.

Figure 4. Left: failure mass by state-diff category, Kimi K3 vs GLM-5.2 — both are broad, with no single dominant failure type. Right: Kimi K3's failure mass by environment. Derived from the frozen field-level grade; used for relative comparison, not as the official pass^5.

A model it outranks, ahead of it here

The cleanest way to feel the gap is head to head with GLM-5.2 — an open-weight model the public index ranks a full tier below Kimi K3 (public #10 vs #4).

Figure 5. Official Arena pass^5 by environment, GLM-5.2 vs Kimi K3. GLM wins five of seven — decisively in logistics (+26) and manufacturing (+19) — while Kimi K3 wins bank HR (+4.7) and essentially ties airlines. Task-weighted across all 834 tasks, GLM-5.2 leads by 8.9 points.

We are not cherry-picking. GLM-5.2 wins on balance and Kimi K3 legitimately takes two environments. But across the enterprise workload as a buyer would meet it, the model the leaderboard ranks lower is the more reliable enterprise agent — by nearly nine points of pass^5.

The panel tracks the public index — except for Kimi

The natural objection is that our benchmark is simply idiosyncratic. It is not. Across the fleet, Arena reliability tracks the public Agentic Index closely: for the sixteen non-Kimi models we can place on both axes, the two agree at Pearson r = 0.90, Spearman rho = 0.90. Strong-public models are strong here; weak ones are weak here.

Figure 6. Public Agentic Index (x) against official Arena pass^5 (y). Line and bands fit on the sixteen non-Kimi reference models. Kimi K3 (red) sits 15.7 points — more than two standard deviations (z = −2.1) — below where its public score predicts. GLM-5.2 (blue) sits above the line at a lower public score. Kimi K2.6 (orange) is below it too.

Against that agreement, Kimi K3 is the standout: it underperforms its public-score-predicted reliability by more than two standard deviations — and it is the only top-ranked public model that falls this far. Two other models sit as low relative to the line (MiniMax-M3, Nemotron 3 Ultra), but both are mid-to-low on the public index already; they are not cases of a celebrated model underdelivering. Kimi K3 is.

And this is not new with K3. Kimi K2.6 sat well below the line too (z = −1.4). The pattern — strong reputation, lower private-enterprise reliability — is a Kimi-family trait that K3 inherits and, because its public score climbed while its Arena reliability did not, actually widens.

What we are not saying

Three disclaimers, all meant.

We say nothing about coding. In our informal internal use, Kimi K3 is a strong coding model, and nothing here touches that. Coding and exact-state enterprise tool use are different capabilities; a model can lead at one and trail at the other. This is only about the second.

We claim nothing about intent. A model that ranks high on public agentic benchmarks and lower on private, exact-state enterprise evaluations is exhibiting a real reliability gap — but we have no evidence about, and make no accusation of, benchmark-gaming by the Moonshot team. We report what two kinds of measurement say and take no position on why they differ.

We claim no universal verdict. Kimi K3 wins two of our seven environments, is a real improvement over K2.6, and is excellent on public agentic tasks and on coding. The bounded claim is this: on private, realistic, exact-state enterprise workflows, Kimi K3 is not yet reliable enough to sit where its public ranking would put it.

So what

If you are choosing a model to stand behind a production agent that touches real systems — filing tickets, updating records, moving inventory, changing HR data — a public leaderboard rank tells you what the model can do, not what it will do every time. Kimi K3 is the current textbook case: a top-four public agent whose single-run competence does not survive the demand for five-in-a-row reliability, beaten on private enterprise workflows by an open-weight model the same leaderboard ranks a tier lower.

The rule this keeps producing: evaluate agents on the work you actually run, on tasks they can't have trained on, graded on the state you actually care about, and demand consistency — not a single lucky pass. Kimi K3 is a strong model. Whether it is reliable enough for your enterprise agent is a question only repeated runs on your own workflows can answer, and on ours the answer is: not yet.

Methods in brief

  • Metric. Composite pass^5 = fraction of tasks solved on all five runs, task-weighted across the seven environments (each environment weighted by its task count). Standard configuration: reasoning medium, max tokens 16,384, temperature 0.6. Headline scores and per-environment values are the official published figures from toloka.ai/arena (updated July 24, 2026).

  • Reliability retention and k-distribution. Computed from the frozen field-level grades over all seven domains; because that grade runs slightly stricter than the live leaderboard, we use retention ratios and cross-model ordering, not absolute levels, and take headline pass^5 from the published board.

  • Fleet concordance. OLS fit of Arena pass^5 on the public Agentic Index across the sixteen non-Kimi models with an unambiguous Arena counterpart; Kimi K3 and K2.6 placed against that reference fit. Public scores retrieved from the Artificial Analysis Agentic Index, 27 July 2026.

  • Failure modes. Deterministic state-diff of every failing trial, categorized by a fixed mutually-exhaustive taxonomy; inconsistency cards drawn from tasks Kimi K3 passed on some runs and failed on others, sanitized to roles.

Subscribe to Toloka news

Case studies, product news, and other articles straight to your inbox.