← Blog
/

From 41.6% to 64.5%: Post-Training Qwen3.5-27B on Our OTS Enterprise RL Data
Toloka Arena is live. See how your model ranks.
TL;DR
We ran RL post-training on Qwen3.5-27B using Toloka’s off-the-shelf (OTS) enterprise tool-use dataset.
The result: One epoch of training (72 steps) lifts Qwen3.5-27B from 41.6% to 64.5% on held-out OTS tasks (+22.8pp, 95% CI [+19.2, +26.6]).
Reliability: The share of tasks the model passes on all three attempts doubles, from 22.9% to 46.4%.
Rubric vs. binary reward: At equal training steps, a reward built on the rubrics that ship with OTS beats a binary pass/fail reward by +9.4pp [+6.1, +12.7].
Cross-domain transfer: On an environment kept out of training entirely, the pass rate rises from 14.0% to 58.0% (+44.0pp [+32.0, +56.0]).
External benchmarks: The model improves on three agent benchmarks it never trained on, including τ³ retail (+8.9pp), AutomationBench Operations (+9.6pp partial credit) and Toolathlon (+4.3pp).

Results
The largest gains are on held-out OTS tasks and the unseen environment. External gains are smaller and concentrated where OTS's environments match the benchmark.
Benchmark (tasks) | Qwen3.5-27B (base) | Qwen3.5-27B-OTS | Gain, pp (95% CI) |
Held-out OTS, pass rate (374) | 41.6% | 64.5% | +22.8 [+19.2, +26.6] |
Held-out OTS, pass^3 (371) | 22.9% | 46.4% | +23.5 [+18.1, +28.8] |
Unseen OTS environment, pass rate (34) | 14.0% | 58.0% | +44.0 [+32.0, +56.0] |
τ³ retail (114) | 77.6% | 86.8% | +8.9 [+4.5, +13.6] |
τ³, all four domains (375) | 69.6% | 70.9% | +1.2 [−1.2, +3.7] |
Toolathlon, pass@1 (108) | 41.7% | 46.0% | +4.3 [+0.0, +9.0] |
AutomationBench, pass rate (600) | 5.6% | 6.8% | +1.2 [+0.0, +2.4] |
AutomationBench, partial credit (600) | 25.0% | 28.4% | +3.4 [+1.5, +5.3] |
— Operations, partial credit (100) | 29.9% | 39.6% | +9.6 [+4.4, +15.0] |
— Support, partial credit (100) | 32.5% | 40.1% | +7.6 [+2.2, +12.8] |
— HR, partial credit (100) | 13.1% | 16.4% | +3.3 [−0.9, +7.6] |
OTS pass rates are final-state checks. Pass^3 counts a task only when all three attempts pass. Both models were served identically. Gains are paired differences over graded trials, so they can differ by a few tenths from subtracting the rounded rates. The unseen environment has 34 tasks, so its interval is wide (about ±12pp).
Most gains are statistically significant. Toolathlon and AutomationBench pass rate sit at the edge, with lower bounds of +0.0, and AutomationBench HR and τ³ across all four domains are within noise.
Dataset and splits
OTS contains 1,990 agentic enterprise tasks across 11 business environments, including customer support, travel, logistics and HR. Each task comes with:
a seeded simulated world the agent works in through tools;
a simulated user with a goal;
a final-state check that decides whether the job got done;
a rubric of task-specific criteria (35,201 across the dataset) that grades how the agent got there.
Split | Tasks | Environments | Purpose |
Training pool | 1,540 | 10 | RL training |
Holdout, trained environments | 340 | 10 (34 each) | Generalization to new tasks |
Holdout, unseen environment | 110 | 1 | Cross-domain transfer |
Reward design: binary vs. gated rubric
We trained the model twice, changing only how each attempt was scored. One run used a simple pass/fail check on the final result. The other used the task rubrics, which also grade how the agent got there. After 30 training steps each, the rubric-trained model passed 54.6% of held-out tasks, against 45.2% for the pass/fail one. That 9.4-point gain comes from the reward alone.
Standard RL setups for agents use a binary final-state check: was the database record updated, yes or no. In multi-step enterprise workflows, that signal has two problems:
Sparse feedback on failure. An agent that completes 8 of 9 steps earns the same 0 as one that fails on step 1.
Passes that hide problems. In our data, 29% of episodes that passed the state check still missed a required rubric criterion, either skipping a required step or breaking a rule.
Each rubric has two kinds of criteria. About 70% are things the agent must do, like looking up the record or following the policy step. The other 30% are things it must not do, like charging without the customer's consent. An agent that does nothing never breaks a rule, so a reward that simply counted criteria met would give it about 30% credit for doing nothing. We score the two kinds separately instead: doing the work adds reward, and breaking a rule takes it away.
The reward then sits in one of two bands:
Pass band [0.6–1.0]: the state check passes and every required criterion is met. The score starts at 1.0, and violations and missed optional steps each cost up to 0.2.
Fail band [0.0–0.4]: anything else. The agent earns partial credit for progress, reduced by violations, capped at 0.4.
The deterministic state check decides the band, and the judge only moves the score within it. A failed attempt can never outscore a pass, while partial credit keeps the model learning on tasks it can't yet solve.
Comparison | Binary reward | Rubric reward | Rubric − binary, pp (95% CI) |
Same 30 steps | 45.2% | 54.6% | +9.4 [+6.1, +12.7] |
End of each run (binary 60 steps, rubric 72) | 49.3% | 64.5% | +15.2 [+11.9, +18.5] |
These are held-out pass rates. The binary run plateaued after step 30, while the rubric run kept improving through step 72.
What the model learned
After one epoch, Qwen3.5-27B-OTS behaves like a more careful support agent: it checks policy and existing records before it changes anything. We measured this from the order of tool calls in every held-out episode (374 tasks × 3 attempts, both models on the same trials).
Behavior on held-out tasks | Base | Qwen3.5-27B-OTS |
Consults the knowledge base before its first write | 49.6% | 81.5% |
Makes a knowledge-base lookup its very first action | 26.0% | 64.3% |
Checks for an existing case before creating a new one | 61.4% | 81.9% |
A write is any action that changes state: create, update, submit, send, amend, cancel and similar.

How this looks in practice
Three scenarios from the held-out evaluation show the shift.
Says no when policy says no. An employee at a food-service chain asks to move her clock-out from 8:00 PM to 9:30 PM.
Base: Submits the correction straight away, routes it for manager approval and tells her it was "successfully processed."
Qwen3.5-27B-OTS: Searches the knowledge base for the time-correction procedure, then pulls her schedule and her punches. The claimed hours don't match the scheduled shift, so it searches the denial rule and checks for an existing case. It then documents a denial and explains the reason to her.
Rubric criteria it now meets: schedule checked before validating the punch; no false claim that the correction was submitted; denial reason communicated.
Validates before charging. A telecom customer asks to pay a $167 past-due balance with a saved payment method.
Base: Takes the $167 payment in all 3 attempts without ever validating the saved payment method.
Qwen3.5-27B-OTS: In all 3 attempts, lists the saved methods, checks payment authorization and validates the default method before charging. In 2 of 3, it starts by searching the payment policy.
Rubric criteria it now meets: saved methods reviewed before validation; method validated before payment; payment policy consulted.
Checks before writing, changes only what's needed. An air-cargo client asks to amend a waybill with revised dimensions.
Base: Creates the same support case twice, skips the route checks and amends the chargeable weight to 1,302.6 kg instead of 1,300.
Qwen3.5-27B-OTS: Checks for existing cases and pulls the contract and both station profiles that decide the route rules. It then amends exactly the two changed fields (1,300 kg, 7.8 cbm) and confirms the revised figures to the client by chat and email.
Rubric criteria it now meets: station profiles retrieved before the amendment; only changed fields amended; accurate confirmation to the client.
Training recipe
Base model: Qwen3.5-27B, LoRA rank 64 / alpha 32 on all linear layers.
Algorithm: GRPO with the Dr. GRPO advantage (reward minus group mean) and token-mean loss.
Batch: 16 tasks × 8 rollouts per step, two optimizer updates per policy sync. Groups where all 8 rollouts score identically are dropped and refilled to keep a learning signal.
Training sampling and optimizer: temperature 1.0, 8,192 tokens per call, constant learning rate 3e-5. Evaluations ran at temperature 0.6.
Length: 72 steps (one epoch), 11,293 training episodes.
Judge and simulated user: gpt-6-luna for both roles. The whole run cost about $26, or roughly 0.2¢ per training episode, which makes rubric-graded rewards practical at scale.
Key takeaways
The rubrics are the most valuable part of the data for RL. They grade what a final-state check can't see, and a reward built on them beats binary pass/fail by +9.4pp at equal steps.
The skills generalize, not just the training tasks. An environment never seen in training improves by +44pp, and external gains are largest where OTS's environments match the benchmark: AutomationBench Operations and Support, and τ³ retail.
A light training run is enough. All of these gains come from a single pass through the training data, with an LLM judge that cost about $26 for the whole run.
Subscribe to our newsletter
Product updates, case studies and the latest news from our team, straight to your inbox