Data your model hasn't seen and can't game.
Expert-built, verified by code, and proven to transfer.
One epoch of RL on our enterprise data lifts models on benchmarks they never trained on.
License our off-the-shelf data for eval, SFT or RL.
Trusted by Leading AI Teams
Proven to move models
One epoch of RL on our enterprise data lifts Qwen3.5-27B on agent benchmarks it never trained on.
Overall
AutomationBench
30.5% → 41.6%
+3.4pp partial credit
Pass@1 5.6% → 6.8% (+1.2pp)
τ³
Retail
77.6% → 86.8%
+8.9pp pass@1
Toolathlon
41.7% → 46.0%
+4.3pp pass@1
Where gains are largest
AutomationBench
Operations
29.9% → 39.6%
+9.6pp partial credit
Pass@1 7.7% → 11.7% (+4.0pp)
AutomationBench
Support
32.5% → 40.1%
+7.6pp partial credit
Why there
Gains land where our environments match the work. Operations and customer-support environments make up most of the training data.
What the model learned
Consults the knowledge base before its first write in 81.5% of held-out tasks, up from 49.6%.
It checks policy before acting, instead of acting first.
Qwen3.5-27B base vs. the same model after 72 GRPO steps on 1,540 enterprise tasks. Identical serving for both; paired 95% confidence intervals don't fall below zero. Evals Sep 25 – Oct 5, 2026.
Off-the-shelf datasets, ready to license
Private, verified data with deterministic rewards, ready for eval, SFT or RL.
Available now
Enterprise Tool Use
Enterprise RL Gyms
Maps to:
τ²
MCP-Atlas
RL environments where agents complete multi-step work across integrated enterprise tools (CRM, ERP, ticketing) in 11 industry environments.
Volume
2,500+ tasks
Domains
11
Grading
End state + tool-call trace + rubric
Uplift after one epoch of RL (Qwen3.5-27B)
AutomationBench Support
30.5% → 41.6%
τ³ retail 77.6% → 86.8%
Available now
Long-Horizon Knowledge Work
WorkBench
Maps to:
GDPval
Agents produce professional deliverables inside full company worlds, graded against real expert answers. Slide decks, financial models and legal memos, built in real tools (.pptx, .xlsx, .docx).
Volume
1,000 tasks
Rubric
~39 expert criteria per task
Harness
Inspect AI
Available now
Coding
Terminal-Bench 2.1 Extension
Maps to:
Terminal-Bench 2.1
Private terminal and CLI agent tasks for real command-line and system-level workflows. TerminalBench-3 compliant and deterministic, validated to eliminate false positives and negatives.
Volume
1,000 tasks
Harness
Harbor
Available now
STEM Reasoning
Frontier STEM
Maps to:
HLE
GPQA
AIME
AMO-Bench
PhD-authored problems across seven domains, weighted toward math and physics. Every answer is recomputed by standalone verification code.
Volume
1,000 problems
Verified
100% reproduced by code
Available now
Scientific Coding
SciCode Extension
Maps to:
SciCode
Expert-authored multi-step scientific coding tasks. Each splits into 2–5 sub-steps with ground-truth Python and assertion-based tests, in SciCode-harness-compatible JSON.
Volume
1,500 tasks
Validation
100% automated + expert review
Available now
Mobile Use
Mobile Use RL Gyms
1:1 UI-fidelity replicas of real mobile apps, each a self-contained RL package. Agents act on real backend state and are graded by deterministic state-diff.
Apps
13
Tasks
~390
Grading
Deterministic state-diff
Harness
Toloka Forge
Available now
Egocentric Video
Physical AI · Household activity
First-person household-activity footage for robotics pre-training and imitation learning, graded into capture-quality tiers with seven layers of time-aligned annotation.
Volume
25,000+ hours
Clips
216,000+
Coverage
12 activity categories, 24 countries
Annotation
7 time-aligned layers
Samples now
Enterprise Knowledge
EnterpriseBench Knowledge
Maps to:
τ³
AutomationBench
Toolathlon
Whole simulated companies: knowledge bases reachable only by search, tools the agent has to discover, and a user who withholds facts and introduces new events mid-conversation.
Knowledge base
500–1,000 docs per company
Per trial
16+ searches, 35+ tool calls
Difficulty
GPT-5.5 medium: under 10% pass^4
Also in the catalog
CharXiv CoT visual reasoning
Reasoning · Hard, deterministic questions on real arXiv charts, each with a gold five-step chain of thought
UMI manipulation
Physical AI · Head and two hand-held gripper cameras, 1080p30, 100 Hz IMU, 6-DoF pose, millimeter gripper aperture
500+ hours
Teleoperation
Physical AI · VR-driven bimanual and single-arm robots, 25 Hz state and actions, wrist and head RGB-D
Exocentric video
Physical AI · Third-person capture in studios, homes and commercial spaces
Coming soon
In production now. Join the waitlist to get samples first.
Terminal-Bench 4.0 Extension
Coding + STEM · Terminal-Bench 4.0 compliant, multi-container setup with a held-out verifier
SWE-Bench Pro Extension
Coding · Fresh, decontaminated, human-verified software-engineering environments mined from real repositories
LiveCodeBench Pro
Coding · Olympiad-authored competitive-programming problems with adversarial test suites and proven-wrong solutions
Lean Math
STEM · Expert-built Lean 4 theorem statements with reference proofs and grading rubrics, Harbor-compatible
Browser Use RL Gyms
Agents · Offline replicas of real web products with deterministic, final-state-graded tasks
Three ways to get started
01 — Ready to ship
Samples in 24 hours.
Fully produced, validated datasets. Request samples today, align on format, and license the full set with no ramp time.
02 — Pipeline-ready
Production live in 48 hours.
We align on volume, format and timeline at kickoff, then run our production pipeline to your spec at scale.
03 — Expand or customize
Tailored to your harness.
Grow any dataset in volume, adapt it to your domain, or turn it into a custom RL environment.
Questions labs ask
What are Toloka's off-the-shelf datasets?
Ready-to-license, expert-authored datasets for evaluating and post-training frontier models. They're private, verified and available now, and each one can be extended.
What kinds of datasets does Toloka offer?
Agentic RL environments (enterprise tool use, enterprise knowledge, mobile use), long-horizon knowledge work, coding, STEM reasoning and scientific coding. The catalog also covers visual reasoning and physical AI data: egocentric and exocentric video, UMI manipulation and teleoperation.
Who creates the datasets?
Domain experts from Toloka's network, including PhD researchers, engineers and industry professionals.
How do you know the data is correct?
Every task has a deterministic check: verification code for STEM, unit tests for coding, and final-state checks plus task-specific rubrics for enterprise tasks, with 35,000+ criteria in the enterprise set alone.
Does training on it actually improve models?
Yes. One epoch of RL on our enterprise data lifted Qwen3.5-27B on external agent benchmarks it never trained on, including +7.6pp on AutomationBench Support and +8.9pp on τ³ retail. Contact us for more details.
How do you keep it uncontaminated?
The data has never been published and Arena has no public test split.
What can we use the datasets for?
Evaluation, SFT and RL with verifiable rewards. Agentic environments ship with deterministic graders that you can set looser for training and stricter for evaluation. Coding tasks run on Harbor, enterprise and mobile RL gyms on Toloka Forge, WorkBench on Inspect AI, and SciCode in harness-compatible JSON. Physical AI data supports robotics pre-training, imitation learning, vision-language-action policy training and world models.
What format is the data delivered in?
Formats vary by dataset: harness-ready JSON for STEM and coding tasks, Docker-packaged RL environments for enterprise and mobile gyms, MCAP plus JSON for UMI, and MP4 plus JSON for video.
Can we review samples before licensing?
Yes. Request samples and we'll send them within 24 hours. Larger trials are available for selected labs.
Can we use the data for commercial training?
Yes. Contact us for licensing terms and pricing.
Can you extend a dataset or build a custom one?
Yes. We can scale volume, adapt a dataset to your domain or harness, or build something new. Pipeline-ready datasets go live within 48 hours.
Trusted by Leading AI Teams
See the data before you commit.
Tell us which capabilities you're working on. We'll send samples, specs and delivery timelines within 24 hours.