Off-the-shelf datasets

Data your model hasn't seen and can't game.

Expert-built, verified by code, and proven to transfer.
One epoch of RL on our enterprise data lifts models on benchmarks they never trained on.
License our off-the-shelf data for eval, SFT or RL.

7K+expert-verified RL and eval tasks
25K+hours of physical AI data
25+domains, from enterprise ops to frontier STEM
24hsample turnaround. No build time

Trusted by Leading AI Teams

Proven to move models

One epoch of RL on our enterprise data lifts Qwen3.5-27B on agent benchmarks it never trained on.

Overall

AutomationBench

30.5% → 41.6%

+3.4pp partial credit

Pass@1 5.6% → 6.8% (+1.2pp)

τ³

Retail

77.6% → 86.8%

+8.9pp pass@1

Toolathlon

41.7% → 46.0%

+4.3pp pass@1

Where gains are largest

AutomationBench

Operations

29.9% → 39.6%

+9.6pp partial credit

Pass@1 7.7% → 11.7% (+4.0pp)

AutomationBench

Support

32.5% → 40.1%

+7.6pp partial credit

Why there

Gains land where our environments match the work. Operations and customer-support environments make up most of the training data.

What the model learned

Consults the knowledge base before its first write in 81.5% of held-out tasks, up from 49.6%.

It checks policy before acting, instead of acting first.

Qwen3.5-27B base vs. the same model after 72 GRPO steps on 1,540 enterprise tasks. Identical serving for both; paired 95% confidence intervals don't fall below zero. Evals Sep 25 – Oct 5, 2026.

Off-the-shelf datasets, ready to license

Private, verified data with deterministic rewards, ready for eval, SFT or RL.

Available now

Enterprise Tool Use

Enterprise RL Gyms

Maps to:

τ²

MCP-Atlas

RL environments where agents complete multi-step work across integrated enterprise tools (CRM, ERP, ticketing) in 11 industry environments.

Volume

2,500+ tasks

Domains

11

Grading

End state + tool-call trace + rubric

Uplift after one epoch of RL (Qwen3.5-27B)

AutomationBench Support
30.5% → 41.6%

τ³ retail 77.6% → 86.8%

Available now

Long-Horizon Knowledge Work

WorkBench

Maps to:

GDPval

Agents produce professional deliverables inside full company worlds, graded against real expert answers. Slide decks, financial models and legal memos, built in real tools (.pptx, .xlsx, .docx).

Volume

1,000 tasks

Rubric

~39 expert criteria per task

Harness

Inspect AI

Available now

Coding

Terminal-Bench 2.1 Extension

Maps to:

Terminal-Bench 2.1

Private terminal and CLI agent tasks for real command-line and system-level workflows. TerminalBench-3 compliant and deterministic, validated to eliminate false positives and negatives.

Volume

1,000 tasks

Harness

Harbor

Available now

STEM Reasoning

Frontier STEM

Maps to:

HLE

GPQA

AIME

AMO-Bench

PhD-authored problems across seven domains, weighted toward math and physics. Every answer is recomputed by standalone verification code.

Volume

1,000 problems

Verified

100% reproduced by code

Available now

Scientific Coding

SciCode Extension

Maps to:

SciCode

Expert-authored multi-step scientific coding tasks. Each splits into 2–5 sub-steps with ground-truth Python and assertion-based tests, in SciCode-harness-compatible JSON.

Volume

1,500 tasks

Validation

100% automated + expert review

Available now

Mobile Use

Mobile Use RL Gyms

1:1 UI-fidelity replicas of real mobile apps, each a self-contained RL package. Agents act on real backend state and are graded by deterministic state-diff.

Apps

13

Tasks

~390

Grading

Deterministic state-diff

Harness

Toloka Forge

Available now

Egocentric Video

Physical AI · Household activity

First-person household-activity footage for robotics pre-training and imitation learning, graded into capture-quality tiers with seven layers of time-aligned annotation.

Volume

25,000+ hours

Clips

216,000+

Coverage

12 activity categories, 24 countries

Annotation

7 time-aligned layers

Samples now

Enterprise Knowledge

EnterpriseBench Knowledge

Maps to:

τ³

AutomationBench

Toolathlon

Whole simulated companies: knowledge bases reachable only by search, tools the agent has to discover, and a user who withholds facts and introduces new events mid-conversation.

Knowledge base

500–1,000 docs per company

Per trial

16+ searches, 35+ tool calls

Difficulty

GPT-5.5 medium: under 10% pass^4

Three ways to get started

01 — Ready to ship

Samples in 24 hours.

Fully produced, validated datasets. Request samples today, align on format, and license the full set with no ramp time.

02 — Pipeline-ready

Production live in 48 hours.

We align on volume, format and timeline at kickoff, then run our production pipeline to your spec at scale.

03 — Expand or customize

Tailored to your harness.

Grow any dataset in volume, adapt it to your domain, or turn it into a custom RL environment.

Questions labs ask

What are Toloka's off-the-shelf datasets?

Ready-to-license, expert-authored datasets for evaluating and post-training frontier models. They're private, verified and available now, and each one can be extended.

What kinds of datasets does Toloka offer?

Agentic RL environments (enterprise tool use, enterprise knowledge, mobile use), long-horizon knowledge work, coding, STEM reasoning and scientific coding. The catalog also covers visual reasoning and physical AI data: egocentric and exocentric video, UMI manipulation and teleoperation.

Who creates the datasets?

Domain experts from Toloka's network, including PhD researchers, engineers and industry professionals.

How do you know the data is correct?

Every task has a deterministic check: verification code for STEM, unit tests for coding, and final-state checks plus task-specific rubrics for enterprise tasks, with 35,000+ criteria in the enterprise set alone.

Does training on it actually improve models?

Yes. One epoch of RL on our enterprise data lifted Qwen3.5-27B on external agent benchmarks it never trained on, including +7.6pp on AutomationBench Support and +8.9pp on τ³ retail. Contact us for more details.

How do you keep it uncontaminated?

The data has never been published and Arena has no public test split.

What can we use the datasets for?

Evaluation, SFT and RL with verifiable rewards. Agentic environments ship with deterministic graders that you can set looser for training and stricter for evaluation. Coding tasks run on Harbor, enterprise and mobile RL gyms on Toloka Forge, WorkBench on Inspect AI, and SciCode in harness-compatible JSON. Physical AI data supports robotics pre-training, imitation learning, vision-language-action policy training and world models.

What format is the data delivered in?

Formats vary by dataset: harness-ready JSON for STEM and coding tasks, Docker-packaged RL environments for enterprise and mobile gyms, MCAP plus JSON for UMI, and MP4 plus JSON for video.

Can we review samples before licensing?

Yes. Request samples and we'll send them within 24 hours. Larger trials are available for selected labs.

Can we use the data for commercial training?

Yes. Contact us for licensing terms and pricing.

Can you extend a dataset or build a custom one?

Yes. We can scale volume, adapt a dataset to your domain or harness, or build something new. Pipeline-ready datasets go live within 48 hours.

Trusted by Leading AI Teams

See the data before you commit.

Tell us which capabilities you're working on. We'll send samples, specs and delivery timelines within 24 hours.