← Blog
/

Synthetic data is now available on the Toloka Platform
Toloka Arena is live. See how your model ranks.
Synthetic data generation is now available on the Toloka Platform. It’s a powerful option when speed matters, when you want to validate a pipeline design before recruiting a human audience, or when you need to quickly generate baseline data for human experts to validate and refine.
Where synthetic data fits in a pipeline
Our platform gives you the flexibility to choose the right approach for each specific project—running synthetic labeling when speed is the priority, or spinning up an expert pipeline for highly complex scenarios. Having both capabilities available in one place means your team can manage its entire data strategy without switching platforms.
Synthetic data is a strong fit for:
Reference-style tasks — anywhere a human would check a website and bring back the information, an LLM can now do the same job directly, and do it well
Search relevance evaluation — the model judges whether a given item matches a query
Audio transcription — the LLM produces a first-pass transcript and a human corrects it, which costs less than transcribing from scratch
These are just a few of the use cases teams are already running on the Toloka Platform.
A hybrid workflow in action
Consider a typical pipeline: a human expert interacts with an on-site agent to produce a conversation, then loads that conversation into Toloka to be labeled for safety, relevance, sentiment, and overall quality. Normally, human experts do that labeling. In the synthetic version, the expert still interacts with the agent and loads the conversation, but LLMs handle the labeling instead.
Quality you can trust
Speed and cost savings only matter if the data holds up. Every data point created synthetically passes through Toloka's LLM QA to confirm it meets quality standards before it reaches you. LLM QA checks for accuracy, task-instruction adherence, and consistency with your labeling guidelines. The same bar your pipeline would hold a human expert to.
Why teams use it
Synthetic data generation costs less than the equivalent human-labeled workflow, and there's no audience to recruit or ramp up before work starts, so results come back faster too. Paired with LLM QA, that means teams get useful data quickly without quality becoming an afterthought.
Pricing
Synthetic generation is billed per input and output token at the selected model's per-million-token rate. The pipeline cost estimate breaks this out as its own line item, labeled with the model and, if set, the reasoning effort.
How to get started
Open the node's Audience & Pricing card and select “synthetic data labeling”. You can also ask the agent to switch a generation node to synthetic labeling for you.
To configure it, you'll specify a few settings:
Choose a generation model from the list
Set a reasoning effort of low, medium, or high, where the selected model supports it
If the task needs guidance beyond the node's own task instruction, add a special LLM instruction; the platform sends both to the model together
A model tends to rate its own output too favorably, so choose a different model for LLM QA than the one handling generation.

Try it today
Open a Generation node in any general project and try synthetic labeling on your next run.
Synthetic data generation is one more way we're helping teams generate high-quality data without sacrificing speed or driving up cost — the same principle behind everything we build at Toloka. If you'd like to talk through your specific use case, connect with our team.
Subscribe to Toloka news
Case studies, product news, and other articles straight to your inbox.