← Blog

/

Insights

Insights

Synthetic data is now available on the Toloka Platform

Toloka Arena is live. See how your model ranks.

Synthetic data generation is now available on the Toloka Platform. It’s a powerful option when speed matters, when you want to validate a pipeline design before recruiting a human audience, or when you need to quickly generate baseline data for human experts to validate and refine.

Where synthetic data fits in a pipeline

Our platform gives you the flexibility to choose the right approach for each specific project—running synthetic labeling when speed is the priority, or spinning up an expert pipeline for highly complex scenarios. Having both capabilities available in one place means your team can manage its entire data strategy without switching platforms.

Synthetic data is a strong fit for:

  • Reference-style tasks — anywhere a human would check a website and bring back the information, an LLM can now do the same job directly, and do it well

  • Search relevance evaluation — the model judges whether a given item matches a query

  • Audio transcription — the LLM produces a first-pass transcript and a human corrects it, which costs less than transcribing from scratch

These are just a few of the use cases teams are already running on the Toloka Platform. 

A hybrid workflow in action

Consider a typical pipeline: a human expert interacts with an on-site agent to produce a conversation, then loads that conversation into Toloka to be labeled for safety, relevance, sentiment, and overall quality. Normally, human experts do that labeling. In the synthetic version, the expert still interacts with the agent and loads the conversation, but LLMs handle the labeling instead.

Quality you can trust

Speed and cost savings only matter if the data holds up. Every data point created synthetically passes through Toloka's LLM QA to confirm it meets quality standards before it reaches you. LLM QA checks for accuracy, task-instruction adherence, and consistency with your labeling guidelines. The same bar your pipeline would hold a human expert to.

Why teams use it

Synthetic data generation costs less than the equivalent human-labeled workflow, and there's no audience to recruit or ramp up before work starts, so results come back faster too. Paired with LLM QA, that means teams get useful data quickly without quality becoming an afterthought.

Pricing

Synthetic generation is billed per input and output token at the selected model's per-million-token rate. The pipeline cost estimate breaks this out as its own line item, labeled with the model and, if set, the reasoning effort.

How to get started

Open the node's Audience & Pricing card and select “synthetic data labeling”. You can also ask the agent to switch a generation node to synthetic labeling for you.

To configure it, you'll specify a few settings:

  • Choose a generation model from the list

  • Set a reasoning effort of low, medium, or high, where the selected model supports it

  • If the task needs guidance beyond the node's own task instruction, add a special LLM instruction; the platform sends both to the model together

A model tends to rate its own output too favorably, so choose a different model for LLM QA than the one handling generation.

Try it today

Open a Generation node in any general project and try synthetic labeling on your next run.

Synthetic data generation is one more way we're helping teams generate high-quality data without sacrificing speed or driving up cost — the same principle behind everything we build at Toloka. If you'd like to talk through your specific use case, connect with our team.

Subscribe to Toloka news

Case studies, product news, and other articles straight to your inbox.