← Blog

/

Arena

Arena

Gemini models explained: Gemini 4 Argon, Pro, Flash, and Flash-Lite

Gemini models explained: Gemini 4 Argon, Pro, Flash, and Flash-Lite

Toloka Arena is live. See how your model ranks.

See how frontier models perform on tasks they have never trained on

See how frontier models perform on tasks they have never trained on

Toloka Arena ranks Gemini, Claude, GPT, and other models by pass^5 reliability and cost per task on private, non-contaminated agentic benchmarks.

Tier names alone do not tell you where a model sits in the Gemini lineup. As of September 2026, the most recently released and most capable generally available Gemini model is a Flash model, not a Pro one. Google describes Gemini 3.8 Flash as its most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows. The Pro tier, meanwhile, is still on version 3.1, and still labelled preview.

And the next frontier model Google announced is not called Pro or Ultra at all. Gemini 4 Argon, unveiled on September 30, 2026, targets real-world software engineering, enterprise knowledge work, and cyber defense, but for now it is available only to trusted cyber defenders, with no public release date.

For teams putting Gemini into production, that inversion changes the default. Flash is no longer the model you settle for when Pro costs too much. It is where Google is shipping its fastest improvements, and it carries the agentic and coding workloads that other vendors reserve for their top tier.

Model lifecycle matters too. Google retires individual versions on published schedules, and it has already shut down Gemini 2.0 Flash, Gemini 2.0 Flash-Lite, and the original Gemini 3 Pro preview. Access to the 2.5 generation is now limited to projects that already used it, with new work pointed at 3.5 Flash-Lite or 3.8 Flash. Choosing a Gemini model is therefore less about picking the highest tier than about matching a capability profile, a price band, and a support horizon to the job.

This guide traces how Gemini’s tiers emerged, explains the differences that matter when choosing and deploying one, compares Gemini with Claude, GPT, Grok, and open-weight alternatives, and looks at where Google may take the family next.

What is Gemini AI?

Gemini is Google’s family of natively multimodal models, and also the name of its consumer assistant and its developer platform. First announced on December 6, 2023, Gemini is used for language, reasoning, coding, analysis, multimodal understanding, and agentic work, with different models offering different balances of capability, speed, and cost.

The models belong to the broader class of foundation models, general-purpose systems trained to support many downstream tasks. Gemini is available through gemini.google.com, the Gemini API and Google AI Studio, and the Gemini Enterprise Agent Platform on Google Cloud.

Google introduced the family with three models ordered by size and deployment target: Ultra for highly complex tasks, Pro for a wide range of tasks, and Nano for on-device work. That ladder did not survive intact. The current lineup is organised around Pro, Flash, and Flash-Lite, with a large set of specialist models around them.

One naming point causes persistent confusion. Gemini Ultra was a model in the 1.0 generation. Today, Google AI Ultra is a consumer subscription plan, not a model you can select in the API. When people ask which Gemini model Ultra is, the answer is that it is a plan that unlocks higher limits and the Deep Think reasoning mode.

The other distinctive feature of Google’s approach is breadth. Where competing families keep a compact ladder of general-purpose tiers, Google publishes a general-purpose ladder plus dedicated models for image generation, video, music, speech, translation, embeddings, robotics, and agents, all reachable through one API with shared grounding entitlements. That is a genuine architectural difference, not just a longer catalogue, and it shapes which projects Gemini fits best. Our overview of multimodal models covers the underlying concepts.

How the Gemini model family evolved

Gemini did not begin with Pro, Flash, and Flash-Lite. At the December 2023 announcement, the family was Ultra, Pro, and Nano, and Gemini Pro reached Google Cloud customers within a week. In February 2024, the Bard chatbot was renamed Gemini and the Duet AI branding across Google Cloud and Workspace was retired in favour of the Gemini name.

The first structural change came with the 1.5 generation, which introduced Flash as a distilled, faster, cheaper sibling of Pro and pushed context windows far beyond what the 1.0 models supported. Ultra quietly stopped being a model line, and Pro became the top tier. The 2.0 generation replaced 1.5 as the workhorse, and 2.5 introduced hybrid reasoning with configurable thinking budgets, which is the direct ancestor of the thinking controls in today’s models.

The Gemini 3 generation arrived with Gemini 3 Pro in November 2025, followed by Gemini 3 Flash in December, which became the default model in the Gemini app. Then the cadence changed. Gemini 3 Deep Think landed in February 2026, Gemini 3.1 Pro entered preview later that month, and the 3.5 family was announced at Google I/O in May. Gemini 3.6 Flash and 3.5 Flash-Lite reached general availability in July, Gemini 3.7 Flash followed in August, and Gemini 3.8 Flash became generally available on September 2, 2026, alongside a restricted cybersecurity variant, 3.8 Flash Cyber. Google called it its third Flash release in six weeks. September also brought Gemini 3.8 Live, a Live Avatar layer for enterprise voice agents, and on September 30 the announcement of Gemini 4 Argon, the first model of the fourth generation. 

The structural change worth noticing is that Flash acquired its own version track and started moving faster than Pro. Between February and September 2026, Google shipped four Flash generations and no new generally available Pro. The tier names still describe cost and latency profiles, but the version numbers now carry capability, which is why a Flash model can sit above a Pro model in the lineup. Gemini 4 then broke the pattern from the other direction, arriving with a new name, Argon, rather than as a Gemini 4 Pro. For the wider chronology behind this shift, see our history of LLMs.

Gemini AI models: Pro, Flash, Flash-Lite, and the specialists

With the lineage established, the useful distinction among current Gemini models is how they spend inference budget. All Gemini 3 models support a 1 million token input context window and up to 64,000 output tokens, and all of them reason dynamically by default. What differs is how much thinking each tier does, how fast it returns, and what it costs.

The Gemini 3 conventions also changed how developers control that behaviour. A string thinking level of low, medium, or high replaced the older numeric thinking budget, and the classic sampling parameters temperature, top_k, and top_p are now ignored by the backend. Teams migrating from 2.5 need to strip those parameters rather than tune them.

Because prices change, the comparisons below use relative cost positions alongside a few published figures. Check Google’s current pricing before deployment, and note that several Flash prices are introductory rates that step up on January 1, 2027.

Gemini Flash-Lite: the high-volume tier

Flash-Lite is the cost floor of the family, optimised for high-volume agentic tasks, translation, classification, and simple data processing where latency and unit cost matter more than depth. Two generations run in parallel: Gemini 3.5 Flash-Lite and Gemini 3.1 Flash-Lite, the latter cheaper still.

Its most instructive role is as a subagent. In systems that need many fast calls rather than one deeply reasoned response, Flash-Lite is the worker model a stronger planner delegates to. It is also the fallback tier inside the Gemini app, where subscribers who exhaust their Pro limits can keep working on Flash-Lite. The same economics that make small language models attractive apply here: for well-scoped tasks, a cheaper model that runs ten times more often often beats a stronger one you can only afford to call once.

Gemini Flash: the tier that moved to the front

Gemini 3.8 Flash is where Google is now concentrating its production work. It supports the 1 million token context window, 64,000 output tokens, and tunable thinking levels, with medium as the default and minimal explicitly unsupported. It is also the default model behind the Antigravity managed agent, which is a clear signal of where Google expects agentic volume to run.

Google is candid that 3.8 Flash uses more tokens by design on long, complex tasks. It takes smaller reasoning steps, calls tools iteratively, and verifies its work along the way, which raises quality on multi-step goals and raises token consumption with it. For everyday work, lowering the thinking level or staying on Gemini 3.7 Flash is the intended alternative. Toloka Arena data bears that out: 3.8 Flash scores 63.5% pass^5 against 62.7% for 3.7 Flash, a gap well inside the error bars, while costing roughly 70% more per task.

The pricing shows how deliberately Google is defending this band. Gemini 3.8 Flash runs at an introductory $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, doubling to $1.50 and $7.50 from January 1, 2027. The older Gemini 3.5 Flash, still live, is priced at $1.50 and $9.00. The newer, more capable Flash model is cheaper than the one it replaced even at full price.

The same release introduced Gemini 3.8 Flash Cyber, built on the same foundation but shipped with more permissive cybersecurity mitigations for vulnerability detection and automated patching. It is available only to trusted defenders through Google’s Fairwind Program. On the external CWE-Bench patching benchmark, Google reports a pass@1 of 47.2% against 47.8% for a leading frontier model, at significantly lower cost.

The practical complication is that four Flash generations are live at once. Pin a specific stable model string in production rather than a latest alias, because the alias is hot-swapped as new releases land.

Gemini Pro: the multimodal and long-context tier

Gemini 3.1 Pro remains Google’s headline model for multimodal understanding, complex problem-solving, and vibe coding, and it is the only current tier with a long-prompt price step: $2.00 per million input tokens and $12.00 per million output tokens for prompts up to 200,000 tokens, rising to $4.00 and $18.00 above that threshold. It is also the text pricing basis for Nano Banana Pro, Google’s highest-quality image model.

Two things follow from that. First, Pro’s differentiator is now less that it is the smartest available model and more that it is the multimodal and very-long-context model, while Flash carries agentic and coding work. The agentic gap is not subtle: on Toloka Arena, Gemini 3.1 Pro scores 10.0% pass^5, against 63.5% for 3.8 Flash at a similar cost per task. Second, at the January 2027 standard rates, the output price gap between 3.8 Flash and 3.1 Pro narrows to less than two times, which makes the choice between them a question of task profile rather than budget tier.

Pro is also still officially in preview. That is not a warning against using it, since Google states preview models may be used in production, but preview models carry more restrictive rate limits and can be deprecated with as little as two weeks’ notice. Controllable inference across both Pro and Flash is part of the broader shift toward reasoning in large language models.

Deep Think: a reasoning mode, not a tier

Gemini 3 Deep Think is not a separate model but a mode that runs on the Pro model and uses Google’s highest thinking level to explore more candidate reasoning paths in parallel. Queries typically take minutes rather than seconds. Google positions the upgraded Deep Think specifically at science, research, and engineering problems.

Choosing Deep Think has an operational consequence that capability tables do not show: it is gated behind the Google AI Ultra subscription in the Gemini app, with API access opened only to selected researchers and enterprises through an early-access programme. The strongest reasoning configuration Google currently ships is therefore not something most teams can call from the API, which matters when a benchmark result you read about was produced in a mode you cannot deploy.

Gemini 4 Argon: the announced frontier tier

Gemini 4 Argon is Google’s new frontier model, announced on September 30, 2026 and built to sustain deep reasoning across long-horizon workflows. Google positions it for three areas: real-world software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense. It is already running internally, where Google reports Argon agents migrating C and C++ codebases to Rust and freeing more than 300 TiB of data-centre memory through automated optimisations.

The headline specification change is output length. Google is raising Argon’s output token limit to 1 million tokens, up from the 64,000 tokens of the Gemini 3 models, so a single trajectory can generate hundreds of thousands of tokens of reasoning and code. Google’s launch results include a state-of-the-art 77.9% on DeepSWE v1.1 for long-horizon software engineering, first place on the Vals Index for economically weighted professional work, 51.3% on Zapier’s AutomationBench, 91.7% on LVBench for long-video understanding, and a tie for first at 68% on CWE-bench v1 for vulnerability remediation. These are vendor-reported figures and should be read as such until independent evaluations catch up.

Argon will launch at an introductory $2 per million input tokens and $10 per million output tokens, with cached input 95% cheaper, before moving to $4 and $20. The post-introductory rate matches Claude Opus 5.5 at $4 and $20 and sits well below GPT-6 Astra at $10 and $50, while the introductory output price is lower than Gemini 3.1 Pro’s.

Access is the real constraint. Argon is rolling out first to trusted cyber defenders in the Fairwind Program, who receive it without cyber guardrails, while Google takes part in the US government’s voluntary pre-release access process and strengthens safeguards against misuse, prompt injection, and misalignment, including monitoring of the model’s chain of thought and internal activations. Broader availability will start with paid API customers and Google AI Ultra subscribers, but Google has not given a date or published a model ID. Until it does, Argon is something to plan for rather than deploy, and teams evaluating it will want their own red teaming and reliability testing alongside Google’s numbers.

The specialist models around the core tiers

The rest of the Gemini catalogue is where the family diverges most sharply from its competitors. Image generation runs through Nano Banana Pro, Nano Banana 2, and Nano Banana 2 Lite, with the older Imagen 4 now shut down. Video covers Veo 3.1 and Veo 3.1 Lite alongside Gemini Omni Flash, which handles generation, editing, keyframe interpolation, and extension with native audio. Music generation runs on Lyria 3.5, Lyria 3 Clip, and Lyria RealTime.

Voice is similarly segmented: Gemini 3.8 Live and 3.8 Live Extended Thinking for audio-to-audio agents, 3.8 Flash TTS across 130 languages for studio-grade synthesis, 3.8 Flash-Lite TTS across 101 languages for high-throughput production, 3.5 Transcribe with speaker diarization and word-level timestamps, and 3.5 Live Translate across more than 70 languages. In late September, Google added Live Avatar to Gemini 3.8 Live, a near-real-time video persona with lip-sync across 97 languages, generally available in Gemini Enterprise for customer service and guided-walkthrough agents.

On the agentic side, Google ships a Computer Use model for driving interfaces, Deep Research and Deep Research Max for multi-source synthesis, and the Antigravity agent, which runs code, manages files, and browses inside an isolated Linux sandbox. If you are evaluating that category, our guide to computer use agents covers how they work and how to deploy them safely. Gemini Embedding 2 extends retrieval into a unified multimodal embedding space across text, images, video, audio, and PDFs.

Robotics is the clearest example of Google building where others have not. Gemini Robotics ER 2 and ER 1.6 are embodied reasoning models for video understanding, spatial reasoning, multi-step task orchestration, and multi-robot collaboration. They sit in the vision-language model lineage but are aimed at physical environments, which brings a different evaluation problem and a different robotics data requirement with it.

One deployment detail applies across the family: grounding with Google Search and Google Maps is built into the API, with 5,000 free requests per month shared across all Gemini 3.x models and $14 per 1,000 requests after that. Grounded retrieval as a first-party, metered capability is something no other major family offers in quite this form.

Gemini Pro vs Flash vs Flash-Lite

The differences among current Gemini models become clearest when the same workload constraints are compared side by side. Flash-Lite minimises latency and cost, Flash now carries the bulk of capable production work, Pro leads on multimodal understanding and very long prompts, and Deep Think extends the reasoning ceiling for a narrow set of hard problems.

Model

Position

Best for

Relative cost

Comparative latency

Gemini 3.5 Flash-Lite

High-volume workhorse

Classification, translation, simple data processing, subagent tasks

Lowest

Fastest

Gemini 3.8 Flash

Most intelligent GA model

Long-horizon coding, autonomous agents, enterprise workflows

Medium

Fast

Gemini 3.1 Pro (preview)

Multimodal and long-context tier

Multimodal understanding, very long prompts, vibe coding

High

Moderate

Gemini 3 Deep Think

Reasoning mode on Pro

Hard multi-step science, research, and engineering problems

Highest

Slowest

Gemini 4 Argon (announced)

Frontier tier, restricted access

Long-horizon coding, legal and finance work, cyber defense

High

Not yet public

How Gemini models score on Toloka Arena

Capability positioning is one thing; reliability on unseen agentic work is another. Toloka Arena evaluates models on private, multi-turn tool-use tasks with strict business rules and reports a composite pass^5 score, the probability that all five independent runs of a task succeed, alongside the average inference cost per task.

Model

Arena rank 

(of 21)

Composite pass^5

Avg cost per task

On cost-reliability frontier

Gemini 3.8 Flash

7

63.5%

~$2.00

Yes

Gemini 3.7 Flash

9

62.7%

~$1.20

Yes

Gemini 3.1 Pro (preview)

20

10.0%

~$1.80

No

Source: Toloka Arena composite leaderboard. Gemini Flash-Lite and Deep Think are not currently on the leaderboard, and Gemini 4 Argon was just announced and is not yet publicly available.

Three things stand out. Both current Flash models sit on the cost-reliability frontier, meaning no model on the board delivers a higher pass^5 score for less money. Gemini 3.7 Flash is the cheapest model on the entire leaderboard to clear 60%. And the Pro tier, which leads on multimodal understanding, ranks 20th of 21 on agentic reliability, which is the clearest single illustration of why tier names should not drive model selection in this family.

The tier table gets you to a shortlist, and the Arena scores show where reliability actually sits. From there, evaluate the shortlisted models on your own representative prompts, data, and edge cases, then weigh the resulting quality against latency and cost. Keep the consistency question in view: single-run benchmark scores show whether a model can succeed, while pass^5 shows whether it succeeds every time, which is what production workflows depend on.

A practical strategy is to start at the cheapest tier that plausibly satisfies the workload and move upward only when evaluation exposes a capability gap. Gemini adds one wrinkle to that rule. Because Flash generations ship every few weeks, the cheapest tier that works is a moving target, and a model that was not good enough three months ago may be good enough now. Build the evaluation once so you can rerun it cheaply, rather than treating model selection as a decision you make a single time.

Evaluate your model on private benchmarks

Test your model on Toloka Arena’s non-contaminated agentic tasks, or license the RL and evaluation datasets behind the leaderboard.

Check our OTS Datasets →

Gemini vs other AI model families

The snapshot below reflects public model lineups in September 2026. Benchmark positions change quickly, so our LLM leaderboard overview is the better place for current rankings, and Toloka Arena tracks how Gemini, Claude, and GPT models perform on private, non-contaminated agentic benchmarks, scored by pass^5 reliability and cost per task. Because Arena tasks are never in any model’s training data, the standings show how these families behave on unseen work rather than on memorised tests.

Model

Developer

Arena rank

(of 21)

Composite pass^5

Avg cost per task

On cost-reliability frontier

Claude Opus 5.5

Anthropic

1

77.3%

~$8.50

Yes

Claude Fable 5.1

Anthropic

2

74.0%

~$20.00

No

GPT-6 Astra

OpenAI

3

70.2%

~$12.80

No

Claude Sonnet 5.5

Anthropic

4

68.7%

~$4.20

Yes

GPT-5.6 Sol

OpenAI

6

64.7%

~$2.60

Yes

Gemini 3.8 Flash

Google

7

63.5%

~$2.00

Yes

GPT-6 Sol

OpenAI

8

63.2%

~$3.20

No

Gemini 3.7 Flash

Google

9

62.7%

~$1.20

Yes

Grok 4.6

xAI

10

62.1%

~$4.20

No

MiMo V2.6 Pro

Xiaomi

11

58.7%

~$0.78

Yes

Qwen 3.8 27B

Alibaba

13

41.0%

~$0.70

Yes

GPT-6 Luna

OpenAI

14

39.1%

~$0.15

Yes

DeepSeek V4.1 Flash

DeepSeek

16

34.0%

~$0.48

No

Gemini 3.1 Pro (preview)

Google

20

10.0%

~$1.80

No

Source: Toloka Arena composite leaderboard. Selected models shown; 21 are ranked in total.

The top of the board belongs to Anthropic and OpenAI’s frontier tiers, and they are priced accordingly. Gemini’s position is in the band just below: Gemini 3.8 Flash trails Claude Opus 5.5 by about 14 points of pass^5 at roughly a quarter of the cost per task, and Gemini 3.7 Flash does so at about one seventh. For workloads where 63% reliability is sufficient, or where a stronger model can be reserved for the hardest cases, that is the part of the market Google is clearly optimising for. Gemini 4 Argon is Google’s bid to contest the top of the board as well, and its launch benchmarks are pitched directly against GPT-6 Astra and Claude, but whether that holds on private, non-contaminated agentic tasks will only be measurable once Argon is publicly available. The comparison below focuses on how the families are packaged, deployed, and controlled.

Gemini vs ChatGPT and GPT models

For teams deploying AI systems, the closer comparison is Gemini versus the GPT model family rather than Gemini versus ChatGPT, which is the consumer-facing app. OpenAI’s GPT-6 lineup, completed in September 2026, keeps a stable three-tier ladder: Astra at the frontier, Sol for complex coding and agentic work, and Luna for focused, high-volume tasks.

The structural difference is version discipline. OpenAI’s tiers move together within a generation, so the tier name reliably tells you the capability order. Gemini’s tiers move independently, which gives Google room to iterate quickly in the price-performance band but means practitioners have to check version numbers rather than trust the tier label. Head to head on Arena, Gemini 3.8 Flash and GPT-6 Sol are effectively level on pass^5, with Gemini cheaper per task, while at the budget end GPT-6 Luna reaches 39.1% for around $0.15 per task, a price point no Gemini model on the board currently matches. Our breakdown of the GPT model family goes through that lineup in detail.

Gemini vs Claude

Anthropic’s Claude lineup is the cleanest counterexample. Haiku, Sonnet, and Opus are persistent capability tiers, with the Mythos-class Fable extending the ladder above Opus. Names stay fixed and version numbers move underneath them, so the ordering is stable even when individual tiers update at different times. Even so, Arena shows the ladder is not strictly monotonic on agentic reliability: Opus 5.5 leads the board at 77.3%, ahead of Fable 5.1 at 74.0%, which costs more than twice as much per task. Argon’s rollout also mirrors Anthropic’s approach to Mythos: both companies now release their most capable cyber-relevant models first to vetted defenders under separate safeguard regimes, before any general release.

Gemini inverts that. Its names describe an operating profile, cost and latency, while the version number carries capability, which is how a Flash model comes to sit above a Pro model. Both families also have an operational catch that raw capability tables miss. With Claude, it is safeguard routing, where some requests to the frontier model are answered by a less capable one. With Gemini, it is version churn and preview status: the model you benchmarked may be superseded or restricted faster than you expect. Our guide to the Claude model family covers that side in full.

Gemini vs Grok

xAI’s Grok lineup is narrower and more concentrated around coding, agentic tasks, and knowledge work. On Arena, Grok 4.6 lands just behind Gemini 3.7 Flash at 62.1% pass^5 but at roughly three and a half times the cost per task, while the newer Grok 4.7 scores 37.5%. The overlap with Gemini’s core tiers is real, but xAI publishes nothing comparable to Google’s media, speech, embedding, and robotics catalogue, so the choice is usually decided by whether a project needs that breadth.

Gemini vs open-weight models

Google is unusual in sitting on both sides of this comparison. Gemma 4, released in April 2026 under the Apache 2.0 licence, ships in E2B, E4B, 26B A4B, and 31B sizes with up to 256,000 tokens of context and support for more than 140 languages. The smallest variants are built to run offline on phones and edge hardware.

That means a team can weigh managed Gemini against self-hosted weights without leaving Google’s ecosystem, as well as against Llama, DeepSeek, Qwen, and GLM. The engineering trade-off is the familiar one. With the Gemini API, model serving, grounding, and vendor-side safeguards stay with Google. Self-hosted models give teams more control over infrastructure, fine-tuning, and data flow, while shifting more responsibility for deployment and maintenance onto the team running them. The exact freedoms depend on the licence, which is why open-weight is usually more precise than open-source. Our guide to open-source vs closed-source LLMs goes deeper into that distinction.

Google AI models: What comes next for Gemini?

The question of what follows Gemini 3.1 Pro now has a partial answer. Google spent most of 2026 shipping Flash generations while Pro stayed in preview, and the larger release turned out to be Gemini 4 Argon. What Google has not said is how Argon relates to the existing tiers: whether it replaces Pro at the top of the API lineup, sits above it the way Deep Think sits above standard Pro, or is the first of several Gemini 4 models with their own Flash and Flash-Lite counterparts. The 3.8 Flash pricing schedule, which doubles at the start of 2027, suggests Google expects the Flash band to stay commercially central whatever Argon becomes.

The direction of travel is agentic. Antigravity, managed agents, Deep Research, Computer Use, and the Robotics ER models all point the same way: the model is becoming one layer of a product rather than the product itself. Deprecation patterns reinforce this, with older general-purpose models retired quickly while specialist models accumulate.

Access is also likely to keep differentiating. Argon’s cyber-first rollout through the Fairwind Program, and 3.8 Flash Cyber before it, make access tiers an explicit part of Google’s release strategy. Deep Think already shows that the strongest configuration can be offered under a different access regime from the rest of the family, and grounding entitlements, priority inference, and enterprise-only throughput all create tiers that sit outside the model names entirely.

How human data supports model training and evaluation

Google publishes evaluation methodology alongside its model releases, but the underlying work still depends on people who design evaluations, investigate failures, and probe models for behaviour that automated tests miss. Techniques like RLHF and its variants automate part of the preference-labelling process, not the judgement that defines what good behaviour looks like in the first place.

Gemini’s breadth creates a specific version of that problem. A family spanning text, image, video, speech, translation, embeddings, and robotics needs evaluation data in each modality, and automated scores travel badly across them. A metric that works for text summarisation says little about whether a generated video is coherent, whether a transcription preserved speaker intent, or whether a robot’s plan was physically sensible.

Human input matters most during training and fine-tuning when the target behaviour cannot be captured by a simple rule or automated score. Domain experts can supply examples, preference judgements, and annotations for specialist tasks, and gaps in that data tend to surface later as uneven model performance. With agents, evaluation often means judging intermediate steps such as plans, tool calls, and error recovery rather than only the final answer.

Human data is therefore most valuable where the target is hard to measure automatically: specialist knowledge, subjective quality, unusual failure modes, and safety boundaries. Toloka provides data for LLM training and fine-tuning, expert model evaluation, and AI safety testing for those parts of the development process.

Train and evaluate your AI with expert human data

Toloka Platform delivers high-quality data for LLM training, RLHF, and model evaluation.

Get started free →

Frequently asked questions

What is the Gemini model family?

Gemini is Google’s family of natively multimodal AI models. The names Pro, Flash, and Flash-Lite identify operating profiles, from highest capability to lowest cost, while numbers such as 3.1 or 3.8 identify generations. Around that core sit specialist models for image, video, music, speech, embeddings, robotics, and agents, all available through the same API.

Which Gemini model is the most capable right now?

It depends on what capable means for your workload. Gemini 3.8 Flash is the most intelligent generally available model and the one Google positions for long-horizon coding and autonomous agents. Gemini 3.1 Pro leads on multimodal understanding and very long prompts but remains in preview. Gemini 4 Argon, announced on September 30, 2026, is positioned above all of them, but it is not yet publicly available. Deep Think, running on Pro, is the strongest reasoning configuration currently shipping, though it is restricted to Google AI Ultra subscribers and a limited API early-access programme. On agentic reliability specifically, Toloka Arena puts 3.8 Flash at 63.5% pass^5 and 3.1 Pro at 10.0%.

Why is Gemini Pro still on version 3.1?

Google does not update every tier on the same schedule. Gemini 3.1 Pro has been the current Pro model since February 2026 while Flash has moved through four generations. The version number tells you when a model was released, not where it sits in the current capability order, which is why comparing a 3.1 Pro with a 3.8 Flash on version number alone gives the wrong answer.

What is the difference between Gemini Flash and Flash-Lite?

Flash is the capable production tier, built for multi-step coding, agentic tool use, and enterprise workflows, with tunable thinking levels. Flash-Lite is the cost floor, optimised for high-volume, low-latency work such as classification, translation, and simple data processing. In practice, many systems use both, with Flash planning and Flash-Lite executing narrow subtasks in parallel.

Which Gemini model is best for agentic tool use?

It depends on how much reliability the workflow needs. Toloka Arena evaluates Gemini models alongside Claude and GPT on private, multi-turn tool-use tasks with strict business rules, scoring each model by pass^5, the probability that all five independent runs of a task succeed. As of September 29, 2026, Gemini 3.8 Flash scores 63.5% pass^5 at about $2.00 per task and Gemini 3.7 Flash scores 62.7% at about $1.20, both on the cost-reliability frontier, while Gemini 3.1 Pro scores 10.0%. For most agentic workloads, 3.7 Flash is the better-value starting point, with 3.8 Flash worth testing where its extra verification steps pay off. Standings also differ across Arena’s domains, from airline and manufacturing scenarios to banking HR and logistics, so it is worth checking the domain closest to your own workload rather than the composite score alone.

What is Gemini 4 Argon and can I use it yet?

Gemini 4 Argon is Google’s new frontier model for long-horizon software engineering, enterprise knowledge work, and cyber defense, with an output limit of 1 million tokens. As of October 1, 2026, it is available only to trusted cyber defenders in Google’s Fairwind Program and to Google’s internal teams. Google says broader access will start with paid API customers and Google AI Ultra subscribers, at an introductory $2 and $10 per million input and output tokens, but it has not announced a date or a public model ID.

What happened to Gemini 2.0 and 2.5?

Gemini 2.0 Flash and 2.0 Flash-Lite have been shut down. The 2.5 models are not deprecated and continue to be served, but access is now limited to projects that already used them, and Google directs new work to Gemini 3.5 Flash-Lite or 3.8 Flash. The original Gemini 3 Pro preview and the 3.1 Flash-Lite preview have also been retired. Check the deprecations page before building on anything outside the current stable set.


From model choice to system design

The lesson from Gemini’s expanding family is that model selection increasingly happens at the task level. A single project may need fast, inexpensive inference in one part of the workflow, deeper multimodal reasoning in another, and a specialist model for speech or video in a third. The practical question is not which Gemini model is best, but how to build the right model architecture for the project.

Gemini makes that pattern unusually visible because its catalogue spans so many operating profiles, but the same logic applies across model families. A project can route routine work to cheaper models and reserve stronger ones for tasks where evaluation shows a measurable gain.

Those routing decisions are only as good as the evaluation data behind them. Teams need representative prompts, edge cases, and human judgements that reflect real failure modes, and ideally independent evaluation on tasks models have never trained on. Otherwise, the model that looks best in testing may not be the one that performs best in production.

Related reading

Multimodal data annotation: the infrastructure layer behind today’s AI

AI agent evaluation: benchmarks and frameworks

How post-training teaches LLMs to reason

Transformer architecture



Subscribe to our newsletter

Product updates, case studies and the latest news from our team, straight to your inbox