Skip to main content
Verified model research

AI model benchmarks for practical decisions

Compare 86 models released since 2024 across 5 task families. Model records include linked official sources, while benchmark results retain their compatible evaluation snapshot.

86 models13 providers5 task familiesDataset 2026-09-09Verified 2026-09-09

Model explorer

Choose one task family, then filter its validated inventory. Charts and comparisons use compatible benchmark snapshots and pricing units; missing values remain visible as Not reported.

Choose a task

Compare models only within one compatible task family.

Decision summary

Facts from one compatible snapshot, measured 2026-06-25.

LiveBench · 2026-06-25

Highest score on LiveBench Overall

GPT-5.5

79.91 points on LiveBench Overall, snapshot 2026-06-25, measured 2026-06-25. Eligible models (4): GPT-5.4, GPT-5.5, Claude Opus 4.8, Claude Sonnet 5. Price unit: USD per 1M input tokens.

Methodology

Cost and quality frontier

3 nondominated models

LiveBench Overall, snapshot 2026-06-25, measured 2026-06-25, compared with USD per 1M input tokens. Eligible models (4): GPT-5.4, GPT-5.5, Claude Opus 4.8, Claude Sonnet 5.

Methodology
More filters
Providers

110 of 63 models

0 of 4 selected

AI model inventory with pricing, selected benchmark, provenance, and comparison controls
GPT-6 AstraCurrentProviderOpenAIRelease2026-09-03AccessAPICapabilitiesreasoningOperational limitsContext 1,050,000 · max output 128,000Price / 1M input tokens$10 / 1M input tokens · standardPrice / 1M output tokens$50 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • OpenAI positions GPT-6 Astra as its frontier model for reasoning, agentic work, and computer use, with a 1,050,000-token context window and 128,000-token maximum output
  • OpenAI reports the strongest published deep-context retention of any frontier model on MRCR v2 (8 needles): 100% up to 512K tokens and 96.3% in the 512K-1M band
  • OpenAI reports saturation-level scores on FrontierMath Tier 4 (97.6%) and, under its own provider adapter harness, ARC-AGI-3 (99.9%)
Cautions
  • Long-context billing is a cliff, not a taper: a request over the long-context threshold is rebilled in full at $20 per million input and $75 per million output, so a prompt that crosses it costs roughly double
  • GPT-6 Astra is the first publicly released model OpenAI classifies as crossing the Critical cybersecurity threshold under its own Preparedness Framework, which carries additional deployment safeguards
  • Vendor-reported benchmark figures come from OpenAI’s own launch material and its own harness; they are not independently reproduced here
Provenance
Gemini 3.8 FlashCurrentProviderGoogleRelease2026-09-02AccessAPICapabilitiesreasoningOperational limitsContext 1,048,576 · max output 65,536Price / 1M input tokens$0.75 / 1M input tokens · standardPrice / 1M output tokens$3.75 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • The cheapest frontier-tier option in this dataset by a wide margin: $0.75 per million input and $3.75 per million output, roughly one thirteenth of GPT-6 Astra on both
  • Accepts text, image, audio, video and PDF input with a 1,048,576-token context window and 65,536-token maximum output
  • Google reports it ahead of Gemini 3.7 Flash on every published benchmark, including DeepSWE v1.1 73.7% vs 65.3% and Terminal-bench 2.1 89.4% vs 85.8%
Cautions
  • The listed price is promotional and ends 31 December 2026; on 1 January 2027 input and output both double, to $1.50 and $7.50, and cached input goes from $0.075 to $0.15
  • Output pricing is billed inclusive of thinking tokens, so reasoning-heavy prompts cost more than the output length suggests
  • Google’s published comparisons are against its own previous Flash release and were not independently reproduced here
Provenance
Muse Spark 1.3CurrentProviderMetaRelease2026-09-02AccessAPICapabilitiesreasoningOperational limitsContext 1,048,576 · max output Not reportedPrice / 1M input tokens$1.25 / 1M input tokens · standardPrice / 1M output tokens$4.25 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • Meta tunes it for agentic workflows (multi-step tool use, browser control, and long-horizon tasks) with a 1,048,576-token context window
  • Meta reports it finishing coding work with roughly 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2, which compounds with the low token price
  • Behavioural changes aimed at real use: it asks clarifying questions on ambiguous prompts and confirms before consequential actions
Cautions
  • Meta’s own documentation states audio understanding is not fully supported and response quality may be degraded on requests containing audio
  • The much cheaper contributor endpoint (about $0.10 / $0.20 per million) is priced that way in exchange for Meta training on your traffic. The standard tier is the rate recorded here.
  • Meta’s scorecard has it trailing Claude Opus 5 on every broader agent-workflow benchmark it published, including GDPVal-AA v2, despite winning the coding rows
Provenance
Claude Fable 5.1CurrentProviderAnthropicRelease2026-09-01AccessAPICapabilitiesreasoningOperational limitsContext 1,000,000 · max output 128,000Price / 1M input tokens$10 / 1M input tokens · standardPrice / 1M output tokens$50 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • Anthropic’s highest-capability generally available model, for demanding reasoning and long-horizon agentic work, with a 1M-token context window and 128,000-token maximum output
  • Cache reads are priced at 0.025x base input ($0.25 per million) rather than the 0.1x every other Claude model uses, which is the single largest cost lever for agentic loops that replay a large prompt
  • Anthropic reports 52.6% on Terminal-Bench-Science 0.1 against 24.7% for Fable 5 and 29.0% for Opus 5
  • Adaptive thinking is always on and the API effort default is `high`
Cautions
  • Base input and output pricing is unchanged from Fable 5 at $10 / $50; the saving is entirely in cache reads, so a workload with little cache reuse sees no discount
  • The 4.7-generation tokenizer produces roughly 30% more tokens for the same text than Claude Sonnet 4.6 and earlier, so per-token rates are not directly comparable across generations
  • Anthropic’s launch comparisons are self-reported and were not reproduced independently here
Provenance
Claude Mythos 5.1CurrentProviderAnthropicRelease2026-09-01AccessAPICapabilitiesreasoningOperational limitsContext Not reported · max output Not reportedPrice / 1M input tokens$10 / 1M input tokens · standardPrice / 1M output tokens$50 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • Shipped alongside Claude Fable 5.1 and listed at the same rates: $10 per million input, $50 per million output, and the same 0.025x cache-read multiplier ($0.25 per million)
  • Carries the same reduced cache-read pricing as Fable 5.1, which no other Claude model has
Cautions
  • Limited availability: access is restricted to organisations vetted through Anthropic’s Glasswing programme, so it cannot be assumed available for a general evaluation
  • Anthropic’s public model comparison table does not list a context window or maximum output for Mythos 5.1, so both are recorded as unknown here rather than inferred from Fable 5.1
Provenance
Gemini 3.7 FlashCurrentProviderGoogleRelease2026-08-13AccessAPICapabilitiesreasoning, coding, agentic-coding, instruction-followingOperational limitsContext 1,048,576 · max output 65,536Price / 1M input tokens$0.75 / 1M input tokens · standardPrice / 1M output tokens$3.75 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • Google positions Gemini 3.7 Flash for higher-accuracy coding and agent workflows with multimodal inputs and a 1,048,576-token context window
Cautions
  • Google publishes the introductory API price through 2026-12-31; standard input and output prices increase on 2027-01-01.
  • No exact compatible public benchmark snapshot is recorded
Provenance
Grok 4.6CurrentProviderxAIRelease2026-08-12AccessAPICapabilitiesreasoningOperational limitsContext 500,000 · max output Not reportedPrice / 1M input tokens$2 / 1M input tokens · standardPrice / 1M output tokens$6 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • Grok 4.6 is positioned for coding, agentic tasks, and knowledge work; its API documentation also lists function calling and structured outputs
Cautions
  • Published token rates increase for all tokens in a request once its prompt reaches xAI's 200,000-token long-context threshold
Provenance
Qwen3.8-MaxCurrentProviderAlibaba QwenRelease2026-08-03AccessAPICapabilitiesreasoningOperational limitsContext 1,000,000 · max output 131,072Price / 1M input tokens$2 / 1M input tokens · standardPrice / 1M output tokens$6 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • A 2.4-trillion-parameter mixture-of-experts model activating about 95 billion parameters per query, with a 1M-token context window and reasoning enabled by default
  • Output is the cheapest of any frontier-tier text model here at $6 per million, an eighth of GPT-6 Astra and Claude Fable 5.1
  • Accepts text, image and video input and returns text
  • The 0902 post-training refresh (2 September 2026) targets coding and agent performance specifically
Cautions
  • Two releases share the name: the original 3 August 2026 model and the 0902 refresh. The rates and limits here are read from the 0902 listing
  • Reported benchmark placements come from Alibaba and from public arena boards rather than a common pinned snapshot, so they are not comparable with the LiveBench results in this dataset
Provenance
Claude Opus 5CurrentProviderAnthropicRelease2026-07-24AccessAPICapabilitiesreasoningOperational limitsContext 1,000,000 · max output 128,000Price / 1M input tokens$5 / 1M input tokens · standardPrice / 1M output tokens$25 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • Anthropic’s recommended default for most workloads, at half Fable 5.1’s token price ($5 / $25) with the same 1M-token context window and 128,000-token maximum output
  • Exposes five effort settings (low, medium, high, xhigh, max), which trade capability against token consumption without changing model
  • One of only two models with a Fast mode research preview, billed at $10 / $50 across the full context window
Cautions
  • Fast mode pricing applies to the whole request, including above 200K input tokens, so it doubles the effective rate rather than adding a surcharge
  • Cache reads are $0.50 per million (the standard 0.1x), five times Fable 5.1’s $0.25, even though Opus 5 is the cheaper model on base tokens. A cache-heavy agent loop can invert the expected cost ordering between the two.
Provenance
Gemini 3.5 Flash-LiteCurrentProviderGoogleRelease2026-07-21AccessAPICapabilitiesreasoningOperational limitsContext 1,048,576 · max output 65,536Price / 1M input tokens$0.3 / 1M input tokens · standardPrice / 1M output tokens$2.5 / 1M output tokens · standardLiveBench Web of Lies V2Not reported
View details
Strengths
  • Gemini 3.5 Flash-Lite is Google's low-latency, cost-efficient 3.5 model for high-throughput subagent and document-processing tasks
Cautions
  • The represented prices are standard API rates; batch, flex, and priority consumption use different rates
Provenance

Decision charts

Compare declared coverage and compatible operational values for the current task.

Capability coverage

Boolean provider-declared coverage for 10 current-task models. Cells are not scores and do not compare capability quality.

Legend: ✓ Supported · × Unsupported

Capability coverage matrixRows are models and columns are task-relevant capabilities. Each status chip explicitly states whether the model supports that capability.

GPT-6 Astra

ReasoningSupported
CodingUnsupported
Agentic CodingUnsupported
Instruction FollowingUnsupported

Gemini 3.8 Flash

ReasoningSupported
CodingUnsupported
Agentic CodingUnsupported
Instruction FollowingUnsupported

Muse Spark 1.3

ReasoningSupported
CodingUnsupported
Agentic CodingUnsupported
Instruction FollowingUnsupported

Claude Fable 5.1

ReasoningSupported
CodingUnsupported
Agentic CodingUnsupported
Instruction FollowingUnsupported

Claude Mythos 5.1

ReasoningSupported
CodingUnsupported
Agentic CodingUnsupported
Instruction FollowingUnsupported

Gemini 3.7 Flash

ReasoningSupported
CodingSupported
Agentic CodingSupported
Instruction FollowingSupported

Grok 4.6

ReasoningSupported
CodingUnsupported
Agentic CodingUnsupported
Instruction FollowingUnsupported

Qwen3.8-Max

ReasoningSupported
CodingUnsupported
Agentic CodingUnsupported
Instruction FollowingUnsupported

Claude Opus 5

ReasoningSupported
CodingUnsupported
Agentic CodingUnsupported
Instruction FollowingUnsupported

Gemini 3.5 Flash-Lite

ReasoningSupported
CodingUnsupported
Agentic CodingUnsupported
Instruction FollowingUnsupported
View exact capability data
Capability coverage data
ModelReasoningCodingAgentic CodingInstruction Following
GPT-6 AstraSupportedUnsupportedUnsupportedUnsupported
Gemini 3.8 FlashSupportedUnsupportedUnsupportedUnsupported
Muse Spark 1.3SupportedUnsupportedUnsupportedUnsupported
Claude Fable 5.1SupportedUnsupportedUnsupportedUnsupported
Claude Mythos 5.1SupportedUnsupportedUnsupportedUnsupported
Gemini 3.7 FlashSupportedSupportedSupportedSupported
Grok 4.6SupportedUnsupportedUnsupportedUnsupported
Qwen3.8-MaxSupportedUnsupportedUnsupportedUnsupported
Claude Opus 5SupportedUnsupportedUnsupportedUnsupported
Gemini 3.5 Flash-LiteSupportedUnsupportedUnsupportedUnsupported

Operational comparisons

Each panel has its own unit, direction, and fixed displayed range. Values are never combined across panels.

LiveBench Overall

points · Higher is better · Range 0100

LiveBench Overall5 models compared in points. Higher is better. Exact values are in the table below.Claude Sonnet 574.85Claude Opus 4.878.93Gemini 3.5 Flash74.64GPT-5.579.91GPT-5.477.97

Price per 1M input tokens

USD / 1M input tokens · Lower is better for cost · Range 010

Price per 1M input tokens28 models compared in USD / 1M input tokens. Lower is better for cost. Exact values are in the table below.GPT-6 Astra10Gemini 3.8 Flash0.75Muse Spark 1.31.25Claude Fable 5.110Claude Mythos 5.110Gemini 3.7 Flash0.75Grok 4.62Qwen3.8-Max2Claude Opus 55Gemini 3.5 Flash-Lite0.3Gemini 3.6 Flash1.5Grok 4.52Kimi K33GPT-5.6 Luna0.2GPT-5.6 Sol4GPT-5.6 Terra2Claude Sonnet 52Claude Fable 510Claude Opus 4.85Gemini 3.1 Flash-Lite0.25GPT-5.55Grok 4.201.25GPT-5.42.5Gemini 3 Flash Preview0.5GPT-4.12GPT-4.1 mini0.4GPT-4.1 nano0.1Grok 4.31.25

Price per 1M output tokens

USD / 1M output tokens · Lower is better for cost · Range 050

Price per 1M output tokens28 models compared in USD / 1M output tokens. Lower is better for cost. Exact values are in the table below.GPT-6 Astra50Gemini 3.8 Flash3.75Muse Spark 1.34.25Claude Fable 5.150Claude Mythos 5.150Gemini 3.7 Flash3.75Grok 4.66Qwen3.8-Max6Claude Opus 525Gemini 3.5 Flash-Lite2.5Gemini 3.6 Flash7.5Grok 4.56Kimi K315GPT-5.6 Luna1.2GPT-5.6 Sol20GPT-5.6 Terra12Claude Sonnet 510Claude Fable 550Claude Opus 4.825Gemini 3.1 Flash-Lite1.5GPT-5.530Grok 4.202.5GPT-5.415Gemini 3 Flash Preview3GPT-4.18GPT-4.1 mini1.6GPT-4.1 nano0.4Grok 4.32.5

Context window

tokens · Higher supports larger inputs · Range 01,050,000

Context window35 models compared in tokens. Higher supports larger inputs. Exact values are in the table below.GPT-6 Astra1,050,000Gemini 3.8 Flash1,048,576Muse Spark 1.31,048,576Claude Fable 5.11,000,000Gemini 3.7 Flash1,048,576Grok 4.6500,000Qwen3.8-Max1,000,000Claude Opus 51,000,000Gemini 3.5 Flash-Lite1,048,576Gemini 3.6 Flash1,048,576Grok 4.5500,000Kimi K31,048,576GPT-5.6 Luna1,050,000GPT-5.6 Sol1,050,000GPT-5.6 Terra1,050,000Claude Sonnet 51,000,000Claude Fable 51,000,000Claude Opus 4.81,000,000Gemini 3.5 Flash1,048,576Gemini 3.1 Flash-Lite1,048,576GPT-5.51,050,000Grok 4.201,000,000GPT-5.41,050,000Gemini 3.1 Pro Preview1,048,576Gemini 3 Flash Preview1,048,576Kimi K2 Base128,000Kimi K2 Instruct128,000GPT-4.11,047,576GPT-4.1 mini1,047,576GPT-4.1 nano1,047,576Kimi-VL-A3B-Instruct128,000Gemini 2.5 Pro1,048,576Qwen2.5-VL 72B Instruct128,000DeepSeek-R1128,000Grok 4.31,000,000
View exact operational data
Operational metric data
MetricModelValueDirection
LiveBench OverallClaude Sonnet 574.85 pointsHigher is better
LiveBench OverallClaude Opus 4.878.93 pointsHigher is better
LiveBench OverallGemini 3.5 Flash74.64 pointsHigher is better
LiveBench OverallGPT-5.579.91 pointsHigher is better
LiveBench OverallGPT-5.477.97 pointsHigher is better
Price per 1M input tokensGPT-6 Astra10 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGemini 3.8 Flash0.75 USD / 1M input tokensLower is better for cost
Price per 1M input tokensMuse Spark 1.31.25 USD / 1M input tokensLower is better for cost
Price per 1M input tokensClaude Fable 5.110 USD / 1M input tokensLower is better for cost
Price per 1M input tokensClaude Mythos 5.110 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGemini 3.7 Flash0.75 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGrok 4.62 USD / 1M input tokensLower is better for cost
Price per 1M input tokensQwen3.8-Max2 USD / 1M input tokensLower is better for cost
Price per 1M input tokensClaude Opus 55 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGemini 3.5 Flash-Lite0.3 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGemini 3.6 Flash1.5 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGrok 4.52 USD / 1M input tokensLower is better for cost
Price per 1M input tokensKimi K33 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGPT-5.6 Luna0.2 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGPT-5.6 Sol4 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGPT-5.6 Terra2 USD / 1M input tokensLower is better for cost
Price per 1M input tokensClaude Sonnet 52 USD / 1M input tokensLower is better for cost
Price per 1M input tokensClaude Fable 510 USD / 1M input tokensLower is better for cost
Price per 1M input tokensClaude Opus 4.85 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGemini 3.1 Flash-Lite0.25 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGPT-5.55 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGrok 4.201.25 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGPT-5.42.5 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGemini 3 Flash Preview0.5 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGPT-4.12 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGPT-4.1 mini0.4 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGPT-4.1 nano0.1 USD / 1M input tokensLower is better for cost
Price per 1M input tokensGrok 4.31.25 USD / 1M input tokensLower is better for cost
Price per 1M output tokensGPT-6 Astra50 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGemini 3.8 Flash3.75 USD / 1M output tokensLower is better for cost
Price per 1M output tokensMuse Spark 1.34.25 USD / 1M output tokensLower is better for cost
Price per 1M output tokensClaude Fable 5.150 USD / 1M output tokensLower is better for cost
Price per 1M output tokensClaude Mythos 5.150 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGemini 3.7 Flash3.75 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGrok 4.66 USD / 1M output tokensLower is better for cost
Price per 1M output tokensQwen3.8-Max6 USD / 1M output tokensLower is better for cost
Price per 1M output tokensClaude Opus 525 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGemini 3.5 Flash-Lite2.5 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGemini 3.6 Flash7.5 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGrok 4.56 USD / 1M output tokensLower is better for cost
Price per 1M output tokensKimi K315 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGPT-5.6 Luna1.2 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGPT-5.6 Sol20 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGPT-5.6 Terra12 USD / 1M output tokensLower is better for cost
Price per 1M output tokensClaude Sonnet 510 USD / 1M output tokensLower is better for cost
Price per 1M output tokensClaude Fable 550 USD / 1M output tokensLower is better for cost
Price per 1M output tokensClaude Opus 4.825 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGemini 3.1 Flash-Lite1.5 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGPT-5.530 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGrok 4.202.5 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGPT-5.415 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGemini 3 Flash Preview3 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGPT-4.18 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGPT-4.1 mini1.6 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGPT-4.1 nano0.4 USD / 1M output tokensLower is better for cost
Price per 1M output tokensGrok 4.32.5 USD / 1M output tokensLower is better for cost
Context windowGPT-6 Astra1,050,000 tokensHigher supports larger inputs
Context windowGemini 3.8 Flash1,048,576 tokensHigher supports larger inputs
Context windowMuse Spark 1.31,048,576 tokensHigher supports larger inputs
Context windowClaude Fable 5.11,000,000 tokensHigher supports larger inputs
Context windowGemini 3.7 Flash1,048,576 tokensHigher supports larger inputs
Context windowGrok 4.6500,000 tokensHigher supports larger inputs
Context windowQwen3.8-Max1,000,000 tokensHigher supports larger inputs
Context windowClaude Opus 51,000,000 tokensHigher supports larger inputs
Context windowGemini 3.5 Flash-Lite1,048,576 tokensHigher supports larger inputs
Context windowGemini 3.6 Flash1,048,576 tokensHigher supports larger inputs
Context windowGrok 4.5500,000 tokensHigher supports larger inputs
Context windowKimi K31,048,576 tokensHigher supports larger inputs
Context windowGPT-5.6 Luna1,050,000 tokensHigher supports larger inputs
Context windowGPT-5.6 Sol1,050,000 tokensHigher supports larger inputs
Context windowGPT-5.6 Terra1,050,000 tokensHigher supports larger inputs
Context windowClaude Sonnet 51,000,000 tokensHigher supports larger inputs
Context windowClaude Fable 51,000,000 tokensHigher supports larger inputs
Context windowClaude Opus 4.81,000,000 tokensHigher supports larger inputs
Context windowGemini 3.5 Flash1,048,576 tokensHigher supports larger inputs
Context windowGemini 3.1 Flash-Lite1,048,576 tokensHigher supports larger inputs
Context windowGPT-5.51,050,000 tokensHigher supports larger inputs
Context windowGrok 4.201,000,000 tokensHigher supports larger inputs
Context windowGPT-5.41,050,000 tokensHigher supports larger inputs
Context windowGemini 3.1 Pro Preview1,048,576 tokensHigher supports larger inputs
Context windowGemini 3 Flash Preview1,048,576 tokensHigher supports larger inputs
Context windowKimi K2 Base128,000 tokensHigher supports larger inputs
Context windowKimi K2 Instruct128,000 tokensHigher supports larger inputs
Context windowGPT-4.11,047,576 tokensHigher supports larger inputs
Context windowGPT-4.1 mini1,047,576 tokensHigher supports larger inputs
Context windowGPT-4.1 nano1,047,576 tokensHigher supports larger inputs
Context windowKimi-VL-A3B-Instruct128,000 tokensHigher supports larger inputs
Context windowGemini 2.5 Pro1,048,576 tokensHigher supports larger inputs
Context windowQwen2.5-VL 72B Instruct128,000 tokensHigher supports larger inputs
Context windowDeepSeek-R1128,000 tokensHigher supports larger inputs
Context windowGrok 4.31,000,000 tokensHigher supports larger inputs

Latest compatible benchmark · Measured 2026-06-25

Benchmark observations by release

LiveBench Overall, snapshot 2026-06-25, measured 2026-06-25. Scores are shown from newest release to oldest. Eligible models (5): Claude Sonnet 5, Claude Opus 4.8, Gemini 3.5 Flash, GPT-5.5, GPT-5.4.

Legend: bars use the fixed range 0100 points; longer is better. Exact values remain visible.

Source snapshot: 2026-06-25, measured 2026-06-25.

LiveBench Overall observations by releaseHorizontal bars show 5 models from newest release to oldest using the 0 to 100 points scale. Values are printed beside every bar.0100Claude Sonnet 5, 74.85 pointsClaude Sonnet 574.85Claude Opus 4.8, 78.93 pointsClaude Opus 4.878.93Gemini 3.5 Flash, 74.64 pointsGemini 3.5 Flash74.64GPT-5.5, 79.91 pointsGPT-5.579.91GPT-5.4, 77.97 pointsGPT-5.477.97
View exact benchmark observation data
Benchmark observation data
ModelScoreMeasuredSource/version
Claude Sonnet 574.85 points2026-06-25
Claude Opus 4.878.93 points2026-06-25
Gemini 3.5 Flash74.64 points2026-06-25
GPT-5.579.91 points2026-06-25
GPT-5.477.97 points2026-06-25

Sources/versions:

Missing scores are omitted.

Value frontier

Takeaway: Claude Sonnet 5, GPT-5.5, GPT-5.4 are not dominated on both published 1M input tokens price and this direct benchmark score among the 4 eligible models. Higher benchmark scores are better.

Frontier modelOther eligible model

Source snapshot: 2026-06-25, measured 2026-06-25. Price unit: 1M input tokens.

The score axis is zoomed to 7382 points; it does not start at 0.

1M input tokens price versus LiveBench Overall scoreScatter plot of 4 models priced per 1M input tokens. Circles identify nondominated frontier models and diamonds identify other eligible models. Higher benchmark scores are better. The score axis covers 73 to 82 points. Every plotted value is also listed in the table below.8280787573$0.00$2.50$5.00Published price (USD per 1M input tokens)LiveBench Overall (points)Claude Sonnet 5: $2 / 1M input tokens · standard, 74.85 pointsClaude Sonnet 5$2 / 1M input tokens · standard · 74.85 pointsClaude Opus 4.8: $5 / 1M input tokens · standard, 78.93 pointsGPT-5.5: $5 / 1M input tokens · standard, 79.91 pointsGPT-5.5$5 / 1M input tokens · standard · 79.91 pointsGPT-5.4: $2.5 / 1M input tokens · standard, 77.97 pointsGPT-5.4$2.5 / 1M input tokens · standard · 77.97 points
View exact value frontier data
Value frontier data
ModelPrice (1M input tokens)ScoreStatusMeasuredSource/version
Claude Sonnet 5$2 / 1M input tokens · standard74.85 pointsFrontier2026-06-25
Claude Opus 4.8$5 / 1M input tokens · standard78.93 pointsEligible2026-06-25
GPT-5.5$5 / 1M input tokens · standard79.91 pointsFrontier2026-06-25
GPT-5.4$2.5 / 1M input tokens · standard77.97 pointsFrontier2026-06-25

Sources/versions:

Prices are provider list prices verified 2026-09-09.

Side-by-side comparison

Compare 2–4 models using directly sourced capability, cost, and access facts.

Select models from the active task. A shared benchmark score is shown only when an audited snapshot covers every selected model.

How to use this comparison

  1. 01Filter or search for the models that fit your workload.
  2. 02Pick them in the selector above, or Select Compare on a row in the inventory.
  3. 03Read the matrix, source evidence, and trade-offs before choosing a default.

0 models selected for comparison

Pick 2 more models above to build the matrix.

Methodology: models with missing scores or prices are omitted from the applicable visual. Prices are provider list prices at the dataset verification date. Benchmark results should not be compared across dataset versions.

Van Data Team analysis

Recommended by use case

Editorial shortlists for evaluation, not automatic winners. Dataset 2026-09-09, last verified 2026-09-09.

AI agents and tool use

GPT-5.6 Sol, Claude Sonnet 5, Gemini 3.6 Flash

Start a current agent-workflow evaluation with GPT-5.6 Sol, Claude Sonnet 5, and Gemini 3.6 Flash.

Review rationale and evidence

Why

  • All three active API records are positioned by their providers for agentic or tool-using workflows.
  • GPT-5.6 Sol and Gemini 3.6 Flash publish context windows above one million tokens, while Claude Sonnet 5 launches with a lower introductory input price than GPT-5.6 Sol.

Watch for

  • The pinned LiveBench snapshot predates GPT-5.6 Sol and Gemini 3.6 Flash and does not measure tool-use success, latency, uptime, or end-to-end agent reliability; Claude Sonnet 5 pricing changes after 2026-08-31.

Fits when: Test the models with the actual tools, permissions, retry rules, and human approval steps used in production.

Evidence

Apply this shortlist

Coding and repository work

GPT-5.6 Sol, Claude Fable 5

Test GPT-5.6 Sol and Claude Fable 5 as current coding-focused API options before choosing a model for repository work.

Review rationale and evidence

Why

  • Both providers position these active records for complex coding or software-engineering workflows.
  • Their published million-token-class context windows support evaluation on large repositories and long-running work.

Watch for

  • The catalog does not include a compatible repository-level coding benchmark for either model, and the pinned LiveBench snapshot predates both models.

Fits when: Evaluate repository navigation, code edits, test execution, and review quality on your own languages and tooling.

Evidence

Apply this shortlist

Complex reasoning and knowledge work

GPT-5.5, Claude Fable 5

Compare GPT-5.5 and Claude Fable 5 for current high-capability analysis, then test both against your review standard.

Review rationale and evidence

Why

  • Both providers position these active API records for demanding reasoning or professional knowledge work.
  • Both records publish context windows of at least 1,000,000 tokens.

Watch for

  • The pinned LiveBench snapshot does not provide a compatible Claude Fable 5 result and does not establish factuality, domain expertise, citation quality, latency, or production reliability.

Fits when: Use representative documents, expected citations, and a defined human review rubric for the production trial.

Evidence

Apply this shortlist

Multimodal workflows

GPT-5.6 Terra, Gemini 3.6 Flash, Grok 4.5

Compare current API options from OpenAI, Google, and SpaceXAI when the workload combines text with visual inputs.

Review rationale and evidence

Why

  • Gemini 3.6 Flash records text, image, audio, and video inputs; GPT-5.6 Terra and Grok 4.5 record text and image inputs.
  • The three current records span published context windows from 500,000 to 1,050,000 tokens and different input-price points.

Watch for

  • The dataset has no compatible multimodal quality benchmark for this shortlist and does not record modality-specific pricing.

Fits when: Build an evaluation set from the image, audio, or video formats and document layouts the system will receive.

Evidence

Apply this shortlist

High-volume, cost-sensitive processing

Gemini 3.5 Flash-Lite, GPT-4.1 nano

Compare current Gemini 3.5 Flash-Lite with GPT-4.1 nano for cost-sensitive API trials using published list prices.

Review rationale and evidence

Why

  • Gemini 3.5 Flash-Lite records $0.30 input and $2.50 output per million tokens, while GPT-4.1 nano records $0.10 input and $0.40 output.
  • Both active records accept multimodal input and publish context windows above one million tokens.

Watch for

  • List price does not include every production cost, and this dataset has no compatible quality score for either record.

Fits when: Estimate full task cost with representative input size, output size, retries, caching, review, and supporting infrastructure.

Evidence

Apply this shortlist
Methods and provenance

Evidence library

Open only the methodology, production guidance, release history, sources, or answers you need. Complete records remain available in the page HTML.

How to read benchmark data

How to read the benchmark data

Use the catalog to screen candidates, then verify them on your workload. Linked official sources support each record as a whole; they are not a field-by-field audit trail. Benchmark results retain their publisher and compatible snapshot.

Keep task families separate

Scores and operational metrics should not be compared across task families. Text reasoning, image quality, video quality, and speech error metrics answer different questions.

Match one benchmark snapshot

Compare results only when the task family, benchmark version, measurement date, and eligible model set match. Different snapshots may use different prompts, evaluators, or datasets.

Follow the score direction

Higher-is-better metrics reward a larger result; lower-is-better metrics such as error rate reward a smaller result. The direction belongs to the named benchmark, not to every task.

Match the pricing unit

A price per million tokens, image, audio minute, video second, or realtime minute is a different cost basis. Compare only the same unit and tier, then add retries, review, tooling, and infrastructure.

Check lifecycle and source status

Current models are active provider offerings; Legacy models remain for historical context. Official records come from providers, while independent results come from benchmark publishers, each with a visible verification date.

Keep missing data missing

Not reported is not zero. The explorer excludes missing values from rankings and does not infer them from a related model.

Scope begins 2024-01-01. Dataset version 2026-09-09. Last verified 2026-09-09. 116 source records.

Production fit

Why a benchmark winner may not fit production

Benchmark rank and production fit measure different things. A result can support a shortlist without deciding the system design.

  • A public score may cover one narrow task while your system combines retrieval, tools, permissions, and review.
  • Production acceptance also depends on latency, reliability, safety, data handling, regional availability, and total cost.
  • Van Data Team recommendations are transparent editorial shortlists. They use published fields and results, apply no hidden weighting, and must be tested on representative work.
Release timeline

Model release timeline since 2024

Generated from exact explorer release dates. Models without a verified exact date remain in the inventory and are omitted here rather than assigned an inferred date.

  1. 2026 Q3

    • GPT-6 Astra2026-09-03
    • Gemini 3.8 Flash2026-09-02
    • Muse Spark 1.32026-09-02
    • Claude Fable 5.12026-09-01
    • Claude Mythos 5.12026-09-01
    • Gemini 3.7 Flash2026-08-13
    • Grok 4.62026-08-12
    • Qwen3.8-Max2026-08-03
    • Claude Opus 52026-07-24
    • Gemini 3.5 Flash-Lite2026-07-21
    • Gemini 3.6 Flash2026-07-21
    • Grok 4.52026-07-16
    • Kimi K32026-07-16
    • GPT-5.6 Luna2026-07-09
    • GPT-5.6 Sol2026-07-09
    • GPT-5.6 Terra2026-07-09
  2. 2026 Q2

    • Claude Sonnet 52026-06-30
    • Gemini 3.1 Flash Lite Image2026-06-30
    • Gemini Omni Flash Preview2026-06-30
    • Claude Fable 52026-06-09
    • Claude Opus 4.82026-05-28
    • Gemini 3 Pro Image2026-05-28
    • Gemini 3.1 Flash Image2026-05-28
    • Gemini 3.5 Flash2026-05-19
    • Gemini 3.1 Flash-Lite2026-05-07
    • GPT-5.52026-04-23
    • Gemini 3.1 Flash TTS Preview2026-04-15
  3. 2026 Q1

    • Gemini 3.1 Flash Live Preview2026-03-26
    • Grok 4.202026-03-10
    • GPT-5.42026-03-05
    • Gemini 3.1 Pro Preview2026-02-19
    • Scribe v22026-01-09
  4. 2025 Q4

    • Gemini 3 Flash Preview2025-12-17
    • Scribe v2 Realtime2025-11-11
    • Veo 3.1 Preview2025-10-15
  5. 2025 Q3

    • Kimi K2 Base2025-07-11
    • Kimi K2 Instruct2025-07-11
  6. 2025 Q2

    • Kimi-Dev-72B2025-06-17
    • GPT-4.12025-04-14
    • GPT-4.1 mini2025-04-14
    • GPT-4.1 nano2025-04-14
    • Kimi-VL-A3B-Instruct2025-04-09
  7. 2025 Q1

    • Gemini 2.5 Pro2025-03-25
    • Claude 3.7 Sonnet2025-02-24
    • Grok 32025-02-19
    • Qwen2.5-VL 72B Instruct2025-01-26
    • DeepSeek-R12025-01-20
  8. 2024 Q4

    • DeepSeek-V32024-12-26
    • Gemini 2.0 Flash2024-12-11
    • Ministral 8B2024-10-16
  9. 2024 Q3

    • Llama 3.2 11B Vision2024-09-25
    • Qwen2.5 32B Instruct2024-09-19
    • Pixtral 12B2024-09-17
    • Grok 22024-08-13
    • Mistral Large 22024-07-24
    • Llama 3.1 405B2024-07-23
  10. 2024 Q2

    • DeepSeek-Coder-V2 Instruct2024-06-17
    • Qwen2 72B Instruct2024-06-07
    • Qwen2 7B Instruct2024-06-07
    • Codestral 22B2024-05-29
    • Gemini 1.5 Flash2024-05-14
    • GPT-4o2024-05-13
    • DeepSeek-V2 Chat2024-05-06
    • Llama 3 70B2024-04-18
    • Llama 3 8B2024-04-18
    • Grok-1.5V2024-04-12
  11. 2024 Q1

    • Grok-1.52024-03-28
    • Claude 3 Haiku2024-03-13
    • Claude 3 Opus2024-03-04
    • Claude 3 Sonnet2024-03-04
    • Gemini 1.5 Pro2024-02-15
Source catalog · 116 records

Source catalog

Official sources come from model providers. Independent sources come from a separate benchmark publisher. Each row shows its publisher, verification date, version, license status, and coverage; missing metadata is not filled or inferred.

Official provider sources

OpenAI
  • GPT-5.4 model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model documentation verifies the dated snapshot, token limits, modalities, and exact standard token prices represented here.

  • Introducing GPT-5.5
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official release announcement verifies model identity and the 2026-04-23 release date.

  • GPT-5.5 model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model documentation verifies token limits, modalities, and exact standard token prices represented here.

  • GPT-5.6: Frontier intelligence that scales with your ambition
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official release announcement verifies the Sol, Terra, and Luna identities and 2026-07-09 general-availability date.

  • OpenAI GPT-5.6 model catalog
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model catalog verifies the Sol, Terra, and Luna token limits, modalities, roles, and exact standard token prices represented here.

  • Hello GPT-4o
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Introducing GPT-4.1 in the API
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • OpenAI API pricing
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    pricing

    Official pricing page used only for the exact token prices represented in model records.

  • GPT-4o model
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official provider documentation used to verify exact model specifications.

  • GPT-4.1 model
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official provider documentation used to verify exact model specifications.

  • GPT-4.1 mini model
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official provider documentation used to verify exact model specifications.

  • GPT-4.1 nano model
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official provider documentation used to verify exact model specifications.

  • GPT Image 2 model
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • GPT-Realtime-1.5 model
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • gpt-audio-1.5 model
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • GPT-4o Transcribe model
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • TTS-1 model
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • GPT-6 Astra model reference
    Verified
    2026-09-09
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • OpenAI API pricing
    Verified
    2026-09-09
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

Anthropic
  • Introducing Claude Opus 4.8
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official release documentation verifies model identity, release date, the published context window, and standard input/output pricing.

  • Introducing Claude Sonnet 5
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official release documentation verifies model identity, release date, API availability, and introductory pricing through 2026-08-31.

  • Claude Fable 5 and Claude Mythos 5
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official release announcement verifies Claude Fable 5 identity, 2026-06-09 release date, API access, standard token pricing, and general availability after redeployment.

  • Claude models overview
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official current-model documentation verifies Fable's text-and-image inputs, 1,000,000-token context window, 128,000-token output limit, and current API availability.

  • Introducing the next generation of Claude
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official provider documentation used to verify model identity and release details.

  • Claude 3.7 Sonnet and Claude Code
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Claude pricing
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    pricing

    Official pricing documentation used only for the exact token prices represented in model records.

  • Claude 3 Haiku
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify exact model specifications.

  • Claude models overview
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify exact model specifications.

  • Claude API pricing
    Verified
    2026-09-09
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • Claude models overview
    Verified
    2026-09-09
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • Claude Mythos limited availability
    Verified
    2026-09-09
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

Google
  • Gemini API release notes
    Verified
    2026-08-18
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Official release notes verify the dated Gemini 3.5 Flash, Gemini 3.1 Flash Live, Gemini 3.1 Flash TTS, Gemini 3.1 Flash Lite Image, Gemini 3 Pro Image, and Gemini Omni Flash releases represented in this catalog.

  • Gemini 3.5 Flash model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Official model documentation verifies model code, modalities, and token limits; exact token prices remain unreported here.

  • Gemini 3.6 Flash model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Official model documentation verifies model code, 2026-07-21 update, modalities, and token limits.

  • Using the latest Gemini models
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity, pricing

    Official latest-model guidance verifies the exact standard input/output prices represented for Gemini 3.6 Flash.

  • Gemini 3.5 Flash-Lite model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Official model documentation verifies the stable model code, multimodal inputs, token limits, capabilities, and July 2026 update.

  • Gemini 3.5 Flash-Lite pricing
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    pricing

    Official pricing table verifies the standard paid-tier input, cached-input, and output rates represented here.

  • Our next-generation model: Gemini 1.5
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Gemini 1.5 Flash
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Introducing Gemini 2.0
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Gemini 2.5: Our most intelligent AI model
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Gemini Developer API pricing
    Verified
    2026-09-09
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    pricing

    Official Gemini Developer API pricing table used only for exact published token, image, audio-minute, and video-second prices represented in model records.

  • Gemini API models
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify exact model specifications.

  • Gemini 2.5 Pro
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify exact model specifications.

  • Gemini 3.1 Flash Image pricing
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party pricing table used to verify Gemini 3.1 Flash Image standard token and resolution-specific image prices.

  • Gemini native image generation capabilities
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    First-party native image guide used to verify Gemini 3.1 Flash Image generation, editing, image input, and mixed text/image output capabilities.

  • Gemini Developer API pricing and native media model identifiers
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • Gemini 3.8 Flash model card
    Verified
    2026-09-09
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • Gemini 3.7 Flash model
    Verified
    2026-08-18
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Official model documentation verifies the stable Gemini 3.7 Flash code, text/image/video/audio inputs, text output, token limits, structured outputs, thinking, and URL context.

  • Introducing Gemini 3.7 Flash
    Verified
    2026-08-18
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official Google announcement verifies the 2026-08-13 release, coding and agent positioning, and introductory token prices through 2026-12-31 before the documented 2027 rate change.

  • Gemini 3.1 Pro Preview
    Verified
    2026-08-18
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Model card verifies the Gemini 3.1 Pro Preview endpoint, multimodal inputs, text output, token limits, and documented capabilities.

  • Gemini 3 Flash Preview
    Verified
    2026-08-18
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Model card verifies the Gemini 3 Flash Preview endpoint, multimodal inputs, text output, token limits, and documented capabilities.

  • Gemini 3.1 Flash-Lite
    Verified
    2026-08-18
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Model card verifies the stable Gemini 3.1 Flash-Lite endpoint, multimodal inputs, text output, token limits, and documented capabilities.

  • Gemini 3.5 Live Translate
    Verified
    2026-08-18
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Model card verifies the Gemini 3.5 Live Translate preview endpoint, audio input, translated-audio and transcript output, and token limits.

  • Gemini Live API
    Verified
    2026-08-18
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Live API guide verifies realtime audio model operation and the preview endpoint lifecycle applies to the live-model records.

  • Gemini 3.1 Flash Live Preview
    Verified
    2026-08-18
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Model card verifies the Gemini 3.1 Flash Live preview endpoint, supported realtime multimodal inputs, text and audio output, token limits, and capabilities.

  • Gemini 3.1 Flash TTS Preview
    Verified
    2026-08-18
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Model card verifies the Gemini 3.1 Flash TTS preview endpoint, text input, audio output, token limits, and audio generation capability.

  • Text-to-speech generation
    Verified
    2026-08-18
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Text-to-Speech guide verifies controllable speech generation and streaming support for the Gemini 3.1 Flash TTS preview endpoint.

  • Gemini 3.1 Flash Lite Image
    Verified
    2026-08-18
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Model card verifies the stable Gemini 3.1 Flash Lite Image endpoint, text and image inputs, text and image output, token limits, and image generation and editing capabilities.

  • Gemini 3 Pro Image
    Verified
    2026-08-18
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Model card verifies the stable Gemini 3 Pro Image endpoint, text and image inputs and outputs, token limits, and image generation capability.

  • Gemini Omni Flash
    Verified
    2026-08-18
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Model card verifies the Gemini Omni Flash preview endpoint, text, image, and video inputs, video output, context window, and 3-to-10-second output range.

  • Generate and edit videos with Gemini Omni Flash
    Verified
    2026-08-18
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Gemini Omni Flash guide verifies preview lifecycle, conversational video editing, native multimodal context, and generated video with audio.

  • Gemini deprecations
    Verified
    2026-08-18
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Lifecycle table verifies published stable or preview status and announced shutdown dates only for the Gemini endpoints that reference it.

SpaceXAI
  • Grok 4.3 model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model documentation verifies Grok 4.3 modalities, context window, reasoning modes, aliases, and standard short-context prices.

  • xAI inference models API reference
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official API reference exposes Grok 4.3 modalities, aliases, and token pricing fields; its object creation timestamp is not treated as a public release date.

  • Skills in web, iOS, and Android
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official 2026-05-18 announcement confirms Grok 4.3 was live across grok.com, iOS, and Android by that date; it is not treated as the model's exact release date.

  • SpaceXAI API release notes
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official release notes verify Grok 4.20 and Grok 4.20 Multi-agent API availability on 2026-03-10.

  • Grok 4.20 System Card
    Verified
    2026-08-07
    Dataset version
    2026-04-07
    License
    Not stated
    Covers
    identity

    Official system card verifies Grok 4.20 identity, single-agent and multi-agent deployment modes, supported input modalities, and intended uses.

  • SpaceXAI model pricing
    Verified
    2026-08-14
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official current pricing table verifies active Grok 4.6, Grok 4.5, and Grok 4.20 variants, their documented context windows, and standard and long-context token rates.

  • Grok model retirement on May 15, 2026
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official lifecycle notice verifies Grok 3 retirement and redirect to Grok 4.3, and identifies Grok 4.3 as the recommended general replacement.

  • Introducing Grok 4.5
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official announcement verifies the model identity, release date, and intended coding, agentic, and knowledge-work scope.

  • Grok 4.5 model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model documentation verifies modalities, context window, and standard short-context token prices; the model caution discloses higher long-context rates.

  • Introducing Grok 4.6
    Verified
    2026-08-14
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official announcement verifies the 2026-08-12 release date and positioning for long-running agents, coding, and knowledge work.

  • Grok 4.6 model
    Verified
    2026-08-14
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model documentation verifies the exact model ID, text and image input, text output, 500,000-token context window, reasoning, function calling, structured outputs, and standard token prices.

  • grok-imagine-image model
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • Video generation
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • Imagine API pricing
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

Meta
  • Introducing Meta Llama 3
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Introducing Llama 3.1
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Llama 3.2
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Meta Model API model list
    Verified
    2026-09-09
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

xAI
  • Announcing Grok-1.5
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Grok-1.5 Vision Preview
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Grok-2 Beta Release
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Grok 3 Beta
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • xAI models and pricing
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model catalog used to verify documented model availability and pricing fields, not historical release dates.

DeepSeek
  • DeepSeek-V2
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • DeepSeek-Coder-V2
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • DeepSeek-V3
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • DeepSeek-R1
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • DeepSeek API pricing
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    pricing

    Official pricing documentation consulted for price provenance; unavailable exact values remain unreported.

Mistral AI
  • Codestral
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Large Enough
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Announcing Pixtral 12B
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Un Ministral, des Ministraux
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Mistral AI pricing
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    pricing

    Official pricing page consulted for price provenance; unavailable exact values remain unreported.

Qwen Team
  • Hello Qwen2
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Qwen2.5: A Party of Foundation Models
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Qwen2.5-VL
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

Moonshot AI
  • Kimi-VL
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Kimi-Dev
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Kimi K2
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Kimi platform model and pricing documentation
    Verified
    2026-09-09
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

Runway
  • Available AI Models
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

ElevenLabs
  • ElevenLabs models
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • Introducing Scribe v2
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • Introducing Scribe v2 Realtime
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

Stability AI
  • Stability AI Developer Platform pricing
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

Black Forest Labs
  • FLUX image generation endpoints
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

  • FLUX API pricing
    Verified
    2026-08-10
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

Alibaba Cloud
  • Alibaba Cloud Model Studio model list
    Verified
    2026-09-09
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    First-party provider documentation used only for the stated model identity, capabilities, lifecycle, access, and pricing.

Independent benchmark sources

LiveBench
  • LiveBench category-weighted averaging implementation
    Verified
    2026-08-07
    Dataset version
    LiveBench-2026-06-25@19e766a5de4de07d672ed5bf9f0a69ceed1d39bf
    License
    Apache-2.0
    Covers
    benchmark

    Pinned implementation averages task scores within each of the seven categories, then averages the category scores; the paired categories_2026_06_25.json file defines category membership.

  • LiveBench 2026-06-25 task results
    Verified
    2026-08-07
    Dataset version
    LiveBench-2026-06-25@19e766a5de4de07d672ed5bf9f0a69ceed1d39bf
    License
    Apache-2.0
    Covers
    benchmark

    Snapshot and measurement date: 2026-06-25. Overall scores are imported only for unambiguous model mappings and are recomputed with the pinned category-weighted averaging implementation.

  • LiveBench datasheet and methodology
    Verified
    2026-08-06
    Dataset version
    LiveBench-2024-11-25
    License
    Apache-2.0
    Covers
    benchmark

    Primary benchmark documentation confirms objective scoring, public distribution, and Apache 2.0 reuse terms.

  • LiveBench 2024-11-25 task results
    Verified
    2026-08-06
    Dataset version
    LiveBench-2024-11-25@347e3d0b6a4916cd66f8bda31ce7ecdd7436ecc3
    License
    Apache-2.0
    Covers
    benchmark

    Snapshot and measurement date: 2024-11-25. Exact published snapshot; only direct Web of Lies V2 percentages with unambiguous model mappings are imported.

OpenRouter
  • OpenRouter models API catalogue
    Verified
    2026-09-09
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Third-party API catalogue used only to cross-check published token prices, context limits, and listing dates against first-party documentation.

FAQ

AI model benchmark FAQ

What does an AI model benchmark score mean?

A benchmark score reports performance on one defined evaluation. Compare scores only when the task family, benchmark version, measurement date, scoring method, and eligible model set match; the result is a screening signal, not a prediction for every workload.

How fresh are the model and source records?

Each source shows a verification date, and the page shows the dataset version and last verified date. A verification date confirms when Van Data Team checked the linked record; it does not promise that a provider page has stayed unchanged since then.

Why are some model values marked Not reported?

Not reported means the catalog does not have a compatible, directly sourced value for that exact model record. Missing values remain empty rather than being estimated, copied from a related model, or treated as zero.

How should I compare AI model pricing?

Compare only the same pricing unit and tier, such as a million tokens, image, audio minute, video second, character, or realtime minute. Prices are provider list prices checked on the source verification date, not live quotes; retries, review, tooling, and infrastructure also affect production cost.

Does open-weight access mean a model is free to run?

No. Open weights means the model parameters are available for deployment under the provider license. This can provide infrastructure control, but compute, serving, monitoring, security, and engineering still carry costs.

How should I choose models for a production evaluation?

Start with the use case, select a small evidence-backed shortlist, and test it on representative work. Set acceptance thresholds for quality, latency, cost, reliability, safety, and human review before comparing results.

How often is this AI model comparison updated?

Van Data Team updates the versioned dataset when a review verifies material model, price, or benchmark changes. The published last verified date is the update record; this is a curated comparison, not a live provider feed.

How can Van Data Team help with model evaluation?

Van Data Team can turn a shortlist into a production evaluation with representative cases, measurable acceptance rules, model routing, workflow controls, observability, and cost tracking tailored to your system.

Production evaluation

Turn a model shortlist into a production decision

We design representative evaluations and the workflow controls needed to operate the selected model with clear quality, cost, and review boundaries.

  • Evaluation cases and acceptance thresholds tied to real work
  • Model routing, tool permissions, observability, and human review design