Skip to main content
Verified model research

AI model benchmarks for practical decisions

Compare 50 models released since 2024, including current GPT-5.6, Claude 5, Gemini 3.6, and Grok 4.5 families, using source-verified access, context, pricing, and compatible evaluation evidence.

50 models9 providersDataset 2026-08-07Verified 2026-08-07

Model explorer

Filter the complete validated inventory by provider, access, use case, price, or one named benchmark. Missing values remain visible as Not reported.

More filters and sorting
Providers

110 of 50 models

0 of 4 selected

AI model inventory with pricing, selected benchmark, provenance, and comparison controls
Claude 3.7 SonnetLegacyProviderAnthropicRelease2025-02-24AccessAPIContext200,000Input price$3Output price$15LiveBench Web of Lies V294 % correct
View details
Strengths
  • Claude 3.7 Sonnet records API access for text and image inputs and a published 200,000-token context window
Cautions
  • This record is marked legacy; confirm current availability and migration requirements before adoption
Provenance
Claude 3 HaikuLegacyProviderAnthropicRelease2024-03-13AccessAPIContext200,000Input price$0.25Output price$1.25LiveBench Web of Lies V2Not reported
View details
Strengths
  • Claude 3 Haiku records API access for text and image inputs and a published 200,000-token context window
Cautions
  • This record is marked legacy; confirm current availability and migration requirements before adoption
Provenance
Claude 3 OpusLegacyProviderAnthropicRelease2024-03-04AccessAPIContext200,000Input price$15Output price$75LiveBench Web of Lies V250 % correct
View details
Strengths
  • Claude 3 Opus records API access for text and image inputs and a published 200,000-token context window
Cautions
  • This record is marked legacy; confirm current availability and migration requirements before adoption
Provenance
Claude 3 SonnetLegacyProviderAnthropicRelease2024-03-04AccessAPIContext200,000Input price$3Output price$15LiveBench Web of Lies V2Not reported
View details
Strengths
  • Claude 3 Sonnet records API access for text and image inputs and a published 200,000-token context window
Cautions
  • This record is marked legacy; confirm current availability and migration requirements before adoption
Provenance
Claude Fable 5CurrentProviderAnthropicRelease2026-06-09AccessAPIContext1,000,000Input price$10Output price$50LiveBench Web of Lies V2Not reported
View details
Strengths
  • Claude Fable 5 is Anthropic's highest-capability widely released model for demanding reasoning and long-horizon agentic work
Cautions
  • Fable requires 30-day data retention, and classifier-triggered requests fall back to Claude Opus 4.8
Provenance
Claude Opus 4.8CurrentProviderAnthropicRelease2026-05-28AccessAPIContext1,000,000Input price$5Output price$25LiveBench Web of Lies V2Not reported
View details
Strengths
  • Claude Opus 4.8 is positioned for coding, agentic tasks, and professional work with a published 1,000,000-token context window
Cautions
  • Maximum output length is not represented because the approved source set does not establish one unambiguously
Provenance
Claude Sonnet 5CurrentProviderAnthropicRelease2026-06-30AccessAPIContextNot reportedInput price$2Output price$10LiveBench Web of Lies V2Not reported
View details
Strengths
  • Claude Sonnet 5 is positioned for agentic execution, tool use, coding, and knowledge work
Cautions
  • The launch price is introductory through 2026-08-31; refresh pricing before evaluating workloads after that date
Provenance
Codestral 22BLegacyProviderMistral AIRelease2024-05-29AccessAPI and open weightsContext32,000Input priceNot reportedOutput priceNot reportedLiveBench Web of Lies V2Not reported
View details
Strengths
  • Codestral 22B records both API and open-weight access for text inputs and a published 32,000-token context window
Cautions
  • This record is marked legacy; confirm current availability and migration requirements before adoption
Provenance
DeepSeek-Coder-V2 InstructLegacyProviderDeepSeekRelease2024-06-17AccessAPI and open weightsContext128,000Input priceNot reportedOutput priceNot reportedLiveBench Web of Lies V2Not reported
View details
Strengths
  • DeepSeek-Coder-V2 Instruct records both API and open-weight access for text inputs and a published 128,000-token context window
Cautions
  • This record is marked legacy; confirm current availability and migration requirements before adoption
Provenance
DeepSeek-R1CurrentProviderDeepSeekRelease2025-01-20AccessAPI and open weightsContext128,000Input priceNot reportedOutput priceNot reportedLiveBench Web of Lies V2100 % correct
View details
Strengths
  • DeepSeek-R1 records both API and open-weight access for text inputs and a published 128,000-token context window
Cautions
  • Hosted token price and infrastructure cost are not represented; measure deployment and operating cost directly
Provenance

Latest compatible benchmark · Measured 2026-06-25

Benchmark rankings

Takeaway: GPT-5.5 ranks first at 79.91 points among 5 models with a compatible score.

Legend: bars use the fixed 0100 points scale; longer is better. Exact values remain visible.

LiveBench Overall rankingHorizontal bars rank 5 models using the 0 to 100 points scale. Values are printed beside every bar.0100Rank 1: GPT-5.5, 79.91 points1. GPT-5.579.91Rank 2: Claude Opus 4.8, 78.93 points2. Claude Opus 4.878.93Rank 3: GPT-5.4, 77.97 points3. GPT-5.477.97Rank 4: Claude Sonnet 5, 74.85 points4. Claude Sonnet 574.85Rank 5: Gemini 3.5 Flash, 74.64 points5. Gemini 3.5 Flash74.64
View exact ranking data
Benchmark ranking data
RankModelScoreMeasuredSource/version
1GPT-5.579.91 points2026-06-25
2Claude Opus 4.878.93 points2026-06-25
3GPT-5.477.97 points2026-06-25
4Claude Sonnet 574.85 points2026-06-25
5Gemini 3.5 Flash74.64 points2026-06-25

Sources/versions:

Missing scores are omitted.

Value frontier

Takeaway: GPT-5.4, GPT-5.5, Claude Sonnet 5 are not dominated on both published input price and this direct benchmark score among the 4 eligible models.

Frontier modelOther eligible model

The score axis is zoomed to 7382 points so the models separate; it does not start at 0. Only frontier models are labelled; hover any mark, or read the table below, for the rest.

Input price versus LiveBench Overall scoreScatter plot of 4 models. Circles identify nondominated frontier models and diamonds identify other eligible models. No trend or sequence line is shown. The score axis covers 73 to 82 points rather than the full benchmark range. Every plotted value is also listed in the table below the chart.8280787573$0.00$2.50$5.00Published input price (USD per 1M tokens)LiveBench Overall (points)GPT-5.4: $2.50, 77.97 pointsGPT-5.4$2.50 · 77.97 pointsGPT-5.5: $5.00, 79.91 pointsGPT-5.5$5.00 · 79.91 pointsClaude Opus 4.8: $5.00, 78.93 pointsClaude Sonnet 5: $2.00, 74.85 pointsClaude Sonnet 5$2.00 · 74.85 points
View exact value frontier data
Value frontier data
ModelInput priceScoreStatusMeasuredSource/version
GPT-5.4$2.50 / 1M77.97 pointsFrontier2026-06-25
GPT-5.5$5.00 / 1M79.91 pointsFrontier2026-06-25
Claude Opus 4.8$5.00 / 1M78.93 pointsEligible2026-06-25
Claude Sonnet 5$2.00 / 1M74.85 pointsFrontier2026-06-25

Sources/versions:

Prices are provider list prices verified 2026-08-07.

Side-by-side comparison

Compare 2–4 models using directly sourced capability, cost, and access facts.

Choose any two to four models without leaving this section.

How to use this comparison

  1. 01Filter or search for the models that fit your workload.
  2. 02Pick them in the selector above, or Select Compare on a row in the inventory.
  3. 03Read the matrix, source evidence, and trade-offs before choosing a default.

0 models selected for comparison

Pick 2 more models above to build the matrix.

Methodology: models with missing scores or prices are omitted from the applicable visual. Prices are provider list prices at the dataset verification date. Benchmark results should not be compared across dataset versions.

Van Data Team analysis

Recommended by use case

Editorial shortlists for evaluation, not automatic winners. Dataset 2026-08-07, last verified 2026-08-07.

AI agents and tool use

GPT-5.6 Sol, Claude Sonnet 5, Gemini 3.6 Flash

Start a current agent-workflow evaluation with GPT-5.6 Sol, Claude Sonnet 5, and Gemini 3.6 Flash.

Review rationale and evidence

Why

  • All three active API records are positioned by their providers for agentic or tool-using workflows.
  • GPT-5.6 Sol and Gemini 3.6 Flash publish context windows above one million tokens, while Claude Sonnet 5 launches with a lower introductory input price than GPT-5.6 Sol.

Watch for

  • The pinned LiveBench snapshot predates GPT-5.6 Sol and Gemini 3.6 Flash and does not measure tool-use success, latency, uptime, or end-to-end agent reliability; Claude Sonnet 5 pricing changes after 2026-08-31.

Fits when: Test the models with the actual tools, permissions, retry rules, and human approval steps used in production.

Evidence

Apply this shortlist

Coding and repository work

GPT-5.6 Sol, Claude Fable 5

Test GPT-5.6 Sol and Claude Fable 5 as current coding-focused API options before choosing a model for repository work.

Review rationale and evidence

Why

  • Both providers position these active records for complex coding or software-engineering workflows.
  • Their published million-token-class context windows support evaluation on large repositories and long-running work.

Watch for

  • The catalog does not include a compatible repository-level coding benchmark for either model, and the pinned LiveBench snapshot predates both models.

Fits when: Evaluate repository navigation, code edits, test execution, and review quality on your own languages and tooling.

Evidence

Apply this shortlist

Complex reasoning and knowledge work

GPT-5.5, Claude Fable 5

Compare GPT-5.5 and Claude Fable 5 for current high-capability analysis, then test both against your review standard.

Review rationale and evidence

Why

  • Both providers position these active API records for demanding reasoning or professional knowledge work.
  • Both records publish context windows of at least 1,000,000 tokens.

Watch for

  • The pinned LiveBench snapshot does not provide a compatible Claude Fable 5 result and does not establish factuality, domain expertise, citation quality, latency, or production reliability.

Fits when: Use representative documents, expected citations, and a defined human review rubric for the production trial.

Evidence

Apply this shortlist

Multimodal workflows

GPT-5.6 Terra, Gemini 3.6 Flash, Grok 4.5

Compare current API options from OpenAI, Google, and SpaceXAI when the workload combines text with visual inputs.

Review rationale and evidence

Why

  • Gemini 3.6 Flash records text, image, audio, and video inputs; GPT-5.6 Terra and Grok 4.5 record text and image inputs.
  • The three current records span published context windows from 500,000 to 1,050,000 tokens and different input-price points.

Watch for

  • The dataset has no compatible multimodal quality benchmark for this shortlist and does not record modality-specific pricing.

Fits when: Build an evaluation set from the image, audio, or video formats and document layouts the system will receive.

Evidence

Apply this shortlist

High-volume, cost-sensitive processing

Gemini 3.5 Flash-Lite, GPT-4.1 nano

Compare current Gemini 3.5 Flash-Lite with GPT-4.1 nano for cost-sensitive API trials using published list prices.

Review rationale and evidence

Why

  • Gemini 3.5 Flash-Lite records $0.30 input and $2.50 output per million tokens, while GPT-4.1 nano records $0.10 input and $0.40 output.
  • Both active records accept multimodal input and publish context windows above one million tokens.

Watch for

  • List price does not include every production cost, and this dataset has no compatible quality score for either record.

Fits when: Estimate full task cost with representative input size, output size, retries, caching, review, and supporting infrastructure.

Evidence

Apply this shortlist
Methods and provenance

Evidence library

Open only the methodology, production guidance, release history, sources, or answers you need. Complete records remain available in the page HTML.

How to read benchmark data

How to read the benchmark data

Use the catalog to screen candidates, then verify them on your workload. Every displayed specification, price, and result resolves to a source record.

Read one metric at a time

A benchmark is a repeatable test of a defined capability. Compare results only inside the same benchmark definition and version.

Check the evidence label

Official means the provider published the record. Independent means a separate benchmark publisher reported it. Both can be useful, but they answer different questions.

Keep missing data missing

Not reported is not zero. The explorer excludes missing values from rankings and does not infer them from a related model.

Model the full cost

Published token prices are inputs to a cost estimate. Retries, caching, review, tooling, and self-hosted infrastructure also affect production cost.

Scope begins 2024-01-01. Dataset version 2026-08-07. Last verified 2026-08-07. 71 source records.

Production fit

Why a benchmark winner may not fit production

Benchmark rank and production fit measure different things. A result can support a shortlist without deciding the system design.

  • A public score may cover one narrow task while your system combines retrieval, tools, permissions, and review.
  • Production acceptance also depends on latency, reliability, safety, data handling, regional availability, and total cost.
  • Van Data Team recommendations are transparent editorial shortlists. They use published fields and results, apply no hidden weighting, and must be tested on representative work.
Release timeline

Model release timeline since 2024

Generated from exact explorer release dates. Models without a verified exact date remain in the inventory and are omitted here rather than assigned an inferred date.

  1. 2024 Q1

    • Gemini 1.5 Pro2024-02-15
    • Claude 3 Opus2024-03-04
    • Claude 3 Sonnet2024-03-04
    • Claude 3 Haiku2024-03-13
    • Grok-1.52024-03-28
  2. 2024 Q2

    • Grok-1.5V2024-04-12
    • Llama 3 8B2024-04-18
    • Llama 3 70B2024-04-18
    • DeepSeek-V2 Chat2024-05-06
    • GPT-4o2024-05-13
    • Gemini 1.5 Flash2024-05-14
    • Codestral 22B2024-05-29
    • Qwen2 7B Instruct2024-06-07
    • Qwen2 72B Instruct2024-06-07
    • DeepSeek-Coder-V2 Instruct2024-06-17
  3. 2024 Q3

    • Llama 3.1 405B2024-07-23
    • Mistral Large 22024-07-24
    • Grok 22024-08-13
    • Pixtral 12B2024-09-17
    • Qwen2.5 32B Instruct2024-09-19
    • Llama 3.2 11B Vision2024-09-25
  4. 2024 Q4

    • Ministral 8B2024-10-16
    • Gemini 2.0 Flash2024-12-11
    • DeepSeek-V32024-12-26
  5. 2025 Q1

    • DeepSeek-R12025-01-20
    • Qwen2.5-VL 72B Instruct2025-01-26
    • Grok 32025-02-19
    • Claude 3.7 Sonnet2025-02-24
    • Gemini 2.5 Pro2025-03-25
  6. 2025 Q2

    • Kimi-VL-A3B-Instruct2025-04-09
    • GPT-4.12025-04-14
    • GPT-4.1 mini2025-04-14
    • GPT-4.1 nano2025-04-14
    • Kimi-Dev-72B2025-06-17
  7. 2025 Q3

    • Kimi K2 Base2025-07-11
    • Kimi K2 Instruct2025-07-11
  8. 2026 Q1

    • GPT-5.42026-03-05
    • Grok 4.202026-03-10
  9. 2026 Q2

    • GPT-5.52026-04-23
    • Gemini 3.5 Flash2026-05-19
    • Claude Opus 4.82026-05-28
    • Claude Fable 52026-06-09
    • Claude Sonnet 52026-06-30
  10. 2026 Q3

    • GPT-5.6 Sol2026-07-09
    • GPT-5.6 Terra2026-07-09
    • GPT-5.6 Luna2026-07-09
    • Grok 4.52026-07-16
    • Gemini 3.6 Flash2026-07-21
    • Gemini 3.5 Flash-Lite2026-07-21
Source catalog

Source catalog

Official sources come from model providers. Independent sources come from a separate benchmark publisher. Missing metadata is not filled or inferred.

Official provider sources

OpenAI
  • GPT-5.4 model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model documentation verifies the dated snapshot, token limits, modalities, and exact standard token prices represented here.

  • Introducing GPT-5.5
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official release announcement verifies model identity and the 2026-04-23 release date.

  • GPT-5.5 model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model documentation verifies token limits, modalities, and exact standard token prices represented here.

  • GPT-5.6: Frontier intelligence that scales with your ambition
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official release announcement verifies the Sol, Terra, and Luna identities and 2026-07-09 general-availability date.

  • OpenAI GPT-5.6 model catalog
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model catalog verifies the Sol, Terra, and Luna token limits, modalities, roles, and exact standard token prices represented here.

  • Hello GPT-4o
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Introducing GPT-4.1 in the API
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • OpenAI API pricing
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    pricing

    Official pricing page used only for the exact token prices represented in model records.

  • GPT-4o model
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official provider documentation used to verify exact model specifications.

  • GPT-4.1 model
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official provider documentation used to verify exact model specifications.

  • GPT-4.1 mini model
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official provider documentation used to verify exact model specifications.

  • GPT-4.1 nano model
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official provider documentation used to verify exact model specifications.

Anthropic
  • Introducing Claude Opus 4.8
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official release documentation verifies model identity, release date, the published context window, and standard input/output pricing.

  • Introducing Claude Sonnet 5
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official release documentation verifies model identity, release date, API availability, and introductory pricing through 2026-08-31.

  • Claude Fable 5 and Claude Mythos 5
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official release announcement verifies Claude Fable 5 identity, 2026-06-09 release date, API access, standard token pricing, and general availability after redeployment.

  • Claude models overview
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official current-model documentation verifies Fable's text-and-image inputs, 1,000,000-token context window, 128,000-token output limit, and current API availability.

  • Introducing the next generation of Claude
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official provider documentation used to verify model identity and release details.

  • Claude 3.7 Sonnet and Claude Code
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Claude pricing
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    pricing

    Official pricing documentation used only for the exact token prices represented in model records.

  • Claude 3 Haiku
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify exact model specifications.

  • Claude models overview
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify exact model specifications.

Google
  • Gemini API release notes
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Official release notes and deprecation schedule verify the 2026-05-19 Gemini 3.5 Flash release date and current availability.

  • Gemini 3.5 Flash model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Official model documentation verifies model code, modalities, and token limits; exact token prices remain unreported here.

  • Gemini 3.6 Flash model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Official model documentation verifies model code, 2026-07-21 update, modalities, and token limits.

  • Using the latest Gemini models
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity, pricing

    Official latest-model guidance verifies the exact standard input/output prices represented for Gemini 3.6 Flash.

  • Gemini 3.5 Flash-Lite model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    identity

    Official model documentation verifies the stable model code, multimodal inputs, token limits, capabilities, and July 2026 update.

  • Gemini 3.5 Flash-Lite pricing
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    CC-BY-4.0
    Covers
    pricing

    Official pricing table verifies the standard paid-tier input, cached-input, and output rates represented here.

  • Our next-generation model: Gemini 1.5
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Gemini 1.5 Flash
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Introducing Gemini 2.0
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Gemini 2.5: Our most intelligent AI model
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Gemini Developer API pricing
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    pricing

    Official pricing documentation consulted for price provenance; tiered values are not collapsed into model records.

  • Gemini API models
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify exact model specifications.

  • Gemini 2.5 Pro
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify exact model specifications.

SpaceXAI
  • Grok 4.3 model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model documentation verifies Grok 4.3 modalities, context window, reasoning modes, aliases, and standard short-context prices.

  • xAI inference models API reference
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official API reference exposes Grok 4.3 modalities, aliases, and token pricing fields; its object creation timestamp is not treated as a public release date.

  • Skills in web, iOS, and Android
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official 2026-05-18 announcement confirms Grok 4.3 was live across grok.com, iOS, and Android by that date; it is not treated as the model's exact release date.

  • SpaceXAI API release notes
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official release notes verify Grok 4.20 and Grok 4.20 Multi-agent API availability on 2026-03-10.

  • Grok 4.20 System Card
    Verified
    2026-08-07
    Dataset version
    2026-04-07
    License
    Not stated
    Covers
    identity

    Official system card verifies Grok 4.20 identity, single-agent and multi-agent deployment modes, supported input modalities, and intended uses.

  • SpaceXAI model pricing
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official current pricing table verifies active Grok 4.20 variants, their 1,000,000-token context window, and standard and long-context token rates.

  • Grok model retirement on May 15, 2026
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official lifecycle notice verifies Grok 3 retirement and redirect to Grok 4.3, and identifies Grok 4.3 as the recommended general replacement.

  • Introducing Grok 4.5
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official announcement verifies the model identity, release date, and intended coding, agentic, and knowledge-work scope.

  • Grok 4.5 model
    Verified
    2026-08-07
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model documentation verifies modalities, context window, and standard short-context token prices; the model caution discloses higher long-context rates.

Meta
  • Introducing Meta Llama 3
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Introducing Llama 3.1
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Llama 3.2
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

xAI
  • Announcing Grok-1.5
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Grok-1.5 Vision Preview
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Grok-2 Beta Release
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Grok 3 Beta
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • xAI models and pricing
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity, pricing

    Official model catalog used to verify documented model availability and pricing fields, not historical release dates.

DeepSeek
  • DeepSeek-V2
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • DeepSeek-Coder-V2
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • DeepSeek-V3
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • DeepSeek-R1
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • DeepSeek API pricing
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    pricing

    Official pricing documentation consulted for price provenance; unavailable exact values remain unreported.

Mistral AI
  • Codestral
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Large Enough
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Announcing Pixtral 12B
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Un Ministral, des Ministraux
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Mistral AI pricing
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    pricing

    Official pricing page consulted for price provenance; unavailable exact values remain unreported.

Qwen Team
  • Hello Qwen2
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Qwen2.5: A Party of Foundation Models
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Qwen2.5-VL
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

Moonshot AI
  • Kimi-VL
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Kimi-Dev
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

  • Kimi K2
    Verified
    2026-08-06
    Dataset version
    Not versioned
    License
    Not stated
    Covers
    identity

    Official provider documentation used to verify model identity and release details.

Independent benchmark sources

LiveBench
  • LiveBench category-weighted averaging implementation
    Verified
    2026-08-07
    Dataset version
    LiveBench-2026-06-25@19e766a5de4de07d672ed5bf9f0a69ceed1d39bf
    License
    Apache-2.0
    Covers
    benchmark

    Pinned implementation averages task scores within each of the seven categories, then averages the category scores; the paired categories_2026_06_25.json file defines category membership.

  • LiveBench 2026-06-25 task results
    Verified
    2026-08-07
    Dataset version
    LiveBench-2026-06-25@19e766a5de4de07d672ed5bf9f0a69ceed1d39bf
    License
    Apache-2.0
    Covers
    benchmark

    Snapshot and measurement date: 2026-06-25. Overall scores are imported only for unambiguous model mappings and are recomputed with the pinned category-weighted averaging implementation.

  • LiveBench datasheet and methodology
    Verified
    2026-08-06
    Dataset version
    LiveBench-2024-11-25
    License
    Apache-2.0
    Covers
    benchmark

    Primary benchmark documentation confirms objective scoring, public distribution, and Apache 2.0 reuse terms.

  • LiveBench 2024-11-25 task results
    Verified
    2026-08-06
    Dataset version
    LiveBench-2024-11-25@347e3d0b6a4916cd66f8bda31ce7ecdd7436ecc3
    License
    Apache-2.0
    Covers
    benchmark

    Snapshot and measurement date: 2024-11-25. Exact published snapshot; only direct Web of Lies V2 percentages with unambiguous model mappings are imported.

FAQ

AI model benchmark FAQ

What does an AI model benchmark score mean?

A benchmark score reports performance on one defined evaluation. Compare scores only when the task, version, scoring method, and test conditions match; the result is a screening signal, not a prediction for every workload.

How fresh are the model and source records?

Each source shows a verification date, and the page shows the dataset version and last verified date. A verification date confirms when Van Data Team checked the linked record; it does not promise that a provider page has stayed unchanged since then.

Why are some model values marked Not reported?

Not reported means the catalog does not have a compatible, directly sourced value for that exact model record. Missing values remain empty rather than being estimated, copied from a related model, or treated as zero.

How should I compare AI model pricing?

Compare the same currency, token unit, and pricing basis, then model the full task cost. Input, cached input, output, retries, review time, and self-hosted infrastructure can change the production total.

Does open-weight access mean a model is free to run?

No. Open weights means the model parameters are available for deployment under the provider license. This can provide infrastructure control, but compute, serving, monitoring, security, and engineering still carry costs.

How should I choose models for a production evaluation?

Start with the use case, select a small evidence-backed shortlist, and test it on representative work. Set acceptance thresholds for quality, latency, cost, reliability, safety, and human review before comparing results.

How often is this AI model comparison updated?

Van Data Team updates the versioned dataset when a review verifies material model, price, or benchmark changes. The published last verified date is the update record; this is a curated comparison, not a live provider feed.

How can Van Data Team help with model evaluation?

Van Data Team can turn a shortlist into a production evaluation with representative cases, measurable acceptance rules, model routing, workflow controls, observability, and cost tracking tailored to your system.

Production evaluation

Turn a model shortlist into a production decision

We design representative evaluations and the workflow controls needed to operate the selected model with clear quality, cost, and review boundaries.

  • Evaluation cases and acceptance thresholds tied to real work
  • Model routing, tool permissions, observability, and human review design