Compare 50 models released since 2024, including current GPT-5.6, Claude 5, Gemini 3.6, and Grok 4.5 families, using source-verified access, context, pricing, and compatible evaluation evidence.
Takeaway: GPT-5.4, GPT-5.5, Claude Sonnet 5 are not dominated on both published input price and this direct benchmark score among the 4 eligible models.
Frontier modelOther eligible model
The score axis is zoomed to 73–82 points so the models separate; it does not start at 0. Only frontier models are labelled; hover any mark, or read the table below, for the rest.
Prices are provider list prices verified 2026-08-07.
Side-by-side comparison
Compare 2–4 models using directly sourced capability, cost, and access facts.
Choose any two to four models without leaving this section.
How to use this comparison
01Filter or search for the models that fit your workload.
02Pick them in the selector above, or Select Compare on a row in the inventory.
03Read the matrix, source evidence, and trade-offs before choosing a default.
0 models selected for comparison
Pick 2 more models above to build the matrix.
Methodology: models with missing scores or prices are omitted from the applicable visual. Prices are provider list prices at the dataset verification date. Benchmark results should not be compared across dataset versions.
Van Data Team analysis
Recommended by use case
Editorial shortlists for evaluation, not automatic winners. Dataset 2026-08-07, last verified 2026-08-07.
AI agents and tool use
GPT-5.6 Sol, Claude Sonnet 5, Gemini 3.6 Flash
Start a current agent-workflow evaluation with GPT-5.6 Sol, Claude Sonnet 5, and Gemini 3.6 Flash.
Review rationale and evidence
Why
All three active API records are positioned by their providers for agentic or tool-using workflows.
GPT-5.6 Sol and Gemini 3.6 Flash publish context windows above one million tokens, while Claude Sonnet 5 launches with a lower introductory input price than GPT-5.6 Sol.
Watch for
The pinned LiveBench snapshot predates GPT-5.6 Sol and Gemini 3.6 Flash and does not measure tool-use success, latency, uptime, or end-to-end agent reliability; Claude Sonnet 5 pricing changes after 2026-08-31.
Fits when: Test the models with the actual tools, permissions, retry rules, and human approval steps used in production.
Compare GPT-5.5 and Claude Fable 5 for current high-capability analysis, then test both against your review standard.
Review rationale and evidence
Why
Both providers position these active API records for demanding reasoning or professional knowledge work.
Both records publish context windows of at least 1,000,000 tokens.
Watch for
The pinned LiveBench snapshot does not provide a compatible Claude Fable 5 result and does not establish factuality, domain expertise, citation quality, latency, or production reliability.
Fits when: Use representative documents, expected citations, and a defined human review rubric for the production trial.
Open only the methodology, production guidance, release history, sources, or answers you need. Complete records remain available in the page HTML.
How to read benchmark data
How to read the benchmark data
Use the catalog to screen candidates, then verify them on your workload. Every displayed specification, price, and result resolves to a source record.
Read one metric at a time
A benchmark is a repeatable test of a defined capability. Compare results only inside the same benchmark definition and version.
Check the evidence label
Official means the provider published the record. Independent means a separate benchmark publisher reported it. Both can be useful, but they answer different questions.
Keep missing data missing
Not reported is not zero. The explorer excludes missing values from rankings and does not infer them from a related model.
Model the full cost
Published token prices are inputs to a cost estimate. Retries, caching, review, tooling, and self-hosted infrastructure also affect production cost.
Scope begins 2024-01-01. Dataset version 2026-08-07. Last verified 2026-08-07. 71 source records.
Production fit
Why a benchmark winner may not fit production
Benchmark rank and production fit measure different things. A result can support a shortlist without deciding the system design.
A public score may cover one narrow task while your system combines retrieval, tools, permissions, and review.
Production acceptance also depends on latency, reliability, safety, data handling, regional availability, and total cost.
Van Data Team recommendations are transparent editorial shortlists. They use published fields and results, apply no hidden weighting, and must be tested on representative work.
Release timeline
Model release timeline since 2024
Generated from exact explorer release dates. Models without a verified exact date remain in the inventory and are omitted here rather than assigned an inferred date.
2024 Q1
Gemini 1.5 Pro2024-02-15
Claude 3 Opus2024-03-04
Claude 3 Sonnet2024-03-04
Claude 3 Haiku2024-03-13
Grok-1.52024-03-28
2024 Q2
Grok-1.5V2024-04-12
Llama 3 8B2024-04-18
Llama 3 70B2024-04-18
DeepSeek-V2 Chat2024-05-06
GPT-4o2024-05-13
Gemini 1.5 Flash2024-05-14
Codestral 22B2024-05-29
Qwen2 7B Instruct2024-06-07
Qwen2 72B Instruct2024-06-07
DeepSeek-Coder-V2 Instruct2024-06-17
2024 Q3
Llama 3.1 405B2024-07-23
Mistral Large 22024-07-24
Grok 22024-08-13
Pixtral 12B2024-09-17
Qwen2.5 32B Instruct2024-09-19
Llama 3.2 11B Vision2024-09-25
2024 Q4
Ministral 8B2024-10-16
Gemini 2.0 Flash2024-12-11
DeepSeek-V32024-12-26
2025 Q1
DeepSeek-R12025-01-20
Qwen2.5-VL 72B Instruct2025-01-26
Grok 32025-02-19
Claude 3.7 Sonnet2025-02-24
Gemini 2.5 Pro2025-03-25
2025 Q2
Kimi-VL-A3B-Instruct2025-04-09
GPT-4.12025-04-14
GPT-4.1 mini2025-04-14
GPT-4.1 nano2025-04-14
Kimi-Dev-72B2025-06-17
2025 Q3
Kimi K2 Base2025-07-11
Kimi K2 Instruct2025-07-11
2026 Q1
GPT-5.42026-03-05
Grok 4.202026-03-10
2026 Q2
GPT-5.52026-04-23
Gemini 3.5 Flash2026-05-19
Claude Opus 4.82026-05-28
Claude Fable 52026-06-09
Claude Sonnet 52026-06-30
2026 Q3
GPT-5.6 Sol2026-07-09
GPT-5.6 Terra2026-07-09
GPT-5.6 Luna2026-07-09
Grok 4.52026-07-16
Gemini 3.6 Flash2026-07-21
Gemini 3.5 Flash-Lite2026-07-21
Source catalog
Source catalog
Official sources come from model providers. Independent sources come from a separate benchmark publisher. Missing metadata is not filled or inferred.
Official release announcement verifies Claude Fable 5 identity, 2026-06-09 release date, API access, standard token pricing, and general availability after redeployment.
Official current-model documentation verifies Fable's text-and-image inputs, 1,000,000-token context window, 128,000-token output limit, and current API availability.
Official API reference exposes Grok 4.3 modalities, aliases, and token pricing fields; its object creation timestamp is not treated as a public release date.
Official 2026-05-18 announcement confirms Grok 4.3 was live across grok.com, iOS, and Android by that date; it is not treated as the model's exact release date.
Official model documentation verifies modalities, context window, and standard short-context token prices; the model caution discloses higher long-context rates.
Pinned implementation averages task scores within each of the seven categories, then averages the category scores; the paired categories_2026_06_25.json file defines category membership.
Snapshot and measurement date: 2026-06-25. Overall scores are imported only for unambiguous model mappings and are recomputed with the pinned category-weighted averaging implementation.
Snapshot and measurement date: 2024-11-25. Exact published snapshot; only direct Web of Lies V2 percentages with unambiguous model mappings are imported.
FAQ
AI model benchmark FAQ
What does an AI model benchmark score mean?
A benchmark score reports performance on one defined evaluation. Compare scores only when the task, version, scoring method, and test conditions match; the result is a screening signal, not a prediction for every workload.
How fresh are the model and source records?
Each source shows a verification date, and the page shows the dataset version and last verified date. A verification date confirms when Van Data Team checked the linked record; it does not promise that a provider page has stayed unchanged since then.
Why are some model values marked Not reported?
Not reported means the catalog does not have a compatible, directly sourced value for that exact model record. Missing values remain empty rather than being estimated, copied from a related model, or treated as zero.
How should I compare AI model pricing?
Compare the same currency, token unit, and pricing basis, then model the full task cost. Input, cached input, output, retries, review time, and self-hosted infrastructure can change the production total.
Does open-weight access mean a model is free to run?
No. Open weights means the model parameters are available for deployment under the provider license. This can provide infrastructure control, but compute, serving, monitoring, security, and engineering still carry costs.
How should I choose models for a production evaluation?
Start with the use case, select a small evidence-backed shortlist, and test it on representative work. Set acceptance thresholds for quality, latency, cost, reliability, safety, and human review before comparing results.
How often is this AI model comparison updated?
Van Data Team updates the versioned dataset when a review verifies material model, price, or benchmark changes. The published last verified date is the update record; this is a curated comparison, not a live provider feed.
How can Van Data Team help with model evaluation?
Van Data Team can turn a shortlist into a production evaluation with representative cases, measurable acceptance rules, model routing, workflow controls, observability, and cost tracking tailored to your system.
Production evaluation
Turn a model shortlist into a production decision
We design representative evaluations and the workflow controls needed to operate the selected model with clear quality, cost, and review boundaries.
Evaluation cases and acceptance thresholds tied to real work
Model routing, tool permissions, observability, and human review design