Benchmarks

Benchmarks & methodology.

How ChatGPT-6 lines up against the current frontier — the full table, the caveats behind every number, and exactly how we put it together.

Last updated: September 2026

Most comparison tables give you a verdict without the receipts. This one shows the numbers and then tells you where they come from, what they can't tell you, and why a free preview changes the maths.

ChatGPT-6 is an early-access preview, so treat its figures as directional rather than final — the honest way to read this page is "frontier-class quality at a price no shipping model can match."

The full table

Nine dimensions, five models, one column that's free.

Metric
ChatGPT-6
GPT-5.6
Opus 5
Gemini 3.7
Grok 4.6
Intelligence Index71*67686665
Agentic coding86*84827876
Context window400K400K1M2M500K
Max output128K128K64K64K64K
Throughput (tps)58*705592120
Price / 1M out$0$10$25$10$6
Multimodal
Open weights
Free to use
Intelligence Index

Composite score across reasoning, knowledge and math benchmarks (higher is better).

Agentic coding

Indicative agentic-coding score — real-world, multi-file tasks with tools (higher is better).

Context window

Maximum tokens the model can attend to in a single conversation.

Max output

Largest response the model will generate in one call.

Throughput (tps)

Median output speed in tokens per second (higher is faster).

Price / 1M out

List price per million output tokens (lower is cheaper).

Multimodal

Accepts image input alongside text.

Open weights

Weights can be downloaded and self-hosted.

Free to use

Usable at no cost right now.

Methodology

Intelligence Index is a composite of public reasoning, knowledge and math benchmarks in the spirit of the Artificial Analysis Intelligence Index. Agentic coding reflects real-world, multi-file tasks executed with tools rather than isolated snippets.

Context window, max output, throughput and price are provider-reported figures. Price is the list rate per one million output tokens; input tokens are usually cheaper and are omitted to keep the table readable.

ChatGPT-6 is in early-access preview, so its figures (marked *) are community-reported estimates rather than a vendor datasheet. They will move as the model is tuned — we refresh this page as better data appears.

Numbers change fast in this space. Use this as a directional guide for picking a model, not as a benchmark leaderboard of record. When in doubt, run your own eval on your own prompts.

See it for yourself

Numbers are one thing. Hand ChatGPT-6 a real task and judge the output.

Open the chat