Kingy AI published a detailed comparison built from the two vendors' launch materials and Artificial Analysis data. The core numbers can be checked against first-party documentation from Anthropic and OpenAI. Some of the independent figures need a caveat, and the details follow.
Same rate card, different bills
On standard API pricing the two models match. Claude Sonnet 5.5 and GPT-6 Sol are priced identically — $2 per million input tokens, $10 per million output tokens — and they arrived six days apart in late September 2026. Anthropic's launch page confirms Sonnet 5.5 costs the same as Sonnet 5, including $0.20 per million tokens for cache reads. Anthropic also claims it usually needs far fewer tokens than its predecessor, which cuts per-task cost by up to 30%. That claim compares Sonnet 5.5 with Sonnet 5, not with Sol.
OpenAI's developer changelog lists Sol's standard rates of $2 input, $0.20 cached input and $10 output. Those rates apply only to prompts with up to 272K input tokens. That limit creates the biggest structural difference between the two. Kingy AI notes that once a Sol request carries more than 272,000 input tokens, the entire request reprices. The higher tier doubles input and cache rates and raises output to $15 per million.
Kingy AI's worked example is a 500,000-token document review with 10,000 tokens of output:
| Model | Input cost | Output cost | Total |
|---|---|---|---|
| Claude Sonnet 5.5 | 500K × $2 = $1.00 | 10K × $10 = $0.10 | $1.10 |
| GPT-6 Sol | 500K × $4 = $2.00 | 10K × $15 = $0.15 | $2.15 |
There is a caveat on the Sonnet side. Anthropic's general pricing docs say Claude models from 4.6 onward bill the whole context window at the standard rate, and Kingy AI is applying that rule to Sonnet 5.5, but its own pricing row had not been published yet. Anthropic's launch page also doesn't give Sonnet 5.5's context window, maximum output or knowledge cutoff. Check Anthropic's model documentation before you build a 500K-token pipeline on it.
For short, cached agent turns, such as 50K cached tokens, 10K fresh input and 4K output, both models cost about seven cents. The main difference is how many tokens each one spends to finish a task. As APIMaster put it, the matching headline token rates do not mean matching bills. A longer prompt, cache writes, reasoning tokens, tool calls and retries all change the cost of a completed task.
Section summary: Both models share the same base rates. Sol's price jump above 272K input tokens is documented. Sonnet 5.5's long-context pricing is inferred from Anthropic's general policy and hasn't been confirmed for this model yet.
Specs: what's documented and what isn't
OpenAI's GPT-6 Sol model page lists:
- Context window: 1,050,000 tokens
- Max output: 128,000 tokens
- Knowledge cutoff: April 20, 2026
- Modalities: text and image in, text out
- Reasoning effort: none, low, medium (default), high, xhigh, max
- Other notes: built-in tools need the Responses API, because Chat Completions supports function calling only when reasoning effort is set to none. EU data residency works only with Standard processing.
APIMaster adds that Sol's published 1.05M context includes a 922K input limit, so don't treat the full 1.05M figure as prompt space.
Sonnet 5.5 is thinner on the spec sheet. Anthropic gives the model ID (claude-sonnet-5-5) and default effort levels: Medium in the Claude apps and Claude Code, High on the Claude Platform. Kingy AI's 1M-token context figure comes from an Artificial Analysis listing, not from Anthropic. Sonnet 5, which Sonnet 5.5 replaces, had a 1M-token window and 128K max output. There's a small cautionary tale here. Before launch, a viral "leak" claimed Sonnet 5.5 was secretly beating Sol, and Startup Fortune found it turns out to be built on Sonnet 5's existing public pricing, not a real leak. Until Anthropic publishes Sonnet 5.5's own specs, don't assume they match Sonnet 5's.
The independent scorecard, with a caveat
Kingy AI's comparison rests on Artificial Analysis's Intelligence Index v4.3.2, a ten-evaluation suite that also records cost per task at list prices. Kingy reports:
| Effort | Sonnet 5.5 score | Sonnet 5.5 $/task | GPT-6 Sol score | GPT-6 Sol $/task |
|---|---|---|---|---|
| Low | 36 | $0.41 | 34 | $0.13 |
| Medium | 41 | $0.59 | 40 | $0.25 |
| High | 47 | $1.08 | 43 | $0.37 |
| Xhigh | 52 | $2.74 | 44 | $0.53 |
| Max | 56 | $7.60 | 48 | $1.06 |
Two points before you quote these. The figures are as Kingy AI reported them, and I couldn't check them directly against Artificial Analysis's own release pages. They also don't fully agree with other coverage. OrcaRouter, citing the same index revision, put it at 56 for Claude Sonnet 5.5 against 47 for GPT-6 Sol, one point below Kingy's 48 for Sol. The gap is small and could reflect a same-day update, but launch-day leaderboards change.
The overall picture holds either way:
- Sonnet 5.5 has the higher ceiling. It leads at every matched effort label, and the lead grows at higher effort.
- Sol is much cheaper per task. By Kingy's figures it costs about 2.4× less at medium and 7.2× less at max. The rate cards are identical, so the difference comes from the number of tokens each model uses.
- At the same spend, they're close. At about $1 per task, Sonnet at High (47) and Sol at Max (48) are roughly tied. Below that level Sol gives more for the money. Above it, Sonnet keeps improving and Sol has no higher setting.
Effort labels aren't calibrated across vendors, so "high" on one model isn't the same amount of work as "high" on the other. Compare by cost per task, not by label.
Artificial Analysis also lists Sonnet 5.5 runs as "Default Fallback." That's Anthropic's new cyber safeguard in action: higher-risk security requests are visibly handed off to Sonnet 5. The index therefore measures the model as customers will actually get it.
Section summary: Sonnet wins on ceiling and Sol wins on price per task. Their cost curves cross near $1 per task, but treat the exact numbers as early.
Head-to-head: the benchmarks both vendors report
Anthropic's launch table includes a GPT-6 Sol column, which is unusual. Anthropic's own charts put Sonnet 5.5 ahead on some measures, including GDPval-AA and AA-Briefcase, while Sol edges out Sonnet 5.5 on FrontierCode's coding benchmark.
| Benchmark | Sonnet 5.5 | GPT-6 Sol |
|---|---|---|
| FrontierCode 1.1 (Main) | 46.2% (Max) | 49.3% |
| GDPval-AA v2.1 (Elo) | 1844 | 1487 |
| AA-Briefcase v1.1 (Elo) | 1811 | 1483 |
| Chartography (no tools) | 61.6% | 53.6% |
The FrontierCode result needs a closer look. Sonnet 5.5 scored lower at Max than at Xhigh. Anthropic says that at Max it ran Claude Code's multi-agent code-review skill more often. In two cases Cognition examined, that caused a timeout or edits outside the task's scope, and FrontierCode penalizes those. Anthropic's chart text also makes a claim that Kingy's summary doesn't mention: at High effort, the default on the Claude Platform, Sonnet 5.5 matches GPT-6 Sol's best score for about a fifth of the cost per task. That points the opposite way from the Artificial Analysis cost data, where Sonnet costs more per task at every matched setting. The two measurements use different benchmarks and harnesses, so neither disproves the other. If you care about mergeable pull requests, test on your own repositories.
Both vendors disclosed bugs that affect these numbers:
- Sonnet 5.5: Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment with a structured-outputs bug. Anthropic says any effect should be small and would understate Sonnet's scores, and the bug has been fixed.
- GPT-6 Sol: OpenAI's September 25 changelog records a fix for an image-encoding bug that degraded image understanding in Sol and Luna. OpenAI recommends rerunning evaluations that use image inputs. Artificial Analysis doesn't expect big changes to the Elo scores, and Anthropic says the bug didn't affect its Chartography testing.
Even with those caveats, gaps of more than 300 Elo on two separate knowledge-work suites are large.
Coding, computer use and knowledge work
Each vendor highlights benchmarks the other doesn't report:
- Terminal-Bench 4.0: Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, a test of agentic command-line work, up from Sonnet 5's 10.3% and actually ahead of Opus 5.5's reported 66.4%. OpenAI hasn't published a Sol score for it.
- CursorBench 4.0: Sonnet 5.5 scores 55.5%, compared with 57.8% for Opus 5.5. No Sol figure is published.
- DeepSWE v1.1: OpenAI reports 68.8% for Sol at max. No Sonnet 5.5 figure is published.
- OSWorld: Sonnet 5.5 scores 80.1% on OSWorld 2.1 (partial reward), and OpenAI reports Sol at 60.5% on OSWorld 2.0 offline. These are different versions run in different harnesses, so don't compare them directly.
- Knowledge work: Anthropic reports 64.5% on Humanity's Last Exam with tools. OpenAI points to AutomationBench and Agents' Last Exam for Sol. Neither vendor reports the other's picks.
Anthropic also says that Opus 5.5 is still clearly stronger at open-ended work that needs sustained judgment, even where Sonnet 5.5 matches it on benchmarks.
Migration notes for Windows developers
If you're moving an existing Sonnet integration to 5.5, check three things:
- Thinking-off configurations change. If you run Sonnet with thinking disabled, switch to the new
between_toolssetting before migrating. Anthropic says this setting keeps up-front thinking off. - Preserved thinking ties reasoning to the account that created it. Moving conversations between accounts now behaves differently, including switching accounts mid-session in Claude Code.
- Security work may be downgraded. Routine bug finding and fixing isn't affected, but higher-risk cybersecurity tasks visibly fall back to Sonnet 5. Anthropic says defenders will soon be able to apply to an expanded Cyber Verification Program. The model also ships with new classifiers meant to block distillation attacks, where large numbers of fake accounts are used to extract a model's capabilities into an unsafeguarded copy.
On the OpenAI side, Sol doesn't support fine-tuning. Built-in tools need the Responses API. In ChatGPT, Sol is available in ChatGPT Work and Codex for paid plans (Plus, Pro, Business, Enterprise and Edu).
Where Azure fits
Anthropic says Sonnet 5.5 is available on all platforms, including Microsoft Azure, alongside AWS and Google Cloud. Organizations that already buy through Azure can try Sonnet 5.5 without adding a new vendor relationship. Kingy AI's table also lists Azure as a route to GPT-6 Sol, but OpenAI's materials here don't confirm that. Check the Azure AI model catalog in your region before you rely on it.
Which should you choose?
Based on the evidence so far:
Sonnet 5.5 fits best when:
- you need the highest ceiling in this price tier
- the work is long-horizon: documents, spreadsheets, slide decks, analysis
- you run terminal-driven coding agents
- prompts regularly exceed 272K tokens, where Sol's pricing jump applies
GPT-6 Sol fits best when:
- you run high volumes and cost per task matters most
- you want a cheap floor (about $0.13 per task at low effort, by Kingy's figures)
- you're already built on OpenAI's Responses API or Codex
For many teams the practical answer is both. Send routine, high-volume work to Sol at low or medium effort, and send the hard, long-running jobs to Sonnet 5.5 at high effort or above. Before committing, run your own representative tasks and measure completed tasks per dollar, latency and the full bill.
Both models are only days old. Both vendors have disclosed deployment bugs affecting third-party scores, independent sources already disagree by a point on the headline index, and Anthropic hasn't published Sonnet 5.5's context and output limits. Treat these numbers as a starting point and expect them to change.
References
- Claude Sonnet 5.5 vs GPT-6 Sol: Benchmarks, Specs, Evals and Pricing Compared - Kingy AI Kingy AI · 2026-09-28T18:58:15+00:00
- GPT-6 Sol Model | OpenAI API developers.openai.com
- Claude Sonnet 5.5 vs GPT-6 Sol: Which $2 Model Wins? orcarouter.ai