Anthropic’s announcement establishes the release, its published pricing and several changes affecting Claude Code users. It also supplies enough detail to distinguish three frequently conflated propositions: lower token prices, lower spending on a particular workload, and better results at a chosen reasoning setting. Those distinctions should drive an evaluation more than the dramatic language surrounding the launch.
Claude Opus 5.5 Moves Beyond the Claude Code Leak
The initial reporting concerned a model identifier. Wccftech reported that claude-opus-5-5 had appeared in the forthcoming Claude Code v2.1.280, citing a social-media post. By the time the publication updated its report, Anthropic had announced Opus 5.5 as the first member of its Claude 5.5 family. The vendor’s dated announcement is the basis for treating this as a product release, rather than interpreting an identifier as proof of availability.
The distinction also resolves apparently contradictory coverage from the preceding day. A September 21 report from Up North AI described Opus 5.5 as unreleased, with discussion centering on a rumored claude-wafer-eap testing codename. That report predates Anthropic’s September 22 announcement; it describes the rumor stage, not the status after launch. There is no reason to keep framing the model as merely imminent now that the company has introduced it.
Anthropic positions Opus 5.5 as performing at roughly the level of Claude Fable 5.1 on most work, while improving efficiency over Opus 5. Its newsroom dates the Fable 5.1 announcement to September 1, placing the releases three weeks apart. The company says Sonnet 5.5 and Haiku 5.5 will follow in the coming weeks. Those are planned follow-on releases, not models that should be counted as available alongside Opus 5.5.
For developers, the release has a direct connection to Microsoft’s tooling. Anthropic’s announcement includes a testimonial from GitHub chief product officer Mario Rodriguez describing tests across GitHub Copilot CLI and Visual Studio Code. Rodriguez says Opus 5.5 used among the fewest tokens and steps GitHub measured, and solved more terminal tasks in VS Code than Opus 5 while taking less than half the steps. This is a partner evaluation published by Anthropic; it establishes that GitHub tested the model, without establishing that every Copilot subscriber can select it at launch.
That availability boundary deserves attention. Anthropic explicitly says that Opus 5.5 Fast mode is available in Claude Code and the Claude Platform, and announces subscription-limit changes for Pro, Max and Team. The material available here does not establish a complete region-by-region or integration-by-integration rollout, nor a Windows-specific installation requirement. A Claude release, a Claude Code feature and a GitHub evaluation are related developments, but each has its own access boundary.
Claude Code’s Reported Interface Changes Need Their Own Confirmation
The early Wccftech report also described a one-million-token context window, a new “Responsive” mode that emits a sentence before reasoning or using tools, and an output style called “Clear, conversational English, with no Claude-isms.” Those details came with the pre-release Claude Code reporting. The inspected Anthropic announcement supports a broader claim of clearer communication, but does not substantiate those exact feature names or provide their activation steps.
Consequently, the reported context size should not become a capacity guarantee for every integration, and the named modes should not become instructions to look for a universally available toggle. There is useful separation here between the model’s behavior and a client’s interface. Anthropic says testers found Opus 5.5’s writing clearer, easier to follow and more likely to put important information first. That is the supported communication change, attributed to the company and its testers.
Clearer output has a practical role in lengthy coding sessions. Anthropic argues that work which is easier to follow is also easier to check. Readers can use that claim as an evaluation criterion: whether a model explains a consequential edit accurately and concisely matters more than whether it avoids a particular stock phrase. A stylistic improvement becomes operationally valuable when it reduces the effort needed to understand and review the work.
Opus 5.5’s Lower Prices Make Workload Accounting More Important
Anthropic’s “40% less” headline describes its estimate of typical workload costs at default settings. It is not the percentage reduction applied to every billing category. The published schedule cuts ordinary input and output prices by 20%, cache-write prices by 20%, and cache-read prices by 60%. Anthropic says those changes, together with fewer tokens used per task, produce the approximately 40% workload saving.
The distinction is visible in the price table. All amounts below are Anthropic’s published prices per one million tokens, the units used to meter model input and output.
| Billing category | Claude Opus 5 | Claude Opus 5.5 | Published unit-price reduction |
|---|---|---|---|
| Input tokens | $5.00 | $4.00 | 20% |
| Output tokens | $25.00 | $20.00 | 20% |
| Cache writes | $6.25 | $5.00 | 20% |
| Cache reads | $0.50 | $0.20 | 60% |
The unusually large cache-read reduction is central to Anthropic’s argument. The company says cache reads account for the majority of agentic and coding-work costs in the workloads it describes. An agentic task is one in which the model takes successive actions—such as making edits or invoking tools—to work toward an outcome. The economics therefore depend on the whole session’s billed activity, including repeated use of context, rather than on the apparent length of the final answer.
A simple calculation shows why different users should expect different savings. At the listed prices, one million ordinary input tokens plus one million output tokens would cost $30 with Opus 5 and $24 with Opus 5.5, before other categories are added. That is exactly a 20% reduction. A million cache-read tokens falls from $0.50 to $0.20, producing the larger 60% reduction in that category. These are illustrations of the rate card, not measurements of a representative coding job.
The 40% saving is a workload estimate, not a universal discount. A session dominated by ordinary input and output, with unchanged token consumption, follows the 20% rate reduction. A session with extensive cache reads has a different billing mix. If Opus 5.5 also finishes the task using fewer tokens, as Anthropic reports in several evaluations, the improvement can exceed what the ordinary input and output rates alone would suggest.
This is why a finance or platform team should compare complete tasks at a specified effort setting. A model that makes fewer attempts, takes fewer steps or requires less rework can change the total even when its individual rates look only moderately better. Conversely, a more demanding reasoning configuration can make a model’s impressive benchmark result a poor guide to the bill from a default production workflow.
Fast Mode Exchanges a Higher Rate for Lower Waiting Time
Anthropic says Opus 5.5 generates output more than 30% faster than Opus 5. Separately, it offers Fast mode in Claude Code and the Claude Platform, advertising up to 2.5 times the speed at $8 per million input tokens and $40 per million output tokens. Those input and output rates are twice Opus 5.5’s standard rates. The announcement does not turn the maximum speed claim into a guaranteed reduction in total task duration.
The rate comparison makes the decision concrete. For the same illustrative one million input tokens and one million output tokens, the listed Fast-mode rates produce a $48 charge, compared with $24 at standard Opus 5.5 rates. That calculation excludes any other billing categories and assumes identical token quantities. It describes the price of those tokens, not the eventual expense of a completed job, which can also depend on how the model proceeds.
Fast mode is therefore a separate purchasing choice from adopting Opus 5.5. A developer waiting interactively for work may value reduced delay differently from a team running a job without continuous supervision. The sensible comparison is the extra billed cost against the reduction in completion time for that particular workflow. Enabling a faster mode everywhere would erase the distinction between the lower standard rate and the premium charged for speed.
Subscription users face another set of boundaries. Anthropic says it is increasing five-hour usage limits for Pro, Max and Team subscribers and providing subscription users with a rate-limit reset that can be saved for later. The available announcement does not quantify the increase for each plan. Those changes should be described as additional allowance, not converted into an invented number of prompts or coding hours, and API token prices should not be presented as subscription fees.
The Artificial Analysis Token Claim Uses a Different Comparison
Wccftech reports that Opus 5.5 reached 58 on the Artificial Analysis Intelligence Index, five points ahead of Fable 5.1 and GPT-6 Astra. It also reports average consumption of 119,000 tokens per task for Opus 5.5 against 17,000 for GPT-6 Astra at xhigh effort. No second inspected source confirms those particular Intelligence Index figures or their full configuration.
Even if taken as reported, those figures answer a different question from Anthropic’s 40% claim. The token comparison is against GPT-6 Astra on that evaluation; the headline workload saving is against Opus 5 at default settings. The original report’s warning about potentially steep real-world costs cannot be derived simply by moving from one comparison to the other. Different opponents, settings and task collections need to remain visible.
The token counts are still a useful lead for buyers interested in that benchmark. They suggest asking how much reasoning and output a model consumed to obtain its score. But neither those counts nor a leaderboard position establishes the cost of an organization’s own repository audit or application change. For procurement, the missing bridge is a priced, reproducible run using the workflow and settings the organization intends to deploy.
Opus 5.5’s Benchmark Lead Comes With Measurable Boundaries
Anthropic’s performance table shows improvements over Opus 5 across several coding and knowledge-work evaluations. The company also supplies an unusually important interpretive warning: at these capability levels, it has found benchmark margins to be a less reliable guide to real-world differences. In its own use, the gap between Opus 5.5 and Fable 5.1 is narrower than the scores suggest.
The published figures are useful when kept with their names, versions and testing conditions. They are Anthropic’s reported results and comparisons, including some figures that the company credits to other organizations.
| Evaluation | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 66.4% | 55.8% | 52.3% | 57.9% |
| FrontierCode v1.1, main set | 54.4% | 50.3% | 48.0% | 53.3% |
| CursorBench 4.0 | 57.8% | 51.8% | 46.6% | Not listed |
| GDPval-AA v2.1 | 1,846 Elo | 1,735 Elo | 1,708 Elo | 1,542 Elo |
| AutomationBench | 40.0% | 31.4% | 26.9% | 41.4% |
| Humanity’s Last Exam, with tools | 67.7% | 65.6% | 63.6% | 57.2% |
| OSWorld 2.0, partial completion | 81.8% | 80.7% | 74.0% | Not listed |
These evaluations cover different work. Anthropic describes Terminal-Bench 4.0 as measuring complex, multistep professional tasks in a command-line interface. FrontierCode measures whether an agent’s changes would be merged, while CursorBench uses ambiguous, multifile coding tasks drawn from real Cursor sessions. A terminal-task improvement is directly relevant to an agent operating through a shell, but it should not be substituted for evidence about every kind of application development.
GDPval-AA v2.1 evaluates professional work across 44 occupations and reports an Elo rating rather than a completion percentage. OSWorld’s figure is explicitly labeled “partial.” Neither should be read as “the model successfully completes this percentage of all business work.” Keeping the original units and labels avoids turning several specialized measurements into a misleading universal intelligence score.
The table also limits the claim that Opus 5.5 beats GPT-6 Astra everywhere. Astra has the higher listed AutomationBench result, 41.4% against 40.0%. Anthropic’s additional Terminal-Bench-Science 0.1 table likewise puts Astra at 64.6% and Opus 5.5 at 58.7%. Opus 5.5 can lead in several relevant categories without sweeping every evaluation, and a buyer’s particular workload may resemble one of the exceptions.
Effort Settings Change Both the Score and the Bill
Most Opus 5.5 results in Anthropic’s headline table use adaptive thinking at maximum effort. Terminal-Bench 4.0 instead uses xhigh effort for Opus 5.5 and a high-effort GPT-6 Astra result reported by OpenAI. Anthropic describes those as each model’s highest score. The comparison is therefore a view of reported peak results under their respective configurations, not a test in which every model has the same effort setting or resource budget.
Anthropic separately publishes cost-versus-performance comparisons at default, medium effort. On FrontierCode, it reports 54.6% for default-effort Opus 5.5, compared with the 54.4% in the main maximum-effort table. On CursorBench, the default figure is 52.5%, against 57.8% at maximum effort. These pairs show why a single setting should not be assumed to improve every task monotonically or to represent the configuration customers will actually use.
The default-effort comparisons are especially relevant to budget decisions. Anthropic says Opus 5.5 beats GPT-6 Astra’s top FrontierCode score at about one-fifth of the cost per task, and matches Astra on Terminal-Bench at about 40% of the cost. Those are vendor-reported relationships for named evaluations. They support trying the cheaper default setting first in a local comparison, rather than automatically choosing the setting used to generate a headline score.
Statistical uncertainty also belongs beside close comparisons. Anthropic gives Terminal-Bench 4.0 a standard error of plus or minus 2.6 points for Opus 5.5 and between 1.6 and 2 points for the other Claude models. It explains that its Opus 5 reproduction, 52.3%, differs slightly from the public leaderboard’s 51.8% but falls within noise. A fraction-of-a-point difference should not be treated with the same confidence as a large, consistently observed improvement.
Safeguard Fallbacks Make Some Scores System Results
Anthropic says it evaluated Opus 5.5 with production safeguards enabled. When those safeguards intervened, cybersecurity tasks were completed by Opus 4.8, while biology and frontier-model-development tasks were completed by Opus 5. It says this likely reduces the reported performance. The important interpretation is that some results describe the deployed arrangement, including fallback behavior, rather than uninterrupted execution by Opus 5.5 alone.
AutomationBench used a different rule. According to Anthropic’s account of Zapier’s evaluation, those runs had no fallback models, so safeguard interventions counted as failures. This difference can affect the relationship between a benchmark score and a production workflow. An administrator needs to know whether a restricted task stops, is handed to another model, or is handled through a verified-access program; a single percentage obscures those operational differences.
There is another comparison boundary in Anthropic’s WANDR results. The company used offline web-search and web-fetch tools, programmatic tool calling, code execution and a 980,000-token task budget. It explicitly says this differs from Perplexity’s published setup, making the scores unsuitable for direct comparison across the two arrangements. That budget also should not be repurposed as confirmation of the model’s context-window specification: a task budget and a context limit describe different constraints.
Opus 5.5’s Coding Examples Reward Reviewable Completion
The most persuasive launch examples concern completed work rather than isolated answers. Anthropic says an early tester completed a 680,000-line code migration in less than a day, work the company describes as taking an engineering team weeks. Another tester reportedly audited and fixed a 200,000-line codebase in under three hours, compared with more than 20 hours for Opus 5. In the latter comparison, Anthropic says the older model used 2.5 times as many tokens.
These are company-reported examples, with insufficient detail to turn them into general productivity multipliers. Repository size alone does not tell a reader what had to change or how acceptance was assessed. Their useful contribution is narrower: they identify migrations and codebase-wide audits as workloads in which Anthropic and early users observed improvements worth testing. They do not justify promising that an organization’s next migration will fit into one working day.
Anthropic’s internal HAProxy exercise supplies a more concrete outcome measure. The company asked Opus 5.5 and Fable 5.1 to translate HAProxy, software used to balance web traffic across servers, from C into Rust. It says both translations passed nearly all of HAProxy’s regression tests. Opus 5.5 completed the exercise in 9.5 hours against Fable 5.1’s 12 hours, with a reported 51% cost reduction.
“Nearly all” is important to the decision. The exercise supports a comparison of the two models under the test’s conditions; it leaves remaining failures to resolve before anyone could treat the translation as accepted work. A reader evaluating a similar migration should record the same categories separately: elapsed time, billed cost, tests passed and outstanding defects. Collapsing those into “finished” would discard the very evidence needed to judge whether the result is usable.
A separate web-application test puts behavioral preservation at the center. Anthropic says Opus 5.5 successfully reduced load times across the application’s pages in 39 of 40 attempts. Opus 5 made smaller improvements but also changed application behavior. This is a more meaningful distinction than simply generating a faster-looking patch: the claimed advantage combines performance improvement with keeping the application’s intended behavior intact.
The GitHub testimonial complements these examples without becoming an independent replication of them. Fewer steps in VS Code and Copilot CLI could reduce both waiting and metered work, but the quoted evaluation does not disclose a complete task set or cost schedule. Its strongest practical use is to suggest what to measure alongside success: terminal actions, model turns and token use, all tied to an accepted outcome.
Knowledge-Work Tests Show Why Acceptance Rules Must Be Explicit
Anthropic also describes an internal research test that required models to write a company-performance report using a copy of the web in which the relevant earnings release was difficult to find. An automated grader checked every figure and quotation. Across effort settings, 16 of 18 Opus 5.5 reports passed a quality bar under which any invented number or quote caused failure; neither Fable 5.1 nor Opus 5 passed in any attempt.
The result is specific to Anthropic’s test, but its acceptance rule is useful. A report with mostly correct prose still failed if it fabricated a figure. For organizations considering model-generated financial or operational summaries, that is a better match to the consequences of an error than judging fluency alone. It also demonstrates why a successful average score should not be interpreted as a guarantee for an individual output: two Opus 5.5 attempts still missed the stated bar.
Another internal test asked Opus 5.5 and Opus 5 to analyze a proposed merger between fictional HR software companies. Each produced an Excel financial model and an executive presentation. Anthropic says they reached the same overall conclusion, but Opus 5.5’s model was more thorough and its presentation easier to read, while Opus 5’s work contained minor errors. The reported completion times were 63 and 93 minutes respectively, with Opus 5.5 costing 50% less.
For Microsoft 365 users, this is evidence about an Excel-related task in Anthropic’s evaluation, not an announcement of a new Microsoft 365 integration or entitlement. Its contribution is to broaden the evaluation beyond code. Time, factual accuracy, completeness and presentation quality can all change the amount of human finishing work required, even when two models reach the same high-level recommendation.
Across these examples, the recurring operational question is the burden left after generation. A migration with failing tests, a report with an invented quotation and a financial model with minor errors create different review obligations. Anthropic’s launch material gives reasons to expect improvements in several of those areas. An organization can turn that into useful evidence by retaining its acceptance criteria when it compares models, rather than relaxing them to match a faster result.
Opus 5.5’s Safeguards Require Deployment-Level Scrutiny
Anthropic describes three layers around autonomous coding: a classifier that screens actions before execution, an open-source sandbox that security teams can audit, and code review intended to catch vulnerabilities before changes merge. It also says Opus 5.5 has stronger defenses against prompt injection across coding, tool use, computer use and web browsing. Prompt injection is the attempt to redirect an agent through content it encounters while performing its assigned task.
The company’s behavioral claims extend beyond malicious prompts. Anthropic says Opus 5.5 performed best among its tested models on its automated behavioral audit, which uses thousands of simulated scenarios. It reports a reduced tendency to take hard-to-reverse actions or act outside assigned boundaries, and says it broadened testing to cover longer tasks, impossible tasks and scenarios modeled on real incidents. These are claims about the company’s evaluations, with acknowledged limits.
Anthropic also reports that Opus 5.5 tied Fable 5.1 for the lowest prompt-injection success rate on a benchmark run by AI security firm Gray Swan. It names Frontier Design and METR as external pre-release evaluators. The available announcement does not supply those organizations’ complete findings, so their involvement should not be compressed into a blanket independent certification of safety.
An autonomous-agent evaluation must include its permissions and safeguards. The model’s behavior is only one part of the arrangement Anthropic describes. A screening layer, a sandbox and a pre-merge review process address different points in the workflow. Before relying on that arrangement, a team needs to establish which of those protections are actually present in its chosen client and integration, rather than assuming a launch-page description applies identically everywhere.
The evidence does not provide a verified Windows-specific setup procedure, universal approval configuration or sandbox deployment sequence. It would be unsafe to invent one. The supported practical approach is to make the proposed deployment’s boundaries part of the evaluation: which actions require approval, what environment the agent can affect, and how its changes are reviewed before acceptance. Those are purchasing and deployment questions, not claims that a particular undocumented toggle exists.
Verified Access and Fallbacks Affect Security Workflows
Anthropic says Opus 5.5 is comparable to Mythos 5.1 in biology and cybersecurity, and that it is deploying safeguards similar to those used with Fable 5.1. Vetted organizations can apply to its Life Sciences Verification Program immediately. Expansion of the Cyber Verification Program is described as coming in the following weeks, allowing verified cybersecurity practitioners to use Opus 5.5 for their work.
Security teams should keep that timing distinct from the model release. Announcing Opus 5.5 does not establish that every specialized cybersecurity use is immediately available under every account. The distinction is particularly material given the benchmark fallback behavior: a task can be subject to a safeguard even where the broader model is accessible. Evaluation results need to record interruptions or substitutions instead of silently treating them as ordinary Opus 5.5 completions.
The same principle applies to long-running work. Anthropic publishes an early-user account of an agent staying on task for more than 18 hours across multiple repositories. That is evidence of one reported use, not a reason to grant an unfamiliar deployment equally broad unattended access. The appropriate progression is to validate accepted outcomes and operating boundaries together before expanding the duration or reach of a job.
The Trump “SI” Claim Has No Established Federal Directive
The political portion of Wccftech’s report requires a different conclusion from the product portion. The publication says Trump officially changed the name of artificial intelligence to “super intelligence,” or “SI,” and wanted government documents to adopt the term. Its immediate supporting material is a social-media post describing what Trump said. The official material available here contains no proclamation, executive order, memorandum or other directive establishing a government-wide terminology change.
That leaves the “officially renamed” formulation unsupported. A reported statement expressing a preference and a binding instruction governing federal documents are different events. Establishing the latter would require the directive itself or an authoritative official account of its scope. The distinction is practical for IT administrators and contractors: terminology requirements should come from the applicable policy instrument, not an inference drawn from a clip description.
The available White House homepage capture still uses “artificial intelligence,” but it contains a June 2026 calendar reference. It therefore cannot resolve what the administration did on September 22. Treating that older page language as conclusive disproof would repeat the same evidentiary mistake in the opposite direction. The supported position is that the claimed federal action remains unverified, not that an older homepage has disproved every possible later statement.
The safety-policy setting also needs careful attribution. Wccftech reports that Dario Amodei and Sam Altman were due to attend a UN Security Council session that week, and describes an Altman open letter calling for U.S. leadership and communication with China on frontier-AI risks. Those are reported forthcoming activity and policy arguments; they do not establish that a session has taken place or that a new government requirement has been adopted.
The report’s reference to Noam Brown discussing thermal communication between air-gapped computers similarly supplies no incident, reproduction or configuration involving Opus 5.5. It should not become an advisory about this release. Administrators already have a more concrete security subject to evaluate in Anthropic’s announcement: the model’s permitted actions, sandbox boundaries, review mechanisms and safeguard behavior.
Nor does the three-week interval after Fable 5.1, by itself, demonstrate that Anthropic violated a defined development-pacing commitment. Anthropic explicitly calls Opus 5.5 its first release since advocating pacing at the frontier and points to pre-release testing. Assessing compliance with a specific promise would require that promise’s criteria. The release cadence is established; an allegation that the cadence breaches an obligation needs separate evidence.
Evaluate Opus 5.5 Against Accepted Work Before Changing Defaults
Teams already using Opus 5 should consider a bounded Opus 5.5 comparison before changing their default model. The published rates offer a clear reason to investigate, and the reported coding results identify promising task categories. Teams satisfied with another model have less reason to switch solely for a leaderboard position, particularly when the available comparisons use different effort settings or do not match their own workflow.
A useful evaluation can be built directly from the distinctions in Anthropic’s evidence. Its purpose is to find the cost and time required to deliver work that meets existing acceptance criteria. It is not necessary to reproduce every benchmark or to start with the most autonomous task an agent can attempt.
- Select representative work with an existing definition of success, such as a migration with regression tests, an audit with reviewable findings, or a report whose figures and quotations can be checked.
- Record the model, integration, effort setting and speed mode for each run, so that a default-effort result is not accidentally compared with a maximum-effort or Fast-mode result.
- Track the billed categories available to the deployment alongside elapsed time, tool activity, retries and any model fallback or safeguard interruption.
- Apply the same acceptance standard to each output, recording unresolved defects and human rework rather than equating a generated answer with completed work.
- Compare the total for accepted results before expanding access, increasing autonomy or changing the team’s default configuration.
This is an evaluation framework, not a vendor-prescribed installation procedure. Its order matters because acceptance criteria determine what the subsequent cost comparison means. A cheap run that fails the required tests and a more expensive run that passes them are not interchangeable deliverables. Similarly, a time saving that leaves substantially more review work needs to be assessed across the whole workflow.
For developers, the most promising first candidates are the tasks Anthropic actually discusses: multifile changes, codebase audits, migrations and terminal-based work. For business teams, the financial-model and sourced-report tests suggest evaluating verifiability and correction effort alongside speed. For security practitioners, access-program status and safeguard behavior must be included in the comparison, because they can affect whether and how a task proceeds.
The immediate takeaways are concrete:
- Anthropic’s September 22 announcement establishes the Opus 5.5 release, while the reported Claude Code v2.1.280 appearance belongs to the pre-release chronology.
- Standard input and output rates fall by 20%, cache-read rates fall by 60%, and the advertised 40% saving reflects Anthropic’s typical-workload estimate.
- Fast mode carries twice the standard input and output rates, so adopting Opus 5.5 and paying for its fastest mode are separate decisions.
- GitHub’s published evaluation supports testing relevance to VS Code and Copilot CLI, but does not establish universal Copilot availability.
- Benchmark comparisons should preserve effort settings, fallback behavior and acceptance criteria, while the claimed federal renaming to “SI” should not be treated as an established policy requirement.
Opus 5.5 gives existing Opus users a specific next step: determine whether lower rates and fewer reported steps translate into cheaper, faster work that still passes review. Anthropic’s announced Sonnet 5.5 and Haiku 5.5 releases will add further choices in the coming weeks. For now, the defensible upgrade decision rests on the work a team can accept, the controls under which it was produced and the complete bill for getting there.
Update: Claude Opus 5.5 rolls out across GitHub Copilot clients (September 22, 2026)
GitHub now confirms that Claude Opus 5.5 is available in GitHub Copilot, subject to a gradual rollout. The model can be selected in Visual Studio Code, Visual Studio, Copilot CLI, the Copilot coding agent, GitHub Copilot app and GitHub Mobile, plus JetBrains IDEs, Xcode and Eclipse.
According to GitHub’s changelog, access is available to Copilot Pro+, Max, Business and Enterprise users. This supersedes the earlier availability caveat: GitHub’s prior evaluation did not establish general customer access, but GitHub now identifies the supported plans and clients. Users who do not yet see the model in the picker may still be awaiting rollout.
For Windows developers, this establishes direct availability in both VS Code and Visual Studio rather than merely evidence from GitHub testing. Copilot Business and Enterprise administrators can manage access through Copilot model policy; newly introduced models are enabled by default unless the global default is disabled or the specific model is turned off.
GitHub also says Opus 5.5 is billed at provider list pricing under usage-based billing and that its text outputs carry an Anthropic watermark which does not alter token usage or cost.