A laptop displaying code is linked by glowing digital streams to an AI chip and cloud servers.
GitHub Copilot is about to get a hybrid brain. Microsoft says Copilot will soon decide for itself whether a coding task runs on a model on your PC or on a cloud model. The caveat is that this is an experimental preview, not a general release.

A laptop displaying code is linked by glowing digital streams to an AI chip and cloud servers. What was announced​

On October 7, 2026, Microsoft's Windows Experience Blog said hybrid intelligence powered by HydraFusion is coming to the GitHub Copilot app, GitHub Copilot CLI and Visual Studio Code in experimental preview later this month. The companion technical post from Microsoft's Patrick Nikoletich and Stuart Schaefer says the change arrives "by the end of the month."

Some coverage turned that into "by the end of October" and implied a firm launch. Microsoft's own wording is "experimental preview." So don't plan a rollout around an October 31 date, and don't assume everyone gets it at once.

HydraFusion: the orchestration layer​

HydraFusion is GitHub's multi-model orchestrator. GitHub introduced it as a research preview that builds an execution plan, choosing from models across multiple providers to draft, critique and revise, or cascade to more powerful models. It followed Auto model selection. Auto chose one model for each workflow, while HydraFusion chooses one or several based on the workflow type.

The cloud-only version has already spread:

  • It started in Copilot CLI.
  • The research preview later became available in Visual Studio Code and the GitHub Copilot app.
  • It is available to Copilot Pro, Pro+, Business, and Enterprise users.
  • For Copilot Business and Enterprise, an administrator must enable preview features.

The new part is that Microsoft is extending HydraFusion to Windows, so it can use models running on the device as well as cloud models. Microsoft's stated rationale is that models are growing faster than cloud budgets, so it wants customers to make "every AI token count."

Microsoft's materials don't say whether the local-model preview will follow the same plan and admin gating as cloud HydraFusion. Enterprise admins should watch for that.

Two ways to use local models​

Microsoft describes two paths across the Copilot CLI, the Copilot app and VS Code:

  1. Auto orchestration. Copilot decides where each task runs. Microsoft says that across a multi-turn session it can weigh task context and cache state while routing work between local and cloud models.
  2. Explicit selection. For workflows that need a specific provider, model or endpoint, you pick it yourself.

For explicit selection, there are two routes:

  • Choose MAI Code 1.1 Flash through the Windows ML provider.
  • Connect Copilot to an OpenAI-compatible local endpoint and pick from the models it exposes.

The second route matters to teams that already run local inference servers. The announcement doesn't say whether every such setup will work, so wait for the preview documentation before committing.

MAI Code 1.1 Flash: the featured local model​

Microsoft's local coding model is a mixture-of-experts design with 137 billion total parameters and 6.8 billion active. The on-device version uses quantization and speculative decoding. Microsoft says the quantized model is 53 GB, about 80% smaller than the bfloat16 cloud variant.

All figures below are Microsoft's own, measured on October 5, 2026. They used mixed-precision quantization at roughly 3.3 bits per weight, DFlash2 speculative decoding and a Windows ARM64 llama.cpp CUDA runtime.

MetricReported result
Peak memory at 256K context75.5 GB
Prompt throughput at 64K context923.5 tokens/sec
Prompt throughput at 128K context769.8 tokens/sec
BenchmarkCloud MAI Code 1.1 FlashGPT-OSS-120B (Unsloth GGUF)Quantized on-device
SWE-Bench Verified (500)72.6%32.0%70.80%
Terminal-Bench 2.1 (89)62.9%23.6%66.29%

Read these with care:

  • They are vendor-run results. Microsoft says actual results vary by device and configuration.
  • The footnote describes the throughput figures as reflecting a synthetic code-generation workload.
  • The GPT-OSS comparison uses a community GGUF build, not an independent head-to-head.
  • The quantized model scored slightly lower than the cloud model on SWE-Bench Verified and higher on Terminal-Bench 2.1. With only 89 tasks in the latter, a gap of a few points shouldn't be over-read.

The hardware reality check​

The featured machine is the Surface Laptop Ultra, built on NVIDIA RTX Spark. It has up to 128 GB of unified memory and up to 1 petaflop of AI compute. Microsoft's Windows post says it is available from October 16 and is on pre-order now.

Microsoft is also clear that memory isn't just model weights. The OS, your apps, the inference runtime and the key-value cache all draw from the same pool. As an agent reads files and gets tool results, its context grows and memory use climbs. That is why the 75.5 GB peak at 256K context matters more than the 53 GB file size.

My read, from general industry knowledge and not from Microsoft's text: a model this size is a workstation-class workload. The 128 GB unified memory configuration is the one built for it. Microsoft says the Surface Laptop Ultra supports models beyond 120 billion parameters locally. It doesn't say what smaller configurations can run. Microsoft's Windows post does mention other models for RTX Spark, including an upcoming NVIDIA Nemotron model quantized to just over 20 GB and DeepSeek V4 Flash.

Local does not mean offline​

Microsoft's post says that model selection, inference and tool execution have different boundaries, and that local inference does not make the session offline. Auto may send work to the cloud by design. If your compliance question is "does this code ever leave the machine," a local-capable Copilot doesn't answer it by default. You'd need to pin an explicit local model and check the other network paths.

Microsoft also doesn't promise that local will be faster, cheaper or the default. It gives no savings figure for Copilot users.

Sandboxing: related, but a separate control​

The same announcement covers Microsoft Execution Containers (MXC), which Microsoft's Windows post says is now generally available on Windows 11. Copilot uses MXC to turn policy into OS-native controls:

  • Windows: the BaseContainer tier of the ProcessContainer backend.
  • macOS: Seatbelt.
  • Linux: bubblewrap.

These backends need no separate VM or container image. Microsoft says it plans to offer more options through MXC later.

The sandbox is independent of model choice. An agent's shell commands normally inherit the access of the account running them, and moving inference on-device doesn't change that. The sandbox applies policy to the processes Copilot launches, whichever model asked for the work.

What's inside the boundary, and what isn't​

  • Inside: shell commands and, by default, local MCP servers and language servers.
  • Not OS-isolated: Copilot's built-in file tools run inside Copilot itself. The harness checks their requests against policy, but Microsoft says this isn't OS-enforced child-process isolation.
  • Outside: remote MCP servers sit outside the local process sandbox. Where MCP sandbox controls apply, Copilot checks their connection policy in process.

Turning it on​

  1. In Copilot CLI, run the /sandbox slash command to configure settings at any time.
  2. In the Copilot app, click the gear icon and select your project in the left menu.
  3. Enable Sandbox new sessions. The project's current working directory is read/write by default, and the rest of the system stays largely read-only or inaccessible to the agent.

Microsoft's worked example is a daily repository dashboard automation, created from the app's Automations section. It pairs the local MAI Code model with a sandboxed project, with the automation set to run at 9 AM.

A caution on maturity​

The public MXC repository describes MXC as a system for running untrusted code such as model output, plugins and tools. Its default backend on Windows 11 is ProcessContainer, and several other backends are marked experimental. The repository's own text, as captured here, doesn't include the "early preview" or "not a security boundary" warnings I'd expected, so I won't claim them. The practical advice stands anyway: a sandbox reduces blast radius. It isn't a reason to give an agent your credentials.

What to do now​

  • Developers: Watch for the experimental preview in the Copilot app, CLI and VS Code later this month. Enable sandboxing before running agents on local repositories, whichever model is selected.
  • IT admins: Expect preview-feature toggles in Copilot Business and Enterprise. Decide your policy on local models and OpenAI-compatible endpoints before developers ask.
  • Hardware buyers: Don't buy a PC for this on the strength of the announcement alone. Microsoft's performance numbers come from the Surface Laptop Ultra, and it hasn't published requirements for other devices.

Bottom line​

The announcement is real and specific: HydraFusion routing across local and cloud models on Windows, MAI Code 1.1 Flash through Windows ML, OpenAI-compatible local endpoints, and MXC sandboxing. It is also an experimental preview with company-supplied benchmarks, tied to high-memory hardware. Hybrid routing could be a sensible way to stretch AI budgets. Whether it works well in daily coding is something only real-world use will show.

 

References

  1. GitHub Copilot Will Soon Switch Between Local and Cloud AI Windows Report 2026-10-08T07:44:41+00:00
  2. (Research Preview) HydraFusion is live in GitHub Copilot CLI: Frontier quality via multi-model orchestration · community · Discussion #206492 github.com
  3. HydraFusion in VS Code and the GitHub Copilot app - GitHub Changelog github.blog