What Microsoft actually announced
The model is a mixture-of-experts design. Microsoft describes it as having 137 billion total and 6.8 billion active parameters, built for real-world coding workloads. For local use, Microsoft says it applies 3-bit precision to cut the model size by nearly 80% while preserving coding quality and supporting a 256K context window locally.
Microsoft AI's own model post, updated on October 7, adds the practical details:
- Microsoft says the quantized build has comparable coding performance to its full-precision counterpart on SWE-Bench Verified and Terminal-Bench 2.1. I found no independent testing of that claim.
- The post says local model calls carry zero inference charges.
- It says the model is available to download and run locally, with more than 120GB of RAM recommended for best performance.
That memory figure is why this is a "monster PC" story. It is a recommendation for best performance, not a published hard minimum. Microsoft's material doesn't spell out what happens on smaller machines.
The hardware: Surface Laptop Ultra and friends
The launch hardware is the Surface Laptop Ultra, built around NVIDIA's RTX Spark. Microsoft's device post says it offers up to 128GB of unified memory and can run AI models exceeding 120 billion parameters locally. Two caveats matter.
- Unified memory isn't all available to the GPU. Microsoft's footnote says the 128GB is shared dynamically between CPU and GPU. It adds that the amount the GPU can address depends on configuration and workload and is less than the total. So 128GB is a capacity figure, not a guarantee of what the model gets.
- Price depends on configuration. Microsoft's device blog lists pre-orders starting at $2,599, with availability beginning October 16. Neowin cites $5,899 for the top configuration. I couldn't confirm that figure in Microsoft's own posts.
Microsoft's Windows post says other RTX Spark PCs from ASUS, Dell, HP, Lenovo and MSI are also open for pre-order and ship October 16. It also says RTX Spark dev boxes and mini desktops arrive later this year. No Microsoft source I reviewed publishes a compatibility list for non-RTX-Spark machines with 128GB of RAM.
How it plugs into GitHub Copilot
The local model is meant to be one option inside Copilot, not a replacement for the cloud. Microsoft AI says it gives Copilot's router an on-device option, so eligible coding work runs locally while cloud models handle tasks needing more capability.
Coverage of the event adds mechanics. Techaeris reports that developers can let Copilot's Auto orchestration route work between local and cloud inference, or explicitly select a local model. It also reports that MAI Code 1.1 Flash is selectable through the Windows ML provider, and OpenAI-compatible local endpoints are supported too. Neowin's separate event report gives the same two paths.
Microsoft's Windows post says the Windows ML runtime now supports llama.cpp, which should give developers more open-source model choices. Microsoft doesn't say this is how MAI Code 1.1 Flash itself runs.
Timing: download now, Copilot integration later
Two things are easy to blur here:
- The model download. Microsoft AI's post says it's available now.
- The Copilot integration. Experimental access in the GitHub Copilot app, Copilot CLI and Visual Studio Code arrives by the end of the month. Microsoft's Windows post puts the experimental preview of hybrid routing "later in October".
So the on-device coding experience inside Copilot is a late-October preview. It is not a finished feature today.
Sandboxing
Local agents that can run commands raise an obvious question about containment. Reporting on the event says GitHub Copilot uses Microsoft Execution Containers, or MXC, to govern access to files, networks and system capabilities. On Windows it uses the BaseContainer tier of the ProcessContainer backend. Microsoft's Windows post says MXC is now generally available on Windows 11, with policies enforced at runtime.
Not a new model
This isn't a brand-new model. The hosted version already shipped in Copilot. GitHub's community post dated August 4 describes it as Microsoft's latest small-tier coding model, adding native vision support and improvements in coding quality, instruction following, tool use and performance. Dates for the August rollout vary between sources, so I'm not pinning an exact day.
GitHub's changelog says the older MAI-Code-1-Flash was deprecated across Copilot on September 10, with 1.1 as the suggested alternative. Microsoft claims 1.1 streams tokens 25% faster and uses 25% fewer tokens per task, at a quarter of the cost of the model launched at Build. Those numbers describe the hosted version in Copilot. They don't show the 3-bit local build performs the same way.
Microsoft's Windows post says MAI Code 1.1 Flash was introduced "at Build". Microsoft AI's own post says the model it launched at Build was the earlier version, not 1.1. I'd treat the Build reference as loose wording.
Analysis: who should care
Developers with 128GB-class machines. Offline agentic coding with a long context window is a real option if you have the memory. Zero per-call charges matter for people who run agents all day.
Everyone else. The recommendation puts it out of reach of typical laptops and most desktops. For you, the practical story is the router. Copilot may quietly send some work to a local model when one exists on your machine, and it still falls back to the cloud for heavier tasks.
IT admins. Local models raise governance questions: which agents run, what they can touch, and how that's audited. Microsoft's answer is MXC plus Entra and Agent 365 integration. Its device post describes the Entra and Agent 365 pieces as upcoming capabilities, not shipping ones.
There's also an obvious commercial angle. A model that wants more than 120GB of RAM launching alongside a 128GB laptop is convenient timing. The model may still be genuinely useful. But the performance claims come from Microsoft, and the independent benchmarks that would test them don't exist yet.
What to watch
- Whether Microsoft publishes a minimum spec and tested performance on smaller memory configurations.
- Real-world token throughput. Microsoft's event charts reportedly show roughly 40 to 63 tokens per second across prompt lengths, but those are Microsoft's own numbers.
- Whether the late-October Copilot preview actually ships on schedule across the app, CLI and VS Code.
- Independent testing of the 3-bit build against the full-precision model on real codebases.
The short version: the local model exists and Microsoft says you can download it now. But the Copilot integration is still weeks away, and it only makes sense on a very well-equipped machine.
Update: Microsoft reportedly expands local AI plans beyond MAI-Code (October 7, 2026)
Brand Icon Image reports that Microsoft also discussed running a version of DeepSeek V4 on Windows machines with at least 60GB of memory. The outlet says Microsoft executive Pavan Davuluri claimed it can outperform OpenAI’s GPT-5 on some coding and reasoning tasks. That is a reported comparison, not independent benchmark evidence, and the 60GB figure is separate from Microsoft’s recommendation of more than 120GB for best performance with MAI-Code-1.1-Flash.
The report adds names to the planned Microsoft Execution Containers ecosystem: Davuluri said Anthropic, OpenAI and Nvidia would use MXC. It also says Meta’s Muse assistant is coming to Windows as a native app and that OpenClaw can work with the security tools. These details could matter to admins assessing which agents may run under Windows’ access controls, though the report does not give deployment dates or specific policy capabilities.
On hardware costs, Brand Icon Image reports the Surface Laptop Ultra’s top configuration at $5,899, with a 20-core processor, 128GB of memory and 1TB of storage. It also says Nvidia raised the DGX Spark’s price by about 75% to $6,950. Those figures underscore that local AI remains costly; the reported memory shortage may add further pressure on buyers.
References
- Microsoft's new coding AI model can run locally, but you'll need a monster PC Neowin · 2026-10-07T19:58:01+00:00
- MAI-Code-1.1-Flash available in GitHub Copilot 🚀 · community · Discussion #203965 github.com
- MAI-Code-1-Flash deprecated - GitHub Changelog github.blog