Windows ML is Microsoft's local AI inferencing framework for Windows. Neowin framed this week's news as Microsoft highlighting "major advancements." Microsoft's own developer post is more cautious. It lists real additions, but it also says the new capabilities are experimental.
What Microsoft announced
Microsoft's Foundry on Windows blog post is dated October 7, 2026. It says Windows adds experimental llama.cpp support to Windows ML so developers can run GGUF models locally through new task-specific APIs. It also calls the experimental Windows-native Runtime API "now in preview" for developers who want more control over how models run and compose.
The update has four parts:
- GGUF and llama.cpp: Developers can pull a new GGUF model from Hugging Face and run it locally through the existing Windows ML stack.
- Task-specific APIs: A Text Generation API and a Speech Recognition API.
- The Windows ML Runtime API: A native inferencing path.
- PyTorch and Triton: Windows on Arm work in both projects.
Microsoft's GitHub release for the package is tagged Windows ML 2.7.2021 Experimental. The release notes call it an early look, with broader workloads, execution engines, performance work and silicon coverage promised as the APIs mature.
GGUF through Windows ML
GGUF is the model file format used by llama.cpp. It packs weights, tokenizer and metadata into one portable file and supports aggressive quantization. That is why many open-source models appear in GGUF form first.
Microsoft's Text Generation API accepts language models in both GGUF and ONNX formats through one surface. Windows ML picks the execution engine automatically, using llama.cpp for GGUF.
Microsoft's example for quick prototyping uses an OpenAI-compatible local endpoint. The steps are:
- Start the local server with a GGUF file, a model ID, a target and a port. Microsoft's example is
WinMLServer.exe model.gguf --model-id qwen2.5-0.5b --target gpu --port 8080. - Point the standard OpenAI Python SDK at
[url]http://127.0.0.1:8080/v1[/url]. - Use the access key that the server prints at startup as the API key.
- Call the chat completions API as usual, with streaming if you want it.
Existing OpenAI-style client code can therefore talk to an on-device model with minimal rewriting. This is Microsoft's example for one small model. It does not show that every GGUF model or device configuration works.
The Speech Recognition API transcribes audio with an ONNX Whisper model. Microsoft says the two APIs can be chained. A developer can transcribe voice input, then pass the text to a GGUF model.
The native Runtime API
Under the task APIs sits the Windows ML Runtime API. Microsoft's description lists several capabilities:
- Feeding images, video frames, audio buffers and text to models through zero-copy paths.
- Building deterministic multi-model pipelines with explicit CPU, GPU or NPU placement for each stage.
- Loading and compiling models ahead of time into ready-to-run artifacts.
The GitHub pull request adds a useful caveat. It says native data integration makes efficient paths possible, but zero-copy and acceleration depend on the actual layout, device and backend. Treat "zero-copy" as a possibility, not a guarantee.
What this means for existing ONNX apps
Nothing is being pulled out from under current developers. Microsoft says the familiar ONNX Runtime APIs stay fully supported for broad compatibility. The post also says deeper Windows-native optimizations will land in the Runtime APIs over time. The two ship side by side.
The practical advice is to keep shipping on what you know and try the native path in a branch. Microsoft says to adopt it when you are ready. Because it is preview software, check the supported scenarios and known limitations first, as the post itself advises.
Performance claims
Microsoft says it worked with NVIDIA and the llama.cpp community on performance. The listed work includes CUDA kernel optimization, kernel fusion, improved CPU–GPU scheduling, weight repacking and CUDA graphs. It also lists:
- Eagle-3, MTP and D-Flash2 speculative decoding.
- Multi-GPU execution.
- NVFP4.
- New model architectures.
- Backend sampling.
The post gives no benchmark numbers for the llama.cpp contributions, so there is no measured speedup to quote. The only timing code in the post is a small PyTorch demo comparing eager mode with torch.compile. It is an illustration, not a published result.
PyTorch and Triton on Windows on Arm
The Arm64 section has three points:
- PyTorch now offers official native Windows Arm64 CPU builds.
- NVIDIA separately publishes CUDA-enabled Windows Arm64 packages for supported hardware. The official CPU builds are not the CUDA path.
- The Windows distribution of Triton brings
triton.jit,torch.compileand custom GPU kernels to supported Windows GPUs.
"Supported" is doing real work in those sentences. Microsoft does not claim every GPU or device qualifies.
The end-to-end example
Microsoft's example workflow runs from training to a deployable model:
- Install Python 3.14 for Arm64 with Microsoft's
pymanager, then create a virtual environment. - Install
torch==2.14.0+cu134andtorchvision==0.29.0+cu134using NVIDIA's extra package index. - Install
triton-windows==3.8.0.post29, plusonnx,onnxscriptandnumpy. - Use
torch.compileso Inductor fuses GPU operations into Triton-generated kernels. - Export a ResNet18 model graph to ONNX.
- Use the Windows ML CLI to prepare it.
The post stresses that the model graph, not the generated kernel, is what gets exported for deployment. That distinction is easy to miss, since the Triton kernels help you work on the model and are not shipped in the app.
These versions are a snapshot from a single day. Check current package compatibility for your GPU before copying them. Microsoft points to a community tracker for Windows Arm64 Python package status.
The CLI step
The Windows ML CLI exposes analyze, optimize, quantize and compile. Microsoft's example installs winml-cli==0.3.1, then runs winml analyze, winml build and winml perf. It targets a GPU with the nv_tensorrt_rtx execution provider. That example is NVIDIA-specific and should not be assumed to apply to other hardware.
Who should care
- Developers building local AI apps on Windows: You get a lower-friction route from a fresh Hugging Face GGUF model to a running local endpoint.
- Arm64 developers: Native PyTorch CPU builds and a documented path for CUDA and Triton on Arm remove some of the usual workarounds.
- IT teams: Nothing here is a deployment mandate. It is a preview aimed at developers. Local inference does keep data on the device and avoids per-token cloud charges, which are benefits Microsoft cites. Neither is guaranteed for every model or machine.
Microsoft ties the update to a new generation of powerful Windows PCs powered by NVIDIA RTX Spark, like Surface Laptop Ultra. That is a hardware tie-in worth noting. The framework itself is described as spanning CPUs, GPUs and NPUs from AMD, Intel, NVIDIA and Qualcomm.
What's next
Microsoft's stated priorities are more native packages, broader kernel coverage, simpler installation, better performance and clearer support across Windows x64 and Windows on Arm. These are goals, not dated commitments.
The bottom line
This is a meaningful widening of Windows' local AI toolbox, especially for GGUF and Arm64. It is also explicitly preview-stage, with no published benchmarks for the headline performance work. Try it on a test machine, keep your ONNX path in production, and tell Microsoft which models and devices you need through the Windows ML GitHub.
References
- Microsoft highlights major advancements in Windows ML - Neowin Neowin · 2026-10-08T10:58:01+00:00
- Windows ML 2.7 Experimental Runtime API by nieubank · Pull Request #23 · microsoft/WindowsML github.com
- AI Development on Windows: from PyTorch and llama.cpp to Windows ML - Microsoft Foundry on Windows Blog devblogs.microsoft.com