# Local coding models in Copilot on Windows: one model, three spec sheets

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/windows-local-coding-models
> date: 2026-10-08
> tags: on-device, agentic-coding, llama-cpp, quantization, speculative-decoding, gguf, security, microsoft

Two posts landed on my timeline within two hours of each other on 7 October. [@windowsdev](https://x.com/windowsdev/status/2107897106237194449) listed six things under the banner "hybrid intelligence": MXC going GA, frontier coding models running locally, GitHub Copilot HydraFusion, llama.cpp plus Windows ML, Surface preorders and new RTX Spark PCs. Then [@MicrosoftAI](https://x.com/MicrosoftAI/status/2107926901150908483) said MAI-Code-1.1-Flash is "going local with on-device model calls inside GitHub Copilot", with "no inference charge for local model calls".

I wanted the boring answers. How big is the model, what does it run on, who decides whether a request goes to my GPU or to Azure, and what is actually downloadable today. The replies under the MicrosoftAI post were asking the same things ("Where can we download the weights?", "how big is the local model?"), and one of them had already noticed something odd: the demo video and the engineering post disagree about the model's size.

They do. So does a third page. That turned out to be the most useful thread to pull.

<Figure
  src="https://ai.thesatyajit.com/articles/windows-local-coding-models/fig1.jpg"
  alt="A grid of announcement tiles: thousands of actions from the taskbar, the most advanced Windows laptops powered by NVIDIA RTX Spark, the most secure platform for agents, the most powerful Surface ever, Windows hybrid intelligence spanning local and cloud, DGX Station for Windows, Copilot's hybrid intelligence toggle, and up to 284 billion parameter models running locally on RTX Spark laptops (MAI Code 1.1 Flash, NVIDIA Nemotron, DeepSeek V4 Flash)."
  caption="The announcement card. The bottom-right tile is the claim this article spends most of its time on: models up to 284 billion parameters running locally on RTX Spark laptops (@windowsdev post image)."
/>

## What exists today, and what is a date

Before any mechanism, the status. Microsoft's posts mix "generally available", "available for pre-order", "experimental preview later in October" and "in the coming months" in adjacent paragraphs, and the tweets flatten all of it into one list. Below is each item with the wording of its own primary source.

| Item | Status on 7 October 2026 | Source |
|---|---|---|
| Microsoft Execution Containers (MXC) | Generally available on Windows 11; open-source SDK | [Windows Developer blog](https://blogs.windows.com/windowsdeveloper/2026/10/07/microsoft-execution-containers-policy-driven-containment-for-ai-agents/) |
| Local sandboxing in GitHub Copilot (built on MXC) | Generally available in Copilot CLI, the Copilot app and VS Code Agent Host sessions | [GitHub changelog](https://github.blog/changelog/2026-10-07-local-sandboxing-for-github-copilot-now-generally-available) |
| Local model discovery in Copilot CLI | Shipped in CLI `1.0.94-0`, for models in a running Ollama instance | [GitHub changelog](https://github.blog/changelog/2026-10-07-discover-local-models-in-github-copilot-cli) |
| llama.cpp in Windows ML | Experimental; Runtime and Text Generation APIs in `2.7.2021-experimental`, not for Store apps | [Foundry on Windows blog](https://devblogs.microsoft.com/foundry-on-windows/build-on-winml-oct-7-26/), [Learn docs](https://learn.microsoft.com/windows/ai/new-windows-ml/runtime/text-generation) |
| MAI-Code-1.1-Flash in the cloud | In production in GitHub Copilot | [Microsoft AI blog](https://microsoft.ai/news/mai-code-1-1-flash-br-better-faster-at-a-quarter-of-the-cost/) |
| MAI-Code-1.1-Flash on the device | "Experimental access" in the Copilot app, CLI and VS Code "by end of the month" | same |
| HydraFusion routing to local models | "Experimental preview later in October" | [Windows Experience blog](https://blogs.windows.com/windowsexperience/2026/10/07/building-windows-for-hybrid-intelligence/) |
| HydraFusion itself (cloud only) | Research preview since 4 September, via `/experimental` in Copilot CLI | [GitHub blog](https://github.blog/ai-and-ml/github-copilot/project-hydrafusion-frontier-quality-via-multi-model-orchestration/) |
| Surface Laptop Ultra, RTX Spark laptops | Pre-order; shipping from 16 October | Windows Experience blog |
| Surface RTX Spark Dev Box | Pre-order; shipping in the U.S. in November | same |
| DGX Station for Windows | "Later this year" | same |
| Copilot (the consumer app) using local models | "Expected to begin rolling out on Copilot+ PCs in the coming months" | same |

So the security half of the story is real software you can install. The local-model half is three weeks away at best, and labelled experimental when it arrives. Neither is a criticism; it is just not what "Frontier coding models running locally" sounds like on a phone screen.

## The model: what Microsoft says it is

MAI-Code-1.1-Flash is Microsoft AI's in-house coding model, the second release of the line that started as [MAI-Code-1-Flash](https://microsoft.ai/news/introducingmai-code-1-flash/) at Build in June. (I covered MAI's earlier image and voice models in [MAI-Image-2.5-Pro and MAI-Voice-2-Flash](/articles/mai-image-2-5-voice-2); the strategy is the same one, applied to code.)

It is a sparse mixture-of-experts transformer. The Windows and Command Line posts both give it 137 billion total and 6.8 billion active parameters, and a 256K-token context that the Windows post and the Microsoft AI post both say is kept on the device. The [model card](https://microsoft.ai/pdf/MAI-Code-1-Flash-Model-Card.PDF) for the June model describes the lineage: trained from MAI-Thinking-1's mid-training checkpoint, then a "mid2" phase of roughly 2 million synthetic agentic tasks and an RL stage over more than 150,000 environments, all inside the Copilot harness. That training-in-the-harness choice is the model's whole pitch; it is tuned to Copilot's tools rather than to a generic benchmark setup.

Version 1.1, per the MAI blog, brings a 22% improvement on Terminal-Bench 2.1 in Copilot CLI, 15% on .NET tasks, 25% faster token streaming, 25% fewer tokens per task, and a price of one quarter of 1.0. The [MAI-Code repository](https://github.com/microsoft/MAI-Code) README adds that 1.1 takes image inputs.

One number does not line up. The June model card lists MAI-Code-1-Flash at "137B total, 5B active". The October posts give 1.1 the same 137B total but 6.8B active. That could be a real architecture change between versions (more experts per token, or a bigger shared expert), or one of the two documents is wrong. There is no config to read, so I can't tell you which.

### Open weights? No.

The model card's licence line reads "Various product and service terms where the model is deployed, such as those for Visual Studio Code." The `microsoft/MAI-Code` repository on GitHub is MIT-licensed, but it holds a README, a security policy and a licence file, nothing else; it exists for issues and feedback. I queried the Hugging Face API for every model under the `microsoft` author with "MAI" in the name and got `MAI-DS-R1`, `MAI-DS-R1-FP8` and the MAIRA radiology models. No MAI-Code.

The MAI blog does say "MAI-Code-1.1-Flash is available to download and run locally", and the demo video shows Copilot offering to download it. The download is a Copilot feature, not a weights release. You get the model as part of using Copilot, under Copilot's terms, and the "no inference charge" applies to Copilot's metering of local calls. The video's install dialog does label the model "MIT", which contradicts the model card; I would not build anything on that label until a model card for 1.1 says the same.

## Three spec sheets

The size question is where the pages part ways.

The [Command Line engineering post](https://commandline.microsoft.com/local-models-sandboxed-tools-github-windows/) by GitHub's Patrick Nikoletich and Windows' Stuart Schaefer is the most specific. It says the on-device build is mixed-precision quantization at "approximately 3.3 bits per weight", comes in at 53GB ("an 80% reduction in size" from the BF16 cloud variant), runs with DFlash2 sliding-window speculative decoding on "a Windows ARM64 llama.cpp CUDA runtime", and peaks at 75.5GB of memory at 256k context on a Surface Laptop Ultra.

The Microsoft AI post says "3-bit quantization", "a full 256K context window", and "we recommend devices with more than 120GB of RAM".

The demo video attached to the MicrosoftAI tweet shows a Copilot CLI model picker in which MAI-Code-1.1-Flash appears under "Local models" with a 128K context, and an install dialog that reads like this:

<Figure
  src="https://ai.thesatyajit.com/articles/windows-local-coding-models/fig8.jpg"
  alt="A dark terminal dialog in GitHub Copilot CLI. It says Copilot will install two things on your machine: llama.cpp from ggml.ai, MIT, 50 MB, the open-source runtime that runs the model; and MAI-Code-1.1-Flash from Microsoft AI, MIT, 22 GB, 128K context, the model itself. Recommended: 32 GB RAM, 12 GB VRAM, 40 GB free. Your machine: 64 GB RAM, 24 GB VRAM, 220 GB free. All stored in ~/.copilot. A highlighted option reads Download and install, 22 GB."
  caption="Eight seconds into the MicrosoftAI demo video: 22 GB, 128K context, 32 GB of RAM and 12 GB of VRAM recommended, and a target machine with a 24 GB discrete GPU. None of this matches the engineering post (frame extracted from the @MicrosoftAI post's video)."
/>

Those cannot all describe the same file. The arithmetic is short. Weight bytes are parameters times bits per weight over eight:

$$
137 \times 10^9 \times 3.3 / 8 \approx 56.5 \times 10^9 \text{ bytes} \approx 52.6 \text{ GiB}
$$

So "53GB" is 3.3 bits per weight measured in binary gigabytes, which is what Windows Explorer shows. BF16 is 16 bits, 274 billion bytes or about 255 GiB, and 53 out of 255 is a 79% cut: "nearly 80%" checks out. Now run it backwards for the video. 22 GB of weights for 137 billion parameters is $22 \times 8 / 137 \approx 1.28$ bits per weight. I know of no agentic coding model shipped at 1.3 bits, and nothing on any of the pages hints at one.

My reading, and it is only a reading: the video is a product-design mock (its model list also includes "GPT-6 Astra" and "Gemini 3.7 Flash") made before the engineering numbers settled, or it shows a different, smaller artifact that has not been described anywhere. The engineering post is dated, has a test date of 5 October, names the runtime and the hardware, and its numbers are internally consistent with each other and with the Microsoft AI post's 120 GB recommendation. I would plan on 53 GB of weights and 75.5 GB at peak. The widget below does the same arithmetic for any bit width and for the other two models the Windows post names.

<WeightBudget />

Two things fall out of playing with it. DeepSeek V4 Flash at 284 billion parameters does not fit in the 110 GB the GPU can see on a 128 GB machine unless you go below about 3.1 bits; the announcement card's "up to 284 billion parameter models" is a 2-to-3-bit claim. And on a 64 GB RTX Spark machine, MAI Code 1.1 Flash's 53 GB of weights leave about 11 GB for everything else, before the GPU-addressable limit even enters. Microsoft's own 75.5 GB peak rules that tier out at long context. The machine you want for this is the 128 GB one.

## What the quantized model costs you in quality

The Command Line post has the only quality numbers for the on-device build:

| Benchmark | Tasks | MAI Code 1.1 Flash (BF16, cloud) | GPT OSS 120B (Unsloth GGUF) | MAI Code 1.1 Flash, quantized on device |
|---|---|---|---|---|
| SWE-Bench Verified | 500 | 72.6% | 32.0% | 70.80% |
| Terminal-Bench 2.1 | 89 | 62.9% | 23.6% | 66.29% |

Convert the percentages back to task counts and they become easier to read. On SWE-Bench Verified the quantized model solves 354 of 500 against the full model's 363: nine tasks fewer, a 1.8-point drop. On Terminal-Bench 2.1, 66.29% of 89 is 59 tasks and 62.9% is 56, so the quantized model solved three more. One task is worth 1.12 points on an 89-task set, and a single agentic run on 89 tasks is noisy at that scale. I would read that row as "no measurable loss", not as quantization helping.

The GPT OSS 120B column is the one I would not quote. 32.0% on SWE-Bench Verified is far below what that model usually scores, and the footnote says the comparison used Unsloth's GGUF. It was evidently run in Microsoft's own harness under conditions the post does not describe. It tells you how a different local model behaved in their setup, nothing more.

## Speed: the number that decides whether you'd use it

<Figure
  src="https://ai.thesatyajit.com/articles/windows-local-coding-models/fig7.png"
  alt="Bar chart titled Decode throughput, MAI Code 1.1 Flash, tokens per second against prompt length: roughly 63 at 2K, 61 at 8K, 57 at 32K, 52 at 64K, 55 at 120K and 39 at 256K. A note says K equals 1,024 input tokens and categories are evenly spaced rather than scaled to length."
  caption="Decode throughput on a Surface Laptop Ultra, with DFlash2 speculative decoding, against prompt length. Microsoft gives no exact values; read off the chart it runs from about 63 tokens per second at 2K to about 39 at 256K (Microsoft Command Line post, decode throughput figure)."
/>

Decode at 40 to 60 tokens per second is fine for an agent. Reading the chart, it goes from about 63 tokens per second at a 2K prompt to about 39 at 256K, with an odd bump at 120K over 64K. With only 6.8 billion active parameters per token the GPU reads a small slice of the 53 GB per step, which is why a 137B model decodes at a usable rate at all, and DFlash2 multiplies that by accepting several drafted tokens per verification pass. (The [DFlash 2 article](/articles/dflash2) explains how its path selector picks among the drafter's candidates.)

The number that matters more for coding agents is in the text, not the chart: prompt processing at 923.5 tokens per second at 64k context and 769.8 at 128k. An agent's context grows with every file it reads and every tool result. At those rates, ingesting a 64K-token context from scratch takes about 71 seconds, and 128K takes about 170 seconds, close to three minutes. Decode speed is what a demo shows; prefill time is what you feel the first time an agent reads your repository.

This also explains a phrase in the routing description that I skipped past at first: Copilot's Auto mode "can consider task context and cache state as it routes work". If the local llama.cpp server already holds 100K tokens of KV cache for this session, sending the next turn to the cloud means paying for those 100K input tokens again, and sending a cloud session's next turn to the laptop means three minutes of prefill. Cache locality is a first-order input to any sane router here. Microsoft says it is one of the inputs. It does not say how it is weighted.

## Memory, and where the NPU went

<Figure
  src="https://ai.thesatyajit.com/articles/windows-local-coding-models/fig5.png"
  alt="Diagram: a CPU box and a GPU box both point down to one bar labelled shared physical memory, which contains three chips: model weights, KV cache, execution plus OS."
  caption="Microsoft's own sketch of the budget: weights, KV cache and everything else share one physical pool (Microsoft Command Line post, memory budget figure)."
/>

RTX Spark is NVIDIA's laptop superchip: a Blackwell RTX GPU, a Grace CPU of up to 20 cores, and up to 128 GB of memory shared between them, with "up to 1 petaflop of AI compute" (the Surface page footnotes that as theoretical FP4 with sparsity). It is Arm, so this is Windows on Arm, and it is the same Grace-plus-Blackwell pairing as the GB10 in the DGX Spark, which I looked at as a [64-user vLLM server](/articles/dgx-spark-batching). Unified memory removes the PCIe copy that sinks large models on discrete GPUs. It does not make all 128 GB available to the model, and Microsoft's Task Manager screenshot shows by how much.

<Figure
  src="https://ai.thesatyajit.com/articles/windows-local-coding-models/fig6.png"
  alt="Windows Task Manager, Performance tab, GPU 0: NVIDIA RTX Spark N1X, 6144-core Blackwell RTX GPU. Utilisation 84%, dedicated GPU memory 62.0 of 74.9 GB, shared GPU memory 0.7 of 35.5 GB, GPU memory 63 of 110 GB. In the sidebar, Memory 25.1 of 51.5 GB, and NPU 0 at 0%."
  caption="Task Manager on a Surface Laptop Ultra while MAI Code 1.1 Flash runs. The GPU can address 110 GB; the NPU sits at 0% (Microsoft Command Line post, Task Manager figure)."
/>

Three readings from that screenshot. The chip identifies itself as "NVIDIA RTX Spark N1X (6144-core Blackwell RTX GPU)". Windows has carved 74.9 GB out as dedicated GPU memory, leaving 51.5 GB visible as system memory, and the GPU can address 110 GB in total once the shared window is counted. The model is using 62.0 GB of the dedicated pool at that moment.

And the NPU is at 0%. The tweet thread says "hybrid intelligence" and Copilot+ PCs are sold on their NPUs, so I went looking for the NPU in this pipeline. It isn't there. The Windows ML Runtime docs say it plainly for GGUF models: the llama.cpp backend "doesn't support NPU targets", and an NPU target returns `ERROR_NOT_SUPPORTED`. Local Copilot inference on these machines is CUDA on the integrated Blackwell GPU. The NPU matters for Windows' own small models (Phi Silica and the like), not for this.

On price, the Surface Laptop Ultra starts at \$2,599.99 and the Surface RTX Spark Dev Box at \$5,999.99. The Surface pages say "up to 128 GB" without telling me which configuration the starting price buys, so I can't say what the 128 GB laptop costs. The performance claim against a MacBook Pro with M5 Pro (2.1x faster time to first token) was measured in llama.cpp on Qwen3.5 27B at Q4K-Medium with an 8,192-token prompt, on 64 GB preproduction machines, in Microsoft-commissioned testing. Time to first token on an 8K prompt is a prefill benchmark, which is compute-bound; a Blackwell GPU should win it. It says nothing about decode.

## How the routing works

There are two ways to use a local model in Copilot, and they are worth keeping apart.

<Figure
  src="https://ai.thesatyajit.com/articles/windows-local-coding-models/fig2.png"
  alt="Diagram with two dashed regions. Connected services: Auto model selection, with an arrow to Cloud models. Windows device: GitHub Copilot as a shared agent runtime for the Copilot app, Copilot CLI and VS Code, with an inference arrow to a local model server for on-device inference and a tool calls arrow to an MXC tool sandbox running scripts and build commands under the effective policy. A dashed arrow labelled model selection runs from Auto model selection down to GitHub Copilot."
  caption="The two paths. Note which box Auto model selection sits in: connected services, not the device (Microsoft Command Line post, local models figure)."
/>

Explicit selection is the simple one. You pick a local model in the picker: MAI Code 1.1 Flash "through the Windows ML provider", or any OpenAI-compatible local endpoint and the models it exposes. The Ollama discovery that shipped in Copilot CLI `1.0.94-0` is the first piece of this you can use now: `/model` lists tool-calling, streaming models from a running Ollama instance alongside your cloud models, and you confirm each one before it is added. It doesn't install a runtime or download a model.

Auto is the HydraFusion extension. To understand it you need HydraFusion as it shipped in September, which had nothing to do with local models.

<Figure
  src="https://ai.thesatyajit.com/articles/windows-local-coding-models/fig3.png"
  alt="HydraFusion: task-aware workflow orchestration in GitHub Copilot. A user request flows to task routing (HyDRA scoring across reasoning, code generation, debugging and tool use, then a local policy that selects a pattern and its constituents from the score and available models), then to orchestration patterns (Single: primary solver; Cascade: draft, isolated judge, optional repair; Critique, Rubber Duck: draft, isolated critic, optional revision), then to execution (solver phases in the shared real workspace with normal tools and permissions; review phases in an isolated context with no tools), then to publish selected results."
  caption="HydraFusion as GitHub described it in September: a capability score per request, then one of three workflow patterns (GitHub blog, Figure 1)."
/>

HydraFusion scores each request on four capabilities (reasoning, code generation, debugging, tool use) and picks one of three execution patterns. Single sends the task to one model. Cascade lets an efficient model draft, runs an isolated judge, and escalates to a stronger model if the draft fails the gate. Critique has a model from a different family review the draft in a tool-less context, and the drafter revises once. GitHub tuned the decision policy with beam search over candidate routing policies rather than hand-set thresholds. Its offline numbers against Claude Opus 5: on Terminal-Bench 2.1, +4.9 points at 67% lower estimated cost; on DeepSWE, -1.5 points at 36% lower cost; on CheckpointBench, its internal benchmark built from real Copilot sessions, -0.1 points at 65% lower cost. Those are GitHub's offline evaluations of the best tuned configuration, priced at list rates. I have no way to rerun them.

What the October posts add is one sentence: HydraFusion can now "tap into models running locally on device". They do not say which pattern a local model plays in, or what "eligible coding work" means. My guess, and I want to be clear it is a guess, is that a local model fits most naturally as the drafter in Cascade: free to call, good enough to pass the gate on routine work, with a cloud model behind it when the judge says no. It is also the only arrangement in which "helping your tokens go further" works out arithmetically, because Critique and Cascade-with-escalation both add cloud calls rather than removing them.

The piece I found most interesting is in the diagram above. Auto model selection is drawn under "Connected services". The decision about where your request runs is made by a GitHub service, not on your laptop. Add the changelog's sentence that "choosing a local model doesn't turn on offline mode or disable GitHub telemetry", and the Command Line post's "local inference does not make the session offline", and the shape is clear: a local model moves the token generation onto your GPU. It does not move the session off GitHub. If you want offline, `COPILOT_OFFLINE=true` is the explicit switch, and the changelog warns that a remote provider "can still receive prompts and code context over the network, even in offline mode".

The widget walks through a session part by part for each mode. The thing to notice is how few rows change when you switch from cloud to local.

<WhereItRuns />

<Figure
  src="https://ai.thesatyajit.com/articles/windows-local-coding-models/fig4.jpg"
  alt="The GitHub Copilot desktop app with a project called Expense Tracker open on a session named Dead code audit. The agent ran dotnet build with 0 warnings and 0 errors, found no unused imports, and lists DiagnosticsLog.Warning as a likely unused API plus intentional TODO markers. The footer of the reply reads 3h ago, Auto, MAI Code 1.1 Flash, with a leaf icon, and the input bar reads Interactive, Auto, Efficiency."
  caption="Auto in the Copilot app picking the local model for a dead-code audit. The reply's footer reads “Auto · MAI Code 1.1 Flash”; the input bar shows an “Auto · Efficiency” setting (Microsoft Command Line post, Copilot app figure)."
/>

That screenshot hints at one more control: an "Auto · Efficiency" setting in the input bar, with a leaf icon that also appears beside the reply that ran locally. Nothing in the text describes it. I'd guess it biases the router toward the cheap path.

## llama.cpp in Windows ML, precisely

"llama.cpp + Windows ML" could mean several things: a new ONNX Runtime execution provider, a DirectML port, a fork. It is none of those. I read the [Windows ML docs](https://learn.microsoft.com/windows/ai/new-windows-ml/runtime/concepts) and the [microsoft/WindowsML](https://github.com/microsoft/WindowsML) repository at `af42ad3`.

Windows ML now has two inference paths side by side: the ONNX Runtime APIs it already had, and a new, experimental, Windows-native Runtime API (C++ and Python only; no C# yet). In the Runtime API, `LoadModelFromFile` picks the backend from the file extension. `.onnx` and `.ort` go to ONNX Runtime and its execution providers. `.gguf` goes to llama.cpp. Your application uses the same objects either way: a model, an execution target, a pipeline builder, stages, tensors, `Run`.

<Figure
  src="https://ai.thesatyajit.com/articles/windows-local-coding-models/fig9.png"
  alt="Diagram: your native or web app with a custom or open-source model in ONNX, PyTorch or GGUF, plus the Windows ML CLI for model prep, feeds Windows ML, which offers ONNX Runtime APIs (existing APIs for reach and compat) and Windows ML Runtime APIs (new Windows-native inferencing APIs), on top of the Windows silicon ecosystem of Intel, AMD, Qualcomm and NVIDIA, down to GPU, NPU and CPU."
  caption="The new Windows ML stack: the Runtime APIs sit beside ONNX Runtime, and GGUF enters through them. The NPU box at the bottom applies to the ONNX path only (Foundry on Windows blog, Windows ML flow figure)."
/>

The llama.cpp copy is not a fork, and it is not shipped with GPU support. The repository's tool README is explicit:

```text
// microsoft/WindowsML  Tools/llama/README.md:5-16
Windows ML ships only the llama.cpp CPU backend, in
`Microsoft.Windows.AI.MachineLearning.LibLlama.Core` and the
`windowsml-llama-core` wheel. It doesn't ship or redistribute GPU backends or
vendor runtimes, including NVIDIA CUDA components. The language samples run on
the CPU with that package alone. To run a GGUF model on a GPU, build a backend
with this folder's script.

`Build-LlamaBackend.ps1` builds a standard upstream `ggml-<family>.dll`, such as
`cuda` for NVIDIA GPUs, `vulkan`, or `opencl`, that matches the Core package.
The output is loaded by `WinMLRuntimeLlama.dll`; it is not a WinML-specific
backend fork.
```

So the GPU path is a stock upstream ggml CUDA backend that you compile yourself, against your own CUDA Toolkit, at the exact llama.cpp commit the NuGet package pins in its manifest; the script reads that commit from the manifest and checks it out. The script is honest about the licensing reason: NVIDIA's redistribution terms are yours to accept, not Microsoft's. Lines 35 to 39 of the same README note the flow "was exercised on Windows ARM64 with Toolkit 13.4", which is the RTX Spark configuration.

The one modification the build makes to llama.cpp is to rename its output DLLs:

```powershell
# microsoft/WindowsML  Tools/llama/Build-LlamaBackend.ps1:596-607
if (WINML_LLAMA_NAMESPACED_CORE)
    set_target_properties(ggml-base PROPERTIES OUTPUT_NAME "WinMLggml-base")
    set_target_properties(ggml PROPERTIES OUTPUT_NAME "WinMLggml")
endif()
...
if (WINML_LLAMA_NAMESPACED_CORE)
    set_target_properties(llama PROPERTIES OUTPUT_NAME "WinMLllama")
endif()
```

A small detail, with a real reason behind it. Every local-AI app on Windows (LM Studio, Ollama, Jan, a game with a chatbot) ships its own `ggml.dll` at its own commit. Loading two different `ggml.dll` builds into one process, or picking up the wrong one from the search path, is a classic Windows DLL failure. Prefixing the names lets Windows ML's copy coexist with whatever else is on the machine, and the script then checks that the CUDA backend it built imports `WinMLggml-base.dll` and no upstream core.

The limits matter if you are deciding whether to build on this today. Per the docs, the llama.cpp backend supports CPU and CUDA only in this release, one decoder stage per GGUF file, and no NPU. The higher-level Text Generation API does greedy decoding only: no sampling, no speculative decoding, no chat templates, no structured output. That last list is interesting next to the Command Line post, whose MAI numbers depend on DFlash2 speculative decoding. Whatever Copilot runs on the device, it is not the public Text Generation API as documented today. The demo video shows Copilot installing llama.cpp itself ("ggml.ai, MIT, 50 MB"), which fits.

DFlash2 support is in upstream llama.cpp, so the speculative half of the stack is public even if the model isn't. At `71ad059` the draft architecture lives in `src/models/dflash.cpp`, and the selector the [DFlash 2 paper](/articles/dflash2) introduced is built as its own graph:

```cpp
// ggml-org/llama.cpp  src/models/dflash.cpp:471-473
// DFlash2 selector: top-k candidates per block position plus the pairwise
// transition scores, packed into the nextn output slot for the CPU-side walk.
static void build_dflash2_selector(llm_graph_context & g, const llama_model & model, ggml_tensor * tokens) {
```

and `--spec-type draft-dflash` is a documented option of `llama-server`. I found no MAI architecture in llama.cpp's model list, which is consistent with MAI's weights not being public: whatever runtime build Copilot downloads either maps MAI onto an existing architecture or carries code that is not upstream. I can't tell which.

The last piece of Windows ML that matters for agents is the server. `WinMLServer.exe model.gguf` serves an OpenAI-compatible Chat Completions endpoint, and its security model is more careful than most local servers:

```text
// microsoft/WindowsML  Samples/Server/README.md:36-45
- It listens only on the loopback addresses `127.0.0.1` and `[::1]`, so other
  computers cannot connect to it.
- Each time it starts, it creates a new random access key. Every request must
  send the key as `Authorization: Bearer <key>`. The server rejects a request
  without it before reading the request body.
- Every request must name a loopback address in its `Host` header and must not
  carry an `Origin` header, so web pages cannot reach the server through a
  browser.
- The server never runs tools. When the model calls a tool, the client runs it
  and sends the result in its next request.
```

The `Origin` check is the one I'd single out. A server on `localhost` with no key can be reached by a web page you happen to visit, through a browser request or DNS rebinding. A random per-start key plus refusing browser-originated requests closes that. Whatever local endpoint you point Copilot at, check it for the same two things.

<RepoCard repo="microsoft/WindowsML" />

## MXC: the part that shipped

Moving inference onto the laptop doesn't change what the agent's shell commands can do. They still run as you. So the sandbox and the local model arrived in the same post, and of the two the sandbox is the more consequential release.

[MXC](https://github.com/microsoft/mxc) (the blog calls it Microsoft Execution Containers, the repository README "Microsoft eXecution Container") is an MIT-licensed SDK, in Rust, .NET and Node, that takes a JSON policy and launches a workload in a platform container that enforces it. The policy covers the containment type, the process, filesystem paths that are writable, read-only or denied, network ingress and egress, and UI access. On Windows the default backend is `processcontainer`; others are WSL containers, a session container that runs the agent under a separate Windows account and desktop, and experimental microVM, Windows Sandbox and Hyperlight backends. On Linux it is bubblewrap, on macOS Seatbelt.

<Figure
  src="https://ai.thesatyajit.com/articles/windows-local-coding-models/fig10.png"
  alt="Diagram: an app manifest of declared needs, an agent request of runtime intent and a user preference of consent and settings feed MXC, which resolves policy, selects primitives and assembles the container, constrained by IT policy for governance. MXC drives OS isolation primitives (security, namespace, virtualization), constrained by OS policy, which fan out to micro-VM, process, session, container, VM and cloud backends."
  caption="MXC sits between the policy sources and the operating system's isolation primitives, and picks a backend (Microsoft Command Line post, sandbox figure)."
/>

Copilot on Windows uses the "BaseContainer tier of the ProcessContainer backend". In the source, a ProcessContainer is an AppContainer launch: `CreateAppContainerProfile`, then a `SECURITY_CAPABILITIES` process attribute, optionally hardened to a less-privileged AppContainer and with Win32k disabled when the policy forbids UI:

```rust
// microsoft/mxc  src/mxc-sdk/src/backends/process_container/common/appcontainer_runner.rs:376-379
/// Always at least 1 (`SECURITY_CAPABILITIES`). LPAC adds one
/// (`ALL_APPLICATION_PACKAGES_POLICY`); UI-disable adds one
/// (`MITIGATION_POLICY` for Win32k disable).
fn compute_attr_count(least_privilege_mode: bool, ui_disable: bool, pipe_mode: bool) -> u32 {
```

AppContainer is the same mechanism that sandboxes UWP apps and Edge's renderers, so this is a well-understood boundary, not a new one. What MXC adds is the policy layer and the authoring tools around it: a Learning mode that enforces and records every denial to a JSON report, and a Permissive mode that records without blocking, so you can find out what a build actually touches before you lock it down. Intune management of MXC policy is "soon"; per the changelog, enterprise-managed settings that developers cannot weaken already work in Copilot.

Read the Command Line post's caveats before you trust it with everything. Shell commands and, by default, local MCP and language servers run inside the boundary. Copilot's built-in file tools do not: they run inside the Copilot process, and the harness checks their requests against policy, "but those checks aren't OS-enforced child-process isolation". Remote MCP servers are outside the sandbox altogether. A prompt-injected agent that wants to write outside the project has to get there through the file tool, and then the question is whether the harness's check is right. The attack surface is smaller than before. It is not zero.

<RepoCard repo="microsoft/mxc" />

## What I'd actually do with this

If you write agents on Windows, turn on Copilot's local sandbox now (`/sandbox` in the CLI, or "Sandbox new sessions" per project in the app). It is GA, costs nothing, and does not care which model you use. If you build your own agent, the MXC SDK's Learning mode is the cheapest way I know to find out what your tool calls really need.

If you want a local model in Copilot today, the supported route is Ollama plus CLI `1.0.94-0`, with any tool-calling model that fits your GPU. The Windows ML server is a reasonable endpoint to experiment with if you are already on Windows and want its loopback protections, but it is experimental, greedy-only, and CPU-only unless you build the CUDA backend yourself.

If you are thinking about buying an RTX Spark machine for local MAI Code 1.1 Flash, wait for two things: the end-of-October preview, so you can see what Auto actually routes locally, and a straight answer on memory. Everything Microsoft has published that is internally consistent says 53 GB of weights and 75.5 GB at 256K, which means the 128 GB configuration, not the 64 GB one. The 22 GB in the demo is the number I would not plan around.

The model is not open. "No inference charge" means Copilot does not meter calls that run on your hardware; you are still paying for Copilot, for the hardware, and for the electricity. The economics only work if the router sends enough routine work local without getting the hard cases wrong, and the router's rules are the one thing nobody has published.

## How I checked

I resolved both posts' short links and read the X threads and replies through the fxtwitter mirror. I read the Windows Experience, Windows Developer (MXC), Foundry on Windows, Microsoft AI and Command Line posts in full, GitHub's HydraFusion post and the two 7 October changelog entries, the Windows ML Runtime, concepts and Text Generation pages on Learn, and the June MAI-Code-1-Flash model card PDF. I shallow-cloned `microsoft/WindowsML` (`af42ad3`), `microsoft/mxc` (`e157fa4`), `microsoft/MAI-Code` (`506a383`) and `ggml-org/llama.cpp` (`71ad059`) and read the files quoted above; I did not build or run any of them. I searched the Hugging Face API for MAI models under the `microsoft` author. I extracted frames from the MicrosoftAI demo video with `ffmpeg` to read the install dialog. The weight sizes, task counts and prefill times are my own arithmetic from Microsoft's numbers; the decode figures are read off Microsoft's chart. I could not check the 6.8B active parameter count, the benchmark scores, the HydraFusion cost numbers or anything about the router's policy, because none of those come with artifacts.
