Microsoft launches on-device coding AI model with new Surface PCs
Microsoft said its locally optimized MAI-Code-1.1-Flash shrinks to 53GB and runs through the new MXC sandbox in Windows, alongside an NVIDIA RTX Spark-based Surface Laptop Ultra.
Microsoft announced an on-device optimized version of its coding AI model, MAI-Code-1.1-Flash, alongside the launch of next-generation PCs including the NVIDIA RTX Spark-based Surface Laptop Ultra, according to aitimes.com.
The optimized model applies quantization and speculative decoding to compress its weights to about 3 bits on average, reducing its size to 53GB — roughly an 80% reduction compared with the BF16 cloud original.
On the SWE-bench Verified test, the locally optimized version scored 70.8%, compared with 72.6% for the cloud original. On Terminal-Bench 2.1, the local version scored 66.3%, above the cloud original's 62.9%.
Microsoft also built the Microsoft Execution Container (MXC), a local sandbox in Windows designed so local AI agents can run safely on the machine. GitHub Copilot can offload tasks to the local model to reduce compute costs, and the Code in Copilot feature lets users build custom software on desktop without external cloud token consumption or app store restrictions.
Satya Nadella said on X that Windows now opens a new chapter by letting agents safely act on behalf of users on any PC.
Earlier model release
In August, Microsoft had already introduced the MAI-Code-1.1-Flash model.
Quick answers
How large is the on-device version of MAI-Code-1.1-Flash?
It is 53GB, about an 80% reduction from the BF16 cloud original, achieved through quantization and speculative decoding that compress weights to about 3 bits on average.
What is the Microsoft Execution Container?
It is a local sandbox built into Windows so local AI agents can run safely.
How does GitHub Copilot use the local model?
GitHub Copilot can offload tasks to the local model to reduce compute costs.