News

Microsoft releases MAI Code1

2 min read
Microsoft releases MAI Code1.1 Flash model, GitHub Copilot will support local and cloud hybrid inference. Microsoft announced at the Windows and Surface launch conference on October 7, 2026 that it will introduce local AI model support for GitHub Copilot by the end of this month, allowing developers to freely switch between cloud and device-side models or implement automatic scheduling. To address the memory bottlenecks and long context resource occupancy issues in client-side inference, Microsoft launched the efficient MAI Code1.1 Flash Hybrid Expert (MoE) model. This model has 137 billion total parameters and 6.8 billion activation parameters. It combines quantification and speculative decoding technology to significantly reduce memory usage while improving device-side encoding response speed. It is the first to be installed on the new Surface Laptop Ultra equipped with NVIDIA RTX Spark hardware. In terms of specific applications, users can choose the inference mode independently through GitHub Copilot CLI, Copilot application and Visual Studio Code. They can rely on the Copilot background to automatically coordinate resources, or specify to call local models such as MAI Code1.1Flash through Windows ML providers or local endpoints compatible with OpenAI. In terms of security, the system integrates the Microsoft Execution Container (MXC) built by the Windows team and uses the native container isolation mechanism of each operating system to ensure the safety of code running. This upgrade marks that the AI ​​coding assistant is accelerating its evolution from pure cloud dependence to a hybrid architecture of "cloud + device-side" collaboration. This move not only simplifies the deployment threshold of local models, but also provides a more efficient implementation solution for high-privacy, low-latency development scenarios. via AI News (author: AI Base)