News

GitHub Copilot is no longer a pure cloud AI model: Surface Laptop Ultra local inference throughput reaches up to 63 words per second

2 min read
Source: ithome.com
IT House reported on October 8 that at a press conference held in San Francisco, USA, at 10 a.m. Pacific time on October 7 (1 a.m. on October 8, Beijing time), Microsoft announced that GitHub Copilot will support local AI model inference by the end of this month. Developers can automatically or manually switch between cloud models and device-side models. GitHub Copilot is an AI programming assistant launched by Microsoft that can be integrated with Visual Studio Code, CLI and other tools to generate completion, explanation or refactoring suggestions based on the code context. GitHub Copilot currently relies on a cloud-hosted model, with a coordinator routing requests based on performance, cost, and accuracy. Microsoft plans to expand this mechanism by the end of this month to allow Copilot to automatically select cloud models and local AI models and coordinate inference tasks in the background. Microsoft offers two modes for GitHub Copilot users: automatic orchestration or forcing the use of device-side models. Users can set preferences in the GitHub Copilot CLI, Copilot apps, and Visual Studio Code. Native model selection supports specifying a provider, model, or endpoint, and developers can select MAI Code 1.1 Flash through Windows ML, or connect to a OpenAI-compliant local endpoint and choose the model it exposes. Microsoft launched the MAI Code 1.1 Flash "Hybrid Expert" model, with a total of 137 billion parameters and 6.8 billion active parameters. It uses quantification and speculative decoding technology to optimize response speed and memory usage. The actual measurement of this model on Surface Laptop Ultra shows that the peak memory usage, token utilization and processing speed in complex tasks are significantly improved. The decoding throughput test shows that when the prompt length is from 2K to 256K Token, the throughput is between 4 per second