Microsoft Lets Developers Run MAI-Code-1.1-Flash Locally, but It Needs 120GB of RAM


MAI Code-1.1-Flash
Image credit: Microsoft

Microsoft released MAI-Code-1.1-Flash last month, and the company has now made MAI-Code-1.1-Flash available to download and run locally.

The move expands Microsoft’s push toward local AI, following its work on hybrid intelligence in Windows with local models.

MAI-Code-1.1-Flash keeps its 256K context window locally

The local version uses 3-bit quantization to reduce its memory requirements while keeping the model’s full 256K context window.

Microsoft says the quantized version delivers coding performance comparable to the full-precision model on SWE-Bench Verified and Terminal-Bench 2.1.

This gives developers access to the model’s large context window while lowering the hardware requirements compared with running the full-precision version.

It still requires more than 120GB of memory

Those reduced requirements do not make MAI-Code-1.1-Flash practical for most PCs.

Microsoft recommends more than 120GB of RAM for the best local performance, which limits the model to high-end workstations and AI-focused systems.

That makes hardware such as the new Microsoft Surface Laptop Ultra a more realistic target for this type of workload. Microsoft has already opened pre-orders for the Surface Laptop Ultra.

GitHub Copilot is getting offline local AI support

MAI-Code-1.1-Flash can also handle compatible GitHub Copilot coding tasks entirely on-device.

Microsoft plans to bring experimental local-model support to GitHub Copilot by the end of October across the Copilot app, Copilot CLI, and Visual Studio Code.

Developers will still have the option to switch to cloud-hosted models when they need more processing power or want to handle more demanding workloads.

MAI-Code-1.1-Flash is faster than MAI-Code-1.0

Microsoft says it already uses MAI-Code-1.1-Flash in production within GitHub Copilot.

Compared with MAI-Code-1.0, the company claims the newer model streams tokens 25% faster while using 25% fewer tokens to complete a task.

The combination of local execution, lower token usage, and faster output gives Microsoft another way to move coding AI workloads from the cloud to high-end developer hardware.

More about the topics: AI, microsoft

Readers help support Windows Report. We may get a commission if you buy through our links. Tooltip Icon

Read our disclosure page to find out how can you help Windows Report sustain the editorial team. Read more

User forum

0 messages