iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
To use Ollama models in Visual Studio Code, install the official Ollama extension, make sure Ollama is running with a model available, then choose that model in VS Code Chat. The extension is the recommended route: Microsoft marks VS Code’s built-in Ollama provider as deprecated and directs users to the official extension.
What you need before connecting Ollama to VS Code
- Visual Studio Code 1.127 or newer.
- Ollama installed and running.
- At least one local or cloud model available through Ollama.
Ollama 0.17.6 or newer is recommended for cloud-model sign-in and richer model metadata. Older Ollama versions may still work with local models, according to the official extension documentation.
Install the official Ollama extension and select a model
- In VS Code, open Extensions, search for Ollama, and install the official Ollama extension from the Visual Studio Code Marketplace.
- Start Ollama and ensure a model is available. The documentation gives
ollama pull qwen3.6as an example for a local model; it is an example, not a recommendation for every system or coding task. - Open Chat in VS Code and open the model picker at the bottom of the chat input.
- Choose a model listed under the Ollama section, then send a prompt in Chat.
By default, the extension discovers models through Ollama’s local endpoint, http://127.0.0.1:11434. The extension documentation and Ollama’s VS Code integration guide describe this setup.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Choose between a local model and a cloud model
Ollama supports both local and cloud models in the VS Code integration, but their setup differs. The sources do not establish a general privacy, cost, or performance advantage for either option.
#1 Best Overall
| Model type | Setup detail | Sign-in |
|---|---|---|
| Local | Make the model available in Ollama; the documentation’s example is ollama pull qwen3.6. |
Not required. Ollama states, “Local models do not require sign-in.” |
| Cloud | The documentation’s example is ollama pull kimi-k2.6:cloud. |
Run ollama signin if authentication is requested. |
These commands are the examples shown in Ollama’s integration documentation; the best model depends on your needs and is not determined by those examples.
Fix Ollama models missing from the VS Code picker
- Confirm Ollama is running and check that it has models available by running
ollama list. - In VS Code, open the Command Palette and run
Ollama: Refresh Models. - If the model still does not appear, run
Ollama: Diagnose Modelsand inspect the Ollama output channel for details. - If a cloud model requests authentication, run
ollama signin, then refresh the model list.
These are the troubleshooting steps documented by Ollama in its extension guide and VS Code integration instructions.
Rank #2
Address context-length needs for local models
VS Code may display a model’s maximum supported context length even when Ollama allocates a smaller context at runtime. For the documented local-model flow, Ollama instructs users to open Ollama Settings, set the context length to at least 64k, reload the VS Code window, and resend the prompt. This setting addresses context needs; it does not guarantee that every model or device will perform well at that size. See the Ollama integration guide for the procedure.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteWhy older built-in Ollama setup instructions may not apply
Microsoft’s current VS Code language-model documentation marks the built-in Ollama provider as deprecated and points users to the official Ollama extension for local models. If an older guide tells you to configure that built-in provider, use the extension-based steps above instead.
Rank #3
What the official setup guidance does not establish
The reviewed Ollama and Microsoft integration documentation does not specify recommended CPU, GPU, memory, or storage configurations, or publish comparative local-inference benchmarks. It is therefore not enough to identify suitable hardware or predict how quickly a particular model will run.
Quick Recap
Rank #4
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

