Running an AI model on your own computer sounds technical, but it has become about as hard as installing a game. It is free, it works without an internet connection once the model is downloaded, and nothing you type leaves your machine. Here is how to do it, step by step.
1. Check what your computer has
What matters most is memory. On a PC, it is the memory of the graphics card (VRAM): on Windows, open the Task Manager, then Performance, then GPU, and read "Dedicated GPU memory". 8 GB is enough for small models, 12 to 16 GB for good mid-size ones, 24 GB for the best that fit at home. On a Mac with an Apple chip (M1 or later), the memory is shared and about 70% of it can go to the model: 16 GB is fine for small models, 32 GB or more opens the mid-size ones. Without a graphics card, small models still run on the processor, only slowly.
The Can my PC run it? tool does the matching for you: pick your graphics card or Mac and it lists the models that fit.
2. Pick an app
- LM Studio is the easiest: a normal app for Windows, Mac and Linux, with a chat window and a built-in model search. Choose it if you have never used a terminal.
- Ollama is a small program you use from the command line, or through other apps built on top of it. Choose it if you want to connect the model to other tools later.
Both are free to download and get their models from the same public sources. You can install both and decide later.
3. Download a model
In LM Studio, open the Discover tab (the magnifying glass), type a model name and pick a version: the app warns you when a file is likely too big for your memory. In Ollama, one line does it: ollama run followed by the model’s name, for example ollama run qwen3, downloads the model and starts a chat in the same window.
4. Choose the right size
Most models come in several sizes, counted in billions of parameters (the "B" in names like 8B or 27B), and in several compression levels. A good default is the version marked Q4_K_M: compressed to about 4 bits, it needs roughly 0.6 GB of memory per billion parameters and loses very little quality. With a normal conversation, an 8B model then needs about 6 GB and a 27B model about 18 GB. If a model does not fit, pick a smaller size before a stronger compression: below 4 bits, quality drops faster.
5. Start chatting, and know the limits
Once loaded, the model works like any chat assistant. Expect it to be slower than the big online services, especially with long documents, and somewhat less capable: the best AI for your graphics card shows how close home models come to the best ones. A local model also knows nothing about recent events unless you paste the text in, and most setups cannot search the web.
6. Use it from your other apps (optional)
Both apps can run a small local server that answers in the same format as the OpenAI API. Start it from the Developer tab in LM Studio; Ollama runs it in the background on its own. Any tool that lets you change the API address can then use your local model instead of a paid service, with no cost per token.
Good to know
- Models are big files, from 2 to 5 GB for small ones to 20 GB or more. Check your free disk space first.
- Download models through the app or from the publisher’s official page, not from random links.
- Check the licence before using a model for work: most open models allow commercial use, some with conditions.