Local DeepSeek Deployment with Ollama and Client Setup Guide
Deploy DeepSeek models on local Apple Silicon Macs, Windows, or Linux hardware using Ollama, exposing a private OpenAI-compatible API endpoint.
On this page
1. VRAM and hardware requirements
Choose an appropriate quantized parameter tier based on your available GPU VRAM or unified memory:
| Model Tag | Parameters | Minimum Memory | Recommended Hardware |
|---|---|---|---|
deepseek-r1:1.5b | 1.5B | 4 GB | Standard Laptops / Raspberry Pi |
deepseek-r1:7b / 8b | 7B ~ 8B | 8 GB | RTX 3060 / 4060 / Mac 8GB+ |
deepseek-r1:14b | 14B | 12 GB | RTX 4070 / Mac 16GB+ |
deepseek-r1:32b | 32B | 20 GB | RTX 3090 / 4090 / Mac 24GB+ |
deepseek-r1:70b | 70B | 48 GB | Dual RTX 3090 / Mac 64GB+ |
Install Ollama and run model
- Download and install Ollama from ollama.com.
- Execute this command in your terminal to pull and run the model:
ollama run deepseek-r1:7bOllama provides an OpenAI-compatible API listening by default on http://127.0.0.1:11434.
Configure network access (Optional)
To allow external web clients or LAN devices to reach your local Ollama server, set environment variables:
# macOS / Linux
export OLLAMA_HOST=0.0.0.0:11434
export OLLAMA_ORIGINS=*
ollama serveConnect tools and coding assistants
In any tool supporting custom OpenAI providers (such as Continue, Cherry Studio, or DSH Custom Provider), set:
- API Base URL:
http://localhost:11434/v1 - API Key: Any placeholder (e.g.
ollama) - Model Name: Your local model tag (e.g.
deepseek-r1:7b)
5. Troubleshooting
Out of memory (OOM) errors during model load
Your GPU VRAM cannot fit the selected model. Switch to a smaller model size (such as 7b or 1.5b).
Slow generation speeds (Low tokens/sec)
Ensure layers are running on the GPU rather than being offloaded to CPU. Run ollama ps to check GPU offloading status.