DeepSeekDSH
Independent community guideNot affiliated with DeepSeek.Download versions and guide references

Local DeepSeek Deployment with Ollama and Client Setup Guide

Deploy DeepSeek models on local Apple Silicon Macs, Windows, or Linux hardware using Ollama, exposing a private OpenAI-compatible API endpoint.

On this page

1. VRAM and hardware requirements

Choose an appropriate quantized parameter tier based on your available GPU VRAM or unified memory:

Model TagParametersMinimum MemoryRecommended Hardware
deepseek-r1:1.5b1.5B4 GBStandard Laptops / Raspberry Pi
deepseek-r1:7b / 8b7B ~ 8B8 GBRTX 3060 / 4060 / Mac 8GB+
deepseek-r1:14b14B12 GBRTX 4070 / Mac 16GB+
deepseek-r1:32b32B20 GBRTX 3090 / 4090 / Mac 24GB+
deepseek-r1:70b70B48 GBDual RTX 3090 / Mac 64GB+
STEP 2

Install Ollama and run model

  1. Download and install Ollama from ollama.com.
  2. Execute this command in your terminal to pull and run the model:
ollama run deepseek-r1:7b

Ollama provides an OpenAI-compatible API listening by default on http://127.0.0.1:11434.

STEP 3

Configure network access (Optional)

To allow external web clients or LAN devices to reach your local Ollama server, set environment variables:

# macOS / Linux export OLLAMA_HOST=0.0.0.0:11434 export OLLAMA_ORIGINS=* ollama serve
STEP 4

Connect tools and coding assistants

In any tool supporting custom OpenAI providers (such as Continue, Cherry Studio, or DSH Custom Provider), set:

  • API Base URL: http://localhost:11434/v1
  • API Key: Any placeholder (e.g. ollama)
  • Model Name: Your local model tag (e.g. deepseek-r1:7b)

5. Troubleshooting

Out of memory (OOM) errors during model load

Your GPU VRAM cannot fit the selected model. Switch to a smaller model size (such as 7b or 1.5b).

Slow generation speeds (Low tokens/sec)

Ensure layers are running on the GPU rather than being offloaded to CPU. Run ollama ps to check GPU offloading status.

References