Run DeepSeek Locally with LM Studio GUI and Local API Server
LM Studio provides a clean graphical desktop experience for local LLMs. Discover, download, and execute quantized DeepSeek GGUF models and expose a local OpenAI-compatible API endpoint.
On this page
Hardware & VRAM recommendations
- 1.5B ~ 7B Distilled models: 8GB RAM / 4GB+ VRAM (Standard laptops / Apple Silicon Macs)
- 8B ~ 14B Distilled models: 16GB RAM / 8GB~12GB VRAM (RTX 3060 / 4060 / Mac 16GB+)
- 32B Distilled models: 32GB RAM / 20GB+ VRAM (RTX 3090 / 4090 / Mac 24GB+)
- 70B Distilled models: 64GB unified memory Mac or dual RTX 3090/4090
STEP 1
Download LM Studio and search DeepSeek models
- Download the installer for your OS (Windows, macOS, or Linux) from lmstudio.ai.
- Open LM Studio and click the Discover (Magnifying Glass 🔍) icon in the left navigation bar.
- Search for
deepseek-r1ordeepseek-v3. - Choose a verified community GGUF quantization (such as packages provided by
unslothorbartowski). For general use, Q4_K_M offers the best balance of speed and reasoning quality.
STEP 2
Enable GPU acceleration and layer offloading
Ensure your hardware handles inference computation effectively:
- macOS (Apple Silicon): Metal acceleration is enabled automatically across unified memory.
- Windows (NVIDIA GPU): Enable GPU Acceleration (CUDA) in the right sidebar and set the GPU Offload Layers slider to maximum.
STEP 3
Launch the local OpenAI-compatible API server
Allow other software tools to query your local DeepSeek model:
- Click the Developer / Local Server (⇄) icon in the left sidebar.
- Select the downloaded DeepSeek model from the top dropdown.
- Click Start Server.
- The default local endpoint is:
http://localhost:1234/v1.
STEP 4
Connect Cursor, VS Code, and other clients
In Cursor, Continue, Cherry Studio, or Zotero, configure a custom provider:
- Base URL:
http://localhost:1234/v1 - API Key: Any placeholder (e.g.
lm-studio) - Model Name: The loaded model identifier
# Verify local server connection via curl
curl http://localhost:1234/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "Hello, verify connection."}]}'Troubleshooting
CUDA out of memory error
The selected parameter size exceeds available physical VRAM. Reduce offload layers or switch to a lighter quantization format (such as Q4_K_M instead of Q8, or 7B instead of 14B).
Connection Refused from other devices on LAN
In Local Server settings, change the listening host from 127.0.0.1 to 0.0.0.0 and enable CORS.