GitHub – antirez/ds4: DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm · GitHub
Local Deepseek-v4-Flash, recommend 128GB memory
Local Deepseek-v4-Flash, recommend 128GB memory
Run GLM 5.2 with only 24GB unified memory. Downside is 0.1 – 1 token per second, so super slow. Could be useful for scheduled overnight reasoning tasks.
Recommended model for DGX Spark