DwarfStar 4 (ds4)
A narrow C inference engine that runs frontier open-weight models locally on high-memory Mac, CUDA and ROCm machines.
At a Glance
Self-hosted DwarfStar 4 (ds4) local inference engine released under the MIT license; no paid tiers are listed.
Engagement
Available On
Listed Oct 2026
About DwarfStar 4 (ds4)
DwarfStar 4 (ds4) is a local inference engine created by Salvatore Sanfilippo (antirez), the creator of Redis. It runs a small set of open-weight model families, including DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next, on high-memory Apple Silicon, NVIDIA CUDA and AMD ROCm machines. The dwarfstar.sh site is a community-maintained resource with docs, hardware guidance and benchmarks around the MIT-licensed engine.
What It Is
ds4 is a C inference engine that is deliberately narrow rather than a generic GGUF runner. It targets project-specific GGUF layouts that are validated end to end against official model outputs. Models use asymmetric 2-bit quantization on the routed experts while keeping critical shared paths precise, which is how the supported builds fit their target machines.
How the Stack Works
One engine exposes three interfaces: ./ds4 for interactive chat, ./ds4-server for local APIs, and ./ds4-agent for persistent coding sessions. The KV cache can be saved to SSD and resumed by prompt hash, so restarts do not require a full re-prefill. The site lists SSD streaming, tensor parallelism, session batching, DSPARK + MTP and vision input among the capabilities.
Agent and API Connectivity
ds4-server speaks OpenAI-style and Anthropic-style APIs, with endpoints such as /v1/chat/completions, /v1/messages and /v1/responses. The docs describe connecting OpenCode, Claude Code, Codex CLI and Pi to the local server via a base URL.
Setup Path
Users clone the repository, download a project GGUF with download_model.sh, and build for their backend (for example make for macOS Metal, or make cuda-spark for DGX Spark). Hardware classes listed include Apple Silicon Macs with 64 GB+, NVIDIA DGX Spark or generic CUDA Linux boxes, and AMD Strix Halo systems. The site publishes benchmark rows, such as 790.2 t/s prefill and 39.4 t/s generation for q2 at 2,048 tokens on an M5 Max with 128 GB.
Project Approach
The About page says upstream treats the project as a working template that users adapt with coding agents, and that there are deliberately no GitHub releases or tags. Upstream credits llama.cpp and GGML for kernels and quantization formats.
Community Discussions
Be the first to start a conversation about DwarfStar 4 (ds4)
Share your experience with DwarfStar 4 (ds4), ask questions, or help others learn from your insights.
Pricing
Open Source (MIT)
Self-hosted DwarfStar 4 (ds4) local inference engine released under the MIT license; no paid tiers are listed.
- Local C inference engine for Metal, CUDA and ROCm
- Supports DeepSeek V4 / V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next
- CLI (./ds4), local server (./ds4-server) and native agent (./ds4-agent)
- OpenAI and Anthropic-style APIs
- SSD KV cache, tensor parallelism, session batching, vision input
Capabilities
Key Features
- Asymmetric 2-bit quantization of routed experts
- KV cache persisted to SSD and resumed by prompt hash
- Interactive CLI (ds4)
- Local server with OpenAI and Anthropic-style APIs (ds4-server)
- Native persistent coding agent (ds4-agent)
- SSD streaming
- Tensor parallelism
- Session batching
- DSPARK + MTP speculative decoding
- Vision input
- Metal, CUDA and ROCm backends
