dstack
Open-source orchestration layer for running AI training and inference workloads across GPU clouds, Kubernetes, Slurm, VMs, and bare-metal clusters.
At a Glance
About dstack
dstack, made by dstack Inc., is an open-source orchestration layer for AI workloads on heterogeneous accelerators. It gives cloud tenants and data-center operators a unified control plane for provisioning compute and running development, training, and inference across GPU clouds, Kubernetes, VMs, and bare-metal. The docs list out-of-the-box support for NVIDIA, AMD, TPU, and Tenstorrent accelerators.
What It Is
dstack is a control plane for GPU provisioning and orchestration that sits between AI frameworks and models on top and compute infrastructure below. Users define configurations as YAML files in their repo and apply them with the dstack apply CLI command or a programmatic API. dstack then handles infrastructure provisioning and job scheduling, plus auto-scaling, port-forwarding, and ingress.
Core Concepts
- Fleets: provision and manage clusters across clouds, Kubernetes, and on-prem.
- Dev environments: launched for access from an IDE or by agents.
- Tasks: training, batch, or other jobs on a single node or across clusters.
- Services: model inference deployed as secure, scalable endpoints, with cache-aware and PD-disaggregated inference.
- Gateways: HTTPS, auto-scaling, domains, and rate limits.
- Presets: an experimental agent-driven inference optimization toolkit.
- Volumes: instance and network volumes for persisting data.
- Projects: tenant isolation and usage metering.
Bring Your Own Compute
SSH fleets connect VMs or bare-metal hosts using SSH credentials. Existing Kubernetes clusters connect through the Kubernetes backend, and existing Slurm clusters through an experimental Slurm backend that submits runs as Slurm jobs over SSH to the login node (requires Pyxis and enroot). dstack also integrates with major GPU clouds and provisions clusters in the user's own cloud account once backends are configured with credentials.
Deployment Options
dstack can be self-hosted as the open-source server (installed with uv, pip, or Docker, then started with dstack server). dstack Factory is described as a multi-tenant stack for AI labs and data centers, and dstack Sky is a hosted offering with unified access to GPU clouds. The site positions dstack as an alternative to, or a layer on top of, Kubernetes and Slurm.
Community Discussions
Be the first to start a conversation about dstack
Share your experience with dstack, ask questions, or help others learn from your insights.
Pricing
dstack (self-hosted open-source)
Open-source orchestration layer for AI workloads, self-hosted with your own compute.
- Install with uv, pip or Docker and run the dstack server
- Bring your own compute: GPU clouds, Kubernetes, VMs, bare-metal, SSH fleets
- Slurm backend (experimental)
- Fleets, tasks, services, gateways, presets, projects
Capabilities
Key Features
- Fleets for cluster provisioning and monitoring
- Dev environments accessible from IDEs or agents
- Tasks for training and batch jobs on single nodes or clusters
- Services for scalable model inference endpoints
- Gateways with HTTPS, auto-scaling, domains, and rate limits
- Presets for agent-driven inference optimization (experimental)
- Volumes for persistent data
- Projects with tenant isolation and usage metering
- SSH fleets for VMs and bare-metal
- YAML configuration applied via CLI or API
- Support for NVIDIA, AMD, TPU, and Tenstorrent accelerators
