# dstack

> Open-source orchestration layer for running AI training and inference workloads across GPU clouds, Kubernetes, Slurm, VMs, and bare-metal clusters.

dstack, made by dstack Inc., is an open-source orchestration layer for AI workloads on heterogeneous accelerators. It gives cloud tenants and data-center operators a unified control plane for provisioning compute and running development, training, and inference across GPU clouds, Kubernetes, VMs, and bare-metal. The docs list out-of-the-box support for NVIDIA, AMD, TPU, and Tenstorrent accelerators.

## What It Is

dstack is a control plane for GPU provisioning and orchestration that sits between AI frameworks and models on top and compute infrastructure below. Users define configurations as YAML files in their repo and apply them with the `dstack apply` CLI command or a programmatic API. dstack then handles infrastructure provisioning and job scheduling, plus auto-scaling, port-forwarding, and ingress.

## Core Concepts

- **Fleets**: provision and manage clusters across clouds, Kubernetes, and on-prem.
- **Dev environments**: launched for access from an IDE or by agents.
- **Tasks**: training, batch, or other jobs on a single node or across clusters.
- **Services**: model inference deployed as secure, scalable endpoints, with cache-aware and PD-disaggregated inference.
- **Gateways**: HTTPS, auto-scaling, domains, and rate limits.
- **Presets**: an experimental agent-driven inference optimization toolkit.
- **Volumes**: instance and network volumes for persisting data.
- **Projects**: tenant isolation and usage metering.

## Bring Your Own Compute

SSH fleets connect VMs or bare-metal hosts using SSH credentials. Existing Kubernetes clusters connect through the Kubernetes backend, and existing Slurm clusters through an experimental Slurm backend that submits runs as Slurm jobs over SSH to the login node (requires Pyxis and enroot). dstack also integrates with major GPU clouds and provisions clusters in the user's own cloud account once backends are configured with credentials.

## Deployment Options

dstack can be self-hosted as the open-source server (installed with uv, pip, or Docker, then started with `dstack server`). dstack Factory is described as a multi-tenant stack for AI labs and data centers, and dstack Sky is a hosted offering with unified access to GPU clouds. The site positions dstack as an alternative to, or a layer on top of, Kubernetes and Slurm.

## Features
- Fleets for cluster provisioning and monitoring
- Dev environments accessible from IDEs or agents
- Tasks for training and batch jobs on single nodes or clusters
- Services for scalable model inference endpoints
- Gateways with HTTPS, auto-scaling, domains, and rate limits
- Presets for agent-driven inference optimization (experimental)
- Volumes for persistent data
- Projects with tenant isolation and usage metering
- SSH fleets for VMs and bare-metal
- YAML configuration applied via CLI or API
- Support for NVIDIA, AMD, TPU, and Tenstorrent accelerators

## Integrations
Kubernetes, Slurm, SSH fleets, GPU clouds, Docker, SGLang

## Platforms
CLI, API, LINUX, MACOS

## Pricing
Open Source

## Links
- Website: https://dstack.ai
- Documentation: https://dstack.ai/docs
- Repository: https://github.com/dstackai/dstack
- EveryDev.ai: https://www.everydev.ai/tools/dstack
