# GPUStack

> Open-source GPU cluster manager for deploying, governing and scaling AI model serving and GPU instances on any hardware.

GPUStack is an open-source GPU cluster manager for AI model serving and on-demand GPU instance provisioning. It configures and orchestrates inference engines such as vLLM, SGLang and TensorRT-LLM across on-premise, Kubernetes and cloud GPU environments, and exposes models through OpenAI-compatible and Anthropic-compatible APIs. The latest release listed on GitHub is v2.2.3, and the project is licensed under Apache 2.0.

## What It Is

GPUStack is a management layer that sits on top of inference engines and GPU hardware. Teams connect GPU workers to a GPUStack server, deploy models from a catalog or from Hugging Face, ModelScope or local files, and serve them through standard APIs. It also launches SSH-accessible GPU instances for development, fine-tuning and interactive workloads, with Jupyter Notebook access and persistent storage.

## How the Workflow Runs

The server can run on a CPU-only machine via a single Docker command, and workers are added from the web UI by running a generated Docker command on each GPU node. Deployment includes automated compatibility checks, and GPUStack maps hardware to a matching inference engine version. Large models can be distributed across nodes and GPUs using tensor and pipeline parallelism. Performance modes (throughput, latency, standard, custom), KV cache extensions such as LMCache and HiCache, and speculative decoding methods like EAGLE3, MTP and N-grams are supported. The vendor reports tuned-deployment gains over unoptimized vLLM baselines in its Performance Lab.

## Hardware and Model Coverage

Supported accelerators include NVIDIA, AMD, Ascend NPU, Hygon DCU, Moore Threads, MetaX, Cambricon MLU, Iluvatar and T-Head PPU. Worker nodes run on Linux only; macOS is not supported for workers. Supported model types include LLM, multimodal, embedding, reranker, image, speech and OCR models.

## Governance and Enterprise Edition

The platform includes user authentication, API key management, token quotas and rate limits, usage analytics, metering, and Prometheus and Grafana monitoring. The Enterprise Edition adds control plane and model service high availability, multi-tenancy with organization isolation, RBAC with LDAP/OIDC/SAML SSO, audit logs, IP allow and block lists, resource topology view, billing reports and white-label branding.

## Features
- Multi-cluster GPU management across on-premise, Kubernetes and cloud
- Pluggable inference engines (vLLM, SGLang, TensorRT-LLM, llama.cpp, MindIE, custom)
- Automatic inference engine selection and compatibility checks
- Distributed inference across multiple nodes and GPUs
- OpenAI-compatible and Anthropic-compatible APIs
- Throughput, latency, standard and custom performance modes
- KV cache extensions and speculative decoding
- SSH-accessible GPU instances with Jupyter and persistent storage
- GPU partitioning and overcommit
- Load balancing, failover and traffic weight routing
- RBAC, API key scoping, token quotas and rate limits
- Usage metering and billing reports
- Prometheus and Grafana monitoring
- Resource topology view
- Enterprise HA, multi-tenancy, SSO, audit logs and white-label branding

## Integrations
vLLM, SGLang, TensorRT-LLM, llama.cpp, MindIE, Hugging Face, ModelScope, Docker, Podman, Kubernetes, Helm, Higress, Prometheus, Grafana, OpenWebUI, LangChain, n8n, Dify, RAGFlow, Claude Code, OpenClaw, OpenAI, Anthropic

## Platforms
WINDOWS, MACOS, LINUX, WEB, API, CLI

## Pricing
Open Source, Free tier available

## Version
v2.2.3

## Links
- Website: https://gpustack.ai
- Documentation: https://docs.gpustack.ai/
- Repository: https://github.com/gpustack/gpustack
- EveryDev.ai: https://www.everydev.ai/tools/gpustack
