GPUStack
GPUStack provides an open-source and enterprise control plane for deploying, governing, and scaling AI models and GPU compute across on-premises, cloud, hybrid, and Kubernetes environments. It unifies GPU cluster management, inference-engine orchestration, model serving, GPU-as-a-Service, observability, access control, and usage metering.
At a Glance
- Development teams
- IT organizations
- Enterprise AI and machine-learning teams
- AI infrastructure and platform teams
- +3 more
AI Tools by GPUStack
(1)GPUStack
Open Source GPU Cluster Manager
Discussions
No discussions yet
Be the first to start a discussion about GPUStack
Latest News
Day 0 Benchmark: Deploying DeepSeek-V4-Flash-DSpark on GPUStack Doubles Throughput
Day 0 Deployment of GLM-5.2-FP8-DSpark on GPUStack: Benchmarking Speculative Decoding
GPUStack Usage brings token, compute, storage, and cost visibility to shared AI infrastructure
Introducing GPUStack — Turn Any GPU Into a Token Factory
Products & Services
Apache-2.0-licensed GPU cluster manager for AI model serving and GPU-instance provisioning. It manages heterogeneous clusters, schedules GPU workloads, configures inference engines, and exposes standard model APIs.
Major platform release focused on tuned inference performance, extended KV cache and speculative decoding support, pluggable engines, multi-cluster operations, and enterprise-grade controls.
Platform release adding T-Head PPU support, vLLM-Omni, unified access to public and private model providers, backend marketplace capabilities, improved operations, and offline installation.
Commercial enterprise edition adding high availability across management and business planes, multi-tenancy and resource isolation, fine-grained RBAC, LDAP/OIDC/SAML SSO, audit and network controls, white-label branding, and dedicated support.
Market Position
GPUStack positions itself as a vendor-neutral AI infrastructure control plane rather than a single inference engine: it manages heterogeneous accelerators, multiple engines, model gateways, GPU instances, governance, and usage metering in one platform. Its competitive set includes DIY Kubernetes plus KServe/Ray Serve, NVIDIA-oriented cluster-management platforms such as Run:ai and Base Command, and standalone serving engines such as vLLM, SGLang, and Triton. GPUStack differentiates through cross-vendor accelerator coverage, automatic engine/configuration selection, day-zero backend versioning, and a combined MaaS/GPUaaS operating model.
Leadership
Founders
George Qin
Co-founder and CEO of GPUStack.ai. His public profile describes more than 10 years working on open-source and cloud solutions; the founding team is described as including former Rancher Labs founders and lead employees.
Peng Jiang
Co-founder of GPUStack.ai. Public professional-profile information lists previous experience at SUSE, Rancher Labs, Microsoft, and Citrix.
Executive Team
George Qin
Co-founder and CEO
Open-source and cloud-solutions professional with more than 10 years of experience according to his public professional profile; part of the former Rancher Labs founding/lead team.
Peng Jiang
Co-founder
Infrastructure and open-source professional with previous roles or experience at SUSE, Rancher Labs, Microsoft, and Citrix.
Founding Story
GPUStack was established in 2022 by a core team with roots in open source, cloud computing, and infrastructure, including former Rancher Labs founders and lead employees. The initial vision was to bring enterprise-grade infrastructure operations to AI workloads: abstract heterogeneous GPUs and inference engines behind one control plane so teams could deploy and operate models without manually stitching together engines, scripts, dashboards, and load balancers.
Business Model
Revenue Model
The core GPUStack software is open source under Apache 2.0. GPUStack monetizes the enterprise edition through commercial licensing, enterprise support, and sales-led architecture/deployment guidance; the platform also provides usage metering and billing capabilities for operators running MaaS or GPUaaS services.
Target Markets
- Development teams
- IT organizations
- Enterprise AI and machine-learning teams
- AI infrastructure and platform teams
- GPU cloud and AI service providers
- Organizations operating on-premises, cloud, or hybrid GPU fleets
- Production LLM, multimodal, embedding, voice, image, and video model serving
- Enterprise Model-as-a-Service platforms
- GPU-as-a-Service and GPU cloud-provider operations
- On-premises and hybrid AI infrastructure
- Multi-GPU and multi-node distributed inference
- Serving newly released models on day zero