EveryDev.ai
Subscribe
Home
Tools

4,218+ AI tools

  • New
  • Trending
  • Featured
  • Rate tools
  • Compare
  • Arena
Categories
  • Agents3186
  • Coding2210
  • Infrastructure966
  • Projects660
  • Marketing625
  • Research579
  • MCP518
  • Design497
  • Analytics490
  • Testing382
  • Security360
  • Data320
  • Integration242
  • Prompts239
  • Communication225
  • Extensions210
  • Voice191
  • Learning188
  • Commerce167
  • DevOps145
  • Web99
  • Finance34
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • Users
  • Rate Tools
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Tools
    3. GPUStack
    GPUStack icon

    GPUStack

    AI Infrastructure
    Featured

    Open-source GPU cluster manager for deploying, governing and scaling AI model serving and GPU instances on any hardware.

    Visit Website

    At a Glance

    Pricing
    Open Source
    Free tier available

    Self-hosted, fully open source GPUStack under the Apache 2.0 license.

    GPUStack Enterprise: Custom/contact

    Engagement

    Available On

    Windows
    macOS
    Linux
    Web
    API

    Resources

    WebsiteDocsGitHubllms.txt

    Topics

    AI InfrastructureLocal InferenceModel Management

    Alternatives

    OpenVINORamaLamacolibri
    Developer
    GPUStackShenzhen, ChinaEst. 2022

    Listed Oct 2026

    About GPUStack

    GPUStack is an open-source GPU cluster manager for AI model serving and on-demand GPU instance provisioning. It configures and orchestrates inference engines such as vLLM, SGLang and TensorRT-LLM across on-premise, Kubernetes and cloud GPU environments, and exposes models through OpenAI-compatible and Anthropic-compatible APIs. The latest release listed on GitHub is v2.2.3, and the project is licensed under Apache 2.0.

    What It Is

    GPUStack is a management layer that sits on top of inference engines and GPU hardware. Teams connect GPU workers to a GPUStack server, deploy models from a catalog or from Hugging Face, ModelScope or local files, and serve them through standard APIs. It also launches SSH-accessible GPU instances for development, fine-tuning and interactive workloads, with Jupyter Notebook access and persistent storage.

    How the Workflow Runs

    The server can run on a CPU-only machine via a single Docker command, and workers are added from the web UI by running a generated Docker command on each GPU node. Deployment includes automated compatibility checks, and GPUStack maps hardware to a matching inference engine version. Large models can be distributed across nodes and GPUs using tensor and pipeline parallelism. Performance modes (throughput, latency, standard, custom), KV cache extensions such as LMCache and HiCache, and speculative decoding methods like EAGLE3, MTP and N-grams are supported. The vendor reports tuned-deployment gains over unoptimized vLLM baselines in its Performance Lab.

    Hardware and Model Coverage

    Supported accelerators include NVIDIA, AMD, Ascend NPU, Hygon DCU, Moore Threads, MetaX, Cambricon MLU, Iluvatar and T-Head PPU. Worker nodes run on Linux only; macOS is not supported for workers. Supported model types include LLM, multimodal, embedding, reranker, image, speech and OCR models.

    Governance and Enterprise Edition

    The platform includes user authentication, API key management, token quotas and rate limits, usage analytics, metering, and Prometheus and Grafana monitoring. The Enterprise Edition adds control plane and model service high availability, multi-tenancy with organization isolation, RBAC with LDAP/OIDC/SAML SSO, audit logs, IP allow and block lists, resource topology view, billing reports and white-label branding.

    GPUStack - 1

    Community Discussions

    Be the first to start a conversation about GPUStack

    Share your experience with GPUStack, ask questions, or help others learn from your insights.

    Pricing

    OPEN SOURCE

    GPUStack Open Source

    Self-hosted, fully open source GPUStack under the Apache 2.0 license.

    • Apache License 2.0
    • Multi-cluster GPU management
    • Pluggable inference engines (vLLM, SGLang, TensorRT-LLM)
    • OpenAI-compatible API
    • Built-in user authentication and access control

    GPUStack Enterprise

    Enterprise edition with HA, multi-tenancy, governance and white-label branding. Contact sales.

    Custom
    contact sales
    • Control plane HA
    • Model service HA
    • Multi-tenancy with organization isolation
    • RBAC & SSO (LDAP, OIDC, SAML)
    • Audit logs and IP access control
    • Token quotas and rate limits
    • Resource topology view
    • Token and GPU time billing
    • White-label branding
    View official pricing

    Capabilities

    Key Features

    • Multi-cluster GPU management across on-premise, Kubernetes and cloud
    • Pluggable inference engines (vLLM, SGLang, TensorRT-LLM, llama.cpp, MindIE, custom)
    • Automatic inference engine selection and compatibility checks
    • Distributed inference across multiple nodes and GPUs
    • OpenAI-compatible and Anthropic-compatible APIs
    • Throughput, latency, standard and custom performance modes
    • KV cache extensions and speculative decoding
    • SSH-accessible GPU instances with Jupyter and persistent storage
    • GPU partitioning and overcommit
    • Load balancing, failover and traffic weight routing
    • RBAC, API key scoping, token quotas and rate limits
    • Usage metering and billing reports
    • Prometheus and Grafana monitoring
    • Resource topology view
    • Enterprise HA, multi-tenancy, SSO, audit logs and white-label branding

    Integrations

    vLLM
    SGLang
    TensorRT-LLM
    llama.cpp
    MindIE
    Hugging Face
    ModelScope
    Docker
    Podman
    Kubernetes
    Helm
    Higress
    Prometheus
    Grafana
    OpenWebUI
    LangChain
    n8n
    Dify
    RAGFlow
    Claude Code
    OpenClaw
    OpenAI
    Anthropic
    API Available
    View Docs

    Ratings & Reviews

    No ratings yet

    Be the first to rate GPUStack and help others make informed decisions.

    Rate other tools you’ve used

    Developer

    GPUStack Team

    GPU cluster manager for optimized AI model deployment

    Founded 2022
    Shenzhen, China
    Read more about GPUStack Team
    WebsiteGitHubLinkedInX / Twitter
    1 tool in directory

    Similar Tools

    OpenVINO icon

    OpenVINO

    Open-source toolkit by Intel for optimizing and deploying deep learning models across CPU, GPU, and NPU hardware targets.

    RamaLama icon

    RamaLama

    An open-source CLI tool that simplifies running and serving AI models locally using OCI containers, with automatic GPU detection and multi-registry support.

    colibri icon

    colibri

    An open-source pure-C inference engine that streams Mixture-of-Experts weights from disk, enabling frontier models up to 2.8 trillion parameters to run on consumer hardware with zero dependencies.

    Browse all tools

    Related Topics

    AI Infrastructure

    Infrastructure designed for deploying and running AI models.

    421 tools

    Local Inference

    Tools and platforms for running AI inference locally without cloud dependence.

    231 tools

    Model Management

    Tools for managing, versioning, and deploying AI models.

    65 tools
    Browse all topics
    Back to all toolsSuggest an edit
    ratings
    discussions