KubeAI
Open-source Kubernetes operator for deploying and autoscaling LLM, embedding, reranking, and speech-to-text models behind an OpenAI-compatible API.
At a Glance
Self-hosted Apache-2.0 licensed AI inference operator for Kubernetes.
Engagement
Available On
Listed Oct 2026
About KubeAI
KubeAI is an open-source AI inference operator for Kubernetes, hosted by the kubeai-project organization on GitHub under the Apache 2.0 license. It deploys and scales machine learning models, including LLMs, embeddings, reranking and speech-to-text, and exposes them through OpenAI-compatible endpoints. The latest listed release is v0.23.5.
What It Is
KubeAI is a Kubernetes-native serving layer for ML models. It operates vLLM and Ollama servers for text generation, FasterWhisper for audio transcription, and Infinity for vector embeddings, and it supports reranking with cross-encoder models. Models are declared through a Model custom resource, and the project ships a catalog of popular models preconfigured for common GPU types.
Architecture and Routing
KubeAI has two main components: a model proxy and a model operator. The proxy provides the OpenAI-compatible API, queues requests while a model scales from zero, retries requests that hit bad backends, and applies a prefix-aware load balancing strategy meant to improve KV cache utilization across vLLM replicas. The operator manages backend server Pods directly, automating model downloads, volume mounts, and loading of dynamic LoRA adapters. Both components are co-located in one deployment.
Compatibility and Setup Path
The API supports /v1/chat/completions, /v1/completions, /v1/embeddings, /v1/rerank, /v1/models and /v1/audio/transcriptions, so existing OpenAI client libraries can be used. The project states it does not require Istio, Knative or a Prometheus metrics adapter, and that it runs on CPU, GPU or TPU. It installs via Helm on any Kubernetes cluster, with guides for AKS, EKS and GKE, and a local quickstart using kind or minikube with a bundled chat UI. Documentation also covers autoscaling, model caching with AWS EFS or GCP Filestore, loading models from OCI images or PVCs, LoRA adapters, multitenancy, and Prometheus observability.
Community Discussions
Be the first to start a conversation about KubeAI
Share your experience with KubeAI, ask questions, or help others learn from your insights.
Pricing
Open Source
Self-hosted Apache-2.0 licensed AI inference operator for Kubernetes.
- Deploy and scale machine learning models on Kubernetes
- Supports VLMs, LLMs, embeddings, and speech-to-text
- OpenAI API compatible
- Catalog of popular models pre-configured for common GPU types
- Licensed under Apache License 2.0
Capabilities
Key Features
- Deploy LLMs via vLLM and Ollama servers
- Speech-to-text with FasterWhisper
- Vector embeddings with Infinity
- Reranking with cross-encoder models
- Autoscaling including scale from zero
- Prefix-aware load balancing for KV cache utilization
- Model caching with EFS and Filestore
- Dynamic LoRA adapter orchestration
- Event streaming with Kafka and PubSub
- OpenAI-compatible API endpoints
- Preconfigured model catalog
- Runs on CPU, GPU, or TPU
- Request queueing and retries
- Prometheus observability
