# KubeAI

> Open-source Kubernetes operator for deploying and autoscaling LLM, embedding, reranking, and speech-to-text models behind an OpenAI-compatible API.

KubeAI is an open-source AI inference operator for Kubernetes, hosted by the kubeai-project organization on GitHub under the Apache 2.0 license. It deploys and scales machine learning models, including LLMs, embeddings, reranking and speech-to-text, and exposes them through OpenAI-compatible endpoints. The latest listed release is v0.23.5.

## What It Is

KubeAI is a Kubernetes-native serving layer for ML models. It operates vLLM and Ollama servers for text generation, FasterWhisper for audio transcription, and Infinity for vector embeddings, and it supports reranking with cross-encoder models. Models are declared through a Model custom resource, and the project ships a catalog of popular models preconfigured for common GPU types.

## Architecture and Routing

KubeAI has two main components: a model proxy and a model operator. The proxy provides the OpenAI-compatible API, queues requests while a model scales from zero, retries requests that hit bad backends, and applies a prefix-aware load balancing strategy meant to improve KV cache utilization across vLLM replicas. The operator manages backend server Pods directly, automating model downloads, volume mounts, and loading of dynamic LoRA adapters. Both components are co-located in one deployment.

## Compatibility and Setup Path

The API supports /v1/chat/completions, /v1/completions, /v1/embeddings, /v1/rerank, /v1/models and /v1/audio/transcriptions, so existing OpenAI client libraries can be used. The project states it does not require Istio, Knative or a Prometheus metrics adapter, and that it runs on CPU, GPU or TPU. It installs via Helm on any Kubernetes cluster, with guides for AKS, EKS and GKE, and a local quickstart using kind or minikube with a bundled chat UI. Documentation also covers autoscaling, model caching with AWS EFS or GCP Filestore, loading models from OCI images or PVCs, LoRA adapters, multitenancy, and Prometheus observability.

## Features
- Deploy LLMs via vLLM and Ollama servers
- Speech-to-text with FasterWhisper
- Vector embeddings with Infinity
- Reranking with cross-encoder models
- Autoscaling including scale from zero
- Prefix-aware load balancing for KV cache utilization
- Model caching with EFS and Filestore
- Dynamic LoRA adapter orchestration
- Event streaming with Kafka and PubSub
- OpenAI-compatible API endpoints
- Preconfigured model catalog
- Runs on CPU, GPU, or TPU
- Request queueing and retries
- Prometheus observability

## Integrations
Kubernetes, vLLM, Ollama, FasterWhisper, Infinity, Helm, OpenAI client libraries, Kafka, PubSub, AWS EFS, GCP Filestore, Prometheus, LangChain, Langtrace, Weaviate, Open WebUI

## Platforms
LINUX, API

## Pricing
Open Source

## Version
v0.23.5

## Links
- Website: https://www.kubeai.org
- Documentation: https://www.kubeai.org
- Repository: https://github.com/kubeai-project/kubeai
- EveryDev.ai: https://www.everydev.ai/tools/kubeai
