kubeai-project
KubeAI is an open-source AI inferencing operator for Kubernetes that deploys, scales, and manages machine-learning models in production. It focuses on LLMs, vision-language models, embeddings, reranking, and speech-to-text, with an OpenAI-compatible interface and cache-aware routing.
At a Glance
- Organizations operating Kubernetes clusters
- AI/ML platform and infrastructure teams
- GPU cloud and managed-Kubernetes providers
- Enterprises running open models in private or sovereign environments
- +2 more
AI Tools by kubeai-project
(1)KubeAI
Kubernetes AI Inference Operator
Discussions
No discussions yet
Be the first to start a discussion about kubeai-project
Latest News
KubeAI 0.23.5 released with Ollama startup-probe handling and OpenTelemetry autoscaling-cardinality fixes.
KubeAI 0.23.1 and matching Helm chart releases published.
KubeAI 0.23.0 and matching Helm chart releases published.
KubeAI 0.22.1 and Helm chart releases published.
Products & Services
The core open-source Kubernetes operator for deploying and scaling model servers, with model catalogs, scale-from-zero, model caching, dynamic LoRA adapters, and hardware-flexible CPU/GPU/TPU operation.
An OpenAI-compatible gateway that provides request queueing, retries, and cache-aware load balancing across model-server replicas.
A CRD-driven controller that manages backend Pods, downloads models, mounts storage, and loads dynamic LoRA adapters.
Helm installation charts and preconfigured model definitions for common GPU and CPU profiles, including text generation, embeddings, reranking, and speech-to-text models.
Market Position
KubeAI positions itself as a vendor-neutral, Kubernetes-native inference layer that is simpler to operate than stacks requiring Istio, Knative, or separate metrics adapters. Its differentiator is cache-aware PrefixHash/CHWBL routing and an OpenAI-compatible proxy, addressing the performance limitations of random Kubernetes Service routing for stateful LLM backends while retaining portability across clouds, edge environments, and hardware types.
Leadership
Founders
Sam Stoelinga
GitHub identifies Stoelinga as the creator of KubeAI and describes him as focused on a Kubernetes operator for serving LLMs; the KubeAI site lists him as a project maintainer/contact.
Nick Stogner
Stogner authored the launch article for KubeAI and describes his background publicly as Cloud & Kubernetes; the launch article identifies him as an early project contributor/co-author.
Executive Team
Sam Stoelinga
Creator and maintainer
Creator of KubeAI and a Kubernetes-focused open-source developer.
Nick Stogner
Early contributor and maintainer/contact
Cloud and Kubernetes practitioner who authored the public KubeAI launch article and is listed among project contacts.
Founding Story
KubeAI was started to make running LLMs, embedding models, and speech-to-text on Kubernetes straightforward. Its early design emphasized an OpenAI-compatible API, direct operation of vLLM and Ollama in isolated Pods, scale-from-zero, and avoiding additional systems such as Istio and Knative.
Target Markets
- Organizations operating Kubernetes clusters
- AI/ML platform and infrastructure teams
- GPU cloud and managed-Kubernetes providers
- Enterprises running open models in private or sovereign environments
- Edge inference deployments
- Teams building LLM, embedding, reranking, or transcription applications
- Production LLM serving on Kubernetes
- Multi-turn chat and agent workloads that benefit from prefix/KV-cache affinity
- Large-scale batch inference
- Multi-region and multi-tenant small-language-model inference
- Vector search embedding generation and reranking
- Speech transcription pipelines
- Telescope
- Google Cloud Distributed Edge
- Lambda AI Developer Cloud
- Vultr Managed Kubernetes