KServe
Open-source Kubernetes platform for serving generative and predictive AI models through a standardized inference API.
At a Glance
Self-hosted KServe under the Apache License 2.0, run on your own Kubernetes cluster.
Engagement
Available On
Listed Oct 2026
About KServe
KServe is an open-source inference platform for deploying generative and predictive machine learning models on Kubernetes. It is a Cloud Native Computing Foundation incubating project and is developed in the open on GitHub under the Apache License 2.0. Users describe a model deployment as a Kubernetes custom resource, and KServe handles autoscaling, networking, health checking, and server configuration.
What It Is
KServe provides a Kubernetes Custom Resource Definition, InferenceService, for serving ML models across frameworks. It unifies two workloads on one platform: generative AI (LLMs) and predictive AI (models built with TensorFlow, PyTorch, scikit-learn, XGBoost, ONNX, and others). Platform and ML engineers write a YAML manifest that points to a model storage location, such as a Hugging Face model URI, and then send inference requests to the resulting endpoint.
How the Architecture Works
The documentation describes a control plane and a data plane. The control plane manages model lifecycle, including revision tracking, canary rollouts, and A/B testing. The data plane defines a standardized inference protocol with request/response APIs for both predictive and generative models. InferenceGraph supports pipelines for pre/post processing, ensembles, and multi-model workflows.
Generative and Predictive Capabilities
For LLMs, KServe supports vLLM and llm-d backends, an OpenAI-compatible protocol, GPU acceleration, model caching, KV cache offloading to CPU or disk, request-based autoscaling, and Hugging Face models. For predictive models, it adds routing between predictor, transformer, and explainer components, scale-to-zero, feature-attribution explainability, and monitoring such as payload logging and outlier, adversarial, and drift detection.
Installation Options
The README lists standard Kubernetes, Knative-based serverless, and ModelMesh installation modes, plus a quick local install. The standard Kubernetes mode is lighter but does not support canary deployment or request-based autoscaling with scale-to-zero. KServe is also available as a Kubeflow add-on.
Community Discussions
Be the first to start a conversation about KServe
Share your experience with KServe, ask questions, or help others learn from your insights.
Pricing
Open Source
Self-hosted KServe under the Apache License 2.0, run on your own Kubernetes cluster.
- Generative AI serving with vLLM and llm-d backends
- OpenAI-compatible inference protocol
- Predictive AI serving across TensorFlow, PyTorch, scikit-learn, XGBoost and ONNX
- Canary rollouts and InferenceGraph
- Autoscaling including scale-to-zero
Capabilities
Key Features
- InferenceService Kubernetes custom resource
- OpenAI-compatible inference protocol
- vLLM and llm-d backends
- GPU acceleration
- Model caching
- KV cache offloading to CPU/disk
- Request-based autoscaling and scale-to-zero
- Hugging Face model support
- Multi-framework predictive serving
- InferenceGraph pipelines and ensembles
- Canary rollouts and A/B testing
- Model explainability
- Payload logging, outlier, adversarial and drift detection
