# KServe

> Open-source Kubernetes platform for serving generative and predictive AI models through a standardized inference API.

KServe is an open-source inference platform for deploying generative and predictive machine learning models on Kubernetes. It is a Cloud Native Computing Foundation incubating project and is developed in the open on GitHub under the Apache License 2.0. Users describe a model deployment as a Kubernetes custom resource, and KServe handles autoscaling, networking, health checking, and server configuration.

## What It Is

KServe provides a Kubernetes Custom Resource Definition, InferenceService, for serving ML models across frameworks. It unifies two workloads on one platform: generative AI (LLMs) and predictive AI (models built with TensorFlow, PyTorch, scikit-learn, XGBoost, ONNX, and others). Platform and ML engineers write a YAML manifest that points to a model storage location, such as a Hugging Face model URI, and then send inference requests to the resulting endpoint.

## How the Architecture Works

The documentation describes a control plane and a data plane. The control plane manages model lifecycle, including revision tracking, canary rollouts, and A/B testing. The data plane defines a standardized inference protocol with request/response APIs for both predictive and generative models. InferenceGraph supports pipelines for pre/post processing, ensembles, and multi-model workflows.

## Generative and Predictive Capabilities

For LLMs, KServe supports vLLM and llm-d backends, an OpenAI-compatible protocol, GPU acceleration, model caching, KV cache offloading to CPU or disk, request-based autoscaling, and Hugging Face models. For predictive models, it adds routing between predictor, transformer, and explainer components, scale-to-zero, feature-attribution explainability, and monitoring such as payload logging and outlier, adversarial, and drift detection.

## Installation Options

The README lists standard Kubernetes, Knative-based serverless, and ModelMesh installation modes, plus a quick local install. The standard Kubernetes mode is lighter but does not support canary deployment or request-based autoscaling with scale-to-zero. KServe is also available as a Kubeflow add-on.

## Features
- InferenceService Kubernetes custom resource
- OpenAI-compatible inference protocol
- vLLM and llm-d backends
- GPU acceleration
- Model caching
- KV cache offloading to CPU/disk
- Request-based autoscaling and scale-to-zero
- Hugging Face model support
- Multi-framework predictive serving
- InferenceGraph pipelines and ensembles
- Canary rollouts and A/B testing
- Model explainability
- Payload logging, outlier, adversarial and drift detection

## Integrations
Kubernetes, Knative, Istio, Kubeflow, ModelMesh, vLLM, llm-d, Hugging Face, TensorFlow, PyTorch, scikit-learn, XGBoost, ONNX, KEDA, Prometheus

## Platforms
API, CLI

## Pricing
Open Source

## Version
v0.21.0

## Links
- Website: https://kserve.github.io/website/
- Documentation: https://kserve.github.io/website/docs/intro
- Repository: https://github.com/kserve/kserve
- EveryDev.ai: https://www.everydev.ai/tools/kserve
