KServe
KServe is an open-source, Kubernetes-native platform for standardized distributed inference of generative and predictive AI models. It provides a unified, cloud-agnostic way to deploy and scale models across frameworks while managing autoscaling, networking, health checks, routing, and serving configuration.
At a Glance
- Data science and ML engineering teams
- DevOps, SRE, and platform engineering teams
- Enterprise organizations operating production AI workloads
- Cloud providers and Kubernetes distributions
- +2 more
AI Tools by KServe
(1)KServe
Kubernetes ML Model Inference Platform
Discussions
No discussions yet
Be the first to start a discussion about KServe
Latest News
KServe v0.20 released with confidential model serving, LLM traffic splitting, raw-deployment canaries, KV-cache offloading, managed DRA, Anthropic Messages routing, OCI ImageVolumes, standalone vLLM, and AutoGluon time-series support.
KServe v0.19 released with static LoRA adapters, model-name routing, graceful vLLM pod draining, scaling status conditions, LocalModelCache integration, dual REST/gRPC routing, and llm-d v0.7 migration support.
KServe v0.18 released with multi-node inference without Ray, LeaderWorkerSet autoscaling, OpenAI Responses API routing, namespace-scoped ModelCache, vLLM 0.19, llm-d 0.6, and security hardening.
KServe v0.17 made LLMInferenceService production-ready with llm-d, KV-cache-aware routing, disaggregated prefill/decode, distributed parallelism, Envoy AI Gateway integration, and modular Helm charts.
Products & Services
Open-source Kubernetes custom-resource platform for deploying, operating, and scaling predictive and generative AI inference services across cloud, on-premises, hybrid, and edge environments.
Core Kubernetes custom resource for predictive model serving, with standardized APIs, framework runtimes, autoscaling including scale-to-zero, health checks, networking, canary rollouts, transformers, explainers, and traffic management.
GenAI-focused custom resource for LLM serving with vLLM and llm-d, distributed inference, KV-cache-aware routing, disaggregated prefill/decode, Gateway Inference Extension, token-based rate limiting, and GPU-aware scaling.
High-scale, high-density model-serving subsystem for frequently changing model fleets and deployments involving very large numbers of concurrently served models.
Market Position
KServe positions itself as the open-source, Kubernetes-native standard that unifies predictive and generative model inference instead of requiring separate serving stacks. Its differentiation is a framework-neutral CRD/API, scale-to-zero and enterprise autoscaling, canary and graph workflows, high-density ModelMesh, and GenAI capabilities such as vLLM/llm-d integration and KV-cache-aware routing. It competes and overlaps with Seldon, NVIDIA Triton Inference Server, Ray Serve, BentoML, TorchServe, MLflow Model Serving, and cloud-provider model-serving services.
Leadership
Founders
Dan Sun
KServe co-founder and project lead; Engineering Team Lead for Cloud Native Compute Services/AI Inference Engineering at Bloomberg. He helped found and lead development of KServe (formerly KFServing).
Animesh Singh
IBM Distinguished Engineer and Director for Watson AI and Data Open Technology; represented IBM, identified by the KServe community as a co-founder and major early champion.
Alexa Nicole Griffith
Bloomberg engineer and KServe community contributor; co-authored the LF AI & Data incubation announcement on behalf of the KServe community.
Executive Team
Dan Sun
KServe Project Lead and Co-founder; Engineering Team Lead, AI Inference Engineering, Bloomberg
Software engineering team lead at Bloomberg and KServe maintainer; co-founder of KServe and co-founder of the Envoy AI Gateway project.
Yuan Tang
KServe Project Lead; Senior Principal Software Engineer, Red Hat AI
Leads AI infrastructure and platform work at Red Hat, including KServe and OpenShift AI; holds leadership roles in KServe, Kubeflow, Argo, Kubernetes, and CNCF communities.
Founding Story
KServe originated in 2019 as KFServing, a collaborative effort by teams from Google, IBM, Bloomberg, NVIDIA, and Seldon under the Kubeflow project. It was created to address the difficulty of deploying and monitoring machine-learning models in production by providing a Kubernetes custom resource, standardized inference interfaces, and built-in scaling and deployment capabilities; KFServing was renamed KServe in September 2021.
Business Model
Revenue Model
KServe is an Apache 2.0 open-source CNCF project rather than a standalone commercial company. The project itself is distributed without a paid subscription or usage-based pricing plan; commercial value is delivered through adopters' products, managed platforms, support, and enterprise distributions.
Target Markets
- Data science and ML engineering teams
- DevOps, SRE, and platform engineering teams
- Enterprise organizations operating production AI workloads
- Cloud providers and Kubernetes distributions
- AI infrastructure and GPU platform providers
- Open-source MLOps and Kubeflow ecosystem users
- Production deployment of traditional machine-learning prediction models
- Large-language-model and generative-AI inference
- Enterprise AI platforms and internal model-serving infrastructure
- High-density serving of hundreds or thousands of frequently changing models
- Multi-node and GPU-accelerated inference
- Progressive model releases, canary testing, and A/B experimentation
- Bloomberg
- IBM
- Red Hat
- NVIDIA