EveryDev.ai
Subscribe
Home
Developers

3,795+ AI companies

  • Radar
  • Trending
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • Users
  • Rate Tools
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Developers
    3. KServe

    KServe

    KServe is an open-source, Kubernetes-native platform for standardized distributed inference of generative and predictive AI models. It provides a unified, cloud-agnostic way to deploy and scale models across frameworks while managing autoscaling, networking, health checks, routing, and serving configuration.

    Visit Website

    At a Glance

    1Tool Listed
    5Products
    12Capabilities
    Discussions
    2019Est.
    Focus Areas
    AI Infrastructure
    Model Management
    Container Orchestration
    Connect
    Latest News
    KServe v0.20 released with confidential model serving, LLM traffic splitting, raw-deployment canaries, KV-cache offloading, managed DRA, Anthropic Messages routing, OCI ImageVolumes, standalone vLLM, and AutoGluon time-series support.Aug 6, 2026
    KServe v0.19 released with static LoRA adapters, model-name routing, graceful vLLM pod draining, scaling status conditions, LocalModelCache integration, dual REST/gRPC routing, and llm-d v0.7 migration support.Jun 14, 2026
    Markets
    • Data science and ML engineering teams
    • DevOps, SRE, and platform engineering teams
    • Enterprise organizations operating production AI workloads
    • Cloud providers and Kubernetes distributions
    • +2 more

    AI Tools by KServe

    (1)
    View KServe
    KServe tool icon

    KServe

    Kubernetes ML Model Inference Platform

    AI InfrastructureModel ManagementContainer Orch.

    Discussions

    No discussions yet

    Be the first to start a discussion about KServe

    Latest News

    08/06/2026

    KServe v0.20 released with confidential model serving, LLM traffic splitting, raw-deployment canaries, KV-cache offloading, managed DRA, Anthropic Messages routing, OCI ImageVolumes, standalone vLLM, and AutoGluon time-series support.

    kserve.github.io
    06/14/2026

    KServe v0.19 released with static LoRA adapters, model-name routing, graceful vLLM pod draining, scaling status conditions, LocalModelCache integration, dual REST/gRPC routing, and llm-d v0.7 migration support.

    kserve.github.io
    04/29/2026

    KServe v0.18 released with multi-node inference without Ray, LeaderWorkerSet autoscaling, OpenAI Responses API routing, namespace-scoped ModelCache, vLLM 0.19, llm-d 0.6, and security hardening.

    kserve.github.io
    03/13/2026

    KServe v0.17 made LLMInferenceService production-ready with llm-d, KV-cache-aware routing, disaggregated prefill/decode, distributed parallelism, Envoy AI Gateway integration, and modular Helm charts.

    kserve.github.io

    Products & Services

    5
    KServe platform
    2019 (as KFServing; renamed KServe in 2021)

    Open-source Kubernetes custom-resource platform for deploying, operating, and scaling predictive and generative AI inference services across cloud, on-premises, hybrid, and edge environments.

    InferenceService
    2019 (as KFServing)

    Core Kubernetes custom resource for predictive model serving, with standardized APIs, framework runtimes, autoscaling including scale-to-zero, health checks, networking, canary rollouts, transformers, explainers, and traffic management.

    LLMInferenceService
    2025 (experimental; production-ready in v0.17 on March 13, 2026)

    GenAI-focused custom resource for LLM serving with vLLM and llm-d, distributed inference, KV-cache-aware routing, disaggregated prefill/decode, Gateway Inference Extension, token-based rate limiting, and GPU-aware scaling.

    ModelMesh
    2022 (included in the LF AI & Data incubation-era KServe project)

    High-scale, high-density model-serving subsystem for frequently changing model fleets and deployments involving very large numbers of concurrently served models.

    Market Position

    KServe positions itself as the open-source, Kubernetes-native standard that unifies predictive and generative model inference instead of requiring separate serving stacks. Its differentiation is a framework-neutral CRD/API, scale-to-zero and enterprise autoscaling, canary and graph workflows, high-density ModelMesh, and GenAI capabilities such as vLLM/llm-d integration and KV-cache-aware routing. It competes and overlaps with Seldon, NVIDIA Triton Inference Server, Ray Serve, BentoML, TorchServe, MLflow Model Serving, and cloud-provider model-serving services.

    Leadership

    Founders

    DS

    Dan Sun

    KServe co-founder and project lead; Engineering Team Lead for Cloud Native Compute Services/AI Inference Engineering at Bloomberg. He helped found and lead development of KServe (formerly KFServing).

    AS

    Animesh Singh

    IBM Distinguished Engineer and Director for Watson AI and Data Open Technology; represented IBM, identified by the KServe community as a co-founder and major early champion.

    AN

    Alexa Nicole Griffith

    Bloomberg engineer and KServe community contributor; co-authored the LF AI & Data incubation announcement on behalf of the KServe community.

    Executive Team

    DS

    Dan Sun

    KServe Project Lead and Co-founder; Engineering Team Lead, AI Inference Engineering, Bloomberg

    Software engineering team lead at Bloomberg and KServe maintainer; co-founder of KServe and co-founder of the Envoy AI Gateway project.

    YT

    Yuan Tang

    KServe Project Lead; Senior Principal Software Engineer, Red Hat AI

    Leads AI infrastructure and platform work at Red Hat, including KServe and OpenShift AI; holds leadership roles in KServe, Kubeflow, Argo, Kubernetes, and CNCF communities.

    Founding Story

    KServe originated in 2019 as KFServing, a collaborative effort by teams from Google, IBM, Bloomberg, NVIDIA, and Seldon under the Kubeflow project. It was created to address the difficulty of deploying and monitoring machine-learning models in production by providing a Kubernetes custom resource, standardized inference interfaces, and built-in scaling and deployment capabilities; KFServing was renamed KServe in September 2021.

    Business Model

    Revenue Model

    KServe is an Apache 2.0 open-source CNCF project rather than a standalone commercial company. The project itself is distributed without a paid subscription or usage-based pricing plan; commercial value is delivered through adopters' products, managed platforms, support, and enterprise distributions.

    Target Markets

    Industries & Segments
    • Data science and ML engineering teams
    • DevOps, SRE, and platform engineering teams
    • Enterprise organizations operating production AI workloads
    • Cloud providers and Kubernetes distributions
    • AI infrastructure and GPU platform providers
    • Open-source MLOps and Kubeflow ecosystem users
    Use Cases
    • Production deployment of traditional machine-learning prediction models
    • Large-language-model and generative-AI inference
    • Enterprise AI platforms and internal model-serving infrastructure
    • High-density serving of hundreds or thousands of frequently changing models
    • Multi-node and GPU-accelerated inference
    • Progressive model releases, canary testing, and A/B experimentation
    Notable Customers
    • Bloomberg
    • IBM
    • Red Hat
    • NVIDIA

    Quick Facts

    Founded
    2019

    History & Milestones

    March 13, 2026

    KServe v0.17 made LLMInferenceService production-ready and added llm-d-based GenAI architecture, KV-cache-aware routing, distributed inference, and Envoy AI Gateway integration.

    August 6, 2026

    KServe v0.20 added confidential model serving with TEEs, traffic splitting for LLMInferenceService, raw-deployment canaries, KV-cache offloading, managed DRA, Anthropic Messages routing, native OCI ImageVolumes, standalone vLLM, and AutoGluon time-series serving.

    September 29, 2025

    KServe was accepted by CNCF at the Incubating maturity level.

    May 15, 2024

    KServe v0.13 expanded generative-AI support with an enhanced Hugging Face runtime, vLLM backend support, and OpenAI protocol support.

    December 23, 2024

    KServe v0.14 introduced a Python client and model cache, promoted OCI model storage to stable, and added direct Hugging Face model deployment.

    Key Capabilities

    12
    Kubernetes Custom Resource Definition and standard Kubernetes API
    Generative and predictive AI inference in one platform
    Multi-framework support including TensorFlow, PyTorch, scikit-learn, XGBoost, ONNX, Hugging Face, vLLM, and custom runtimes
    OpenAI-compatible APIs for LLM chat completion, streaming, and embeddings
    Autoscaling based on requests, tokens, queue depth, GPU utilization, HPA, KEDA, or KPA, including scale-to-zero
    Canary deployments, A/B testing, traffic splitting, model revision tracking, and intelligent routing

    Integrations & Partnerships

    Platform Integrations

    • Kubernetes via KServe CRDs and Helm charts
    • Kubeflow
    • Knative
    • ModelMesh
    • vLLM
    • llm-d
    • Envoy AI Gateway and Gateway API Inference Extension
    • NVIDIA NIM and NVIDIA Triton Inference Server

    Key Partnerships

    Kubeflow ecosystem integration
    CNCF incubation and Linux Foundation governance
    Original collaboration among Google, IBM, Bloomberg, NVIDIA, and Seldon

    Connect

    Website
    kserve.github.io/website/
    GitHub
    kserve
    LinkedIn
    kserve-project

    AI Topics

    3

    KServe focuses on these topics:

    AI Infrastructure(1)
    Model Management(1)
    Container Orchestration(1)
    Back to all developersSuggest an edit