# ML Drift

> Cross-platform GPU-accelerated inference engine for running ML and large generative models on-device across mobile, web, and servers.

ML Drift is an open-source GPU inference engine from the google-ai-edge organization, written in C++ and licensed under Apache 2.0. It runs machine learning workloads, including large generative models, on mobile phones, web browsers, and servers. The README describes it as the successor to the GPU delegate in TensorFlow Lite.

## What It Is

ML Drift is a runtime for executing ML models on device GPUs. It targets developers who want low latency and low power use, plus privacy and offline operation, for apps on Android, iOS, the web, and GPU-equipped edge devices. Models are represented as a backend-agnostic `GpuModel` compute graph, and individual steps are `GpuOperation` units containing shader code. The README states that it achieves an order-of-magnitude performance improvement over existing open-source GPU inference engines.

## How It Works

Kernels are written once in the Unified Compute Language (UCL), an abstraction over GLSL, MSL, and WGSL, and shaders are generated at runtime for the specific model, inputs, and GPU. Supported backends are OpenCL (primary for Android), Metal, WebGPU via Dawn, and OpenGL ES 3.1+. Runtime optimizations include operator fusion, layout adjustments, weight rearrangement, workgroup tuning, and memory reuse with algorithms such as GREEDY_BY_SIZE. Tensor virtualization decouples logical tensor views from physical GPU storage.

## Large Generative Models

For LLM workloads, ML Drift provides stage-aware execution separating prefill from decode, efficient KV cache management, and FP16/INT8/INT4 quantization. `GpuModel` graphs can also be hand-tailored for custom models.

## Getting Started

The repository includes hello-world guides for the OpenCL and WebGPU APIs and a samples directory. Public API headers live in `api/`, with backend implementations in `cl/`, `gl/`, `metal/`, and `webgpu/`.

## Features
- Cross-platform GPU inference for mobile, web, desktop, and servers
- OpenCL, Metal, WebGPU (Dawn), and OpenGL ES 3.1+ backends
- Unified Compute Language (UCL) for write-once GPU kernels
- Runtime dynamic shader generation
- Advanced GPU memory reuse
- Stage-aware prefill/decode execution for LLMs
- FP16/INT8/INT4 quantization
- Operator fusion and tensor virtualization
- Custom GpuModel graphs
- KV cache management

## Integrations
OpenCL, Metal, WebGPU, Dawn, OpenGL ES, TensorFlow Lite

## Platforms
WINDOWS, MACOS, LINUX, ANDROID, IOS, WEB, API, DEVELOPER_SDK

## Pricing
Open Source

## Links
- Website: https://github.com/google-ai-edge/ml-drift
- Repository: https://github.com/google-ai-edge/ml-drift
- EveryDev.ai: https://www.everydev.ai/tools/ml-drift
