ML Drift
Cross-platform GPU-accelerated inference engine for running ML and large generative models on-device across mobile, web, and servers.
At a Glance
ML Drift core inference engine released under the Apache License 2.0 with no-charge, royalty-free terms.
Engagement
Available On
Listed Oct 2026
About ML Drift
ML Drift is an open-source GPU inference engine from the google-ai-edge organization, written in C++ and licensed under Apache 2.0. It runs machine learning workloads, including large generative models, on mobile phones, web browsers, and servers. The README describes it as the successor to the GPU delegate in TensorFlow Lite.
What It Is
ML Drift is a runtime for executing ML models on device GPUs. It targets developers who want low latency and low power use, plus privacy and offline operation, for apps on Android, iOS, the web, and GPU-equipped edge devices. Models are represented as a backend-agnostic GpuModel compute graph, and individual steps are GpuOperation units containing shader code. The README states that it achieves an order-of-magnitude performance improvement over existing open-source GPU inference engines.
How It Works
Kernels are written once in the Unified Compute Language (UCL), an abstraction over GLSL, MSL, and WGSL, and shaders are generated at runtime for the specific model, inputs, and GPU. Supported backends are OpenCL (primary for Android), Metal, WebGPU via Dawn, and OpenGL ES 3.1+. Runtime optimizations include operator fusion, layout adjustments, weight rearrangement, workgroup tuning, and memory reuse with algorithms such as GREEDY_BY_SIZE. Tensor virtualization decouples logical tensor views from physical GPU storage.
Large Generative Models
For LLM workloads, ML Drift provides stage-aware execution separating prefill from decode, efficient KV cache management, and FP16/INT8/INT4 quantization. GpuModel graphs can also be hand-tailored for custom models.
Getting Started
The repository includes hello-world guides for the OpenCL and WebGPU APIs and a samples directory. Public API headers live in api/, with backend implementations in cl/, gl/, metal/, and webgpu/.
Community Discussions
Be the first to start a conversation about ML Drift
Share your experience with ML Drift, ask questions, or help others learn from your insights.
Pricing
Open Source
ML Drift core inference engine released under the Apache License 2.0 with no-charge, royalty-free terms.
- Apache License 2.0
- GPU backends: OpenCL, Metal, WebGPU (Dawn), OpenGL ES 3.1+
- Unified Compute Language (UCL) for write-once GPU kernels
- FP16/INT8/INT4 quantization support
- Cross-platform: Android, iOS, web browsers, servers
Capabilities
Key Features
- Cross-platform GPU inference for mobile, web, desktop, and servers
- OpenCL, Metal, WebGPU (Dawn), and OpenGL ES 3.1+ backends
- Unified Compute Language (UCL) for write-once GPU kernels
- Runtime dynamic shader generation
- Advanced GPU memory reuse
- Stage-aware prefill/decode execution for LLMs
- FP16/INT8/INT4 quantization
- Operator fusion and tensor virtualization
- Custom GpuModel graphs
- KV cache management
