OpenLake Project
OpenLake provides high-performance storage infrastructure to shatter LLM inference boundaries by saturating GPUs and offloading KV caches using RDMA and thread-per-core designs.
At a Glance
- AI Research Labs
- Enterprise AI Infrastructure
- Cloud Service Providers
- Companies using large-scale LLM deployments
AI Tools by OpenLake Project
(1)OpenLake
GPU Storage Engine for AI
Discussions
No discussions yet
Be the first to start a discussion about OpenLake Project
Latest News
Introducing OpenLake: Fast, Durable Storage for LLM Inference and Training
OpenLake presence at India Electronics Week (IEW) 2026
OpenLake v0.4 release with enhanced CLI and node reporting
OpenLake v0.1 release
Products & Services
A high-performance object store built for GPU workloads, featuring RDMA, GPUDirect, and io_uring to bypass the CPU for direct drive-to-VRAM data transfer.
A management interface to inspect, pin, and reuse KV cache entries across GPU workers.
Market Position
Claims 8x higher throughput and significantly lower latency compared to conventional object stores like MinIO, RustFS, and Ceph, specifically optimized for the small random I/O patterns of AI workloads.
Leadership
Founders
Arnav Balyan
Previously a Software Engineer at Uber and a committer for the Apache Gluten project. An alumnus of Netaji Subhas University of Technology (NSUT).
Executive Team
Arnav Balyan
Founder
Expert in native engine offloading and distributed systems; former committer at Apache Gluten (Uber).
Board of Directors
Founding Story
Started with the vision that while compute is abundant, storage bottlenecks prevent GPUs from being fully saturated. The project aims to provide a storage path that moves data directly from NVMe to GPU memory, bypassing the CPU to reduce latency and increase throughput for AI workloads.
Business Model
Revenue Model
Likely subscription or usage-based model for its cloud-managed storage and enterprise features, though the core engine is open source (Apache 2.0).
Target Markets
- AI Research Labs
- Enterprise AI Infrastructure
- Cloud Service Providers
- Companies using large-scale LLM deployments
- LLM Inference (Long Context)
- GPU Training & Checkpointing
- Vector Search and RAG
- Data Pipelines (Spark, Ray, Flink)