InstinctFlash
A high-performance, open-source serving framework for robotics models that accelerates inference from perception to action on edge hardware like NVIDIA Jetson Thor.
At a Glance
Fully open-source under AGPL-3.0. Free to use, modify, and distribute.
Engagement
Available On
Alternatives
Listed Oct 2026
About InstinctFlash
InstinctFlash is an open-source inference runtime built by General Instinct specifically for Physical AI and robotics models. Released under the AGPL-3.0 license, it targets the challenge of deploying frontier robot models on the hardware that ships with real machines — from NVIDIA Jetson Thor edge devices to RTX 4090 and RTX 5090 workstations. The project is backed by Y Combinator and the NVIDIA Inception Program.
What It Is
InstinctFlash is a serving framework that sits between a trained robotics model checkpoint and the hardware executing it. Rather than requiring users to hand-tune inference pipelines, it reads a short checkpoint declaration, determines which optimizations are provably valid for those weights without needing a GPU or the weights themselves, applies them in a structured six-layer optimization stack, and shows its work. The result is a single instinctflash serve /path/to/checkpoint command that handles model loading, optimization planning, and WebSocket-based serving for robot-side clients.
Six-Layer Optimization Architecture
The runtime organizes optimization into six distinct layers, each targeting a different aspect of execution:
- MODEL — what is computed (distillation, step reduction, checkpoint compression via companion tools InstinctCompress and instinct-pdd)
- GRAPH — when work is issued (prefill extraction, CUDA-graph capture, memory planning)
- CACHE — what is recomputed (KV reuse, cross-attention and episode caches)
- ATTENTION — how tokens mix (FlashAttention, hybrid and linear attention)
- KERNEL — how a kernel is written (backend and layout dispatch, fusion)
- HARDWARE — what it executes on (FP8/INT8, TensorRT, Jetson-class edge devices)
The planner decides which optimizations are valid before touching weights or a GPU, then measures where time actually goes and starts there — rather than applying a fixed priority order.
Supported Model Families and Benchmarks
InstinctFlash ships with adapters for eight robotics model families out of the box, including LingBot-VA, LingBot-VLA-4B, LingBot-VLA-V2-6B, pi0.5, GR00T-N1.7-3B, Cosmos3 Edge and Nano policies, and DreamZero. The README publishes Jetson Thor benchmark results measured September 15, 2026, showing prediction latency improvements ranging from 1.19× (GR00T N1.7) to 7.88× (pi0.5) in single-configuration comparisons. The headline 33.78× figure combines FP8 precision and a 25V/50A → 2V/4A schedule reduction for LingBot-VA.
Deployment and Integration Path
Users install the Python 3.10+ core, then bootstrap a pinned per-model-family environment using scripts/bootstrap_vendor.py. The core can inspect checkpoints and plan optimizations without PyTorch or a GPU. Serving exposes a msgpack-over-WebSocket protocol compatible with the existing openpi-client ecosystem, so robot-side clients written for π0/π0.5 connect without modification. Four serving flags cover dry-run preflight, smoke testing, seeded native execution for paired comparisons, and live observation streaming to a Rerun viewer.
Update: thor-2026-09-15 Release
The latest GitHub release, tagged thor-2026-09-15 and published September 15, 2026, marks the full-source release of InstinctFlash. This release includes all eight robotics model family adapters, acceleration kernels, and Python/WebSocket serving through a unified Runtime API. Subsequent point updates added RTX 4090 support (September 16) and RTX 5090 support (September 17), extending the same Runtime API from Jetson Thor to desktop workstations. The repository was last pushed September 23, 2026, indicating active development. The roadmap lists attention upgrades, few-step distillation gated on edge control budgets, and device-specific serving defaults as upcoming work.
Community Discussions
Be the first to start a conversation about InstinctFlash
Share your experience with InstinctFlash, ask questions, or help others learn from your insights.
Pricing
Open Source
Fully open-source under AGPL-3.0. Free to use, modify, and distribute.
- Full source code access
- Eight robotics model family adapters
- Jetson Thor, RTX 4090, RTX 5090 support
- Python and WebSocket serving
- Benchmark and evaluation tooling
Capabilities
Key Features
- High-performance inference serving for robotics models
- Six-layer optimization stack (MODEL, GRAPH, CACHE, ATTENTION, KERNEL, HARDWARE)
- FP8 and INT8 quantization support
- CUDA-graph capture and memory planning
- KV reuse and cross-attention caching
- FlashAttention integration
- TensorRT backend support
- Jetson Thor edge deployment
- RTX 4090 and RTX 5090 workstation support
- WebSocket serving via msgpack protocol
- Compatible with openpi-client ecosystem
- Checkpoint-driven automatic optimization planning
- Support for 8 robotics model families (LingBot-VA, LingBot-VLA, pi0.5, GR00T-N1.7, Cosmos3, DreamZero)
- Hugging Face Hub model loading
- Benchmark and evaluation tooling (LIBERO, RoboTwin)
- Dynamic step cache for DreamZero
- Plugin system via instinctflash.adapters entry points
- Dry-run preflight, smoke test, and Rerun visualization flags
- Non-inferiority certification via instinctflash validate
