# InstinctFlash

> A high-performance, open-source serving framework for robotics models that accelerates inference from perception to action on edge hardware like NVIDIA Jetson Thor.

InstinctFlash is an open-source inference runtime built by General Instinct specifically for Physical AI and robotics models. Released under the AGPL-3.0 license, it targets the challenge of deploying frontier robot models on the hardware that ships with real machines — from NVIDIA Jetson Thor edge devices to RTX 4090 and RTX 5090 workstations. The project is backed by Y Combinator and the NVIDIA Inception Program.

## What It Is

InstinctFlash is a serving framework that sits between a trained robotics model checkpoint and the hardware executing it. Rather than requiring users to hand-tune inference pipelines, it reads a short checkpoint declaration, determines which optimizations are provably valid for those weights without needing a GPU or the weights themselves, applies them in a structured six-layer optimization stack, and shows its work. The result is a single `instinctflash serve /path/to/checkpoint` command that handles model loading, optimization planning, and WebSocket-based serving for robot-side clients.

## Six-Layer Optimization Architecture

The runtime organizes optimization into six distinct layers, each targeting a different aspect of execution:

- **MODEL** — what is computed (distillation, step reduction, checkpoint compression via companion tools InstinctCompress and instinct-pdd)
- **GRAPH** — when work is issued (prefill extraction, CUDA-graph capture, memory planning)
- **CACHE** — what is recomputed (KV reuse, cross-attention and episode caches)
- **ATTENTION** — how tokens mix (FlashAttention, hybrid and linear attention)
- **KERNEL** — how a kernel is written (backend and layout dispatch, fusion)
- **HARDWARE** — what it executes on (FP8/INT8, TensorRT, Jetson-class edge devices)

The planner decides which optimizations are valid before touching weights or a GPU, then measures where time actually goes and starts there — rather than applying a fixed priority order.

## Supported Model Families and Benchmarks

InstinctFlash ships with adapters for eight robotics model families out of the box, including LingBot-VA, LingBot-VLA-4B, LingBot-VLA-V2-6B, pi0.5, GR00T-N1.7-3B, Cosmos3 Edge and Nano policies, and DreamZero. The README publishes Jetson Thor benchmark results measured September 15, 2026, showing prediction latency improvements ranging from 1.19× (GR00T N1.7) to 7.88× (pi0.5) in single-configuration comparisons. The headline 33.78× figure combines FP8 precision and a 25V/50A → 2V/4A schedule reduction for LingBot-VA.

## Deployment and Integration Path

Users install the Python 3.10+ core, then bootstrap a pinned per-model-family environment using `scripts/bootstrap_vendor.py`. The core can inspect checkpoints and plan optimizations without PyTorch or a GPU. Serving exposes a msgpack-over-WebSocket protocol compatible with the existing openpi-client ecosystem, so robot-side clients written for π0/π0.5 connect without modification. Four serving flags cover dry-run preflight, smoke testing, seeded native execution for paired comparisons, and live observation streaming to a Rerun viewer.

## Update: thor-2026-09-15 Release

The latest GitHub release, tagged `thor-2026-09-15` and published September 15, 2026, marks the full-source release of InstinctFlash. This release includes all eight robotics model family adapters, acceleration kernels, and Python/WebSocket serving through a unified Runtime API. Subsequent point updates added RTX 4090 support (September 16) and RTX 5090 support (September 17), extending the same Runtime API from Jetson Thor to desktop workstations. The repository was last pushed September 23, 2026, indicating active development. The roadmap lists attention upgrades, few-step distillation gated on edge control budgets, and device-specific serving defaults as upcoming work.

## Features
- High-performance inference serving for robotics models
- Six-layer optimization stack (MODEL, GRAPH, CACHE, ATTENTION, KERNEL, HARDWARE)
- FP8 and INT8 quantization support
- CUDA-graph capture and memory planning
- KV reuse and cross-attention caching
- FlashAttention integration
- TensorRT backend support
- Jetson Thor edge deployment
- RTX 4090 and RTX 5090 workstation support
- WebSocket serving via msgpack protocol
- Compatible with openpi-client ecosystem
- Checkpoint-driven automatic optimization planning
- Support for 8 robotics model families (LingBot-VA, LingBot-VLA, pi0.5, GR00T-N1.7, Cosmos3, DreamZero)
- Hugging Face Hub model loading
- Benchmark and evaluation tooling (LIBERO, RoboTwin)
- Dynamic step cache for DreamZero
- Plugin system via instinctflash.adapters entry points
- Dry-run preflight, smoke test, and Rerun visualization flags
- Non-inferiority certification via instinctflash validate

## Integrations
NVIDIA Jetson Thor, RTX 4090, RTX 5090, PyTorch, TensorRT, FlashAttention, NVIDIA CUTLASS, Hugging Face Hub, LeRobot, OpenPI / openpi-client, Rerun (visualization), LIBERO simulator, RoboTwin simulator, Triton compiler, msgpack-numpy, CUDA, InstinctCompress, instinct-pdd

## Platforms
LINUX, WEB, API, DEVELOPER_SDK, CLI

## Pricing
Open Source

## Version
thor-2026-09-15

## Links
- Website: https://general-instinct.com/
- Repository: https://github.com/General-Instinct/InstinctFlash
- EveryDev.ai: https://www.everydev.ai/tools/instinctflash
