# Awesome-ML-SYS-Tutorial

> A curated collection of learning notes and deep-dive tutorials on ML systems, covering RL infrastructure, online/offline inference, SGLang, CUDA, and distributed training.

Awesome-ML-SYS-Tutorial is an open-source GitHub repository of in-depth learning notes on machine learning systems (ML SYS), maintained by Chenyang Zhao, a researcher now working full-time at RadixArk. The project started in August 2024 as personal notes taken while working with the SGLang inference framework, and has grown into a broad technical reference covering RL infrastructure, inference system design, and AI infrastructure fundamentals. The repository is licensed under Apache 2.0 and is freely available to the community.

## What It Is

Awesome-ML-SYS-Tutorial is a structured knowledge base — not a software tool — that collects detailed technical write-ups, source code walkthroughs, and engineering notes on the internals of modern ML systems. Topics span reinforcement learning from human feedback (RLHF) infrastructure, LLM inference engines (especially SGLang), distributed training, CUDA/GPU programming, and quantization. Articles are written primarily in Chinese with English translations available for most major pieces, and many are cross-posted to Zhihu (a Chinese technical blogging platform).

## Coverage and Content Areas

The repository is organized into several major sections:

- **RLHF System Development**: Deep dives into frameworks including slime, verl, AReal, and OpenRLHF — covering PPO, GRPO, weight update mechanisms, FSDP training backends, multi-turn RL, and speculative decoding in RL sampling.
- **SGLang Learning Notes**: Architecture walkthroughs of the SGLang inference engine, including KV cache management, scheduler design, constraint decoding, quantization, diffusion model support, and omni model inference.
- **Omni Model Inference**: Notes on multi-stage generative model inference, TTS optimization (achieving 1.9–3.4× speedups on SGLang-Omni per the author's benchmarks), and CPU resource management in speech serving.
- **ML System Fundamentals**: Coverage of CUDA Graph, NCCL, tensor parallelism, expert parallelism, PyTorch distributed training, and transformer architecture internals.
- **Developer Guide**: Practical engineering notes on Docker, development environment setup, and CI/CD for Jupyter notebooks.

## Audience and Purpose

The primary audience is researchers and engineers entering or working in AI infrastructure — particularly those building or studying RL training pipelines and LLM inference systems. The author explicitly frames the project as a response to a concern about flawed RL infrastructure leading to unreliable research conclusions, and positions the notes as a way to help the community build on a correct foundation. The repository also serves as an onboarding resource for contributors to the SGLang open-source community.

## Update: Active and Rapidly Expanding

The repository was created in November 2024 and, according to the GitHub metadata, was last pushed to in September 2026, indicating sustained and active development. The README lists numerous articles marked "Pending Review," signaling ongoing content production. Recent additions include notes on Qwen3-Omni inference, INT4 QAT RL end-to-end practice, FSDP2 as a training backend for slime, unified FP8 for RL sampling and training, and CPU topology-aware allocation for speech model serving. The project has accumulated over 7,000 GitHub stars and 500 forks, reflecting significant community interest as reported by the repository's own metadata.

## Why It Stands Out

Unlike typical "awesome list" repositories that aggregate links, this project consists primarily of original long-form technical writing with source code analysis, benchmark results, and engineering lessons learned from real production and research work. The author's dual role — as both an active contributor to SGLang and a reviewer at venues like ICLR — gives the notes a practitioner's perspective grounded in hands-on infrastructure development rather than purely academic exposition.

## Features
- In-depth RL infrastructure notes (slime, verl, AReal, OpenRLHF)
- SGLang inference engine walkthroughs and code analysis
- Omni model and TTS inference optimization notes
- CUDA Graph, NCCL, and distributed training fundamentals
- Quantization deep dives (AWQ, FP8, INT4 QAT)
- Multi-turn RL and tool-calling implementation guides
- KV cache and scheduler architecture analysis
- Bilingual content (Chinese and English)
- Cross-posted articles on Zhihu
- Developer guides for Docker and environment setup

## Integrations
SGLang, verl, OpenRLHF, slime, PyTorch, NCCL, Megatron, FSDP, HuggingFace, Ray, CUDA, Docker

## Platforms
WEB, API

## Pricing
Open Source

## Links
- Website: https://github.com/zhaochenyang20/Awesome-ML-SYS-Tutorial
- Repository: https://github.com/zhaochenyang20/Awesome-ML-SYS-Tutorial
- EveryDev.ai: https://www.everydev.ai/tools/awesome-ml-sys-tutorial
