# club-3090

> Community recipes and Docker Compose configs for serving LLMs locally on RTX 3090s and other NVIDIA GPUs, with an OpenAI-compatible API via vLLM, SGLang, llama.cpp, and more.

Club-3090 is an open-source collection of working Docker Compose configurations, patches, and benchmark data for running large language models locally on consumer NVIDIA GPUs — primarily the RTX 3090, but also tested on 4090s, 5090s, and A-series cards. Maintained by the GitHub user noonghunna and licensed under Apache 2.0, the project lets users pick a model, launch it with a single command, and get an OpenAI-compatible API endpoint on their own hardware.

## What It Is

Club-3090 is a local inference recipe library: a structured repository of per-model, per-engine, per-topology Docker Compose files that abstract away the complexity of running quantized LLMs on consumer GPUs. It is not a new inference engine itself — it sits on top of established engines (vLLM, SGLang, llama.cpp, ik_llama, exllamav3) and provides the glue: verified compose configs, vendored patches, setup scripts, and a benchmark table so users know what performance to expect before they download anything.

## How the Workflow Works

The setup path is designed to go from clone to first reply in about five minutes:

- `scripts/setup.sh` — picks a model, downloads weights, and SHA-verifies them
- `scripts/launch.sh` — selects a config for the detected GPU topology and boots it with a health check
- `scripts/switch.sh` — lists all configs the current machine can run and switches between them
- `scripts/update.sh` — pulls the latest stack and re-runs setup

A terminal cockpit called `c3` (installable via `uv pip install -e tools/serve-cockpit`) provides a screen-based UI for browsing the catalog, launching configs, watching GPU and container state, and running health checks — all without touching the CLI directly.

## Multi-Engine and Multi-GPU Architecture

Configs are organized under `models/<model>/<engine>/compose/<topology>/<quant>/<serving>.yml`, making the structure explicit: one compose file per "slug" (a unique combination of model, engine, topology, and quantization). The repository currently ships configs for:

- **Single-card (1× RTX 3090)** — documented in `docs/SINGLE_CARD.md`
- **Dual-card (2× RTX 3090)** — documented in `docs/DUAL_CARD.md`
- **Multi-card (4× and 8× GPUs)** — documented in `docs/MULTI_CARD.md`

Supported engines include vLLM, SGLang, llama.cpp, ik_llama, and exllamav3. The configs are hardware-class aware, and contributors have validated them on 4090s, 5090s, and mixed-architecture rigs (e.g., 5090 + 3090 Ti in one box).

## Beyond Chat: AI Studio

The repository includes a **Club 3090 AI Studio** extension (`docs/ai-studio/`) that adds open-weight image, video, and audio generation alongside a chat model, driven from Open WebUI. The image bundle launches with a single script: `bash scripts/setup-image-studio.sh`.

## Update: v0.12.0

The latest release is **v0.12.0**, published on 2026-09-29. The repository was created in April 2026 and has seen active development since, with the GitHub About description noting current configs for Qwen3.6-27B, Qwen3.6 35B, Gemma 4 26B, and Gemma 4 31B across 1× and 2× card topologies. The repo's Announcements discussion category is the canonical source for the newest slugs and benchmark numbers, with the docs catching up afterward. A community project, `VykosX/club-3090-server`, adds a browser admin panel, OpenAI-compatible reverse proxy, and GPU-aware multi-instance orchestration on top of the base repo.

## Community and Contribution Model

The project uses GitHub Discussions for async benchmark sharing and cross-rig comparisons, GitHub Issues for bug reports, and a Discord server for synchronous Q&A and hardware questions. Benchmark results from community contributors across different GPU rigs are collected in `BENCHMARKS.md`, with a standardized Results Card format (including per-pack dispersion and p50/p95 latency) adopted from community input.

## Features
- Working Docker Compose configs for vLLM, SGLang, llama.cpp, ik_llama, and exllamav3
- OpenAI-compatible API endpoint out of the box
- Single-command setup, launch, and switch scripts
- SHA-verified model weight downloads
- Hardware-class-aware configs for 1×, 2×, 4×, and 8× GPU topologies
- Terminal cockpit (c3) for browsing catalog and managing serving
- Benchmark table with measured TPS and latency numbers
- AI Studio extension for image, video, and audio generation via Open WebUI
- Bring-your-own-model support with VRAM pre-check
- Quality and health check eval scripts
- Troubleshooting guide, FAQ, glossary, and quantization reference docs

## Integrations
vLLM, SGLang, llama.cpp, ik_llama, exllamav3, Docker, NVIDIA Container Toolkit, Open WebUI, Hugging Face, WSL2

## Platforms
WINDOWS, LINUX, API, CLI

## Pricing
Open Source

## Version
v0.12.0

## Links
- Website: https://github.com/noonghunna/club-3090
- Documentation: https://github.com/noonghunna/club-3090/blob/master/docs/README.md
- Repository: https://github.com/noonghunna/club-3090
- EveryDev.ai: https://www.everydev.ai/tools/club-3090
