EveryDev.ai
Subscribe
Home
Tools

4,153+ AI tools

  • New
  • Trending
  • Featured
  • Compare
  • Arena
Categories
  • Agents2782
  • Coding1973
  • Infrastructure825
  • Projects603
  • Marketing598
  • Research520
  • Analytics468
  • Design462
  • MCP419
  • Testing346
  • Security323
  • Data305
  • Integration224
  • Prompts220
  • Communication210
  • Extensions196
  • Learning179
  • Voice175
  • Commerce160
  • DevOps135
  • Web95
  • Finance31
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Tools
    3. club-3090
    club-3090 icon

    club-3090

    Local Inference

    Community recipes and Docker Compose configs for serving LLMs locally on RTX 3090s and other NVIDIA GPUs, with an OpenAI-compatible API via vLLM, SGLang, llama.cpp, and more.

    Visit Website

    At a Glance

    Pricing
    Open Source

    Fully free and open-source under Apache 2.0. Clone, use, modify, and distribute freely.

    Engagement

    Available On

    Windows
    Linux
    API
    CLI

    Resources

    WebsiteDocsGitHubllms.txt

    Topics

    Local InferenceLLM OrchestrationAI Infrastructure

    Alternatives

    SyntheticBodega Inference EngineFreeLLMAPI
    Developer
    noonghunnanoonghunna maintains club-3090, an open-source community pro…

    Listed Oct 2026

    About club-3090

    Club-3090 is an open-source collection of working Docker Compose configurations, patches, and benchmark data for running large language models locally on consumer NVIDIA GPUs — primarily the RTX 3090, but also tested on 4090s, 5090s, and A-series cards. Maintained by the GitHub user noonghunna and licensed under Apache 2.0, the project lets users pick a model, launch it with a single command, and get an OpenAI-compatible API endpoint on their own hardware.

    What It Is

    Club-3090 is a local inference recipe library: a structured repository of per-model, per-engine, per-topology Docker Compose files that abstract away the complexity of running quantized LLMs on consumer GPUs. It is not a new inference engine itself — it sits on top of established engines (vLLM, SGLang, llama.cpp, ik_llama, exllamav3) and provides the glue: verified compose configs, vendored patches, setup scripts, and a benchmark table so users know what performance to expect before they download anything.

    How the Workflow Works

    The setup path is designed to go from clone to first reply in about five minutes:

    • scripts/setup.sh — picks a model, downloads weights, and SHA-verifies them
    • scripts/launch.sh — selects a config for the detected GPU topology and boots it with a health check
    • scripts/switch.sh — lists all configs the current machine can run and switches between them
    • scripts/update.sh — pulls the latest stack and re-runs setup

    A terminal cockpit called c3 (installable via uv pip install -e tools/serve-cockpit) provides a screen-based UI for browsing the catalog, launching configs, watching GPU and container state, and running health checks — all without touching the CLI directly.

    Multi-Engine and Multi-GPU Architecture

    Configs are organized under models/<model>/<engine>/compose/<topology>/<quant>/<serving>.yml, making the structure explicit: one compose file per "slug" (a unique combination of model, engine, topology, and quantization). The repository currently ships configs for:

    • Single-card (1× RTX 3090) — documented in docs/SINGLE_CARD.md
    • Dual-card (2× RTX 3090) — documented in docs/DUAL_CARD.md
    • Multi-card (4× and 8× GPUs) — documented in docs/MULTI_CARD.md

    Supported engines include vLLM, SGLang, llama.cpp, ik_llama, and exllamav3. The configs are hardware-class aware, and contributors have validated them on 4090s, 5090s, and mixed-architecture rigs (e.g., 5090 + 3090 Ti in one box).

    Beyond Chat: AI Studio

    The repository includes a Club 3090 AI Studio extension (docs/ai-studio/) that adds open-weight image, video, and audio generation alongside a chat model, driven from Open WebUI. The image bundle launches with a single script: bash scripts/setup-image-studio.sh.

    Update: v0.12.0

    The latest release is v0.12.0, published on 2026-09-29. The repository was created in April 2026 and has seen active development since, with the GitHub About description noting current configs for Qwen3.6-27B, Qwen3.6 35B, Gemma 4 26B, and Gemma 4 31B across 1× and 2× card topologies. The repo's Announcements discussion category is the canonical source for the newest slugs and benchmark numbers, with the docs catching up afterward. A community project, VykosX/club-3090-server, adds a browser admin panel, OpenAI-compatible reverse proxy, and GPU-aware multi-instance orchestration on top of the base repo.

    Community and Contribution Model

    The project uses GitHub Discussions for async benchmark sharing and cross-rig comparisons, GitHub Issues for bug reports, and a Discord server for synchronous Q&A and hardware questions. Benchmark results from community contributors across different GPU rigs are collected in BENCHMARKS.md, with a standardized Results Card format (including per-pack dispersion and p50/p95 latency) adopted from community input.

    club-3090 - 1

    Community Discussions

    Be the first to start a conversation about club-3090

    Share your experience with club-3090, ask questions, or help others learn from your insights.

    Pricing

    OPEN SOURCE

    Open Source

    Fully free and open-source under Apache 2.0. Clone, use, modify, and distribute freely.

    • All Docker Compose configs for vLLM, SGLang, llama.cpp, ik_llama, exllamav3
    • Setup, launch, switch, and update scripts
    • Benchmark data and results cards
    • Terminal cockpit (c3)
    • AI Studio for image/video/audio generation

    Capabilities

    Key Features

    • Working Docker Compose configs for vLLM, SGLang, llama.cpp, ik_llama, and exllamav3
    • OpenAI-compatible API endpoint out of the box
    • Single-command setup, launch, and switch scripts
    • SHA-verified model weight downloads
    • Hardware-class-aware configs for 1×, 2×, 4×, and 8× GPU topologies
    • Terminal cockpit (c3) for browsing catalog and managing serving
    • Benchmark table with measured TPS and latency numbers
    • AI Studio extension for image, video, and audio generation via Open WebUI
    • Bring-your-own-model support with VRAM pre-check
    • Quality and health check eval scripts
    • Troubleshooting guide, FAQ, glossary, and quantization reference docs

    Integrations

    vLLM
    SGLang
    llama.cpp
    ik_llama
    exllamav3
    Docker
    NVIDIA Container Toolkit
    Open WebUI
    Hugging Face
    WSL2
    API Available
    View Docs

    Ratings & Reviews

    No ratings yet

    Be the first to rate club-3090 and help others make informed decisions.

    Developer

    noonghunna

    noonghunna maintains club-3090, an open-source community project providing recipes and Docker Compose configs for serving large language models locally on consumer NVIDIA GPUs. The project ships verified multi-engine configs (vLLM, llama.cpp, SGLang, and more), benchmark data, and tooling to get an OpenAI-compatible API running on RTX 3090s and similar hardware with a single command. Development is community-driven, with contributors submitting benchmark results, patches, and new model configs via GitHub.

    Read more about noonghunna
    WebsiteGitHub
    1 tool in directory

    Similar Tools

    Synthetic icon

    Synthetic

    AI platform providing access to multiple LLMs with subscription or usage-based pricing, offering both UI and API access.

    Bodega Inference Engine icon

    Bodega Inference Engine

    Enterprise-grade local LLM inference engine built specifically for Apple Silicon, featuring a multi-model registry, OpenAI-compatible API, and high-throughput continuous batching.

    FreeLLMAPI icon

    FreeLLMAPI

    An open-source, self-hosted router that aggregates free LLM tiers from 34 providers behind a single OpenAI-compatible endpoint with smart routing and automatic failover.

    Browse all tools

    Related Topics

    Local Inference

    Tools and platforms for running AI inference locally without cloud dependence.

    221 tools

    LLM Orchestration

    Platforms and frameworks for designing, managing, and deploying complex LLM workflows with visual interfaces, allowing for the coordination of multiple AI models and services.

    247 tools

    AI Infrastructure

    Infrastructure designed for deploying and running AI models.

    416 tools
    Browse all topics
    Back to all toolsSuggest an edit
    ratings
    discussions