EveryDev.ai
Subscribe
Home
Developers

3,289+ AI companies

  • Radar
  • Trending
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Developers
    3. HungryGPU

    HungryGPU

    HungryGPU is a hardware-aware open-source inference catalogue: it helps people discover what they can run on their machine, with community recipes and patches for local AI. It reads merged pull requests, model releases, community setups and papers, and tags them with the machines they help or leave out.

    Visit Website

    At a Glance

    1Tool Listed
    4Products
    9Capabilities
    Discussions
    Focus Areas
    Local Inference
    Model Management
    MCP Servers
    Connect
    Latest News
    HungryGPU catalogue reports 2,102 community setups, 1,069 open models and 1,276 merged PRs readSep 13, 2026
    GLM-5.3-Flash DGX Spark repository adds stock-baseline and configuration-tuning measurementsSep 13, 2026
    Markets
    • Local AI and open-weight model users
    • Developers and researchers running inference on consumer GPUs, workstations, Apple Silicon, AMD systems and data-center accelerators
    • AI infrastructure and inference-engine developers
    • AI agents needing machine-aware model and patch discovery

    AI Tools by HungryGPU

    (1)
    View HungryGPU
    HungryGPU tool icon

    HungryGPU

    GPU Filtered AI Model Tracker

    Local InferenceModel ManagementMCP Servers

    Discussions

    No discussions yet

    Be the first to start a discussion about HungryGPU

    Latest News

    09/13/2026

    HungryGPU catalogue reports 2,102 community setups, 1,069 open models and 1,276 merged PRs read

    hungrygpu.com
    09/13/2026

    GLM-5.3-Flash DGX Spark repository adds stock-baseline and configuration-tuning measurements

    github.com
    09/12/2026

    Published one-file vLLM overlay and runbook for GLM-5.3-Flash NVFP4 on DGX Spark

    github.com
    09/01/2026

    Show HN launch: HungryGPU – Track local AI models, patches and recipes by hardware

    news.ycombinator.com

    Products & Services

    4
    HungryGPU hardware-aware inference catalogue
    September 2026

    A web catalogue for discovering local AI models, community inference recipes, hardware compatibility, measured throughput, papers, creators and software changes. It currently covers 25 machine profiles and distinguishes measured runs from arithmetic memory-fit estimates.

    HungryGPU community recipes
    September 2026

    Published setups pairing models, engines, quantisation and hardware with reported speeds, conditions and source repositories; the projects page reported 2,102 recipes and 9,410 runs across all machines on 2026-09-13.

    HungryGPU patches and compatibility exclusions
    September 2026

    A feed of merged upstream changes, corrections, improvements and new capabilities, including evidence from diffs and explicit machines that a change leaves out.

    GLM-5.3-Flash NVFP4 on DGX Spark overlay
    2026-09-12

    MIT-licensed open-source repository containing a drop-in Python overlay, serving scripts and runbook for running NVIDIA's official GLM-5.3-Flash NVFP4 checkpoint on two DGX Spark GB10 nodes using vLLM, Ray and a direct link.

    Market Position

    HungryGPU positions itself as a hardware-aware complement to general model and code repositories: instead of only cataloguing models or source code, it connects a model or software change to concrete machines, published run conditions, measured throughput and explicit compatibility exclusions. Its differentiation is the machine-specific evidence trail and the deliberate separation of measured results, estimates and unknowns.

    Founding Story

    The site presents itself as a beta built around a practical gap in local AI: model and software announcements often do not say whether a particular machine can run them. HungryGPU was created to make the answer hardware-specific by collecting runnable recipes, measured results and compatibility exclusions, while keeping estimates and unknowns explicitly separated from actual runs.

    Business Model

    Revenue Model

    The public site provides catalogue browsing and MCP access without an account; an optional token adds setup synchronisation and queueing features. No paid plan or other monetisation mechanism was identified on the public pages.

    Pricing Tiers

    Public catalogue
    Free

    Browsing and MCP access are available without an account.

    Target Markets

    Industries & Segments
    • Local AI and open-weight model users
    • Developers and researchers running inference on consumer GPUs, workstations, Apple Silicon, AMD systems and data-center accelerators
    • AI infrastructure and inference-engine developers
    • AI agents needing machine-aware model and patch discovery
    Use Cases
    • Finding open models that fit and run on a user's specific GPU or Apple/AMD system
    • Reproducing local LLM inference configurations from community-published recipes
    • Comparing measured throughput across hardware, engines, quantisation and concurrency
    • Assessing whether an upstream patch or software change helps or excludes a particular machine
    • Running GLM-5.3-Flash NVFP4 on DGX Spark GB10 hardware
    • Agent-driven discovery through the site's MCP interface

    History & Milestones

    September 2026

    HungryGPU launched as a Show HN project for tracking local AI models, patches and recipes by hardware; the site described itself as beta with feedback welcome.

    2026-09-13

    The catalogue reported 25 tracked machines, 2,102 community setups, 1,069 open-model releases in its collection window and 1,276 merged pull requests read across eight repositories over 30 days.

    2026-09-12

    The associated tenhkspark repository published a one-file vLLM overlay that makes NVIDIA's GLM-5.3-Flash NVFP4 checkpoint run on DGX Spark (GB10/sm_121), with a documented no-rope sparse-MLA fix and two-node serving runbook.

    2026-09-13

    The repository added measured stock-baseline and configuration-tuning results for GLM-5.3-Flash NVFP4 on DGX Spark.

    Key Capabilities

    9
    Machine-specific model and inference compatibility tagging across 25 hardware profiles
    Community recipes with engines, quantisation, reported throughput, context length, concurrency and conditions
    Measured results kept distinct from estimates; memory-fit estimates are arithmetic rather than benchmark claims
    Compatibility exclusions showing changes that leave a machine out, with evidence from source diffs
    Tracking of merged pull requests, open-weight releases, papers, third-party scores and OpenRouter applications
    Plain-text search plus model facets for attributes such as quantisation and MoE

    Integrations & Partnerships

    Platform Integrations

    • Public MCP interface at https://hungrygpu.com/api/mcp
    • Atom feed at https://hungrygpu.com/feed.xml
    • GitHub source repositories and creator profiles
    • Hugging Face model links and release metadata
    • Supported inference stacks represented in recipes include vLLM, llama.cpp, SGLang, MLX, Ollama and TensorRT-LLM
    • The supplied GLM project integrates vLLM's official GLM-5.3-Flash arm64 CUDA 13 image, FlashInfer, Ray tensor parallelism and direct QSFP networking between DGX Spark nodes

    Key Partnerships

    The catalogue ingests or links to upstream work from eight software repositories and links recipes to their source GitHub repositories.
    The site's public methodology cites Artificial Analysis Intelligence Index v4.3 for the open-versus-proprietary trend visualization.
    The supplied GLM-5.3-Flash overlay works with NVIDIA's official checkpoint/image, vLLM and Ray; the repository acknowledges Z.ai and NVIDIA as upstream model/image publishers.

    Connect

    Website
    hungrygpu.com
    GitHub
    tenhkspark

    AI Topics

    3

    HungryGPU focuses on these topics:

    Local Inference(1)
    Model Management(1)
    MCP Servers(1)
    Back to all developersSuggest an edit