HungryGPU
HungryGPU is a hardware-aware open-source inference catalogue: it helps people discover what they can run on their machine, with community recipes and patches for local AI. It reads merged pull requests, model releases, community setups and papers, and tags them with the machines they help or leave out.
At a Glance
- Local AI and open-weight model users
- Developers and researchers running inference on consumer GPUs, workstations, Apple Silicon, AMD systems and data-center accelerators
- AI infrastructure and inference-engine developers
- AI agents needing machine-aware model and patch discovery
AI Tools by HungryGPU
(1)HungryGPU
GPU Filtered AI Model Tracker
Discussions
No discussions yet
Be the first to start a discussion about HungryGPU
Latest News
HungryGPU catalogue reports 2,102 community setups, 1,069 open models and 1,276 merged PRs read
GLM-5.3-Flash DGX Spark repository adds stock-baseline and configuration-tuning measurements
Published one-file vLLM overlay and runbook for GLM-5.3-Flash NVFP4 on DGX Spark
Show HN launch: HungryGPU – Track local AI models, patches and recipes by hardware
Products & Services
A web catalogue for discovering local AI models, community inference recipes, hardware compatibility, measured throughput, papers, creators and software changes. It currently covers 25 machine profiles and distinguishes measured runs from arithmetic memory-fit estimates.
Published setups pairing models, engines, quantisation and hardware with reported speeds, conditions and source repositories; the projects page reported 2,102 recipes and 9,410 runs across all machines on 2026-09-13.
A feed of merged upstream changes, corrections, improvements and new capabilities, including evidence from diffs and explicit machines that a change leaves out.
MIT-licensed open-source repository containing a drop-in Python overlay, serving scripts and runbook for running NVIDIA's official GLM-5.3-Flash NVFP4 checkpoint on two DGX Spark GB10 nodes using vLLM, Ray and a direct link.
Market Position
HungryGPU positions itself as a hardware-aware complement to general model and code repositories: instead of only cataloguing models or source code, it connects a model or software change to concrete machines, published run conditions, measured throughput and explicit compatibility exclusions. Its differentiation is the machine-specific evidence trail and the deliberate separation of measured results, estimates and unknowns.
Founding Story
The site presents itself as a beta built around a practical gap in local AI: model and software announcements often do not say whether a particular machine can run them. HungryGPU was created to make the answer hardware-specific by collecting runnable recipes, measured results and compatibility exclusions, while keeping estimates and unknowns explicitly separated from actual runs.
Business Model
Revenue Model
The public site provides catalogue browsing and MCP access without an account; an optional token adds setup synchronisation and queueing features. No paid plan or other monetisation mechanism was identified on the public pages.
Pricing Tiers
Browsing and MCP access are available without an account.
Target Markets
- Local AI and open-weight model users
- Developers and researchers running inference on consumer GPUs, workstations, Apple Silicon, AMD systems and data-center accelerators
- AI infrastructure and inference-engine developers
- AI agents needing machine-aware model and patch discovery
- Finding open models that fit and run on a user's specific GPU or Apple/AMD system
- Reproducing local LLM inference configurations from community-published recipes
- Comparing measured throughput across hardware, engines, quantisation and concurrency
- Assessing whether an upstream patch or software change helps or excludes a particular machine
- Running GLM-5.3-Flash NVFP4 on DGX Spark GB10 hardware
- Agent-driven discovery through the site's MCP interface