llmrun
llmrun is a public web tool for answering whether a particular large language model can run locally on a user's GPU or Apple Silicon device. It compares model VRAM requirements, hardware fit, and estimated decode performance, helping users choose models and local inference setups.
At a Glance
- Developers running open-weight LLMs locally
- AI hobbyists and home-lab users
- Researchers and engineers evaluating local inference hardware
- Users of NVIDIA, AMD, Intel, and Apple Silicon systems
- +1 more
AI Tools by llmrun
(1)llmrun
Local LLM GPU Compatibility Checker
Discussions
No discussions yet
Be the first to start a discussion about llmrun
Latest News
Why averaging LLM benchmarks gives the wrong leaderboard — llmrun describes the failure of its first composite-ranking methodology and its active benchmark-panel fix.
Estimating tokens/s for Mixture-of-Experts models: active parameters, plus a routing term — llmrun documents its current decode-speed estimator and calibration limits.
Products & Services
A browser-based catalog that ranks open-weight language models for selected GPUs and Apple Silicon devices. It shows model parameter counts, quantization, VRAM requirements, context length, estimated tokens per second, quality score, and a hardware-fit grade; users can then install models through Ollama or LM Studio or download GGUF files.
A composite open-model leaderboard covering categories such as reasoning, coding, and mathematics. The methodology evolved from equal-weight category averages to an active benchmark panel and is documented on the llmrun benchmark pages.
Market Position
llmrun positions itself as a hardware-aware local-LLM decision tool rather than an inference host: it combines a model catalog, VRAM-fit analysis, speed estimates, hardware pages, and benchmark rankings in one browser experience. Its closest adjacent alternatives identified in web search are CanIRun.ai and WillItRunAI, which also match open models to local hardware, while llmfit is a related command-line/TUI approach.
Founding Story
The project was started to answer the practical question of which open-weight LLM a person's existing hardware can actually run and how fast it will be. Its initial vision was a straightforward hardware-aware model ranking; the first version launched in June 2026 and was later revised after the creators found that simple averaging produced misleading benchmark rankings.
Target Markets
- Developers running open-weight LLMs locally
- AI hobbyists and home-lab users
- Researchers and engineers evaluating local inference hardware
- Users of NVIDIA, AMD, Intel, and Apple Silicon systems
- Teams seeking private or offline model deployment
- Choosing an open-weight LLM that fits a user's GPU or Apple Silicon memory
- Planning local, private, or offline LLM inference
- Selecting quantization levels and estimating whether a model will leave usable context headroom
- Comparing expected local generation speed before downloading a model
- Evaluating local chat, coding assistance, reasoning, summarization, visual question answering, and document workloads
- Comparing open-model benchmark performance through the llmrun Score leaderboard