EveryDev.ai
Subscribe
Home
Developers

3,669+ AI companies

  • Radar
  • Trending
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • Users
  • Rate Tools
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Developers
    3. llmrun

    llmrun

    llmrun is a public web tool for answering whether a particular large language model can run locally on a user's GPU or Apple Silicon device. It compares model VRAM requirements, hardware fit, and estimated decode performance, helping users choose models and local inference setups.

    Visit Website

    At a Glance

    1Tool Listed
    2Products
    10Capabilities
    Discussions
    Focus Areas
    Local Inference
    LLM Evaluations
    Performance Metrics
    Latest News
    Why averaging LLM benchmarks gives the wrong leaderboard — llmrun describes the failure of its first composite-ranking methodology and its active benchmark-panel fix.Oct 5, 2026
    Estimating tokens/s for Mixture-of-Experts models: active parameters, plus a routing term — llmrun documents its current decode-speed estimator and calibration limits.Oct 5, 2026
    Markets
    • Developers running open-weight LLMs locally
    • AI hobbyists and home-lab users
    • Researchers and engineers evaluating local inference hardware
    • Users of NVIDIA, AMD, Intel, and Apple Silicon systems
    • +1 more

    AI Tools by llmrun

    (1)
    View llmrun
    llmrun tool icon

    llmrun

    Local LLM GPU Compatibility Checker

    Local InferenceLLM EvaluationsPerformance Metrics

    Discussions

    No discussions yet

    Be the first to start a discussion about llmrun

    Latest News

    10/05/2026

    Why averaging LLM benchmarks gives the wrong leaderboard — llmrun describes the failure of its first composite-ranking methodology and its active benchmark-panel fix.

    dev.to
    10/05/2026

    Estimating tokens/s for Mixture-of-Experts models: active parameters, plus a routing term — llmrun documents its current decode-speed estimator and calibration limits.

    dev.to

    Products & Services

    2
    llmrun local LLM hardware checker and model browser
    June 2026

    A browser-based catalog that ranks open-weight language models for selected GPUs and Apple Silicon devices. It shows model parameter counts, quantization, VRAM requirements, context length, estimated tokens per second, quality score, and a hardware-fit grade; users can then install models through Ollama or LM Studio or download GGUF files.

    llmrun Score benchmark leaderboard
    June 2026

    A composite open-model leaderboard covering categories such as reasoning, coding, and mathematics. The methodology evolved from equal-weight category averages to an active benchmark panel and is documented on the llmrun benchmark pages.

    Market Position

    llmrun positions itself as a hardware-aware local-LLM decision tool rather than an inference host: it combines a model catalog, VRAM-fit analysis, speed estimates, hardware pages, and benchmark rankings in one browser experience. Its closest adjacent alternatives identified in web search are CanIRun.ai and WillItRunAI, which also match open models to local hardware, while llmfit is a related command-line/TUI approach.

    Founding Story

    The project was started to answer the practical question of which open-weight LLM a person's existing hardware can actually run and how fast it will be. Its initial vision was a straightforward hardware-aware model ranking; the first version launched in June 2026 and was later revised after the creators found that simple averaging produced misleading benchmark rankings.

    Target Markets

    Industries & Segments
    • Developers running open-weight LLMs locally
    • AI hobbyists and home-lab users
    • Researchers and engineers evaluating local inference hardware
    • Users of NVIDIA, AMD, Intel, and Apple Silicon systems
    • Teams seeking private or offline model deployment
    Use Cases
    • Choosing an open-weight LLM that fits a user's GPU or Apple Silicon memory
    • Planning local, private, or offline LLM inference
    • Selecting quantization levels and estimating whether a model will leave usable context headroom
    • Comparing expected local generation speed before downloading a model
    • Evaluating local chat, coding assistance, reasoning, summarization, visual question answering, and document workloads
    • Comparing open-model benchmark performance through the llmrun Score leaderboard

    History & Milestones

    June 2026

    The first version of the llmrun composite model-ranking system launched, using normalized benchmark results and category averages.

    September 20, 2026

    The project identified a failure in its original leaderboard methodology: models evaluated on only a few easy benchmarks could outrank models evaluated on many harder, newer benchmarks.

    October 5, 2026

    llmrun documented its revised active-panel methodology, in which benchmarks count only while their newest result comes from a model released within the prior 365 days; the article reports that Kimi K3 ranked first of 67 under the revised system.

    October 5, 2026

    The project published an updated Mixture-of-Experts speed estimator using active-parameter scaling plus a per-layer routing-overhead term, calibrated against llama.cpp benchmarks.

    Key Capabilities

    10
    Hardware-aware model compatibility rankings with S-to-F fit grades
    VRAM estimation from quantized model weights, KV cache, and framework overhead
    Support for discrete GPUs, Apple Silicon, MacBooks, desktops, mini PCs, and AI development kits
    Quantization comparisons including F16, Q8_0, Q6_K, Q4_K_M, Q4_K_S, and Q2_K
    Estimated decode tokens per second calibrated by hardware platform
    Separate prefill and time-to-first-token estimates

    Integrations & Partnerships

    Platform Integrations

    • Ollama
    • LM Studio
    • GGUF files
    • llama.cpp benchmark and inference ecosystem
    • Hugging Face model downloads

    Connect

    Website
    llmrun.dev

    AI Topics

    3

    llmrun focuses on these topics:

    Local Inference(1)
    LLM Evaluations(1)
    Performance Metrics(1)
    Back to all developersSuggest an edit