EveryDev.ai
Subscribe
Home
Tools

3,569+ AI tools

  • New
  • Trending
  • Featured
  • Compare
  • Arena
Categories
  • Agents2189
  • Coding1574
  • Infrastructure698
  • Marketing534
  • Projects498
  • Research456
  • Design416
  • Analytics389
  • Testing296
  • MCP290
  • Security286
  • Data262
  • Integration197
  • Prompts189
  • Communication183
  • Extensions173
  • Learning170
  • Voice151
  • Commerce135
  • DevOps123
  • Web86
  • Finance26
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Tools
    3. inferock-bench
    inferock-bench icon

    inferock-bench

    LLM Evaluations
    Featured

    A local LLM cost-tracking proxy that records per-call receipts for OpenAI, Anthropic, Gemini, and OpenRouter calls with token usage, failure evidence, and billing-integrity signals.

    Visit Website

    At a Glance

    Pricing
    Open Source
    Free tier available

    Keep your own provider accounts and pay providers directly. Inferock measures eligible calls in path and reports which calls failed and what they cost on your provider bill. No Inferock service credit applies.

    Managed: Custom/contact

    Engagement

    Available On

    CLI
    Web
    API

    Resources

    WebsiteDocsGitHubllms.txt

    Topics

    LLM EvaluationsObservability PlatformsAI Infrastructure

    Alternatives

    FoglampAtla AIFuture AGI
    Developer
    OpiusAISan Francisco, CAEst. 2024

    Listed Aug 2026

    About inferock-bench

    inferock-bench is an open-source local proxy built by OpiusAI that sits between your application and AI provider APIs, recording independent per-call receipts with spend, bill-bounded money loss, time loss, and invoice-check exposure. It supports OpenAI, Anthropic (Claude), Gemini Developer API, and pinned OpenRouter endpoints, and is distributed as an npm package runnable via npx inferock-bench. The project launched in July 2026 and has accumulated 127 GitHub stars and 49 forks as of mid-August 2026.

    What It Is

    inferock-bench is a local diagnostic measurement proxy for metered LLM API traffic. It intercepts provider calls routed through localhost, attaches your provider API key only to outbound provider requests, and writes local event records that are graded by the @inferock/measure library against The Inferock Standard — a published, versioned rulebook for what counts as a billable failure, what is provider-recognized recoverable, and what stays as invoice-check exposure. The tool addresses a specific gap: AI providers report usage totals but do not give customers per-call receipts that cross-check delivery evidence against charges.

    How the Receipt Works

    Every proxied call produces a structured receipt with four labeled headline numbers:

    • spent — provider spend observed by the run for priced calls the proxy saw
    • money loss — bill-bounded dollar loss tied to observed spend or charge evidence under The Inferock Standard
    • time loss — real wait or downtime measured as time, never added to dollars
    • invoice-check exposure — amounts like cache-discount-at-risk, labeled "verify your invoice" and never summed into money loss

    The receipt also reports provider-recognized recovery, a bill-bounded recognition gap, and a coverage line such as surfaces watched 10/12 | signals 3 | not-openable 2. A zero only counts for a watched surface; unopened surfaces are named as coverage debt rather than silently claimed clean.

    Supported Providers and Integrations

    The tool measures four provider planes: OpenAI, Anthropic, Gemini Developer API, and pinned OpenRouter endpoints covering meta-llama, deepseek, mistral, moonshot/kimi, z-ai/glm, and qwen on observed hosts. Integration guides exist for Claude Code, the OpenAI SDK, the Gemini Developer API, OpenRouter, and CI/headless usage. Pointing an existing SDK at the local proxy requires changing only two settings — apiKey (to the local ibl_ bench key) and baseURL (to http://127.0.0.1:4318).

    Architecture and Key Boundary

    inferock-bench runs as a local Node.js process with a browser dashboard. Provider keys are saved locally under ~/.inferock-bench/ with owner-only file permissions, are shown back only in masked form, and are attached only to outbound provider requests — they are not sent to Inferock's hosted service. The generated local ibl_ bench key is a local-only credential. The @inferock/measure grading library is Apache-2.0 licensed and ships separately on npm; the benchmark CLI uses FSL-1.1-ALv2 (converting to Apache-2.0 after two years); The Inferock Standard documents are CC-BY-4.0.

    Update: Public Run Card 2026-08-05

    The current cumulative public ledger (as of the 2026-08-05 addendum) contains 1,303 measured calls across OpenAI, Anthropic, Gemini, and pinned OpenRouter coverage, with $8.43 provider spend observed, $0.03 bill-bounded money loss (0.3%), approximately 2.9 minutes of time loss, and $18.88 cache-discount-at-risk invoice-check exposure kept separate from money loss. The current receipt watches 12 of 13 surfaces. The public receipt presentation was introduced in version 0.1.10 and re-rendered in the 0.2.4 Nominal Light product UI. The project also includes a pre-launch Reliability Index feature that lets users opt in locally to preview an anonymized aggregate payload before any data is sent to a public backend.

    Why It Matters

    The README cites a third-party audit firm report from June 2026 claiming it reviewed approximately $34M of AI invoices and found approximately $1.7M in overbilling — the source attribution and caveats are linked in the spec annex. The tool's stated purpose is to give customers independent, per-call evidence rather than relying solely on provider-reported totals. Providers including Anthropic and OpenAI have issued denials of widespread overbilling; the README includes these as scope boundaries and treats them as claims to test against local per-call evidence rather than as admissions.

    inferock-bench - 1

    Community Discussions

    Be the first to start a conversation about inferock-bench

    Share your experience with inferock-bench, ask questions, or help others learn from your insights.

    Pricing

    FREE

    Bring Your Own Keys

    Keep your own provider accounts and pay providers directly. Inferock measures eligible calls in path and reports which calls failed and what they cost on your provider bill. No Inferock service credit applies.

    • Provider bill remains between you and your providers
    • Loss report shows objective failures, thresholded signals, and review-only flags
    • No Inferock service credit in this mode

    Managed

    Inferock operates the provider relationship for selected traffic. Eligible objective failures can receive bounded service credits under agreed terms. Final platform fee and service credit terms set at onboarding.

    Custom
    contact sales
    • Inferock-operated provider relationship for selected traffic
    • Bounded service credits for eligible objective failures
    • Per-call receipts with spend, money loss, time loss, and invoice-check exposure
    • Objective failure coverage: downtime, timeouts, empty output, truncated output, invalid structured output, billing anomalies
    • Latency, drift, refusals, factuality, and policy signals with thresholds or review
    View official pricing

    Capabilities

    Key Features

    • Local LLM cost-tracking proxy via localhost
    • Per-call receipts with spend, money loss, time loss, and invoice-check exposure
    • Supports OpenAI, Anthropic, Gemini Developer API, and pinned OpenRouter endpoints
    • Independent billing-integrity signals and token cross-check
    • Browser dashboard with provider key setup and receipt viewer
    • CLI commands: start, setup, test, receipt, status, key reveal/copy, init
    • Built-in test battery with spend estimate and consent gate before provider calls
    • Coverage state per surface: watched-clean, signal, or not-openable
    • Agent test mode for OpenAI and Anthropic runs
    • CI/headless usage support
    • FSL-1.1-ALv2 license converting to Apache-2.0 after 2 years
    • Open-source @inferock/measure grading library (Apache-2.0)
    • Published Inferock Standard rulebook (CC-BY-4.0)
    • Pre-launch Reliability Index with opt-in anonymized aggregate payload

    Integrations

    OpenAI SDK
    Anthropic Claude Code
    Gemini Developer API
    OpenRouter
    meta-llama
    deepseek
    mistral
    moonshot/kimi
    z-ai/glm
    qwen
    CI/CD pipelines
    API Available
    View Docs

    Ratings & Reviews

    No ratings yet

    Be the first to rate inferock-bench and help others make informed decisions.

    Developer

    OpiusAI

    OpiusAI builds Inferock, an AI-provider accountability platform that provides reliable and accountable LLM inference with independent per-call receipts. The company was founded by Bharath Koneti and Himashwetha Gowda, who created the open-source inferock-bench local proxy and The Inferock Standard — a published, versioned rulebook for AI billing integrity. OpiusAI operates both a hosted managed inference product and an open-source measurement tool, positioning itself as an independent third party between AI customers and providers.

    Founded 2024
    San Francisco, CA
    10 employees

    Used by

    Billboard.com (mentioned in similar…
    Early stage AI startups and developers…
    Read more about OpiusAI
    WebsiteGitHubLinkedInX / Twitter
    1 tool in directory

    Similar Tools

    Foglamp icon

    Foglamp

    Open-source observability layer for AI agents built on the Vercel AI SDK — tracks cost, latency, token usage, distributed traces, and evals with two lines of code.

    Atla AI icon

    Atla AI

    Atla AI is an AI evaluation platform that helps teams assess and improve the quality of large language model outputs.

    Future AGI icon

    Future AGI

    An AI lifecycle platform for building, evaluating, monitoring, and securing generative AI agents with hallucination detection, simulations, and real-time guardrails.

    Browse all tools

    Related Topics

    LLM Evaluations

    Platforms and frameworks for evaluating, testing, and benchmarking LLM systems and AI applications. These tools provide evaluators and evaluation models to score AI outputs, measure hallucinations, assess RAG quality, detect failures, and optimize model performance. Features include automated testing with LLM-as-a-judge metrics, component-level evaluation with tracing, regression testing in CI/CD pipelines, custom evaluator creation, dataset curation, and real-time monitoring of production systems. Teams use these solutions to validate prompt effectiveness, compare models side-by-side, ensure answer correctness and relevance, identify bias and toxicity, prevent PII leakage, and continuously improve AI product quality through experiments, benchmarks, and performance analytics.

    116 tools

    Observability Platforms

    Comprehensive platforms that combine metrics, logs, and traces with AI-powered analytics to provide deep insights into complex distributed systems and application behavior.

    114 tools

    AI Infrastructure

    Infrastructure designed for deploying and running AI models.

    347 tools
    Browse all topics
    Back to all toolsSuggest an edit
    ratings
    discussions