EveryDev.ai
Subscribe
Home
Tools

4,072+ AI tools

  • New
  • Trending
  • Featured
  • Compare
  • Arena
Categories
  • Agents2782
  • Coding1973
  • Infrastructure825
  • Projects603
  • Marketing598
  • Research520
  • Analytics468
  • Design462
  • MCP419
  • Testing346
  • Security323
  • Data305
  • Integration224
  • Prompts220
  • Communication210
  • Extensions196
  • Learning179
  • Voice175
  • Commerce160
  • DevOps135
  • Web95
  • Finance31
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Tools
    3. jevals
    jevals icon

    jevals

    LLM Evaluations
    Featured

    Agent evals and guardrails using Jev-style decision models instead of LLM judges — one request per trace, a fraction of a cent, fast enough to run inside the agent loop.

    Visit Website

    At a Glance

    Pricing
    Open Source

    Fully free and open-source under the MIT license. Use, modify, and distribute freely.

    Engagement

    Available On

    macOS
    Web
    API
    SDK
    CLI

    Resources

    WebsiteDocsGitHubllms.txt

    Topics

    LLM EvaluationsAgent FrameworksAutonomous Systems

    Alternatives

    terminal-benchAgentBenchAshr
    Developer
    openlayer-aiopenlayer-ai builds open-source tooling for evaluating and m…

    Listed Sep 2026

    About jevals

    jevals is an open-source Python library for evaluating and guarding AI agents using Jev-style decision models rather than traditional LLM judges. Built under the MIT license by the openlayer-ai organization, it packages all evals for a trace into a single HTTP request that the README reports costs a few thousandths of a cent and returns in a few hundred milliseconds — fast enough to run on every trace and inside the agent loop itself.

    What It Is

    jevals is an eval and guardrails framework for AI agents. Instead of routing judgment calls through a frontier LLM that generates text token-by-token, it sends typed questions (yes/no, pick-one, rubric) to a decision model — Jev, Kev, or Laya — that returns calibrated probabilities in a single forward pass. The library ships 37 built-in evals across three modules (jevals.agent, jevals.security, jevals.quality), a YAML gate format, framework adapters, an MCP server, and a CLI. It is designed to replace or complement Ragas-style LLM judges for agent-specific checks that those libraries don't cover.

    The Core Architecture

    Each eval is a Python class with three methods: state() selects what the model should look at, questions() defines what to ask, and reduce() turns the returned probabilities into a score. When multiple evals are passed to evaluate(), their states are merged and their questions are packed into one request. Plain code handles what code is good at — sentence splitting, tool-call matching, regex for secrets, Presidio for entity detection — and typed questions handle the judgment calls.

    The README benchmarks this approach against Ragas on four equivalent metrics over 20 rows:

    • Ragas + gpt-4.1-mini: 6 LLM calls + embeddings per sample, ~$2.60/1k samples, 22–35s for 20 rows
    • jevals emulating Jev with gpt-4.1-mini: 1 call per sample, ~$0.46/1k samples, 4s for 20 rows
    • jevals + Jev via Vercel AI Gateway: 1 call per sample, ~$0.03/1k samples, 0.8s for 20 rows
    • jevals + Kev or Laya locally: 1 call per sample, $0/1k samples, ~1–6s (author-published estimates)

    Backend Flexibility

    jevals resolves its backend from environment variables in priority order: TYPESAFE_API_KEY for Jev direct, AI_GATEWAY_API_KEY for Jev via Vercel AI Gateway (no waitlist required per the README), KEV_BASE_URL for a self-hosted Kev instance on a 32 GB Mac, JEVALS_BACKEND=laya for fully local Apple Silicon inference, or OPENROUTER_API_KEY for any chat LLM in emulation mode. Backends can also be passed explicitly per call, and anything that speaks the System One wire format can be subclassed as a custom backend.

    Guardrails and Gates

    The same eval definitions that score traces offline can run as gates inside the request path — before a tool call executes or before a tool result reaches the model. A gate pairs an eval with a policy that maps its answers to allow, escalate, or block decisions. Gates can be defined in Python or YAML; the README shows a tool_call_risk.yaml example that scores lookup_order, issue_refund, send_email, and run_sql calls and routes irreversible or financially significant actions to a human reviewer. Adapters are provided for the OpenAI Agents SDK, LangGraph, and the Claude Agent SDK.

    What's Included

    • jevals.agent: ToolChoice, ArgumentValidity, UsedToolResult, Grounded, StayedInScope, StepProgress, LoopDetection, GoalCompletion, PlanAdherence, Quality, ToolCallRisk, TrajectoryMatch, ToolCallF1
    • jevals.security: PromptInjection, IndirectInjection, Jailbreak, GoalHijacking, SystemPromptLeakage, ExcessiveAgency, PII, PHI, SecretsExposure, Toxicity, Bias, NonAdvice, TopicAdherence
    • jevals.quality: Faithfulness, AnswerRelevancy, ContextPrecision, ContextRecall, Hallucination, Correctness, Completeness, Coherence, InstructionFollowing, Refusal, CustomRubric
    • MCP server: exposes list_evals, describe_eval, evaluate, evaluate_file, gate, validate_eval, author_eval, schema, and docs so coding agents (Cursor, Claude Code, Copilot) can write and run evals without guessing at the API
    • CLI: jevals run, jevals calibrate, jevals bench, jevals mcp, jevals hook, jevals validate, jevals schema, jevals docs

    Current Status

    The README describes jevals as alpha, created in September 2026 and last updated September 27, 2026 (version 0.1.4 referenced in the benchmark section). All 37 evals, gates, the CLI, and the benchmark have been run end-to-end against Jev through Vercel's AI Gateway and against gpt-4.1-mini through OpenRouter. The TypeSafe direct backend and framework adapters are written to documented wire formats and tested against mocks but not yet run live. A TypeScript package is noted as next on the roadmap. The project is MIT-licensed and accepts contributions, with calibration data noted as the most valuable contribution.

    jevals - 1

    Community Discussions

    Be the first to start a conversation about jevals

    Share your experience with jevals, ask questions, or help others learn from your insights.

    Pricing

    OPEN SOURCE

    Open Source

    Fully free and open-source under the MIT license. Use, modify, and distribute freely.

    • 37 built-in evals (agent, security, quality)
    • YAML and Python gate definitions
    • MCP server for coding agent integration
    • CLI for batch runs, calibration, and benchmarking
    • Adapters for OpenAI Agents SDK, LangGraph, Claude Agent SDK

    Capabilities

    Key Features

    • 37 built-in evals across agent, security, and quality modules
    • Single HTTP request for all evals on a trace
    • Calibrated probability outputs instead of text generation
    • YAML and Python gate definitions with allow/escalate/block policies
    • PII and PHI detection with Presidio integration and redaction support
    • Indirect injection detection in tool results and retrieved documents
    • Secrets exposure detection via regex + model verification
    • MCP server for Cursor, Claude Code, and Copilot integration
    • CLI for batch runs, calibration, benchmarking, and hook installation
    • Adapters for OpenAI Agents SDK, LangGraph, and Claude Agent SDK
    • Sync and async APIs (evaluate/aevaluate, check/acheck)
    • Multiple backend support: Jev, Kev, Laya, any chat LLM
    • YAML eval authoring with JSON Schema validation
    • Dataset batch evaluation with jevals run CLI
    • Calibration tooling with threshold/error-rate tradeoff tables
    • Mock backend for testing
    • Ragas-equivalent metrics in one request per sample
    • Custom eval authoring in Python or YAML

    Integrations

    Jev (TypeSafe)
    Vercel AI Gateway
    Kev (self-hosted, Qwen3)
    Laya (Apple Silicon, ModernBERT)
    OpenRouter
    OpenAI Agents SDK
    LangGraph
    Claude Agent SDK
    Presidio (PII/PHI entity detection)
    Cursor (via MCP)
    Claude Code (via MCP)
    GitHub Copilot (via MCP)
    OpenAI chat format (user/assistant/tool messages)
    Anthropic content blocks
    LangChain message objects
    API Available
    View Docs

    Ratings & Reviews

    No ratings yet

    Be the first to rate jevals and help others make informed decisions.

    Developer

    openlayer-ai

    openlayer-ai builds open-source tooling for evaluating and monitoring AI agents and LLM pipelines. The organization develops jevals, a Python library that replaces expensive LLM judges with fast, calibrated decision models for agent evals and guardrails. Their work focuses on making production-grade evaluation practical at scale — cheap enough to run on every trace and fast enough to run inside the agent loop.

    Read more about openlayer-ai
    WebsiteGitHubX / Twitter
    1 tool in directory

    Similar Tools

    terminal-bench icon

    terminal-bench

    Terminal-Bench is an open-source benchmark suite for evaluating AI agents' ability to complete complex tasks in terminal environments, built on the Harbor framework.

    AgentBench icon

    AgentBench

    AgentBench is an open-source benchmark framework for evaluating LLMs as autonomous agents across 8 diverse environments including OS, database, web, and knowledge graph tasks.

    Ashr icon

    Ashr

    Ashr is an AI agent evaluation platform that mimics production environments and user behavior to catch agent failures before they reach real users.

    Browse all tools

    Related Topics

    LLM Evaluations

    Platforms and frameworks for evaluating, testing, and benchmarking LLM systems and AI applications. These tools provide evaluators and evaluation models to score AI outputs, measure hallucinations, assess RAG quality, detect failures, and optimize model performance. Features include automated testing with LLM-as-a-judge metrics, component-level evaluation with tracing, regression testing in CI/CD pipelines, custom evaluator creation, dataset curation, and real-time monitoring of production systems. Teams use these solutions to validate prompt effectiveness, compare models side-by-side, ensure answer correctness and relevance, identify bias and toxicity, prevent PII leakage, and continuously improve AI product quality through experiments, benchmarks, and performance analytics.

    129 tools

    Agent Frameworks

    Tools and platforms for building and deploying custom AI agents.

    754 tools

    Autonomous Systems

    AI agents that can perform complex tasks with minimal human guidance.

    445 tools
    Browse all topics
    Back to all toolsSuggest an edit
    ratings
    discussions