jevals
Agent evals and guardrails using Jev-style decision models instead of LLM judges — one request per trace, a fraction of a cent, fast enough to run inside the agent loop.
At a Glance
Fully free and open-source under the MIT license. Use, modify, and distribute freely.
Engagement
Available On
Alternatives
Listed Sep 2026
About jevals
jevals is an open-source Python library for evaluating and guarding AI agents using Jev-style decision models rather than traditional LLM judges. Built under the MIT license by the openlayer-ai organization, it packages all evals for a trace into a single HTTP request that the README reports costs a few thousandths of a cent and returns in a few hundred milliseconds — fast enough to run on every trace and inside the agent loop itself.
What It Is
jevals is an eval and guardrails framework for AI agents. Instead of routing judgment calls through a frontier LLM that generates text token-by-token, it sends typed questions (yes/no, pick-one, rubric) to a decision model — Jev, Kev, or Laya — that returns calibrated probabilities in a single forward pass. The library ships 37 built-in evals across three modules (jevals.agent, jevals.security, jevals.quality), a YAML gate format, framework adapters, an MCP server, and a CLI. It is designed to replace or complement Ragas-style LLM judges for agent-specific checks that those libraries don't cover.
The Core Architecture
Each eval is a Python class with three methods: state() selects what the model should look at, questions() defines what to ask, and reduce() turns the returned probabilities into a score. When multiple evals are passed to evaluate(), their states are merged and their questions are packed into one request. Plain code handles what code is good at — sentence splitting, tool-call matching, regex for secrets, Presidio for entity detection — and typed questions handle the judgment calls.
The README benchmarks this approach against Ragas on four equivalent metrics over 20 rows:
- Ragas + gpt-4.1-mini: 6 LLM calls + embeddings per sample, ~$2.60/1k samples, 22–35s for 20 rows
- jevals emulating Jev with gpt-4.1-mini: 1 call per sample, ~$0.46/1k samples, 4s for 20 rows
- jevals + Jev via Vercel AI Gateway: 1 call per sample, ~$0.03/1k samples, 0.8s for 20 rows
- jevals + Kev or Laya locally: 1 call per sample, $0/1k samples, ~1–6s (author-published estimates)
Backend Flexibility
jevals resolves its backend from environment variables in priority order: TYPESAFE_API_KEY for Jev direct, AI_GATEWAY_API_KEY for Jev via Vercel AI Gateway (no waitlist required per the README), KEV_BASE_URL for a self-hosted Kev instance on a 32 GB Mac, JEVALS_BACKEND=laya for fully local Apple Silicon inference, or OPENROUTER_API_KEY for any chat LLM in emulation mode. Backends can also be passed explicitly per call, and anything that speaks the System One wire format can be subclassed as a custom backend.
Guardrails and Gates
The same eval definitions that score traces offline can run as gates inside the request path — before a tool call executes or before a tool result reaches the model. A gate pairs an eval with a policy that maps its answers to allow, escalate, or block decisions. Gates can be defined in Python or YAML; the README shows a tool_call_risk.yaml example that scores lookup_order, issue_refund, send_email, and run_sql calls and routes irreversible or financially significant actions to a human reviewer. Adapters are provided for the OpenAI Agents SDK, LangGraph, and the Claude Agent SDK.
What's Included
jevals.agent: ToolChoice, ArgumentValidity, UsedToolResult, Grounded, StayedInScope, StepProgress, LoopDetection, GoalCompletion, PlanAdherence, Quality, ToolCallRisk, TrajectoryMatch, ToolCallF1jevals.security: PromptInjection, IndirectInjection, Jailbreak, GoalHijacking, SystemPromptLeakage, ExcessiveAgency, PII, PHI, SecretsExposure, Toxicity, Bias, NonAdvice, TopicAdherencejevals.quality: Faithfulness, AnswerRelevancy, ContextPrecision, ContextRecall, Hallucination, Correctness, Completeness, Coherence, InstructionFollowing, Refusal, CustomRubric- MCP server: exposes
list_evals,describe_eval,evaluate,evaluate_file,gate,validate_eval,author_eval,schema, anddocsso coding agents (Cursor, Claude Code, Copilot) can write and run evals without guessing at the API - CLI:
jevals run,jevals calibrate,jevals bench,jevals mcp,jevals hook,jevals validate,jevals schema,jevals docs
Current Status
The README describes jevals as alpha, created in September 2026 and last updated September 27, 2026 (version 0.1.4 referenced in the benchmark section). All 37 evals, gates, the CLI, and the benchmark have been run end-to-end against Jev through Vercel's AI Gateway and against gpt-4.1-mini through OpenRouter. The TypeSafe direct backend and framework adapters are written to documented wire formats and tested against mocks but not yet run live. A TypeScript package is noted as next on the roadmap. The project is MIT-licensed and accepts contributions, with calibration data noted as the most valuable contribution.
Community Discussions
Be the first to start a conversation about jevals
Share your experience with jevals, ask questions, or help others learn from your insights.
Pricing
Open Source
Fully free and open-source under the MIT license. Use, modify, and distribute freely.
- 37 built-in evals (agent, security, quality)
- YAML and Python gate definitions
- MCP server for coding agent integration
- CLI for batch runs, calibration, and benchmarking
- Adapters for OpenAI Agents SDK, LangGraph, Claude Agent SDK
Capabilities
Key Features
- 37 built-in evals across agent, security, and quality modules
- Single HTTP request for all evals on a trace
- Calibrated probability outputs instead of text generation
- YAML and Python gate definitions with allow/escalate/block policies
- PII and PHI detection with Presidio integration and redaction support
- Indirect injection detection in tool results and retrieved documents
- Secrets exposure detection via regex + model verification
- MCP server for Cursor, Claude Code, and Copilot integration
- CLI for batch runs, calibration, benchmarking, and hook installation
- Adapters for OpenAI Agents SDK, LangGraph, and Claude Agent SDK
- Sync and async APIs (evaluate/aevaluate, check/acheck)
- Multiple backend support: Jev, Kev, Laya, any chat LLM
- YAML eval authoring with JSON Schema validation
- Dataset batch evaluation with jevals run CLI
- Calibration tooling with threshold/error-rate tradeoff tables
- Mock backend for testing
- Ragas-equivalent metrics in one request per sample
- Custom eval authoring in Python or YAML
