# jevals

> Agent evals and guardrails using Jev-style decision models instead of LLM judges — one request per trace, a fraction of a cent, fast enough to run inside the agent loop.

jevals is an open-source Python library for evaluating and guarding AI agents using Jev-style decision models rather than traditional LLM judges. Built under the MIT license by the openlayer-ai organization, it packages all evals for a trace into a single HTTP request that the README reports costs a few thousandths of a cent and returns in a few hundred milliseconds — fast enough to run on every trace and inside the agent loop itself.

## What It Is

jevals is an eval and guardrails framework for AI agents. Instead of routing judgment calls through a frontier LLM that generates text token-by-token, it sends typed questions (yes/no, pick-one, rubric) to a decision model — Jev, Kev, or Laya — that returns calibrated probabilities in a single forward pass. The library ships 37 built-in evals across three modules (`jevals.agent`, `jevals.security`, `jevals.quality`), a YAML gate format, framework adapters, an MCP server, and a CLI. It is designed to replace or complement Ragas-style LLM judges for agent-specific checks that those libraries don't cover.

## The Core Architecture

Each eval is a Python class with three methods: `state()` selects what the model should look at, `questions()` defines what to ask, and `reduce()` turns the returned probabilities into a score. When multiple evals are passed to `evaluate()`, their states are merged and their questions are packed into one request. Plain code handles what code is good at — sentence splitting, tool-call matching, regex for secrets, Presidio for entity detection — and typed questions handle the judgment calls.

The README benchmarks this approach against Ragas on four equivalent metrics over 20 rows:

- **Ragas + gpt-4.1-mini**: 6 LLM calls + embeddings per sample, ~$2.60/1k samples, 22–35s for 20 rows
- **jevals emulating Jev with gpt-4.1-mini**: 1 call per sample, ~$0.46/1k samples, 4s for 20 rows
- **jevals + Jev via Vercel AI Gateway**: 1 call per sample, ~$0.03/1k samples, 0.8s for 20 rows
- **jevals + Kev or Laya locally**: 1 call per sample, $0/1k samples, ~1–6s (author-published estimates)

## Backend Flexibility

jevals resolves its backend from environment variables in priority order: `TYPESAFE_API_KEY` for Jev direct, `AI_GATEWAY_API_KEY` for Jev via Vercel AI Gateway (no waitlist required per the README), `KEV_BASE_URL` for a self-hosted Kev instance on a 32 GB Mac, `JEVALS_BACKEND=laya` for fully local Apple Silicon inference, or `OPENROUTER_API_KEY` for any chat LLM in emulation mode. Backends can also be passed explicitly per call, and anything that speaks the System One wire format can be subclassed as a custom backend.

## Guardrails and Gates

The same eval definitions that score traces offline can run as gates inside the request path — before a tool call executes or before a tool result reaches the model. A gate pairs an eval with a policy that maps its answers to allow, escalate, or block decisions. Gates can be defined in Python or YAML; the README shows a `tool_call_risk.yaml` example that scores `lookup_order`, `issue_refund`, `send_email`, and `run_sql` calls and routes irreversible or financially significant actions to a human reviewer. Adapters are provided for the OpenAI Agents SDK, LangGraph, and the Claude Agent SDK.

## What's Included

- **`jevals.agent`**: ToolChoice, ArgumentValidity, UsedToolResult, Grounded, StayedInScope, StepProgress, LoopDetection, GoalCompletion, PlanAdherence, Quality, ToolCallRisk, TrajectoryMatch, ToolCallF1
- **`jevals.security`**: PromptInjection, IndirectInjection, Jailbreak, GoalHijacking, SystemPromptLeakage, ExcessiveAgency, PII, PHI, SecretsExposure, Toxicity, Bias, NonAdvice, TopicAdherence
- **`jevals.quality`**: Faithfulness, AnswerRelevancy, ContextPrecision, ContextRecall, Hallucination, Correctness, Completeness, Coherence, InstructionFollowing, Refusal, CustomRubric
- **MCP server**: exposes `list_evals`, `describe_eval`, `evaluate`, `evaluate_file`, `gate`, `validate_eval`, `author_eval`, `schema`, and `docs` so coding agents (Cursor, Claude Code, Copilot) can write and run evals without guessing at the API
- **CLI**: `jevals run`, `jevals calibrate`, `jevals bench`, `jevals mcp`, `jevals hook`, `jevals validate`, `jevals schema`, `jevals docs`

## Current Status

The README describes jevals as alpha, created in September 2026 and last updated September 27, 2026 (version 0.1.4 referenced in the benchmark section). All 37 evals, gates, the CLI, and the benchmark have been run end-to-end against Jev through Vercel's AI Gateway and against gpt-4.1-mini through OpenRouter. The TypeSafe direct backend and framework adapters are written to documented wire formats and tested against mocks but not yet run live. A TypeScript package is noted as next on the roadmap. The project is MIT-licensed and accepts contributions, with calibration data noted as the most valuable contribution.

## Features
- 37 built-in evals across agent, security, and quality modules
- Single HTTP request for all evals on a trace
- Calibrated probability outputs instead of text generation
- YAML and Python gate definitions with allow/escalate/block policies
- PII and PHI detection with Presidio integration and redaction support
- Indirect injection detection in tool results and retrieved documents
- Secrets exposure detection via regex + model verification
- MCP server for Cursor, Claude Code, and Copilot integration
- CLI for batch runs, calibration, benchmarking, and hook installation
- Adapters for OpenAI Agents SDK, LangGraph, and Claude Agent SDK
- Sync and async APIs (evaluate/aevaluate, check/acheck)
- Multiple backend support: Jev, Kev, Laya, any chat LLM
- YAML eval authoring with JSON Schema validation
- Dataset batch evaluation with jevals run CLI
- Calibration tooling with threshold/error-rate tradeoff tables
- Mock backend for testing
- Ragas-equivalent metrics in one request per sample
- Custom eval authoring in Python or YAML

## Integrations
Jev (TypeSafe), Vercel AI Gateway, Kev (self-hosted, Qwen3), Laya (Apple Silicon, ModernBERT), OpenRouter, OpenAI Agents SDK, LangGraph, Claude Agent SDK, Presidio (PII/PHI entity detection), Cursor (via MCP), Claude Code (via MCP), GitHub Copilot (via MCP), OpenAI chat format (user/assistant/tool messages), Anthropic content blocks, LangChain message objects

## Platforms
MACOS, WEB, API, DEVELOPER_SDK, CLI

## Pricing
Open Source

## Version
0.1.4

## Links
- Website: https://github.com/openlayer-ai/jevals
- Documentation: https://github.com/openlayer-ai/jevals#readme
- Repository: https://github.com/openlayer-ai/jevals
- EveryDev.ai: https://www.everydev.ai/tools/jevals
