# Jev Cookbook

> Runnable question sets (recipes) for TypeSafe's Jev structured-decision API, with an eval harness, CLINC150 benchmark results, and a linter for common request-shape mistakes.

Jev Cookbook is a community-maintained, MIT-licensed GitHub repository by chr-kelly that provides copy-paste-ready question sets ("recipes") for TypeSafe's Jev structured-decision API. It ships alongside an evaluation harness with reproducible CLINC150 benchmark results and a linter that catches silent API mis-reads before they reach production.

## What It Is

Jev is a typed, low-latency API that answers structured questions about a piece of text — classifying intent, scoring urgency, or flagging attributes — without generating free-form prose. The Jev Cookbook fills the gap between the API's three primitives (`choice`, `score`, `noul`) and real-world usage: it shows practitioners *what to ask* and *why each question is shaped that way*, including what was tried first, what broke, and where each recipe still gets things wrong.

## Recipes Included

Each recipe lives in its own folder with a `questions.json` ready to paste into a request and a README explaining the design decisions:

- **customer-support-routing** — category, refund intent, repeat contact, urgency
- **roleplay-state** — tone drift, stalled tension, lore breaks as live sensors
- **agent-tool-guardrail** — whether an agent's tool call matches its own stated plan
- **context-pruning** — which old tool calls in a long transcript still earn their place
- **llm-router** — which model tier should handle a request, judged on the work not the model
- **content-qa** — whether a generated page is thin, and which of five ways

## Four Design Rules

The cookbook enforces four rules that exist because skipping any one of them produces wrong answers on real data:

- **Every `choice` has a fallback option** — Jev picks from a closed set and will choose confidently even on off-topic input; always include `OTHER`/`UNRESOLVED` and review what lands there.
- **Ask atomic questions, combine in code** — separate urgency, refund intent, and repeat-contact into individual questions, then weight them with your own formula so changing priorities means editing a coefficient, not rewriting a prompt.
- **Multi-label means several `noul`s, not one `choice`** — questions in one call are evaluated independently and in parallel, so a dozen `noul`s cost roughly what one does.
- **Split judgement from extraction** — Jev generates nothing; use it to gate, then pass survivors to an LLM only where prose is needed.

## The Two-Pass Pattern

The cookbook documents a two-pass architecture for large corpora: a cheap gate of 5–7 questions runs on everything, roughly 10% survive to a full 40–60 question checklist, and an LLM is invoked only on that filtered set. Because questions inside a single call are nearly free but *state* is what you pay for, this is where the order-of-magnitude cost savings come from.

## Eval Harness and Benchmarks

The `eval/` directory contains a harness that runs any recipe against a labelled dataset and reports agreement, a confidence-vs-error curve, escalation rate at a given threshold, and fallback-option behavior. The README publishes one measured run: CLINC150 test set plus 1,000 out-of-scope items, one 151-way `choice`, untuned option names, run on 2026-09-21 against `jev-1.13.0`. The run reported 89.4% agreement, 90.6% in-scope accuracy, 83.8% OOS recall, and an ECE of 0.025 at approximately $0.63 total cost. Full per-item results and the exact command are included in the repo. The project explicitly states: "No numbers are published here that you cannot reproduce."

## Provider Compatibility

The same request body works against TypeSafe's own endpoint (`api.typesafe.ai`), OpenRouter (`openrouter.ai/api/v1/decisions`), and NanoGPT (`nano-gpt.com/api/v1/decisions`). The `--base-url` flag on the eval harness points it at any compatible server. At the time of the README, TypeSafe's direct API was gated behind a waitlist; the other two providers were open.

## Features
- Runnable question sets (recipes) for TypeSafe's Jev API
- Recipes for customer support routing, roleplay state, agent tool guardrails, context pruning, LLM routing, and content QA
- Eval harness with CLINC150 benchmark results
- Linter (eval/lint.py) for catching silent API request-shape mistakes
- Two-pass pattern documentation for cost-efficient large-corpus processing
- Multi-provider support: TypeSafe, OpenRouter, NanoGPT
- MIT licensed and fully reproducible benchmarks
- CONTRIBUTING guide for adding new recipes

## Integrations
TypeSafe Jev API, OpenRouter, NanoGPT, CLINC150 dataset

## Platforms
API, CLI

## Pricing
Open Source

## Version
jev-1.13.0 (eval reference)

## Links
- Website: https://github.com/chr-kelly/jev-cookbook
- Documentation: https://github.com/chr-kelly/jev-cookbook/blob/main/eval/README.md
- Repository: https://github.com/chr-kelly/jev-cookbook
- EveryDev.ai: https://www.everydev.ai/tools/jev-cookbook
