EveryDev.ai
Subscribe
Home
Tools

4,072+ AI tools

  • New
  • Trending
  • Featured
  • Compare
  • Arena
Categories
  • Agents2782
  • Coding1973
  • Infrastructure825
  • Projects603
  • Marketing598
  • Research520
  • Analytics468
  • Design462
  • MCP419
  • Testing346
  • Security323
  • Data305
  • Integration224
  • Prompts220
  • Communication210
  • Extensions196
  • Learning179
  • Voice175
  • Commerce160
  • DevOps135
  • Web95
  • Finance31
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Tools
    3. Jev Cookbook
    Jev Cookbook icon

    Jev Cookbook

    LLM Evaluations

    Runnable question sets (recipes) for TypeSafe's Jev structured-decision API, with an eval harness, CLINC150 benchmark results, and a linter for common request-shape mistakes.

    Visit Website

    At a Glance

    Pricing
    Open Source

    Fully free and open-source under the MIT License. Clone, use, and contribute with no cost.

    Engagement

    Available On

    API
    CLI

    Resources

    WebsiteDocsGitHubllms.txt

    Topics

    LLM EvaluationsPrompt EngineeringAgent Harness

    Alternatives

    VerifiersSupabase EvalsExploitBench
    Developer
    chr-kellychr-kelly is an independent developer building open-source t…

    Listed Sep 2026

    About Jev Cookbook

    Jev Cookbook is a community-maintained, MIT-licensed GitHub repository by chr-kelly that provides copy-paste-ready question sets ("recipes") for TypeSafe's Jev structured-decision API. It ships alongside an evaluation harness with reproducible CLINC150 benchmark results and a linter that catches silent API mis-reads before they reach production.

    What It Is

    Jev is a typed, low-latency API that answers structured questions about a piece of text — classifying intent, scoring urgency, or flagging attributes — without generating free-form prose. The Jev Cookbook fills the gap between the API's three primitives (choice, score, noul) and real-world usage: it shows practitioners what to ask and why each question is shaped that way, including what was tried first, what broke, and where each recipe still gets things wrong.

    Recipes Included

    Each recipe lives in its own folder with a questions.json ready to paste into a request and a README explaining the design decisions:

    • customer-support-routing — category, refund intent, repeat contact, urgency
    • roleplay-state — tone drift, stalled tension, lore breaks as live sensors
    • agent-tool-guardrail — whether an agent's tool call matches its own stated plan
    • context-pruning — which old tool calls in a long transcript still earn their place
    • llm-router — which model tier should handle a request, judged on the work not the model
    • content-qa — whether a generated page is thin, and which of five ways

    Four Design Rules

    The cookbook enforces four rules that exist because skipping any one of them produces wrong answers on real data:

    • Every choice has a fallback option — Jev picks from a closed set and will choose confidently even on off-topic input; always include OTHER/UNRESOLVED and review what lands there.
    • Ask atomic questions, combine in code — separate urgency, refund intent, and repeat-contact into individual questions, then weight them with your own formula so changing priorities means editing a coefficient, not rewriting a prompt.
    • Multi-label means several nouls, not one choice — questions in one call are evaluated independently and in parallel, so a dozen nouls cost roughly what one does.
    • Split judgement from extraction — Jev generates nothing; use it to gate, then pass survivors to an LLM only where prose is needed.

    The Two-Pass Pattern

    The cookbook documents a two-pass architecture for large corpora: a cheap gate of 5–7 questions runs on everything, roughly 10% survive to a full 40–60 question checklist, and an LLM is invoked only on that filtered set. Because questions inside a single call are nearly free but state is what you pay for, this is where the order-of-magnitude cost savings come from.

    Eval Harness and Benchmarks

    The eval/ directory contains a harness that runs any recipe against a labelled dataset and reports agreement, a confidence-vs-error curve, escalation rate at a given threshold, and fallback-option behavior. The README publishes one measured run: CLINC150 test set plus 1,000 out-of-scope items, one 151-way choice, untuned option names, run on 2026-09-21 against jev-1.13.0. The run reported 89.4% agreement, 90.6% in-scope accuracy, 83.8% OOS recall, and an ECE of 0.025 at approximately $0.63 total cost. Full per-item results and the exact command are included in the repo. The project explicitly states: "No numbers are published here that you cannot reproduce."

    Provider Compatibility

    The same request body works against TypeSafe's own endpoint (api.typesafe.ai), OpenRouter (openrouter.ai/api/v1/decisions), and NanoGPT (nano-gpt.com/api/v1/decisions). The --base-url flag on the eval harness points it at any compatible server. At the time of the README, TypeSafe's direct API was gated behind a waitlist; the other two providers were open.

    Jev Cookbook - 1

    Community Discussions

    Be the first to start a conversation about Jev Cookbook

    Share your experience with Jev Cookbook, ask questions, or help others learn from your insights.

    Pricing

    OPEN SOURCE

    Open Source

    Fully free and open-source under the MIT License. Clone, use, and contribute with no cost.

    • All recipes (questions.json files)
    • Eval harness
    • Linter (eval/lint.py)
    • CLINC150 benchmark results
    • CONTRIBUTING guide

    Capabilities

    Key Features

    • Runnable question sets (recipes) for TypeSafe's Jev API
    • Recipes for customer support routing, roleplay state, agent tool guardrails, context pruning, LLM routing, and content QA
    • Eval harness with CLINC150 benchmark results
    • Linter (eval/lint.py) for catching silent API request-shape mistakes
    • Two-pass pattern documentation for cost-efficient large-corpus processing
    • Multi-provider support: TypeSafe, OpenRouter, NanoGPT
    • MIT licensed and fully reproducible benchmarks
    • CONTRIBUTING guide for adding new recipes

    Integrations

    TypeSafe Jev API
    OpenRouter
    NanoGPT
    CLINC150 dataset
    API Available
    View Docs

    Ratings & Reviews

    No ratings yet

    Be the first to rate Jev Cookbook and help others make informed decisions.

    Developer

    chr-kelly

    chr-kelly is an independent developer building open-source tooling for TypeSafe's Jev structured-decision API. The jev-cookbook project provides runnable question sets, an evaluation harness, and a linter to help practitioners move from API primitives to production-ready decision pipelines. The project is MIT-licensed and community-contribution-friendly.

    Read more about chr-kelly
    WebsiteGitHub
    1 tool in directory

    Similar Tools

    Verifiers icon

    Verifiers

    An open-source Python library by Prime Intellect for creating environments to train and evaluate LLMs using reinforcement learning.

    Supabase Evals icon

    Supabase Evals

    An open-source evaluation framework that benchmarks how well AI agents perform across the Supabase developer journey, from building and deploying to investigating and resolving production issues.

    ExploitBench icon

    ExploitBench

    ExploitBench measures how far AI agents can climb the exploitation ladder, from reaching vulnerable code to achieving arbitrary code execution, using a five-tier grading system against real CVEs.

    Browse all tools

    Related Topics

    LLM Evaluations

    Platforms and frameworks for evaluating, testing, and benchmarking LLM systems and AI applications. These tools provide evaluators and evaluation models to score AI outputs, measure hallucinations, assess RAG quality, detect failures, and optimize model performance. Features include automated testing with LLM-as-a-judge metrics, component-level evaluation with tracing, regression testing in CI/CD pipelines, custom evaluator creation, dataset curation, and real-time monitoring of production systems. Teams use these solutions to validate prompt effectiveness, compare models side-by-side, ensure answer correctness and relevance, identify bias and toxicity, prevent PII leakage, and continuously improve AI product quality through experiments, benchmarks, and performance analytics.

    129 tools

    Prompt Engineering

    Tools for creating and refining effective AI prompts.

    84 tools

    Agent Harness

    Infrastructure, orchestrators, and task runners that wrap around LLM coding agents — covering session management, context delivery, worktree isolation, architecture enforcement, and issue-to-PR pipelines.

    182 tools
    Browse all topics
    Back to all toolsSuggest an edit
    ratings
    discussions