# jev-effort

> A proxy tool that uses Jev to dynamically select Claude Code's reasoning effort per step without breaking the prompt cache, reducing costs at max effort by up to 55%.

jev-effort is an open-source research project by GitHub user ifoster01 that runs Claude Code behind a local proxy to dynamically adjust reasoning effort on every agent step using Jev, TypeSafe's small decision model. The project was created to answer a concrete question: does letting Jev choose Claude Code's reasoning effort step by step actually save money, and by how much?

## What It Is

jev-effort is a CLI proxy tool written in JavaScript that intercepts Claude Code's API requests via `ANTHROPIC_BASE_URL`. Before each model request, it sends Jev a trimmed view of the conversation and asks which effort level the next step needs and for how many steps to hold it. It then applies the answer by appending an effort-only system message — using Anthropic's per-message effort beta — without resetting the prompt cache. The tool is unofficial and not affiliated with Anthropic or TypeSafe.

## How the Proxy Works

The proxy sits between Claude Code and the Anthropic API. On each step it:
- Sends Jev a condensed view of the conversation (prompts, Claude's visible replies, the last six tool calls)
- Receives Jev's effort recommendation and duration
- Appends an effort-only system message as the last message in the request (required for the beta per-turn control to take effect)
- Re-inserts earlier effort messages at their original positions so the cached prefix remains intact

Jev may lower effort but never exceed the session's own ceiling setting. The tool also supports a shadow mode (`--jev-shadow`) that logs Jev's choices without changing anything, enabling safe observation on real work.

## Cache Preservation: The Key Technical Finding

A central concern was whether changing effort mid-session would invalidate Claude Code's prompt cache. The project confirmed it does not have to: Anthropic's per-message effort beta allows an effort-only system message to change the level while leaving the cached prefix intact. In a 470-step, 92-minute real session, the tool recorded **0 unexpected cache misses in 411 checked steps** and a 99.1% cache hit rate, even with effort changing on every step.

## Benchmark Results

The project ran 48 headless Claude Code sessions across six small coding tasks, comparing fixed effort against Jev-chosen effort under the same ceiling:

- **At `high` effort:** Cost fell about 1.1%, thinking tokens dropped 46%, all hidden tests passed. Opus 5.5 at `high` already thinks little on routine steps, leaving little to cut.
- **At `max` effort:** Cost fell 55% (21% excluding one outlier task), thinking tokens dropped 98.5%, wall time fell 58%, all hidden tests passed. Jev moved routine steps to `low` or `medium`, where `max` would otherwise think heavily even on simple work.
- **Real session at `max` (shadow mode):** Estimated saving of 4–9%, because cached context reads accounted for 82% of spend and effort does not affect those costs.

## Tradeoffs and Limitations

The README is explicit about where the approach falls short:

- At `high` effort, the saving is marginal (~1%) because thinking is already a small share of the bill
- In long sessions, context re-reading dominates cost regardless of effort level
- Jev adds ~601 ms median latency per decision, though at `max` the reduced thinking more than compensates
- The benchmark tasks are small and well-specified; quality impact on complex real work is untested
- Using any custom `ANTHROPIC_BASE_URL` disables Claude Code's MCP tool search by default (workaround: `ENABLE_TOOL_SEARCH=true`)
- One real session at one effort level is the only real-world measurement

## Current Status

The repository was created and last updated in late September 2026. The code is described as working and results as reproducible. It is an MIT-licensed research project available via `npx jev-effort`, with commands for shadow mode, spend analysis, and controlled benchmarking. Related projects in the same space include jev-opus, jev-model-router, and effort-router, none of which (as of the README's writing) reported total cost against a fixed-effort baseline.

## Features
- Per-step reasoning effort selection via Jev decision model
- Local proxy via ANTHROPIC_BASE_URL without breaking prompt cache
- Shadow mode to log Jev's choices without changing behavior
- Spend analysis with jev-effort stats command
- Controlled benchmark runner (jev-effort bench)
- Cache-safe effort injection using Anthropic per-message effort beta
- Effort ceiling enforcement (Jev never exceeds session setting)
- Support for low, medium, high, xhigh, and max effort levels
- MCP tool search compatibility via ENABLE_TOOL_SEARCH=true
- MIT licensed and reproducible results

## Integrations
Claude Code, Anthropic API, Jev (TypeSafe decision model), OpenRouter (Jev model access), Claude Opus 5.5, Claude Fable 5.1, Claude Mythos 5.1, Claude Opus 5

## Platforms
API, CLI

## Pricing
Open Source

## Links
- Website: https://github.com/ifoster01/jev-effort
- Documentation: https://github.com/ifoster01/jev-effort/blob/main/docs/usage.md
- Repository: https://github.com/ifoster01/jev-effort
- EveryDev.ai: https://www.everydev.ai/tools/jev-effort
