jev-effort
A proxy tool that uses Jev to dynamically select Claude Code's reasoning effort per step without breaking the prompt cache, reducing costs at max effort by up to 55%.
At a Glance
Free and open-source under the MIT License. Run via npx with no cost beyond API usage.
Engagement
Available On
Alternatives
Listed Sep 2026
About jev-effort
jev-effort is an open-source research project by GitHub user ifoster01 that runs Claude Code behind a local proxy to dynamically adjust reasoning effort on every agent step using Jev, TypeSafe's small decision model. The project was created to answer a concrete question: does letting Jev choose Claude Code's reasoning effort step by step actually save money, and by how much?
What It Is
jev-effort is a CLI proxy tool written in JavaScript that intercepts Claude Code's API requests via ANTHROPIC_BASE_URL. Before each model request, it sends Jev a trimmed view of the conversation and asks which effort level the next step needs and for how many steps to hold it. It then applies the answer by appending an effort-only system message — using Anthropic's per-message effort beta — without resetting the prompt cache. The tool is unofficial and not affiliated with Anthropic or TypeSafe.
How the Proxy Works
The proxy sits between Claude Code and the Anthropic API. On each step it:
- Sends Jev a condensed view of the conversation (prompts, Claude's visible replies, the last six tool calls)
- Receives Jev's effort recommendation and duration
- Appends an effort-only system message as the last message in the request (required for the beta per-turn control to take effect)
- Re-inserts earlier effort messages at their original positions so the cached prefix remains intact
Jev may lower effort but never exceed the session's own ceiling setting. The tool also supports a shadow mode (--jev-shadow) that logs Jev's choices without changing anything, enabling safe observation on real work.
Cache Preservation: The Key Technical Finding
A central concern was whether changing effort mid-session would invalidate Claude Code's prompt cache. The project confirmed it does not have to: Anthropic's per-message effort beta allows an effort-only system message to change the level while leaving the cached prefix intact. In a 470-step, 92-minute real session, the tool recorded 0 unexpected cache misses in 411 checked steps and a 99.1% cache hit rate, even with effort changing on every step.
Benchmark Results
The project ran 48 headless Claude Code sessions across six small coding tasks, comparing fixed effort against Jev-chosen effort under the same ceiling:
- At
higheffort: Cost fell about 1.1%, thinking tokens dropped 46%, all hidden tests passed. Opus 5.5 athighalready thinks little on routine steps, leaving little to cut. - At
maxeffort: Cost fell 55% (21% excluding one outlier task), thinking tokens dropped 98.5%, wall time fell 58%, all hidden tests passed. Jev moved routine steps tolowormedium, wheremaxwould otherwise think heavily even on simple work. - Real session at
max(shadow mode): Estimated saving of 4–9%, because cached context reads accounted for 82% of spend and effort does not affect those costs.
Tradeoffs and Limitations
The README is explicit about where the approach falls short:
- At
higheffort, the saving is marginal (~1%) because thinking is already a small share of the bill - In long sessions, context re-reading dominates cost regardless of effort level
- Jev adds ~601 ms median latency per decision, though at
maxthe reduced thinking more than compensates - The benchmark tasks are small and well-specified; quality impact on complex real work is untested
- Using any custom
ANTHROPIC_BASE_URLdisables Claude Code's MCP tool search by default (workaround:ENABLE_TOOL_SEARCH=true) - One real session at one effort level is the only real-world measurement
Current Status
The repository was created and last updated in late September 2026. The code is described as working and results as reproducible. It is an MIT-licensed research project available via npx jev-effort, with commands for shadow mode, spend analysis, and controlled benchmarking. Related projects in the same space include jev-opus, jev-model-router, and effort-router, none of which (as of the README's writing) reported total cost against a fixed-effort baseline.
Community Discussions
Be the first to start a conversation about jev-effort
Share your experience with jev-effort, ask questions, or help others learn from your insights.
Pricing
Open Source
Free and open-source under the MIT License. Run via npx with no cost beyond API usage.
- Per-step effort selection via Jev
- Shadow mode logging
- Spend analysis (jev-effort stats)
- Controlled benchmark runner
- Cache-safe effort injection
Capabilities
Key Features
- Per-step reasoning effort selection via Jev decision model
- Local proxy via ANTHROPIC_BASE_URL without breaking prompt cache
- Shadow mode to log Jev's choices without changing behavior
- Spend analysis with jev-effort stats command
- Controlled benchmark runner (jev-effort bench)
- Cache-safe effort injection using Anthropic per-message effort beta
- Effort ceiling enforcement (Jev never exceeds session setting)
- Support for low, medium, high, xhigh, and max effort levels
- MCP tool search compatibility via ENABLE_TOOL_SEARCH=true
- MIT licensed and reproducible results
