# AnyJev

> Open-source Python library that turns an open LLM into a Jev-style decision model returning typed decisions with calibrated probabilities.

AnyJev is an Apache-2.0 Python library published by Nokia applied research authors, with the latest release being 0.2.0 on PyPI and GitHub. According to its README, it lets you ask an open LLM a typed question, such as a choice, a yes/no or a score, and get a decision with a probability you can threshold. The PyPI classifiers list it as Pre-Alpha, and the project states it is not affiliated with TypeSafe AI or Jev.

## What It Is

AnyJev is a decision-readout layer for open LLMs. The README says each decision is read from one prefill of the model's next-token distribution, so nothing is generated and nothing is parsed. The project's stated aim is to fix two problems with raw logits: answers that change when options are reordered, and confidence values that cannot be trusted.

## How the levels work

The README describes a ladder of readout levels, and every Decision carries its level so downstream code can require a minimum one.

- Raw is a restricted softmax over label tokens.
- L0 needs no labels. It averages out position bias over option rotations and divides out the label prior.
- L1 adds temperature scaling using 100 to 500 labels per question.
- L2 fits a closed-form head on a hidden state partway down the model, using 100 to 300 labels per question.

The project says L2 heads are per question and per model, and do not transfer.

## Reported results

The maintainers report, for Qwen3-8B on BANKING77 with 20 classes, that the order-flip rate falls from 0.230 to 0.073 with zero labels. They also report that the share of traffic auto-decidable at 5% error rises from 7.7% to 52.0% with L1. For L2 on the typed-decisions set, they report accuracies of 0.730 to 0.799 across five Qwen3 models. The README says every number is regenerated from committed JSON, and the figures for Jev and Laya were published by their authors and not rerun. The README also notes that accuracy on typed-decisions is agreement with a teacher LLM.

## Serving with vLLM

The README says vLLM can serve every level, including L2, through an embed server's pooler. Its documented path is to install the hf extra, truncate a model to fewer blocks with anyjev.truncate, and serve it with vllm. The documentation presents an L2 deployment as a pooling server plus a few kilobytes of head. A pipeline command converts, serves and measures accuracy, ECE and latency in one step. An opt-in adaptive rotation budget is described as using 7.2 rotations instead of 18 at a certified 1% disagreement rate.

## Tradeoffs to know

The README lists several limitations:

- Calibration cannot fix a model that cannot answer.
- L0 can cost accuracy when one label dominates.
- Only Qwen3 heads ship, and the headline tables are Qwen models.
- The letter readout supports at most 26 options.
- Decisions are scored in isolation, not inside an agent loop.

The roadmap lists agent-loop evaluation, more model families and SGLang support as not yet done.

## Current Status

PyPI shows releases 0.0.1 and 0.0.2 on Sep 21, 0.1.0 on Sep 26 and 0.2.0 on Sep 28, 2026. The GitHub repository is public and still under active updates.

## Features
- Typed decisions: choice, yes/no (noul) and score from one prefill
- L0 zero-label position-bias correction over option rotations
- L1 temperature-scaled calibration with 100-500 labels
- L2 closed-form heads on mid-depth hidden states with 100-300 labels
- Adaptive rotation budget for fewer prefills per decision
- vLLM serving via embed server pooler or truncated checkpoint
- Model truncation tool and one-command pipeline benchmark
- Label-free head adaptation and online label collection via observe
- Shipped heads for five Qwen3 models
- Transformers (Hugging Face) backend

## Integrations
vLLM, Hugging Face Transformers, Qwen3, Qwen2.5

## Platforms
WEB, API, DEVELOPER_SDK, CLI

## Pricing
Open Source

## Version
0.2.0

## Links
- Website: https://pypi.org/project/anyjev/
- Documentation: https://github.com/nokia-applied-research/AnyJev/blob/main/docs/levels.md
- Repository: https://github.com/nokia-applied-research/AnyJev
- EveryDev.ai: https://www.everydev.ai/tools/anyjev
