EveryDev.ai
Subscribe
Home
Tools

4,176+ AI tools

  • New
  • Trending
  • Featured
  • Rate tools
  • Compare
  • Arena
Categories
  • Agents2782
  • Coding1973
  • Infrastructure825
  • Projects603
  • Marketing598
  • Research520
  • Analytics468
  • Design462
  • MCP419
  • Testing346
  • Security323
  • Data305
  • Integration224
  • Prompts220
  • Communication210
  • Extensions196
  • Learning179
  • Voice175
  • Commerce160
  • DevOps135
  • Web95
  • Finance31
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • Users
  • Rate Tools
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Tools
    3. AnyJev
    AnyJev icon

    AnyJev

    AI Decision Models

    Open-source Python library that turns an open LLM into a Jev-style decision model returning typed decisions with calibrated probabilities.

    Visit Website

    At a Glance

    Pricing
    Open Source

    Self-hosted open-source Python library under Apache License 2.0; free to use, modify, and distribute.

    Engagement

    Available On

    Web
    API
    SDK
    CLI

    Resources

    WebsiteDocsGitHubllms.txt

    Topics

    AI Decision ModelsAI Development LibrariesLLM Evaluations

    Alternatives

    JevlikeVonJulia 1
    Developer
    nokia-applied-research

    Listed Oct 2026

    About AnyJev

    AnyJev is an Apache-2.0 Python library published by Nokia applied research authors, with the latest release being 0.2.0 on PyPI and GitHub. According to its README, it lets you ask an open LLM a typed question, such as a choice, a yes/no or a score, and get a decision with a probability you can threshold. The PyPI classifiers list it as Pre-Alpha, and the project states it is not affiliated with TypeSafe AI or Jev.

    What It Is

    AnyJev is a decision-readout layer for open LLMs. The README says each decision is read from one prefill of the model's next-token distribution, so nothing is generated and nothing is parsed. The project's stated aim is to fix two problems with raw logits: answers that change when options are reordered, and confidence values that cannot be trusted.

    How the levels work

    The README describes a ladder of readout levels, and every Decision carries its level so downstream code can require a minimum one.

    • Raw is a restricted softmax over label tokens.
    • L0 needs no labels. It averages out position bias over option rotations and divides out the label prior.
    • L1 adds temperature scaling using 100 to 500 labels per question.
    • L2 fits a closed-form head on a hidden state partway down the model, using 100 to 300 labels per question.

    The project says L2 heads are per question and per model, and do not transfer.

    Reported results

    The maintainers report, for Qwen3-8B on BANKING77 with 20 classes, that the order-flip rate falls from 0.230 to 0.073 with zero labels. They also report that the share of traffic auto-decidable at 5% error rises from 7.7% to 52.0% with L1. For L2 on the typed-decisions set, they report accuracies of 0.730 to 0.799 across five Qwen3 models. The README says every number is regenerated from committed JSON, and the figures for Jev and Laya were published by their authors and not rerun. The README also notes that accuracy on typed-decisions is agreement with a teacher LLM.

    Serving with vLLM

    The README says vLLM can serve every level, including L2, through an embed server's pooler. Its documented path is to install the hf extra, truncate a model to fewer blocks with anyjev.truncate, and serve it with vllm. The documentation presents an L2 deployment as a pooling server plus a few kilobytes of head. A pipeline command converts, serves and measures accuracy, ECE and latency in one step. An opt-in adaptive rotation budget is described as using 7.2 rotations instead of 18 at a certified 1% disagreement rate.

    Tradeoffs to know

    The README lists several limitations:

    • Calibration cannot fix a model that cannot answer.
    • L0 can cost accuracy when one label dominates.
    • Only Qwen3 heads ship, and the headline tables are Qwen models.
    • The letter readout supports at most 26 options.
    • Decisions are scored in isolation, not inside an agent loop.

    The roadmap lists agent-loop evaluation, more model families and SGLang support as not yet done.

    Current Status

    PyPI shows releases 0.0.1 and 0.0.2 on Sep 21, 0.1.0 on Sep 26 and 0.2.0 on Sep 28, 2026. The GitHub repository is public and still under active updates.

    AnyJev - 1

    Community Discussions

    Be the first to start a conversation about AnyJev

    Share your experience with AnyJev, ask questions, or help others learn from your insights.

    Pricing

    OPEN SOURCE

    Open Source (Apache-2.0)

    Self-hosted open-source Python library under Apache License 2.0; free to use, modify, and distribute.

    • Typed decisions with calibrated probabilities
    • Training-free L0 with zero labels
    • L2 closed-form heads per question
    • vLLM serving support
    • Apache-2.0 license

    Capabilities

    Key Features

    • Typed decisions: choice, yes/no (noul) and score from one prefill
    • L0 zero-label position-bias correction over option rotations
    • L1 temperature-scaled calibration with 100-500 labels
    • L2 closed-form heads on mid-depth hidden states with 100-300 labels
    • Adaptive rotation budget for fewer prefills per decision
    • vLLM serving via embed server pooler or truncated checkpoint
    • Model truncation tool and one-command pipeline benchmark
    • Label-free head adaptation and online label collection via observe
    • Shipped heads for five Qwen3 models
    • Transformers (Hugging Face) backend

    Integrations

    vLLM
    Hugging Face Transformers
    Qwen3
    Qwen2.5
    API Available
    View Docs

    Ratings & Reviews

    No ratings yet

    Be the first to rate AnyJev and help others make informed decisions.

    Rate other tools you’ve used

    Developer

    nokia-applied-research

    Read more about nokia-applied-research
    WebsiteGitHub
    1 tool in directory

    Similar Tools

    Jevlike icon

    Jevlike

    An open-source Python library for training small models that select among a changing list of text options in a single forward pass, inspired by TypeSafe's Jev model.

    Von icon

    Von

    An open-source, non-autoregressive System One decision model delivering calibrated discrete, probabilistic, and ordinal inference in sub-25ms, running locally without autoregressive text generation.

    Julia 1 icon

    Julia 1

    A compact 144.3M-parameter multilingual decision model from Supersonic Labs that classifies, ranks, and answers yes/no questions by choosing among supplied options, running on CPU.

    Browse all tools

    Related Topics

    AI Decision Models

    Models that make decisions instead of generating text. Given context and predefined options, they return choices, scores, or probabilities that applications can act on.

    11 tools

    AI Development Libraries

    Programming libraries and frameworks that provide machine learning capabilities, model integration, and AI functionality for developers.

    337 tools

    LLM Evaluations

    Platforms and frameworks for evaluating, testing, and benchmarking LLM systems and AI applications. These tools provide evaluators and evaluation models to score AI outputs, measure hallucinations, assess RAG quality, detect failures, and optimize model performance. Features include automated testing with LLM-as-a-judge metrics, component-level evaluation with tracing, regression testing in CI/CD pipelines, custom evaluator creation, dataset curation, and real-time monitoring of production systems. Teams use these solutions to validate prompt effectiveness, compare models side-by-side, ensure answer correctness and relevance, identify bias and toxicity, prevent PII leakage, and continuously improve AI product quality through experiments, benchmarks, and performance analytics.

    134 tools
    Browse all topics
    Back to all toolsSuggest an edit
    ratings
    discussions