# LocalJev

> A local Jev-compatible POST /v1/systemone API in TypeScript for Bun, backed by DiffusionGemma via an OpenAI-compatible Chat Completions endpoint.

LocalJev is an open-source, MIT-licensed project from the githubnext organization on GitHub. It provides a local, Jev-compatible `POST /v1/systemone` API written in TypeScript for Bun. The README says it is backed by DiffusionGemma through an OpenAI-compatible Chat Completions endpoint.

## What It Is

LocalJev is a bridge server that exposes the typed decision API used by Jev while sending the actual inference to a local model server. Clients send a `state` and a set of typed questions (choice, score, or noul), and receive the normal Jev response shape. The README lists defaults of an inference server at 127.0.0.1:8000, the model `diffusiongemma-26B-A4B-it-4bit`, and the LocalJev API at 127.0.0.1:8080.

## Why a bridge is needed

The README explains that Jev uses a typed decision API rather than an OpenAI chat API. OpenJev implements the Jev wire protocol and reads probabilities with a one-step DiffusionGemma structured read, which depends on unmerged vLLM request extensions. The normal oMLX API does not expose those primitives, so LocalJev takes a portable approach:

- It translates state and typed questions into a classification prompt.
- It asks DiffusionGemma for a JSON probability scalar or vector.
- It validates the result and retries malformed output.
- It normalizes vectors and computes choices, expected scores, and entropy-based confidence.

## Tradeoffs to know

The README states that this is wire-compatible but not mathematically equivalent to OpenJev's logit read. Probabilities are generated or self-reported by the model rather than read from its logits. It advises evaluating calibration on your own workload before relying on results for consequential decisions.

## Setup path

The README says LocalJev requires Bun 1.2+ and a running oMLX server. Setup involves running `bun install`, copying `.env.example` to `.env`, setting the upstream API key, and running `bun run start`. A `/ready` endpoint checks that the configured model is available. The TypeSafe Python SDK can point at LocalJev through environment variables, and the aliases `jev-latest` and `jev-preview` are accepted so SDK defaults work unchanged. Configuration covers upstream URL, model, timeouts, concurrency limits, queue size, retries, and chunking.

## Model evaluation and runner notes

The repository includes a repeatable bake-off using AG News, BoolQ, and SST-5 to compare installed models on quality, calibration, retries, and latency. The README reports a first completed run of 1,200 requests on an M5 Max, in which Gemma 4 26B-A4B and Qwen3.6 were the strongest overall candidates in a small screening sample, with caveats. It also says that, as of September 18, 2026, DiffusionGemma support in LM Studio was still tracked as open, making oMLX the better runner for this model.

## Features
- Jev-compatible POST /v1/systemone API
- Supports choice, score, and noul question types
- Backed by DiffusionGemma via OpenAI-compatible Chat Completions
- Validation and corrective retries for malformed model JSON
- Normalized probabilities, expected scores, and entropy-based confidence
- Accepts jev-latest and jev-preview model aliases
- Optional Bearer API key for clients
- Configurable concurrency, queueing, timeouts, and chunking
- /ready endpoint for model availability checks
- Repeatable multi-model evaluation bake-off

## Integrations
Bun, oMLX, DiffusionGemma, TypeSafe SDK, OpenJev, OpenAI-compatible Chat Completions

## Platforms
MACOS, API, CLI

## Pricing
Open Source

## Links
- Website: https://github.com/githubnext/localjev
- Documentation: https://github.com/githubnext/localjev/blob/main/docs/evaluation.md
- Repository: https://github.com/githubnext/localjev
- EveryDev.ai: https://www.everydev.ai/tools/localjev
