# decider

> Open-source family of language models that return calibrated probability distributions for typed questions in a single forward pass.

decider is an open-source project by Mapika (Mark Marosi per the license and citation) that provides language models which do not generate text. Instead, each model reads a state and a set of typed questions and returns a probability distribution for every question from one forward pass. The code is Apache-2.0 licensed and is installed as the decider-ai Python package, with weights published on Hugging Face.

## What It Is

The README describes decider as an open reproduction of the "System One" model class and says it is independent and not affiliated with or endorsed by TypeSafe AI. A typed decision is a question with a fixed answer set: Choice over 2 to 255 options, Score over 2 to 10 described levels, or Noul, the probability of yes. There is no decoding or parsing, and no output outside the options you define. The README says the 2B, 4B and 35B mixture-of-experts models are built on Qwen3.5 base models, with training data drawn from public datasets plus data labelled by a local Qwen3.5-27B teacher.

## How the one-pass readout works

A request is rendered as text with one answer slot per question. The model reads the hidden state at each slot, projects it onto one label token per option, and applies a softmax over the valid options at a fitted temperature. Because letters are never generated, all slots come out of a single pass. Two prompt layouts are trained: state-first, the default, and schema-first, which caches a state-independent prefix for speed at some cost in accuracy. Per-type temperatures are configured in decider_config.json.

## Model lineup and hardware fit

The README maps hardware to models. According to it, the 2B and 4B models have GGUF builds for CPU use through llama.cpp, with Q4_K_M files of 1.3 GB and 2.7 GB. It lists the 4B at 8.4 GB in bf16 and decider-12b at 24 GB in bf16. The 35B-A3B needs about 65 GB in bf16 or 19.6 GB with NVFP4. It also lists two decider-chat repositories that read stock Gemma-4-31B and Qwen3.6-27B with a configured temperature, plus a 0.8B model and a vision variant.

## Standing and measured results

The README reports third-party leaderboard positions that the author says were read on 2026-09-29 and not run by the author. On JevBench it lists decider-4b v2 at #7 with a score of 71.3. On the Decision Index it lists decider-chat-gemma4-31b at #2 with 57.33. It also notes that Jev leads on knowledge-heavy areas such as GPQA, GSM8K, CRUXEval and MMLU.

## Runtime options

The package runs on CUDA, Apple Silicon via MPS, CPU in float32, and llama.cpp for GGUF files. It includes an HTTP server with a /v1/systemone endpoint that follows TypeSafe's wire format and a /decide plain form, plus a vLLM server for large checkpoints. A community port runs decider-12b on the Snapdragon X Elite NPU. Training scripts and data builders for about 95 public datasets are included.

## Limits to know

The README states the limits plainly: one pass cannot do multi-step arithmetic, calibration on hard items is the weak axis, and the models are English only. It also says knowledge-heavy multiple choice improves little on smaller models and that the RL stage is not yet in the package. The vision variant is still on v5 text weights and is being retrained.

## Current Status

The README's latest entry is decider-ai 1.8.1 dated 2026-09-30, which bounds the shared-prefix forward for Gemma-4 and long per-question suffixes and adds the DECIDER_SHARED_SUFFIX_TOKENS setting. The repository was last pushed 2026-10-02.

## Features
- One forward pass returns a probability distribution for every typed question
- Choice, Score and Noul question types
- Calibrated probabilities with per-type temperatures in decider_config.json
- HTTP server with /v1/systemone (TypeSafe wire format) and /decide endpoints
- vLLM server for large stock or chat-layout checkpoints
- GGUF loading through llama.cpp for CPU, CUDA or Metal
- CUDA graphs, torch.compile and optional FP8 linears
- Apple Silicon MPS acceleration
- Schema-first prompt layout with cached prefixes
- Training recipe, data builders and calibration fitting tools
- Open weights for 0.8B, 2B, 4B, 12B and 35B-A3B variants

## Integrations
Hugging Face, llama.cpp, vLLM, PyTorch, ONNX Runtime QNN, TypeSafe SDKs

## Platforms
LINUX, MACOS, WINDOWS, API, DEVELOPER_SDK, CLI

## Pricing
Open Source

## Version
1.8.1

## Links
- Website: https://github.com/Mapika/decider
- Repository: https://github.com/Mapika/decider
- EveryDev.ai: https://www.everydev.ai/tools/decider
