# Strands Decider

> A small, fast open source decision model that picks between options or rates on a scale, with a calibrated confidence on every decision.

Strands Decider is an open source decision model from Strands Labs, the experimental arm of Strands Agents. Instead of generating text like an LLM, it picks between sets of options and rates things on a scale, aimed at the decisions inside agentic workflows built with the Strands Agents SDK. The reference model, strands-decider-2B (v19), has 1.9 billion parameters and ships with a CLI and an HTTP server.

## What It Is

Strands Decider is a "system one" decision model: a general-purpose classification and scoring model that, per the README, responds faster than an LLM and needs less time and expertise than training a traditional classifier. It takes a state (text, optionally images) plus questions, and returns answers with confidence values. Three question types are supported: `noul` (yes/no), `choice` (one of N options) and `score` (an ordered rubric).

## How It Works

The model starts from a pretrained Qwen3.5-2B-Base decoder, discards its language-modelling head, and replaces it with a small pointer head of about a million parameters. The head scores each option by comparing the hidden state at the answer position with the hidden state at each option's last token. This takes one forward pass with no generation loop, and label sets are defined per request rather than baked into the weights. Multiple questions about the same text can be combined in one call so the state is read only once.

## Using It

Install with `pip install strands-decider`, then use the `strands-decider ask` command with `--state` and `--choice`, `--noul` or `--score` arguments. `strands-decider serve` runs an HTTP server exposing a `/v1/systemone` endpoint. An MLX device option is documented for Apple-silicon Macs, and a `--vision` option lets requests carry base64 images. A worked example in the repository gates a Strands agent's tool call with a `before_tool_call` intervention based on two yes/no decisions.

## Intended Uses and Measurements

The README lists model routing, tool selection, argument checking, triage, guardrails, evaluations and hybrid agents as target uses. It reports 0.723 accuracy (167 of 231 tasks) on the JevBench v1 public set and a median latency of 115 ms on an RTX 3090, and states that answers at confidence 0.9 or above are right about 95% of the time on short unseen classification tasks. It also cautions that differences under about 10 tasks between single runs are unresolved. Training is documented at about 11 hours on one RTX 3090, and the repository includes a preregistered research log.

## Features
- Choice questions selecting one of N options
- Yes/no (noul) questions with a probability output
- Score questions over ordered rubrics
- Calibrated confidence on every decision
- Multiple questions per request over the same state
- CLI ask command and HTTP server endpoint
- Optional vision input with images
- Apple-silicon MLX, CUDA, MPS and CPU serving
- Reproducible training recipe

## Integrations
Strands Agents SDK, Hugging Face, Qwen3.5-2B-Base, MLX

## Platforms
MACOS, LINUX, WEB, API, CLI

## Pricing
Open Source

## Links
- Website: https://strandsagents.com/docs/labs/
- Repository: https://github.com/strands-labs/strands-decider
- EveryDev.ai: https://www.everydev.ai/tools/strands-decider
