# Drex DLM

> A decision model from Nace.AI that returns probabilities for typed questions about a given document using an 8B diffusion language model with a pointer head.

Drex DLM is a decision model from Nace.AI that answers typed questions about a supplied context. You pass the context as `state` along with one or more named questions, and it returns a probability for every option. The repository provides Python inference and server code, and the weights are published on Hugging Face in BF16 and Q8_0 GGUF forms.

## What It Is

Drex DLM is built on NVIDIA Efficient-DLM-8B, a diffusion language model, with a decision adapter merged into its weights. One forward pass plus a shared pointer head produces the probability distribution. The pointer head projects the hidden state at a decision marker into a query and the hidden states at option-ending markers into keys. Scaled dot products, temperature scaling and a softmax then give per-question probabilities. The shared context uses bidirectional attention, and question branches attend to that context but not to one another.

## Request and Output Format

A request has a `state` (a string, object or list) and a dictionary of named questions. Three question types are supported:

- `choice`: named options, returning the chosen option, a probability per option and a confidence value.
- `noul`: a yes/no question, returning the probability of yes.
- `score`: an ordered scale, returning a probability-weighted score, a legend, per-level probabilities and a confidence value.

Responses also include token usage and scoring latency. One decision per request is the stated release contract.

## Serving Options

The model can be run through three local runners that each expose `POST /v1/systemone`: a Python server, a llama-server build from the `edlm` branch of Nace's llama.cpp fork, and a custom Ollama fork. The maximum context window is 32,768 tokens, while the local runners default to 16,384 tokens. A separate Drex agent skill lets coding agents and other harnesses call a self-hosted server without a hosted API key.

## Requirements and Licensing

Python 3.12 is the validated interpreter. Local inference was tested on an Apple M5 Max with 128 GiB of unified memory; CUDA and CPU-only inference have not been validated. BF16 weights take about 16 GB. The model weights are released under CC BY-NC 4.0, while the original Nace.AI code in the repository is MIT licensed.

## Features
- Typed question answering over a shared document context
- Choice, yes/no (noul) and score question types
- Per-option probabilities with confidence values
- Diffusion language model backbone with pointer head
- Single forward pass for packed questions
- Python, llama-server and Ollama runners exposing /v1/systemone
- BF16 and Q8_0 GGUF checkpoints
- 32,768-token maximum context window

## Integrations
Hugging Face, llama.cpp, Ollama, NVIDIA Efficient-DLM-8B, Drex agent skill

## Platforms
MACOS, LINUX, API, CLI

## Pricing
Open Source

## Links
- Website: https://github.com/nace-ai/drex-dlm
- Documentation: https://github.com/nace-ai/drex-dlm
- Repository: https://github.com/nace-ai/drex-dlm
- EveryDev.ai: https://www.everydev.ai/tools/drex-dlm
