Drex DLM
A decision model from Nace.AI that returns probabilities for typed questions about a given document using an 8B diffusion language model with a pointer head.
At a Glance
Self-hosted Drex DLM repository code under MIT; model weights under CC BY-NC 4.0 (non-commercial).
Engagement
Available On
Listed Oct 2026
About Drex DLM
Drex DLM is a decision model from Nace.AI that answers typed questions about a supplied context. You pass the context as state along with one or more named questions, and it returns a probability for every option. The repository provides Python inference and server code, and the weights are published on Hugging Face in BF16 and Q8_0 GGUF forms.
What It Is
Drex DLM is built on NVIDIA Efficient-DLM-8B, a diffusion language model, with a decision adapter merged into its weights. One forward pass plus a shared pointer head produces the probability distribution. The pointer head projects the hidden state at a decision marker into a query and the hidden states at option-ending markers into keys. Scaled dot products, temperature scaling and a softmax then give per-question probabilities. The shared context uses bidirectional attention, and question branches attend to that context but not to one another.
Request and Output Format
A request has a state (a string, object or list) and a dictionary of named questions. Three question types are supported:
choice: named options, returning the chosen option, a probability per option and a confidence value.noul: a yes/no question, returning the probability of yes.score: an ordered scale, returning a probability-weighted score, a legend, per-level probabilities and a confidence value.
Responses also include token usage and scoring latency. One decision per request is the stated release contract.
Serving Options
The model can be run through three local runners that each expose POST /v1/systemone: a Python server, a llama-server build from the edlm branch of Nace's llama.cpp fork, and a custom Ollama fork. The maximum context window is 32,768 tokens, while the local runners default to 16,384 tokens. A separate Drex agent skill lets coding agents and other harnesses call a self-hosted server without a hosted API key.
Requirements and Licensing
Python 3.12 is the validated interpreter. Local inference was tested on an Apple M5 Max with 128 GiB of unified memory; CUDA and CPU-only inference have not been validated. BF16 weights take about 16 GB. The model weights are released under CC BY-NC 4.0, while the original Nace.AI code in the repository is MIT licensed.
Community Discussions
Be the first to start a conversation about Drex DLM
Share your experience with Drex DLM, ask questions, or help others learn from your insights.
Pricing
Open Source (Self-hosted)
Self-hosted Drex DLM repository code under MIT; model weights under CC BY-NC 4.0 (non-commercial).
- Original Nace.AI code in the repository is under the MIT License
- Model weights and GGUF files are released under CC BY-NC 4.0
- Local Python, llama-server, and Ollama runners
- Maximum context window of 32,768 tokens (16,384 by default)
- No hosted API key needed for self-hosted use
Capabilities
Key Features
- Typed question answering over a shared document context
- Choice, yes/no (noul) and score question types
- Per-option probabilities with confidence values
- Diffusion language model backbone with pointer head
- Single forward pass for packed questions
- Python, llama-server and Ollama runners exposing /v1/systemone
- BF16 and Q8_0 GGUF checkpoints
- 32,768-token maximum context window
