# Ollaya

> Run open decision models locally — pull and serve typed, calibrated classification models via a TypeSafe-compatible API, on your own CPU or NVIDIA/Apple GPU.

Ollaya is an independent open-source project (Apache-2.0) that lets you download and run decision models locally, the same way Ollama runs large language models. It serves a local HTTP daemon with a TypeSafe-compatible `/v1/systemone` API, so existing Jev clients work by changing a single environment variable. The project is written in Rust and currently in beta, with its latest release at v0.7.2 published in late September 2026.

## What It Is

A decision model is a specialized neural network that reads a *state* — a message, email, ticket, or JSON blob — plus typed questions (`choice`, `score`, `noul`/yes-no), and returns calibrated probabilities in a single forward pass without generating any text. Ollaya acts as the local runtime and model manager for this class of models, pulling weights from their authors' Hugging Face repositories, verifying them by sha256, and serving them through a unified API. It is not affiliated with Ollama or TypeSafe.

## How the Runtime Works

Ollaya runs one binary (`ollaya serve`) that starts a local daemon. The CLI mirrors Ollama's command surface: `run`, `pull`, `list`, `ps`, `show`, `rm`, `cp`, `stop`, and `create`. Inference uses ONNX Runtime for encoder models (CPU and NVIDIA CUDA) and llama.cpp for GGUF-based decoder models. On Apple silicon, `laya` and `nli` run on the Apple GPU through MLX. Precision is chosen automatically: fp16 on GPU, fp32 on CPU.

Key runtime details:
- Default server address is `127.0.0.1:11435`; data never leaves the machine by default
- NVIDIA GPUs require driver R580 or newer; CUDA libraries are fetched only when a GPU is detected
- Docker images are published to GHCR for both CPU (`amd64`/`arm64`) and CUDA (`amd64`) targets
- A desktop app (menu bar on macOS, windowed on Windows/Linux) wraps the daemon for non-CLI users
- Modelfiles let users bake a question set, calibration, and precision into a named model

## Model Library

Ollaya ships a curated set of open-weight decision models, each kept under its author's own license:

- **laya** — Router that dispatches to `laya:en` (ModernBERT-large, 421M) or `laya:multilingual` (mmBERT-base, 322M); the homepage reports 8–10 ms for five questions on an RTX 4090
- **decider** — Mapika's Qwen3.5 decoder family (0.8B, 2B, 4B); `decider:4b` scores 0.680 on the typed-decisions benchmark
- **kev** — Jared Palmer's LoRA + pointer-head on Qwen3.5 (0.8B default, 4B, 9B); `kev:9b` scores 0.722 on typed-decisions
- **nli** — Moritz Laurer's zero-shot NLI classifiers (DeBERTa-v3-large, ModernBERT-large)
- **gliclass** — Knowledgator's instruction-following zero-shot classifier (DeBERTa-v3-large)
- **von** — Victor Hugo Panisa's Von 1.1 (ModernBERT-large) with 8k-token context
- **qwen3guard** — Qwen team's safety guard with built-in questions, 119 languages
- **decision**, **winnow**, **jevk5** — additional community fine-tunes

Weights are never re-hosted by Ollaya; the project publishes only small ONNX graphs (~3 MB each) that reference the author's Hugging Face files.

## Agent and MCP Integration

Ollaya includes first-class support for AI agent workflows. `ollaya mcp` starts an MCP server that exposes local decision models to Claude Code, Claude Desktop, Cursor, and other MCP clients. An `ollaya-decisions` agent skill (installable via `npx skills add`) teaches agents when and how to call the local API. The TypeSafe-compatible wire format means any code already written against the TypeSafe SDK works against a local Ollaya server by setting `TYPESAFE_BASE_URL=http://localhost:11435`.

## Update: v0.7.2

The latest release is v0.7.2, published 2026-09-26. The repository was created 2026-09-23 and has accumulated 425 stars and 16 forks in its first days. The project is actively iterating, with the GitHub README and docs covering CLI reference, API reference, Modelfile format, TypeSafe compatibility, and agent/MCP integration as distinct documentation sections. The beta label on the homepage reflects its early-access status.

## Features
- Run decision models locally on CPU or GPU
- TypeSafe-compatible /v1/systemone and /v1/models API
- Single forward pass inference — no token generation
- Calibrated probability outputs for choice, score, and yes/no questions
- Pull models by name from authors' Hugging Face repositories
- SHA256 verification of all model weights
- ONNX Runtime for encoder models; llama.cpp for GGUF decoder models
- Apple GPU (MLX) support for laya and nli on Apple silicon
- NVIDIA CUDA support on Linux, Windows, and Docker
- Desktop app for macOS, Windows, and Linux
- CLI mirroring Ollama's command surface (run, pull, list, ps, show, rm, cp, stop, create)
- Modelfiles to bake question sets, calibration, and precision into named models
- MCP server for Claude Code, Claude Desktop, Cursor, and other MCP clients
- ollaya-decisions agent skill for AI agent workflows
- Language and script detection router (laya)
- Docker images on GHCR for CPU and CUDA targets
- No per-token fees or metering
- Private by default — server listens on 127.0.0.1

## Integrations
TypeSafe SDK, Hugging Face, Claude Code, Claude Desktop, Cursor, ONNX Runtime, llama.cpp, MLX (Apple), NVIDIA CUDA, Docker, MCP (Model Context Protocol)

## Platforms
WINDOWS, MACOS, LINUX, API, CLI

## Pricing
Open Source

## Version
v0.7.2

## Links
- Website: https://ollaya.dev
- Documentation: https://ollaya.dev/docs
- Repository: https://github.com/ollaya-dev/ollaya
- EveryDev.ai: https://www.everydev.ai/tools/ollaya
