ollaya-dev
Ollaya is an independent open-source runtime that runs open decision models locally. It accepts text, JSON or other state plus typed questions and returns calibrated, structured answers in a single forward pass, keeping sensitive data on the user's hardware.
At a Glance
- Developers building AI applications
- Enterprises requiring private/on-premises inference
- Customer-support and operations teams processing tickets, emails and messages
- AI-agent and automation developers
- +1 more
AI Tools by ollaya-dev
(1)Ollaya
Local Decision Model Runner
Discussions
No discussions yet
Be the first to start a discussion about ollaya-dev
Latest News
Website and documentation updated to recommend winnow:e4b, with a 0.722 typed-decisions score and 89 ms five-question RTX 4090 benchmark.
v0.7.2 added RTX 50-series/Blackwell GPU support and the JevK5 decision model.
v0.7.1 added Apple-silicon GPU execution through MLX for laya and nli:modernbert-large.
v0.7.0 added native NVIDIA GPU support on Windows, larger Kev models, the Decision model family and GGUF execution through llama.cpp.
Products & Services
Apache-2.0 Rust binary and daemon for pulling, serving and managing open decision models locally. It provides Ollama-like commands such as run, pull, serve, list, ps, show, stop, rm, cp and create, plus native /api endpoints and TypeSafe-compatible /v1/systemone, /v1/decisions and /v1/models endpoints.
Desktop application for macOS, Windows and Linux to start and stop the server, download models and try them in one window.
The ollaya mcp command exposes local decision models to MCP clients, and the ollaya-decisions skill teaches agents when and how to use the local models.
A registry and runtime for open model families including winnow, laya, decider, kev, nli, gliclass, qwen3guard, decision, von and jevk5. Models are pulled by name; Ollaya publishes small derived graphs/manifests and does not re-host authors' weights.
Market Position
Ollaya positions itself as the Ollama-like local runtime for open decision models, rather than a generative LLM server. Its differentiators are local privacy, calibrated typed outputs, single-pass multi-question inference, and drop-in TypeSafe/Jev API compatibility. It competes conceptually with TypeSafe's hosted Jev and with general LLM APIs or local LLM runtimes such as Ollama; the project's own benchmark places winnow:e4b and kev:9b near Jev on typed-decision accuracy while running locally.
Leadership
Founders
Mert Cobanov
Identified in the project's Hacker News launch discussion as its developer and as the maintainer of the ollaya-dev GitHub organization. His GitHub profile describes him as a Senior AI Engineer and Generative Artist at Refik Anadol Studio, working on Dataland, AI agents, developer-experience tools, backend systems, vector databases and generative-AI pipelines.
Founding Story
Ollaya began as an open-source answer to hosted Jev-style decision inference: the first release describes the goal as running open decision models locally, the way Ollama runs language models. The initial vision was a single binary and local daemon that pulls open models by name, serves them privately on user hardware, and provides a TypeSafe-compatible API so existing clients can switch from hosted inference.
Business Model
Revenue Model
The project is free and open source under Apache-2.0, with no per-token fees or API metering. Users run it on their own hardware; models are downloaded locally and the project distributes binaries, a desktop app and Docker images.
Target Markets
- Developers building AI applications
- Enterprises requiring private/on-premises inference
- Customer-support and operations teams processing tickets, emails and messages
- AI-agent and automation developers
- Users running local or edge inference on macOS, Windows, Linux and Docker
- Ticket and email intent routing
- Refund, urgency, frustration and churn-risk decisions
- Safety screening and content moderation
- Structured classification of text, JSON and other sensitive local data
- Low-latency, privacy-preserving AI routing on servers, desktops, developer machines and edge hardware
- Agent workflows that need local typed decisions through MCP