# LM-Kit One

> A private AI application server that runs models, agents, RAG, search, and document intelligence on your own hardware, compatible with OpenAI, Anthropic, Ollama, and MCP clients.

LM-Kit One is a private AI application server from LM-Kit, designed to run models, agents, RAG, search, and document intelligence entirely on infrastructure you control. It serves the OpenAI, Anthropic, Ollama, and MCP API dialects, so existing clients and coding agents can connect by changing a base URL. The product launched with version 2026.9.8, signed for Windows, Linux, and macOS.

## What It Is

LM-Kit One is a self-hosted AI backend that consolidates what would otherwise require twelve separate components — a model runner, API gateway, authentication, vector database, RAG framework, OCR and document stack, agent runtime, MCP host, admin console, observability, fine-tuning stack, and workload queue — into a single versioned server. It is positioned as a "private AI application server" for teams that cannot or will not send documents and data to hosted AI services due to policy, contract, regulatory, cost, or connectivity constraints. The server is not open source; LM-Kit publishes the free tier without activation, account, or runtime license checks.

## Architecture: One Engine, Twelve Jobs

Rather than gluing independently versioned open-source services together, LM-Kit One is built and released as a single product by one team. The product page states it has shipped 150+ releases since 2024. Key architectural properties include:

- **Four API dialects**: Native REST, OpenAI (chat completions, embeddings, files, vector stores, Responses API), Anthropic (Messages with thinking), and Ollama (full lifecycle including pull, push, create, copy, delete, blob upload)
- **Storage flexibility**: Fully local, 100% on PostgreSQL, or MySQL/SQL Server paired with Qdrant for vectors
- **Hardware backends**: CPU, CUDA, Vulkan, and Metal — no GPU required to start
- **Horizontal scaling**: Nodes share state so any node serves any request; KEDA-ready on Kubernetes
- **Security posture**: Loopback-only by default, API tokens hashed at rest, egress governed by allowlist, air-gapped operation supported

## Document Intelligence and Search

The server includes a production-grade document intelligence stack covering OCR, layout understanding, PDF/A conversion, digital signatures, redaction, and structured data extraction with per-field confidence scores and grammar-constrained decoding. The search engine provides BM25 with language-aware analyzers (including Snowball stemming and CJK bigram handling), vector search, hybrid retrieval with reciprocal rank fusion or convex combination, reranking, facets, filters, MMR diversity, recency decay, and multi-tenant collection management — capabilities the product page compares to what teams typically run a dedicated OpenSearch or Elasticsearch cluster for, while noting it is not wire-compatible with those APIs.

## Agent Runtime and MCP Integration

LM-Kit One includes a server-side agent runtime with 18 pre-built agent templates, six planning/reasoning strategies, multi-agent workflow patterns (pipeline, parallel, router, supervisor), graph orchestration, persistent memory, and 70+ built-in tools. Agents are defined server-side and adopted by clients with a single field on a chat request. The server acts as a full MCP host, exposing document and knowledge tools over the Model Context Protocol so external AI assistants (such as Claude Code, Continue, Cline, or Zed) can call governed local tools without source documents leaving the network.

## Deployment and Migration Path

LM-Kit One installs as a single package on Windows, Linux (x64 and ARM64), and macOS, and can run as a Windows service or desktop session. The product provides explicit migration guides for teams coming from OpenAI, Ollama, and Anthropic — in most cases, the migration is a base URL change. The compatibility matrix documents every endpoint served and every gap. The admin console covers model management, access tokens, skills, collections, request history, telemetry, and logs.

## Update: Launch of LM-Kit One

LM-Kit One was announced as a new product with a dedicated launch post ("LM-Kit One is here"). The current version at launch is 2026.9.8. The companion product LM-Kit.NET (the embeddable .NET SDK) has been available since at least 2024 and the NuGet packages report 360,000+ downloads according to the LM-Kit homepage. LM-Kit One extends the same engine to a standalone server deployment, adding the four API dialects, authentication, policies, audit, console, telemetry, and fine-tuning job management that the embedded SDK does not expose.

## Features
- OpenAI-compatible API (chat completions, embeddings, files, vector stores, Responses API)
- Anthropic Messages API with thinking support
- Ollama API with full model lifecycle (pull, push, create, copy, delete)
- MCP host with governed tool catalog
- Native REST API for documents, agents, search, and training
- OCR and layout understanding for scans and photographs
- Structured data extraction with per-field confidence scores
- PDF/A conversion and validation (ISO 19005)
- PDF digital signatures (PAdES sign, verify, timestamp, LTV)
- Smart and reviewed PDF redaction
- BM25 full-text search with language-aware analyzers
- Vector and hybrid search with rank fusion
- Built-in vector database (no separate provisioning)
- RAG with citations (document, page, passage, retrieval score)
- 18 pre-built agent templates
- Six agent reasoning/planning strategies
- Multi-agent workflows (pipeline, parallel, router, supervisor)
- Graph orchestration with composable workflow nodes
- Agent memory (persistent, retrieval-grounded)
- 70+ built-in tools plus custom function calling
- Fine-tuning as managed jobs (submit, poll, cancel)
- LoRA integration with hot-swap adapters
- Model quantization (30+ formats, FP32 to 1-bit)
- Hardware backends: CPU, CUDA, Vulkan, Metal
- Multi-GPU and tensor overrides
- Horizontal scaling with KEDA-ready Kubernetes support
- Admin console with request history, telemetry, and logs
- OpenTelemetry GenAI tracing
- Air-gapped operation support
- Loopback-only default network posture
- API tokens hashed at rest
- Egress governed by allowlist
- Real-time speech transcription (100+ languages)
- Context hibernation (pause and resume conversations)
- Encrypted model loading

## Integrations
OpenAI SDK, Anthropic SDK, Ollama CLI, MCP (Model Context Protocol), Open WebUI, Claude Code, Continue, Cline, Zed, Microsoft.Extensions.AI, Semantic Kernel, Amazon Bedrock, PostgreSQL, MySQL, SQL Server, Qdrant, pgvector, Kubernetes / KEDA, NuGet

## Platforms
WINDOWS, MACOS, LINUX, API, CLI

## Pricing
Freemium — Free tier available with paid upgrades

## Version
2026.9.8

## Links
- Website: https://lm-kit.com/products/lm-kit-one/
- Documentation: https://docs.lm-kit.com/lm-kit-one/guides/index.html
- Repository: https://github.com/LM-Kit
- EveryDev.ai: https://www.everydev.ai/tools/lm-kit-one
