LM-Kit One
A private AI application server that runs models, agents, RAG, search, and document intelligence on your own hardware, compatible with OpenAI, Anthropic, Ollama, and MCP clients.
At a Glance
Complete SDK and server including commercial use and redistribution for small companies below published thresholds. Evaluation and development free at any size.
Engagement
Available On
Alternatives
Listed Oct 2026
About LM-Kit One
LM-Kit One is a private AI application server from LM-Kit, designed to run models, agents, RAG, search, and document intelligence entirely on infrastructure you control. It serves the OpenAI, Anthropic, Ollama, and MCP API dialects, so existing clients and coding agents can connect by changing a base URL. The product launched with version 2026.9.8, signed for Windows, Linux, and macOS.
What It Is
LM-Kit One is a self-hosted AI backend that consolidates what would otherwise require twelve separate components — a model runner, API gateway, authentication, vector database, RAG framework, OCR and document stack, agent runtime, MCP host, admin console, observability, fine-tuning stack, and workload queue — into a single versioned server. It is positioned as a "private AI application server" for teams that cannot or will not send documents and data to hosted AI services due to policy, contract, regulatory, cost, or connectivity constraints. The server is not open source; LM-Kit publishes the free tier without activation, account, or runtime license checks.
Architecture: One Engine, Twelve Jobs
Rather than gluing independently versioned open-source services together, LM-Kit One is built and released as a single product by one team. The product page states it has shipped 150+ releases since 2024. Key architectural properties include:
- Four API dialects: Native REST, OpenAI (chat completions, embeddings, files, vector stores, Responses API), Anthropic (Messages with thinking), and Ollama (full lifecycle including pull, push, create, copy, delete, blob upload)
- Storage flexibility: Fully local, 100% on PostgreSQL, or MySQL/SQL Server paired with Qdrant for vectors
- Hardware backends: CPU, CUDA, Vulkan, and Metal — no GPU required to start
- Horizontal scaling: Nodes share state so any node serves any request; KEDA-ready on Kubernetes
- Security posture: Loopback-only by default, API tokens hashed at rest, egress governed by allowlist, air-gapped operation supported
Document Intelligence and Search
The server includes a production-grade document intelligence stack covering OCR, layout understanding, PDF/A conversion, digital signatures, redaction, and structured data extraction with per-field confidence scores and grammar-constrained decoding. The search engine provides BM25 with language-aware analyzers (including Snowball stemming and CJK bigram handling), vector search, hybrid retrieval with reciprocal rank fusion or convex combination, reranking, facets, filters, MMR diversity, recency decay, and multi-tenant collection management — capabilities the product page compares to what teams typically run a dedicated OpenSearch or Elasticsearch cluster for, while noting it is not wire-compatible with those APIs.
Agent Runtime and MCP Integration
LM-Kit One includes a server-side agent runtime with 18 pre-built agent templates, six planning/reasoning strategies, multi-agent workflow patterns (pipeline, parallel, router, supervisor), graph orchestration, persistent memory, and 70+ built-in tools. Agents are defined server-side and adopted by clients with a single field on a chat request. The server acts as a full MCP host, exposing document and knowledge tools over the Model Context Protocol so external AI assistants (such as Claude Code, Continue, Cline, or Zed) can call governed local tools without source documents leaving the network.
Deployment and Migration Path
LM-Kit One installs as a single package on Windows, Linux (x64 and ARM64), and macOS, and can run as a Windows service or desktop session. The product provides explicit migration guides for teams coming from OpenAI, Ollama, and Anthropic — in most cases, the migration is a base URL change. The compatibility matrix documents every endpoint served and every gap. The admin console covers model management, access tokens, skills, collections, request history, telemetry, and logs.
Update: Launch of LM-Kit One
LM-Kit One was announced as a new product with a dedicated launch post ("LM-Kit One is here"). The current version at launch is 2026.9.8. The companion product LM-Kit.NET (the embeddable .NET SDK) has been available since at least 2024 and the NuGet packages report 360,000+ downloads according to the LM-Kit homepage. LM-Kit One extends the same engine to a standalone server deployment, adding the four API dialects, authentication, policies, audit, console, telemetry, and fine-tuning job management that the embedded SDK does not expose.
Community Discussions
Be the first to start a conversation about LM-Kit One
Share your experience with LM-Kit One, ask questions, or help others learn from your insights.
Pricing
Free
Complete SDK and server including commercial use and redistribution for small companies below published thresholds. Evaluation and development free at any size.
- Complete LM-Kit One server
- Complete LM-Kit.NET SDK
- Commercial use and redistribution
- No activation key or runtime license check
- For companies under $1M USD annual gross revenue, 10 or fewer employees, and no more than $3M USD raised from outside investors
Professional
Required above the free thresholds. Scoped per application (LM-Kit.NET) or per deployment (LM-Kit One). Not metered by tokens or seats.
- Production and redistribution rights set in your order
- Long-term support builds and security patches
- Support with response-time commitments
- Priced on scope, not on developers or end users
Capabilities
Key Features
- OpenAI-compatible API (chat completions, embeddings, files, vector stores, Responses API)
- Anthropic Messages API with thinking support
- Ollama API with full model lifecycle (pull, push, create, copy, delete)
- MCP host with governed tool catalog
- Native REST API for documents, agents, search, and training
- OCR and layout understanding for scans and photographs
- Structured data extraction with per-field confidence scores
- PDF/A conversion and validation (ISO 19005)
- PDF digital signatures (PAdES sign, verify, timestamp, LTV)
- Smart and reviewed PDF redaction
- BM25 full-text search with language-aware analyzers
- Vector and hybrid search with rank fusion
- Built-in vector database (no separate provisioning)
- RAG with citations (document, page, passage, retrieval score)
- 18 pre-built agent templates
- Six agent reasoning/planning strategies
- Multi-agent workflows (pipeline, parallel, router, supervisor)
- Graph orchestration with composable workflow nodes
- Agent memory (persistent, retrieval-grounded)
- 70+ built-in tools plus custom function calling
- Fine-tuning as managed jobs (submit, poll, cancel)
- LoRA integration with hot-swap adapters
- Model quantization (30+ formats, FP32 to 1-bit)
- Hardware backends: CPU, CUDA, Vulkan, Metal
- Multi-GPU and tensor overrides
- Horizontal scaling with KEDA-ready Kubernetes support
- Admin console with request history, telemetry, and logs
- OpenTelemetry GenAI tracing
- Air-gapped operation support
- Loopback-only default network posture
- API tokens hashed at rest
- Egress governed by allowlist
- Real-time speech transcription (100+ languages)
- Context hibernation (pause and resume conversations)
- Encrypted model loading
