Kottos AI, Inc.
Kottos AI builds exchange-grade infrastructure for LLM inference: it records request metadata such as cost, cache usage, timing and errors, then uses that evidence to provide intelligent routing across model providers and venues. Its open-source llmbridge gateway supplies the low-latency translation and proxy layer.
At a Glance
- AI application developers and teams operating LLM inference at scale
- Enterprises requiring managed or customer-hosted inference infrastructure
- Latency-sensitive AI applications
- Teams using multiple model providers or open-weight model venues
- +2 more
AI Tools by Kottos AI, Inc.
(1)llmbridge
Open Source C++ LLM Gateway
Discussions
No discussions yet
Be the first to start a discussion about Kottos AI, Inc.
Latest News
llmbridge v0.59.1: Bedrock credential handling fix
llmbridge v0.59.0 and v0.58.0 add request sequencing and optional prefix hashing metadata
llmbridge v0.57.0 adds the upstream venue request ID to request records
llmbridge v0.56.0 adds recording of OpenAI cache-write counts
Products & Services
Commercial routing intelligence that observes served requests, provider quotes, cache behavior, price, timing and venue health, predicts expected cost/performance, and routes requests according to customer constraints.
Managed OpenAI-compatible endpoint that records venue, cost, timing and failures, provides routing, observability and team-management capabilities, and supports BYOK credentials passed per request.
Apache-2.0 open-source C++ sub-millisecond LLM gateway. It accepts OpenAI-compatible requests, translates to provider dialects such as Anthropic, Gemini and Cohere, proxies responses, supports streaming SSE, tool calling, prompt-cache forwarding, timing headers, TLS and credential passthrough.
No-signup trial that places Kottos between Claude Code and Anthropic and exposes request metadata such as model, venue, time to first token, cache reads, cost and status in a shared public dashboard; prompt and response text are not stored.
Market Position
Kottos positions llmbridge against general-purpose LLM gateways such as LiteLLM, Bifrost, Portkey, GoModel and similar proxies by emphasizing C++ implementation, microsecond translation overhead, single-core efficiency and high concurrency. It differentiates the commercial Kottos layer from a basic gateway through an inference tape, live provider price/latency data, cache-aware expected-cost routing, verification of expected versus actual cost, observability and managed deployment. The company explicitly describes LiteLLM, Bifrost and Helicone as adjacent gateway/observability work and published benchmark comparisons with several of them.
Leadership
Founders
Lluís Antoni Jiménez Rugama
Founder; PhD in applied mathematics with nine years in high-frequency and electronic trading. His LinkedIn profile identifies him as Kottos AI's founder and places him in Austin; he also maintains the kottos-ai GitHub organization.
Executive Team
Lluís Antoni Jiménez Rugama
Founder
PhD in applied mathematics and nine years in high-frequency and electronic trading; listed by LinkedIn as Kottos AI founder and described by Kottos as the person who founded and built the company.
Founding Story
The founder describes an analogy between LLM inference and pre-TRACE fixed-income markets: providers quote prices, but users need an auditable record of what was actually delivered, including effective cost, latency and failures. Kottos AI was started to build an inference 'tape' and use it to improve routing decisions; llmbridge was built as the exchange-grade, microsecond-overhead gateway underneath it.
Business Model
Revenue Model
The commercial layer is managed/private-beta hosted infrastructure and enterprise or customer-hosted routing intelligence, observability and integrations. The open-source llmbridge core is Apache 2.0; hosted use is currently private beta and BYOK, with customers pointing an OpenAI-compatible client at Kottos and Kottos managing the gateway and routing.
Pricing Tiers
Shared public trial account/dashboard; provider quotas apply and the trial page says nothing is billed to Kottos.
Managed infrastructure, routing, monitoring and founder support; prospective customers apply by email.
Customer runs the gateway; Kottos supplies routing intelligence and receives routing/usage metadata.
Target Markets
- AI application developers and teams operating LLM inference at scale
- Enterprises requiring managed or customer-hosted inference infrastructure
- Latency-sensitive AI applications
- Teams using multiple model providers or open-weight model venues
- Organizations needing production observability, cost attribution and SSO/enterprise controls
- Design partners evaluating broad provider comparisons
- Production LLM inference routing and provider selection
- Reducing inference cost while preserving prompt-cache benefits
- Latency-sensitive agent loops
- Voice agents and high-concurrency streaming responses
- Trading agents and other workloads where request-path latency is critical
- Cost attribution, monitoring and observability by user, team and project