Baseten
Inference platform for serving open-source, custom, and fine-tuned AI models with dedicated deployments, Model APIs, and training.
At a Glance
Deploy custom, fine-tuned, and open-source models; pay as you go with $0 per month base.
Engagement
Available On
Alternatives
Listed Oct 2026
About Baseten
Baseten is an AI inference platform for running open-source, custom, and fine-tuned models in production. It offers dedicated deployments, pre-optimized Model APIs, and a training product, all built on what it calls the Baseten Inference Stack. The company says it was founded in 2019 by engineers.
What It Is
Baseten provides infrastructure for serving AI models at scale. Teams deploy a custom or proprietary model, or use hosted open models through Model APIs, and the platform handles runtime optimization, autoscaling, and deployment management. Supported workloads named on the site include LLMs, transcription, text-to-speech, image generation (including ComfyUI workflows), embeddings, and compound AI via Baseten Chains.
Deployment Options
Workloads can run on Baseten Cloud, a multi-cloud, multi-region network with optional single-tenant clusters, or self-hosted inside a customer's own VPC. A hybrid option adds on-demand flex capacity on Baseten Cloud, and the self-hosted setup can fail over to Baseten Cloud. Baseten also offers forward deployed engineers who work with customers to tune deployments toward latency, throughput, and cost targets.
Performance and Runtimes
Baseten says its runtimes include custom kernels, speculative decoding techniques, and advanced caching. It claims 99.99% uptime and fast cold starts, and says Baseten Embeddings Inference delivers higher throughput and lower latency than other solutions. These are vendor claims.
Security and Enterprise Controls
The enterprise page states Baseten does not store inference inputs or outputs, is SOC 2 Type II certified and HIPAA compliant, and supports region-restricted deployments for data residency. It lists RBAC, team-level access controls, SSO, logging, custom metrics, and tracing.
Training and Model Labs
The Training product uses the Loops SDK for reinforcement-learning-style training, with deployment to production inference on the same stack. Baseten for Model Labs lets model creators distribute and monetize models on Baseten infrastructure.
Community Discussions
Be the first to start a conversation about Baseten
Share your experience with Baseten, ask questions, or help others learn from your insights.
Pricing
Basic
Deploy custom, fine-tuned, and open-source models; pay as you go with $0 per month base.
- Dedicated deployments
- Model APIs
- Training
- Fast cold starts
- SOC 2 Type II and HIPAA compliant
Pro
Unlimited autoscaling and priority compute access; volume discounts available.
- Everything in Basic
- Priority access to high-demand GPUs
- Dedicated compute
- Higher Model API rate limits
- Hands-on engineering expertise
- Dedicated support on Slack and Zoom
- Volume discounts available
Enterprise
Full control in your cloud and ours; custom pricing via quote.
- Everything in Pro
- Custom SLAs
- Self-host deployments
- On-demand flex compute
- Use existing cloud commitments
- Full control over data residency
- Advanced security and compliance
- Custom global regions
- Advanced RBAC with Teams
- Deployment options: Baseten, Your VPC, Hybrid
GLM-5.3 Fast
Model API metered pricing.
- Input: $2.10 per 1M tokens
- Cache Input: $0.21 per 1M tokens
- Output: $6.60 per 1M tokens
GLM-5.3
Model API metered pricing.
- Input: $1.40 per 1M tokens
- Cache Input: $0.14 per 1M tokens
- Output: $4.40 per 1M tokens
Kimi K3
Model API metered pricing.
- Input: $3.00 per 1M tokens
- Cache Input: $0.30 per 1M tokens
- Output: $15.00 per 1M tokens
DeepSeek V4 Pro 0813
Model API metered pricing.
- Input: $1.32 per 1M tokens
- Cache Input: $0.132 per 1M tokens
- Output: $3.96 per 1M tokens
GPT OSS 120B
Model API metered pricing.
- Input: $0.10 per 1M tokens
- Output: $0.50 per 1M tokens
Capabilities
Key Features
- Dedicated inference deployments for custom and fine-tuned models
- Pre-optimized Model APIs for open models
- Model library
- Training with the Loops SDK
- Baseten Chains for compound AI
- Baseten Embeddings Inference
- Multi-cloud capacity management and autoscaling
- Self-hosted, cloud, and hybrid deployment
- Forward deployed engineers
- RBAC, SSO, logging, and tracing
- SOC 2 Type II and HIPAA compliance
