Baseten
Baseten builds an inference cloud and infrastructure stack for bringing AI products to production. It provides fast model runtimes, multi-cloud GPU capacity, deployment tooling, training, observability, and support for mission-critical inference.
At a Glance
- AI-native startups
- Enterprise software companies
- Model labs and model creators
- Healthcare and clinical AI
- +4 more
AI Tools by Baseten
(1)Baseten
AI Model Inference Platform
Discussions
No discussions yet
Be the first to start a discussion about Baseten
Latest News
Agentic inference optimization: 50–90% faster engines
Baseten partners with OpenAI to serve open models through Codex and the Responses API
Baseten joins NVIDIA OpenShell/Open Agent Safety Platform effort and supports Blaxel Carbon sandboxes
Sheila Vashee joins Baseten as Chief Marketing Officer
Products & Services
Deploy open-source, custom, and fine-tuned models on dedicated GPUs with configurable hardware, autoscaling, environments, release lifecycle, multi-cloud capacity, and performance engineering.
Hosted, pre-optimized open-model APIs with OpenAI-compatible and Anthropic Messages-compatible interfaces, designed for rapid prototyping and production workloads and billed by token.
Train or fine-tune models using Loops or Training Jobs on dedicated GPU clusters, then deploy resulting checkpoints on the same inference stack.
Training and post-training tooling for supervised fine-tuning and reinforcement learning, including workflows from LangSmith traces.
Market Position
Baseten positions itself as an inference-first, production-grade alternative to generic cloud GPU infrastructure and API-only model hosts: it combines optimized runtimes and performance research with dedicated, multi-cloud, hybrid, and enterprise deployment control. Commonly cited alternatives include Modal, Replicate, RunPod, Together AI, Fireworks AI, Anyscale, and AWS SageMaker; Baseten differentiates on high-performance production inference, customization, and hands-on forward-deployed engineering.
Leadership
Founders
Tuhin Srivastava
CEO and co-founder; previously worked on machine learning and co-founded Sutro Health.
Amir Haghighat
CTO and co-founder; previously Head of Engineering at Gumroad and worked in engineering at Clover Health.
Philip (Phil) Howes
Co-founder and Chief Scientist; holds a PhD in mathematics from the University of Sydney and previously co-founded Sutro Health.
Pankaj Gupta
Co-founder and model-performance leader; previously a software engineer at Uber.
Executive Team
Tuhin Srivastava
CEO and Co-Founder
Co-founded Baseten in 2019 and previously worked in machine learning and co-founded Sutro Health.
Amir Haghighat
CTO and Co-Founder
Previously Head of Engineering at Gumroad and an engineer at Clover Health.
Board of Directors
Founding Story
The founders started Baseten after repeatedly seeing strong ML models get stuck in deployment hell: productionization took weeks, infrastructure was fragile, and training, serving, scaling, and hardware orchestration were disconnected. They set out to build the integrated platform they wanted to use themselves so builders could bring AI into products and operate it reliably at scale.
Business Model
Revenue Model
Usage-based infrastructure and inference: Model APIs are billed per million input and output tokens; dedicated deployments and training are billed for GPU/CPU time, generally per minute. Enterprise customers can buy dedicated support and cloud, self-hosted, or hybrid deployments.
Pricing Tiers
Prices vary by hosted model; the public catalog lists model-specific token rates.
Price varies by selected CPU/GPU hardware and deployment configuration; scale-to-zero can avoid idle GPU charges.
Price varies by hardware and training job configuration.
Dedicated support on Slack and Zoom plus enterprise hosting, networking, security, and capacity options.
Target Markets
- AI-native startups
- Enterprise software companies
- Model labs and model creators
- Healthcare and clinical AI
- Legal AI
- Developer tools and coding agents
- High-scale, low-latency LLM inference
- Agentic coding and compound AI applications
- Speech transcription, diarization, and text-to-speech
- Image and video generation
- Embeddings, reranking, and classification
- Fine-tuning and reinforcement-learning post-training
- Abridge
- Cursor
- Lovable
- Notion