# Vast.ai

> A GPU cloud marketplace that lets developers and AI teams rent on-demand, interruptible, or reserved GPU compute across 20,000+ GPUs and 40+ data centers at market-driven prices.

Vast.ai is a GPU cloud marketplace founded in 2016 by Jake Cannell and Christian Horne, built on the thesis that underutilized GPU hardware worldwide could be pooled into a competitive, low-cost alternative to hyperscaler pricing. The platform is SOC 2 certified, operates across 40+ data centers, and offers over 68 GPU types with per-second billing and no long-term contracts required.

## What It Is

Vast.ai operates as a two-sided marketplace and infrastructure platform for GPU compute. On one side, GPU owners — from independent hosts with gaming rigs to professional data centers — list their hardware. On the other, developers, researchers, and enterprises search, filter, and deploy instances in seconds via a web console, CLI, Python SDK, or REST API. Prices are set by supply and demand rather than fixed by Vast, making the platform structurally competitive with major cloud providers. The company positions itself as "agent-ready AI infrastructure," meaning its API-native provisioning model is designed for AI agents to autonomously procure and optimize compute without human intervention.

## Three Deployment Modes

Vast.ai offers three distinct ways to run GPU workloads:

- **GPU Cloud**: On-demand instances across 40+ data centers and 20,000+ GPUs. Deployable in seconds via CLI, SDK, or API. Best for production workloads requiring guaranteed uptime.
- **Serverless**: Deploy models as endpoints with automatic benchmarking and optimization across GPU types. Autoscales to zero; users pay only for compute time consumed.
- **Clusters**: Dedicated multi-node GPU clusters with InfiniBand networking, designed for large-scale distributed training jobs.

## Developer Tooling and API-Native Design

The platform is built for programmatic access from the ground up. Developers can interact through three interfaces:

- **CLI**: Install with a single curl command; search, deploy, and manage instances from the terminal with no Python required.
- **Python SDK**: `pip install vastai` provides programmatic compute provisioning in a few lines of code.
- **REST API**: Full HTTP access to every platform operation, with an OpenAPI spec available.

This API-native architecture is central to Vast.ai's positioning for agentic workloads, where AI agents call the provisioning API directly to spin up, scale, and tear down compute without human involvement.

## Supported Workloads and Use Cases

Vast.ai supports a broad range of GPU workloads across its platform:

- AI/ML framework execution (PyTorch, TensorFlow, JAX)
- LLM inference and text generation using open-source models via vLLM, TGI, and similar frameworks
- AI image and video generation with Stable Diffusion, FLUX, and diffusion models
- AI agent deployment and scaling
- Batch data processing
- Audio-to-text transcription
- AI fine-tuning
- GPU programming and HPC
- 3D graphics rendering
- Virtual computing / GPU-enabled VMs

A Model Library provides pre-configured templates for popular open-source models including Qwen, MiniMax, and others, enabling deployment without manual setup.

## Enterprise and Compliance Features

For enterprise customers, Vast.ai offers dedicated infrastructure with single-tenant isolation, data sovereignty controls, optional private VPNs, persistent audit logging, and custom security configurations to support HIPAA, GDPR, and regional compliance requirements. The platform achieved SOC 2 Type I certification in 2025. Enterprise tiers include volume discounts, reserved GPU contracts, white-glove onboarding, and SLA-backed support with priority escalation. The company's enterprise page cites case studies including organizations that scaled to 200K monthly active users and achieved significant infrastructure cost reductions, though these are vendor-published claims.

## Platform Scale and Current Status

According to Vast.ai's own published figures, the platform processes 700,000+ transactions per month, hosts 20,000+ GPUs across 40+ data centers, and supports 68+ GPU types spanning architectures from Pascal through NVIDIA's latest Blackwell generation (including H100, H200, B200, B300, and RTX 5090). The company reported 310% growth in 2024. Vast.ai opened its Los Angeles headquarters in July 2024 and a San Francisco engineering office in 2025, growing to 40+ employees across both locations. The platform's GPU pricing is real-time and updates hourly, with instance types including on-demand, interruptible (preemptible), and reserved options.

## Features
- On-demand GPU instances across 20,000+ GPUs
- 68+ GPU types from Pascal to Blackwell
- Per-second billing with no minimum hours
- CLI, Python SDK, and REST API access
- Serverless inference endpoints with autoscaling to zero
- Multi-node GPU clusters with InfiniBand networking
- Real-time market-driven pricing
- Model Library with pre-configured open-source model templates
- Interruptible (preemptible) instances for batch workloads
- Reserved instances with volume discounts
- SOC 2 Type I certified
- Single-tenant isolation for enterprise
- Private networking and VPN support
- Docker-based instance management
- GPU search and filtering by model, VRAM, price, and availability
- Earnings calculator for GPU hosts
- Startup program
- Enterprise white-glove support and SLA-backed response

## Integrations
vLLM, Text Generation Inference (TGI), Stable Diffusion, FLUX, PyTorch, TensorFlow, JAX, Jupyter, Docker, SSH, OpenAPI

## Platforms
WEB, API, CLI, DEVELOPER_SDK

## Pricing
Freemium — Free tier available with paid upgrades

## Links
- Website: https://vast.ai
- Documentation: https://docs.vast.ai
- Repository: https://github.com/vast-ai
- EveryDev.ai: https://www.everydev.ai/tools/vast-ai
