Artificial Analysis
Independent AI benchmarking platform that evaluates and compares AI models across intelligence, speed, cost, and capabilities to help users choose the best model and provider for their use case.
At a Glance
About Artificial Analysis
Artificial Analysis is an independent AI benchmarking and analysis company that publishes continuously updated evaluations of language models, coding agents, image, video, and speech models. The platform covers over 575 models and tracks performance across intelligence, speed, cost, and specialized capabilities, giving developers, researchers, and businesses a neutral reference point for AI model selection.
What It Is
Artificial Analysis operates as a third-party benchmarking service — not affiliated with any AI provider — that independently runs evaluations and publishes results publicly. The core product is a web-based dashboard where users can compare models side-by-side on the Artificial Analysis Intelligence Index (a composite of nine evaluations including GDPval-AA v2, τ³-Banking, Terminal-Bench, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR), as well as on output speed (tokens per second), cost per task, and latency. A personalized model recommender lets users weight their own priorities across intelligence, speed, and cost.
What Gets Benchmarked
The platform covers a wide range of AI modalities and use cases:
- Language models: Intelligence Index, agentic index, coding index, capability indices (agentic, coding, finance, legal, healthcare, engineering, economics, strategy)
- Coding agents: End-to-end software engineering tasks via DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA
- Image and video: Text-to-image, image editing, text-to-video, image-to-video, and video editing leaderboards powered by arena-style blind preference votes
- Speech: Text-to-speech arena Elo, speech-to-text (AA-WER Index), and speech-to-speech evaluations
- Hardware: GPU and inference hardware benchmarks
Evaluations are run independently; the FAQ states that providers cannot pay for results, methodology changes, or listing on the website.
Proprietary Benchmarks and Methodology
Artificial Analysis develops and maintains several proprietary evaluations that appear in its Intelligence Index and capability indices:
- GDPval-AA v2 — agentic real-world work tasks scored via Elo
- AA-Briefcase — a new frontier benchmark for long-horizon knowledge work requiring deliverables such as spreadsheets, presentations, and memos
- AA-Omniscience — knowledge accuracy and non-hallucination rate
- AA-LCR — long context reasoning
- τ³-Banking — agentic tool use in a banking context
- Terminal-Bench — agentic coding and terminal use
- AutomationBench-AA, Harvey LAB-AA, EnterpriseOps-Gym-AA, APEX-Agents-AA, ITBench-AA — domain-specific agentic evaluations
The site also incorporates established third-party benchmarks such as GPQA Diamond, SciCode, Humanity's Last Exam, MMMU-Pro, and IFBench.
Update: Intelligence Index v4.1
The changelog shows active and frequent product movement. As of July 2025, the Intelligence Index is at version 4.1, which incorporates updates to GDPval-AA V2, τ³-Banking, and Terminal-Bench v2.1. New evaluations AA-Briefcase, AutomationBench-AA, Harvey LAB-AA, EnterpriseOps-Gym-AA, and APEX-Agents-AA were recently added. The platform typically aims to benchmark leading models and providers within 24 hours of their release, according to the FAQ. The changelog shows new model evaluations published multiple times per week.
Independence and Audience
The platform positions itself as a neutral reference for AI decision-making. The pricing page states that Artificial Analysis describes itself as "the leading independent AI benchmarking company." The site's FAQ explicitly states that providers cannot pay for results or methodology changes. The platform serves developers choosing APIs, researchers tracking model progress, and enterprise teams making AI strategy decisions. A premium API and data export tier is available for organizations that need programmatic access to the underlying benchmark data, custom visualizations, and industry reports. An Enterprise tier offers custom benchmarking, compute market forecasts, workshops, and AI strategy advisory for larger organizations.
Community Discussions
Be the first to start a conversation about Artificial Analysis
Share your experience with Artificial Analysis, ask questions, or help others learn from your insights.
Pricing
Free
Public access to benchmarks, leaderboards, and model comparisons on the website.
- Access to Intelligence Index leaderboards
- Speed and cost comparisons
- Image, video, and speech leaderboards
- Coding agent index
- Personalized model recommender
Pro
Data and insights for individuals and small organizations, including API access and data export.
- Build with our API
- Export data for local analysis
- Build custom charts and tables
- Access industry reports and guides
- Email support
- Language model data
- Media model data (image, video, speech, music)
- Language model provider data
- State of AI Report
- AI Adoption Survey
- Model Deployment Report
- Leaders AI Strategy Guide
Enterprise
Tailored solutions for large organizations with custom benchmarking, advisory, and dedicated support.
- Everything in Pro
- Benchmark models, inference, and hardware for your use cases
- Highest API rate limits
- Compute market model access with forecasts
- Workshops and education
- AI strategy advisory
- Personalized support
- For organizations with 150+ employees
Capabilities
Key Features
- Artificial Analysis Intelligence Index (composite of 9 evaluations)
- Speed benchmarks (output tokens per second)
- Cost per task analysis
- Coding agent leaderboard (DeepSWE, Terminal-Bench, SWE-Atlas-QnA)
- Image and video arena leaderboards
- Text-to-speech arena Elo
- Speech-to-text (AA-WER Index)
- Hardware benchmarks
- Personalized model recommender
- Capability indices (agentic, coding, finance, legal, healthcare, engineering, economics)
- Proprietary benchmarks: GDPval-AA, AA-Briefcase, AA-Omniscience, AA-LCR, τ³-Banking
- AI Trends tracking
- Frontier model intelligence over time charts
- Data API and export
- Custom chart and table builder
- Industry reports and guides
- Custom benchmarking services (Enterprise)
