Armature
Armature provides MCP analytics and evals for product teams, capturing agent sessions from Claude, ChatGPT, and other AI clients to surface user intent, issues, and regressions.
At a Glance
About Armature
Armature is a product analytics and evaluation platform built specifically for MCP servers, Claude Connectors, and ChatGPT Apps — the agent-facing interfaces that live inside AI clients rather than traditional UIs. Founded by Theodore Otzenberger and Louis Scremin, who previously shipped secure infrastructure at Palantir, built observability at Tsuga, and deployed AI agents to six million users at Joko, Armature is backed by Y Combinator. The platform is actively available with a free tier and a custom enterprise plan.
What It Is
Armature addresses a fundamental visibility gap: when users delegate tasks to Claude, ChatGPT, Cursor, or other AI agents, the session happens inside the AI client — not in the product's own UI. Traditional analytics tools like PostHog, Amplitude, or Mixpanel track human clicks on a UI that agents never open. Armature captures exactly those agent sessions instead, rebuilding each one with user intent, agent reasoning, and every tool call, then scoring whether the user's goal was achieved.
The platform has two core modules that can be used independently or together:
- MCP Analytics — captures and reconstructs agent sessions, groups them into use cases by volume and success rate (including unsupported use cases), and surfaces issues like agent loops, missing auth scopes, and pagination truncation even when every API response returned 200 OK.
- MCP & CLI Evals — runs real agents end-to-end across major models and harnesses (Claude, ChatGPT, Codex, Claude Code, and more), scores each run against defined pass criteria, and catches regressions before shipping.
How the Workflow Fits Together
Armature's setup path is designed to be minimal: developers add the SDK to an existing MCP server, Claude Connector, or ChatGPT App backend with a few lines of code and one deploy. No changes to the server's behavior are required. Once deployed, sessions flow in automatically.
Analytics and evals are designed to close a loop: real usage data from MCP Analytics surfaces top use cases and issues, which Armature can draft directly into eval suites. Those evals then run against the MCP across models and harnesses to catch regressions before release. The platform also supports writing evals from scratch without analytics data.
Privacy and Data Handling
Because agent sessions can carry personal information and secrets, Armature runs detection models that scan and redact PII and secrets by default before anything reaches storage. Users control retention and can delete their data at any time.
Target Audience
Armature is positioned for product teams — not just the engineers who build agents — who need to understand how their users' agents experience what they ship. The platform explicitly distinguishes itself from LangSmith and Langfuse, which the company describes as tools for observing agents you build yourself, serving the engineers who build them. Armature's stated focus is on the product team's perspective: what users ask for, how agents deliver, and where the experience breaks.
Update: MCP & CLI Evals Launch
The most recent announced addition is MCP & CLI Evals, highlighted as "New" on the homepage and detailed in a dedicated blog post ("Armature launch: evals for MCPs and CLIs"). This extends the platform from pure analytics into end-to-end synthetic testing across the full matrix of models and agent harnesses, completing what Armature describes as a closed loop between real usage signal and pre-release quality assurance.
Community Discussions
Be the first to start a conversation about Armature
Share your experience with Armature, ask questions, or help others learn from your insights.
Pricing
Free
Session analytics and evals with 1,000 credits per month, unlimited projects and users, 7-day retention.
- 1,000 credits/month (1,000 sessions, 100 eval runs, or any mix)
- Unlimited projects (launch offer)
- Unlimited users
- 7-day retention
- Dashboard access
Extra Credits
Additional credits beyond the free tier at $50 per 1,000 credits.
- 1,000 additional credits (1,000 sessions, 100 eval runs, or any mix)
Custom
Custom plan for teams with serious traffic, including priority support, SSO/SAML, audit logs, custom retention, and dedicated onboarding.
- Everything in Free
- Priority support & SLA
- SSO / SAML
- Audit logs
- Custom retention
- Dedicated onboarding
Capabilities
Key Features
- MCP session analytics
- Agent session replay
- User intent capture
- Use-case grouping and ranking
- Issue identification and root-cause grouping
- Session scoring
- MCP & CLI evals
- End-to-end eval runs across major models and harnesses
- Eval generation from top use cases and issues
- PII and secret redaction by default
- SDK for MCP servers and Claude Connectors
- Dashboard access
- Unlimited projects (launch offer)
- Unlimited users
