EveryDev.ai
Subscribe
Home
Tools

3,484+ AI tools

  • New
  • Trending
  • Featured
  • Compare
  • Arena
Categories
  • Agents2189
  • Coding1574
  • Infrastructure698
  • Marketing534
  • Projects498
  • Research456
  • Design416
  • Analytics389
  • Testing296
  • MCP290
  • Security286
  • Data262
  • Integration197
  • Prompts189
  • Communication183
  • Extensions173
  • Learning170
  • Voice151
  • Commerce135
  • DevOps123
  • Web86
  • Finance26
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Tools
    3. Armature
    Armature icon

    Armature

    LLM Evaluations
    Featured

    Armature provides MCP analytics and evals for product teams, capturing agent sessions from Claude, ChatGPT, and other AI clients to surface user intent, issues, and regressions.

    Visit Website

    At a Glance

    Pricing
    Free tier available

    Session analytics and evals with 1,000 credits per month, unlimited projects and users, 7-day retention.

    Extra Credits: $50 usage-based
    Custom: Custom/contact

    Engagement

    Available On

    Web
    API
    SDK
    CLI

    Resources

    WebsiteDocsllms.txt

    Topics

    LLM EvaluationsMCP ToolsObservability Platforms

    Alternatives

    LunaryLangfuseHoneyHive
    Developer
    Armature, Inc.San Francisco, CAEst. 2026$500000 raised

    Listed Aug 2026

    About Armature

    Armature is a product analytics and evaluation platform built specifically for MCP servers, Claude Connectors, and ChatGPT Apps — the agent-facing interfaces that live inside AI clients rather than traditional UIs. Founded by Theodore Otzenberger and Louis Scremin, who previously shipped secure infrastructure at Palantir, built observability at Tsuga, and deployed AI agents to six million users at Joko, Armature is backed by Y Combinator. The platform is actively available with a free tier and a custom enterprise plan.

    What It Is

    Armature addresses a fundamental visibility gap: when users delegate tasks to Claude, ChatGPT, Cursor, or other AI agents, the session happens inside the AI client — not in the product's own UI. Traditional analytics tools like PostHog, Amplitude, or Mixpanel track human clicks on a UI that agents never open. Armature captures exactly those agent sessions instead, rebuilding each one with user intent, agent reasoning, and every tool call, then scoring whether the user's goal was achieved.

    The platform has two core modules that can be used independently or together:

    • MCP Analytics — captures and reconstructs agent sessions, groups them into use cases by volume and success rate (including unsupported use cases), and surfaces issues like agent loops, missing auth scopes, and pagination truncation even when every API response returned 200 OK.
    • MCP & CLI Evals — runs real agents end-to-end across major models and harnesses (Claude, ChatGPT, Codex, Claude Code, and more), scores each run against defined pass criteria, and catches regressions before shipping.

    How the Workflow Fits Together

    Armature's setup path is designed to be minimal: developers add the SDK to an existing MCP server, Claude Connector, or ChatGPT App backend with a few lines of code and one deploy. No changes to the server's behavior are required. Once deployed, sessions flow in automatically.

    Analytics and evals are designed to close a loop: real usage data from MCP Analytics surfaces top use cases and issues, which Armature can draft directly into eval suites. Those evals then run against the MCP across models and harnesses to catch regressions before release. The platform also supports writing evals from scratch without analytics data.

    Privacy and Data Handling

    Because agent sessions can carry personal information and secrets, Armature runs detection models that scan and redact PII and secrets by default before anything reaches storage. Users control retention and can delete their data at any time.

    Target Audience

    Armature is positioned for product teams — not just the engineers who build agents — who need to understand how their users' agents experience what they ship. The platform explicitly distinguishes itself from LangSmith and Langfuse, which the company describes as tools for observing agents you build yourself, serving the engineers who build them. Armature's stated focus is on the product team's perspective: what users ask for, how agents deliver, and where the experience breaks.

    Update: MCP & CLI Evals Launch

    The most recent announced addition is MCP & CLI Evals, highlighted as "New" on the homepage and detailed in a dedicated blog post ("Armature launch: evals for MCPs and CLIs"). This extends the platform from pure analytics into end-to-end synthetic testing across the full matrix of models and agent harnesses, completing what Armature describes as a closed loop between real usage signal and pre-release quality assurance.

    Armature - 1

    Community Discussions

    Be the first to start a conversation about Armature

    Share your experience with Armature, ask questions, or help others learn from your insights.

    Pricing

    FREE

    Free

    Session analytics and evals with 1,000 credits per month, unlimited projects and users, 7-day retention.

    • 1,000 credits/month (1,000 sessions, 100 eval runs, or any mix)
    • Unlimited projects (launch offer)
    • Unlimited users
    • 7-day retention
    • Dashboard access

    Extra Credits

    Additional credits beyond the free tier at $50 per 1,000 credits.

    $50
    usage based
    • 1,000 additional credits (1,000 sessions, 100 eval runs, or any mix)

    Custom

    Custom plan for teams with serious traffic, including priority support, SSO/SAML, audit logs, custom retention, and dedicated onboarding.

    Custom
    contact sales
    • Everything in Free
    • Priority support & SLA
    • SSO / SAML
    • Audit logs
    • Custom retention
    • Dedicated onboarding
    View official pricing

    Capabilities

    Key Features

    • MCP session analytics
    • Agent session replay
    • User intent capture
    • Use-case grouping and ranking
    • Issue identification and root-cause grouping
    • Session scoring
    • MCP & CLI evals
    • End-to-end eval runs across major models and harnesses
    • Eval generation from top use cases and issues
    • PII and secret redaction by default
    • SDK for MCP servers and Claude Connectors
    • Dashboard access
    • Unlimited projects (launch offer)
    • Unlimited users

    Integrations

    Claude
    ChatGPT
    Claude Code
    Cursor
    Codex
    Gemini CLI
    TypeScript SDK
    Python SDK
    Go SDK
    API Available
    View Docs

    Ratings & Reviews

    No ratings yet

    Be the first to rate Armature and help others make informed decisions.

    Developer

    Armature, Inc.

    Armature builds product analytics and evaluation tooling for agent-facing interfaces — MCP servers, Claude Connectors, and ChatGPT Apps. Co-founders Theodore Otzenberger and Louis Scremin bring experience from Palantir (secure infrastructure), Tsuga (bring-your-own-cloud observability), and Joko (AI agents for six million consumers). Backed by Y Combinator, Armature focuses on giving product teams visibility into how users' AI agents experience their products.

    Founded 2026
    San Francisco, CA
    $500000 raised
    3 employees
    Read more about Armature, Inc.
    WebsiteLinkedInX / Twitter
    1 tool in directory

    Similar Tools

    Lunary icon

    Lunary

    Open-source platform to monitor, improve, and secure AI chatbots with observability, prompt management, evaluations, and analytics.

    Langfuse icon

    Langfuse

    Open source LLM engineering platform for observability, prompt management, evaluation, and debugging of AI applications and agents.

    HoneyHive icon

    HoneyHive

    AI observability and evaluation platform to monitor, evaluate, and govern AI agents and applications across any model, framework, or agent runtime.

    Browse all tools

    Related Topics

    LLM Evaluations

    Platforms and frameworks for evaluating, testing, and benchmarking LLM systems and AI applications. These tools provide evaluators and evaluation models to score AI outputs, measure hallucinations, assess RAG quality, detect failures, and optimize model performance. Features include automated testing with LLM-as-a-judge metrics, component-level evaluation with tracing, regression testing in CI/CD pipelines, custom evaluator creation, dataset curation, and real-time monitoring of production systems. Teams use these solutions to validate prompt effectiveness, compare models side-by-side, ensure answer correctness and relevance, identify bias and toxicity, prevent PII leakage, and continuously improve AI product quality through experiments, benchmarks, and performance analytics.

    114 tools

    MCP Tools

    Tools built with the Model Context Protocol for specific tasks.

    73 tools

    Observability Platforms

    Comprehensive platforms that combine metrics, logs, and traces with AI-powered analytics to provide deep insights into complex distributed systems and application behavior.

    113 tools
    Browse all topics
    Back to all toolsSuggest an edit
    ratings
    discussions