Oqoqo
Oqoqo is a cloud-based platform for building evals and custom benchmarks that test how well AI agents can use real-world products, tools, and interfaces at scale.
At a Glance
100 plan runs per month. Unused plan runs do not carry over.
Engagement
Available On
Listed Aug 2026
About Oqoqo
Oqoqo is a managed cloud platform that lets teams build private benchmarks and run evaluation experiments to measure how well AI agents perform on real-world tasks. It supports testing across skills, MCP servers, CLIs, SDKs, APIs, and documentation, with each trial running in its own isolated sandbox environment. The platform is accessible via a web app, CLI, and MCP server, making it usable both by human developers and by agents themselves.
What It Is
Oqoqo sits in the agent evaluation and benchmarking category. Its core job is to answer the question: "Can any agent actually use this product?" Teams define task sets and rubrics in plain language, configure which agents and treatments to test, and launch experiments that run in parallel on fully managed cloud infrastructure. The platform captures full trajectories — every tool call, command, file diff, and error — and reports pass rates, token usage, friction counts, and lift between treatments.
How the Experiment Loop Works
The workflow follows a structured input-sandbox-output-loop model:
- Input: Define tasks, rubrics, agents, treatments, files, and instructions held constant across the grid.
- Sandbox: Each run gets its own isolated, reproducible environment with project state, credentials, and the tools the agent needs. Setup, execution, and teardown are timed separately.
- Output: Pass/fail on each requirement with a written reason, the full trajectory, token counts, friction counts, and experiment-level pass rates.
- Loop: Read the trajectory, fix the interface or prompt, and relaunch — with results staying comparable across iterations because tasks are versioned.
Agent and Model Support
Oqoqo supports Claude Code, Codex, Cursor, GitHub Copilot, OpenCode, OpenClaw, Pi, and Hermes, with more agents described as on the way. Where supported, users can choose model and effort level per agent. The platform uses a bring-your-own-keys model: Oqoqo manages the infrastructure and orchestration, while users supply their own model API keys and subscriptions. Model inference costs are separate from platform run costs.
What Can Be Evaluated
The platform is designed to evaluate any interface an agent might use:
- Skills and MCP servers — test whether an agent can correctly invoke tools exposed via MCP
- CLIs and SDKs — run agents against command-line interfaces and software development kits
- APIs and docs — measure whether agents can navigate API references and documentation to complete tasks
- Workflows — compare raw agent performance against augmented treatments (e.g., adding an MCP server) to measure lift
The comparison feature lets teams cross tasks, treatments, agents, and repeated trials, changing one variable at a time to isolate what actually improves pass rates.
CI Integration and Automation
Oqoqo supports triggering experiments from CI pipelines, so teams can run agent evals automatically when a change might break agent workflows. Agents themselves can launch and inspect experiments through the CLI or MCP interface, enabling autonomous evaluation loops where Codex, Claude Code, Cursor, and similar agents can self-test.
Current Status
Oqoqo is actively available with a live web app at app.oqoqo.ai, a public pricing page, and documentation at docs.oqoqo.ai. The platform offers a free plan with monthly runs and paid subscription tiers for higher volume, plus one-time top-up run purchases that never expire. A demo booking option is available via the founders' calendar link.
Community Discussions
Be the first to start a conversation about Oqoqo
Share your experience with Oqoqo, ask questions, or help others learn from your insights.
Pricing
Free
100 plan runs per month. Unused plan runs do not carry over.
- 100 runs each month
- Unlimited team members and projects
- Unlimited experiments, tasks, treatments, and assets
- Full web app, CLI, and MCP access
Pro
300 plan runs per month. Unused plan runs do not carry over.
- 300 runs each month
- Unlimited team members and projects
- Unlimited experiments, tasks, treatments, and assets
- Full web app, CLI, and MCP access
Ultra
1,000 plan runs per month. Unused plan runs do not carry over.
- 1,000 runs each month
- Unlimited team members and projects
- Unlimited experiments, tasks, treatments, and assets
- Full web app, CLI, and MCP access
Custom
Higher volume plan with greater discounts. Contact for pricing.
- Custom run volume
- Greater discounts on higher volumes
- All features included
Capabilities
Key Features
- Build private benchmarks with custom task sets and rubrics
- Run eval experiments at scale on managed cloud infrastructure
- Isolated sandboxes per run with project state, credentials, and tools
- Full trajectory capture: tool calls, commands, file diffs, errors
- Compare agents, models, and treatments on the same tasks
- Pass/fail reporting with written reasons and friction analysis
- Token usage and efficiency metrics per run
- CI integration to trigger experiments on code changes
- Web app, CLI, and MCP access
- Bring your own model API keys and subscriptions
- Versioned tasks for comparable results across iterations
- Support for MCP servers, CLIs, SDKs, APIs, and docs evaluation
- Unlimited team members and projects on all plans
- Top-up runs that never expire
