AQuA (Ambient Quality Agent)
An open-source ADK recipe that runs beside a production agent in Google Cloud, clustering and diagnosing failures from production trajectories.
At a Glance
AQuA (Ambient Quality Agent) is released in the open in the google/adk-recipes repository under the Apache License 2.0. Running it in your own Google Cloud project incurs separate Google Cloud and Gemini model costs.
Engagement
Available On
Alternatives
Listed Oct 2026
About AQuA (Ambient Quality Agent)
AQuA (Ambient Quality Agent) is a reference implementation from Google, published in the adk-recipes repository, that runs unattended beside a production agent in your Google Cloud project. It sweeps production trajectories from Cloud Trace, Cloud Logging, or BigQuery, then clusters failures and diagnoses their root causes. Google describes it as released in the open as composable building blocks to run, adapt, and help shape.
What It Is
AQuA is an outer-loop agent quality tool for agents built with the Agent Development Kit (ADK). It works on live production traffic, not pre-launch test cases, and it never sits in the request path or writes back to the observed agent. Transcripts, source snapshots, and BigQuery tables stay inside the user's project. The repository README describes the recipes in adk-recipes as demonstration starting points, not production use, and as not an officially supported Google product.
How a Sweep Works
Each run follows a five-stage pipeline:
- Sample: pulls a random sample of up to 1,000 recent sessions.
- Review: grades each session against a nine-point checklist, with an optional plain-English
goal.mdsteering the review, and writes structured actual/expected findings. - Cluster: groups findings that share a failure mechanism.
- Verify: a separate model checks each cluster against up to three full transcripts and discards unsupported ones.
- Track: matches surviving clusters against open insights in BigQuery as NEW, RECURRING, or auto-RESOLVED after 14 days unseen.
By default a single-pass session_review judge is used, and Gemini platform trajectory AutoRaters can be enabled optionally. Deterministic Python custom metrics in eval_config.yaml can run alongside the judge.
Root-Cause Diagnosis
From the dashboard chat or agents-cli aqua run, a diagnosis agent reads failing trajectories against the immutable source snapshot captured at deploy time. When the defect is in the repository, it cites file and line ranges and proposes an anchored edit. When the fault is upstream, it attributes the failure to a trajectory step without proposing a diff. It never applies edits or opens pull requests itself. Insights can be fetched with agents-cli aqua get-insight so a coding agent can apply a fix and verify it by replaying sessions.
Trust and Limits
The blog states that insights are backed by citations to session IDs and validated line ranges rather than confidence scores, and that skipped or failed work is recorded on the run. Users can permanently dismiss by-design findings. Google notes deliberate trade-offs, including random session sampling, capped verification at 50 clusters per run, and static replay that does not reconstruct external environment state. It also lists directions it is exploring, such as attaching across a fleet of agents.
Community Discussions
Be the first to start a conversation about AQuA (Ambient Quality Agent)
Share your experience with AQuA (Ambient Quality Agent), ask questions, or help others learn from your insights.
Pricing
Open Source (Apache 2.0)
AQuA (Ambient Quality Agent) is released in the open in the google/adk-recipes repository under the Apache License 2.0. Running it in your own Google Cloud project incurs separate Google Cloud and Gemini model costs.
- Apache License 2.0 (declared for the adk-recipes repository)
- Runs beside your agent in your own Google Cloud project
- Local dashboard exploration with no cloud project, no credentials, and no model calls
- Sweeps up to 1,000 sessions per run
- External costs: Gemini model usage at standard Gemini platform pricing and Google Cloud resources; example sweep cost $0.70 for 96 sessions and $3.76 for 32 multi-agent sessions
Capabilities
Key Features
- Scheduled, post-deployment, or on-demand sweeps of production sessions
- Random sampling of up to 1,000 sessions per run
- Nine-point checklist session review with actual/expected findings
- Failure clustering and transcript-based verification
- Insight tracking in BigQuery as NEW, RECURRING, or RESOLVED
- Developer goal (goal.md) and custom Python metrics
- Optional managed trajectory AutoRaters
- Root-cause diagnosis anchored to deploy-time source snapshots
- Dashboard on Cloud Run behind Identity-Aware Proxy
- agents-cli aqua commands and agents-cli-aqua skill for coding agents
- Local dashboard demo with synthetic data
