# CodeVetter

> Local-first, execution-backed verification of coding-agent changes that produces pass, fail, or unverified verdicts with retained evidence.

CodeVetter is an open-source, local-first verification system for coding-agent changes, built by Sarthak Agrawal. It binds a requested task to the exact agent change, runs repository-owned checks, and records a pass, fail, or unverified verdict along with the evidence and limitations behind it. The current public build is a macOS Apple-silicon desktop app that bundles a CLI and a local MCP sidecar.

## What It Is

CodeVetter is an execution-backed verification and evaluation tool aimed at engineers and teams who supervise or compare coding-agent output. Instead of treating a second model's opinion as the result, it treats task-relevant execution evidence, such as tests, type checks, builds, and browser or API checks, as the verdict boundary. Model-backed review is optional and can suggest risks and checks, but it does not decide the verdict.

## How the Verification Loop Works

A run follows a chain: task, change, execute, evidence, verdict. The requested behavior and exact revision range are recorded, repository-owned checks are executed locally, and the result is exported with commands, bounded output, artifacts, provenance, and explicit gaps. Failures are classified separately from existing failures, environment problems, timeouts, and missing coverage. Missing or irrelevant evidence yields an unverified verdict rather than a confidence score. The bundled CLI is invoked with a revision range, a task description, and a JSON output option.

## Local-First Model and Providers

Product state is stored locally in SQLite, and no CodeVetter account or hosted verification backend is required. If you start a provider-backed review, selected prompt and code context go directly to the provider you configured: Anthropic, OpenAI, or OpenRouter. Local checks and the desktop viewer can work offline, while provider-backed review needs network access.

## Current Status

The latest listed release is v1.15.0, published as an Apple-silicon macOS DMG with an updater archive; other platform installers are not currently published. The project also publishes a benchmark of 27 synthetic cases and 29 labeled expected findings, which the site describes as measuring a narrow recognition task rather than production-wide accuracy.

## Features
- Binds requested task to the exact agent change
- Runs repository-owned checks (tests, builds, browser and API checks)
- Pass, fail, or unverified verdicts
- Portable machine-readable evidence bundles with explicit limitations
- Failure classification separating agent regressions from existing failures and environment issues
- Bundled CLI and local MCP sidecar
- Native macOS desktop review workbench
- Optional provider-backed code review
- Local SQLite state with no hosted verifier
- Public benchmark with scorer and documented limitations

## Integrations
Anthropic, OpenAI, OpenRouter, MCP, Claude Code, Codex

## Platforms
MACOS, WEB, API, VSC_EXTENSION, CLI

## Pricing
Open Source

## Version
v1.15.0

## Links
- Website: https://codevetter.com
- Documentation: https://codevetter.com/docs
- Repository: https://github.com/Codevetter/codevetter
- EveryDev.ai: https://www.everydev.ai/tools/codevetter
