# gutcheck

> A local CLI tool that applies plain-English semantic questions to lines, logs, diffs, and JSON records — like grep, but for meaning — running entirely on your CPU with no API key required.

gutcheck is an open-source command-line tool written in Rust that brings semantic filtering to text files, logs, diffs, and JSON records. It lets you ask plain-English questions about every line or paragraph and explore the scored answers in a live terminal view — all without sending data to any external service. The project is MIT-licensed and published by sfmqrb on GitHub, with its first release (v0.5.0) appearing in September 2026.

## What It Is

gutcheck is best described as "grep for meaning." Where grep matches patterns, gutcheck matches intent: you write a natural-language question like "is this a bug report?" or "does this log line describe an error?" and the tool scores every line against that question using a local ONNX model running on your CPU. The result is a ranked, filterable list of matches — no cloud, no API key, no data leaving your machine. The first run downloads a 1.3 GB model to a local cache directory.

## How the Workflow Feels

The tool is designed to carry over grep muscle memory while adding semantic depth:

- **Grep-style flags** — `-n`, `-H`, `-l`, `-q`, `-m`, `-v`, `-r`, `-C` all work as expected; exit codes follow grep conventions (0 = match, 1 = none, 2 = error).
- **Interactive mode (`-i`)** — scores appear as they arrive, best first; you can slide the threshold with arrow keys, press `/` to re-ask, and press `enter` to print matches.
- **`--auto`** — automatically finds the score cut-off from the distribution, removing the need to guess a threshold.
- **Named questions** — built-in presets like `@secret`, `@pii`, `@security`, `@error`, `@bug`, and `@spam`, plus user-defined questions stored in `~/.config/gutcheck/questions`.
- **Structured input** — `--field` targets a specific JSON key; `--csv` handles CSV columns; `--diff` judges each git diff hunk independently.
- **Streaming** — `tail -f app.log | gutcheck -v "is this routine noise?"` works on live streams.

## Built for Log Analysis

The README highlights a fuzzy-deduplication strategy that makes gutcheck practical on large log files. Identical lines are answered once; with `-f` (fuzzy mode), lines that differ only in numbers, timestamps, or IDs share a single model call. Benchmarks on Loghub samples show that 2,000-line Apache logs contain only 12 distinct shapes, reducing 500-line runs from 57 seconds (exact) to 4 seconds (fuzzy) while maintaining 99.8% decision agreement. HDFS and Linux log samples show similar compression ratios.

## Three Local Models

gutcheck ships with three selectable models, all running via ONNX on CPU:

- **multilingual** — supports 100+ languages, fastest inference
- **english** — balanced accuracy and speed for English text
- **typed** — highest accuracy on English; used automatically for named questions

The models are Apache-2.0 licensed and downloaded at runtime. A `--estimate` flag previews the number of model calls before committing to a run.

## Update: v0.5.0

The initial public release, v0.5.0, was published on 2026-09-26. The repository was created the same day, making this a brand-new project. The README is already comprehensive, covering a command reference, model accuracy docs, benchmarks, tips, and an explanation of how the tool works internally. The project is independent and explicitly states it is not affiliated with the Laya or Jev model authors.

## Features
- Plain-English semantic filtering of lines, logs, diffs, and JSON records
- Fully local inference — no API key, no data leaves the machine
- Interactive live terminal view with adjustable threshold and re-ask
- grep-compatible flags and exit codes
- Named question presets (@secret, @pii, @security, @error, @bug, @spam)
- Fuzzy deduplication for efficient log analysis
- Three selectable ONNX models (multilingual, english, typed)
- JSON field and CSV column targeting
- Git diff hunk evaluation with --diff
- Streaming input support (tail -f)
- -auto threshold detection
- -top N ranking with heat bars
- -tally for labelled histograms
- Run receipt showing record count, model calls, cache hits, and elapsed time
- User-defined named questions via config file
- -estimate flag to preview model call count before running

## Integrations
Git (diff hunk analysis), JSON (--field flag), CSV (--csv flag), Homebrew (brew install), Cargo (cargo install), ONNX Runtime (local model inference)

## Platforms
MACOS, LINUX, API, CLI

## Pricing
Open Source

## Version
v0.5.0

## Links
- Website: https://github.com/sfmqrb/gutcheck
- Documentation: https://github.com/sfmqrb/gutcheck/blob/main/docs/cli.md
- Repository: https://github.com/sfmqrb/gutcheck
- EveryDev.ai: https://www.everydev.ai/tools/gutcheck
