# Mingbird

> A local-first agent harness that runs 2–9B Ollama models on ordinary laptops to complete real tasks with offline-first operation.

Mingbird is an open-source, local-first agent harness published by the Mingbird organization on GitHub under Apache-2.0. It is designed to run small 2–9B models through Ollama on laptops with integrated graphics. The latest release listed is v1.9.2, published 2026-10-02, and the README says the project is actively maintained.

The README argues that small-model failures in mainstream agent frameworks are harness defects rather than model defects, and it builds its mechanisms around that claim.

## What It Is

Mingbird is an agent harness: the layer that sits between a local model served by Ollama and the work the model is asked to do. According to the README, it handles tasks such as writing code and running tests, organizing files, web research, data analysis, and calling MCP tools. It ships as a Windows desktop client, with experimental Linux and macOS builds, a source-based GUI and CLI, and an experimental local web UI bound to 127.0.0.1.

## How the harness compensates for small models

The README describes each mechanism as the harness doing what a small model cannot. It lists these:

- The harness runs tests itself and feeds back exact failures with file and line.
- Every edit is auto-backed up with a .bak file, and rollback is one command.
- Anti-loop logic escalates from a nudge to a hard reset to a graceful exit.
- Tool prefill is loaded by category, and the README states the factory prefill is 797 tokens, with CI failing any change that grows it.
- A delivery self-check re-reads the original task before claiming completion.
- Chunked write feedback helps when large file output is truncated.

## Benchmark claims

The README reports a 288-cell benchmark (four harnesses, four open models, 18 tasks) with every cell published. It states that the same 2B model scores 0.821 in Mingbird versus 0.017 to 0.271 in the other harnesses, and that the overall score is 0.886 versus goose at 0.631, opencode at 0.479 and agent-mini at 0.405. The benchmark, LRAB, was designed by the authors themselves, and the README says so openly and publishes tasks, scoring code and raw data. It also reports results on Sierra Research's τ²-bench, run by the project authors.

## Privacy and offline operation

The README says there is no account system, telemetry, crash reporting, update checks or analytics. A toolbar toggle enables offline mode, in which web tools and URL-based MCP servers are not assembled into the prompt and the optional cloud provider is disabled. In online mode, the README says network access occurs only for web search and for MCP servers the user configures. An optional OpenAI-compatible cloud model can be configured in config.json.

## Safety net

The README describes a five-ring safety design: sandboxed sub-agents, refusal of irreversible operations such as format and diskpart, confirmation tiers for uninstalls and environment changes, deletes confined to the working directory, and rollback via .bak files and a .mingbird_trash folder. An AGENT_UNSAFE=1 environment variable disables the safeguards.

## Skills, voice and setup

The README lists 17 built-in skills (11 coding and 6 universal), user-added Markdown skills, MCP servers configured as JSON, drag-and-drop attachments, bilingual English and Chinese UI, and bundled local speech-to-text for Chinese and English voice input. Setup requires installing Ollama, pulling a model such as gemma4:e2b or qwen3.5:4b, then installing the Windows setup executable or running python agent_gui.py from source. Python 3.12 is recommended.

## Honest limits

The README notes that a 2B model will not rewrite an entire codebase in one shot, that 35B on integrated graphics runs end to end but slowly, and that Linux and macOS packages are experimental while Windows is the primary platform.

## Features
- Local-first agent harness for 2–9B Ollama models
- One-click offline mode with web tools removed from the prompt
- Harness-run tests with file:line failure feedback
- Automatic .bak backups and rollback
- Anti-loop escalation ladder
- Fixed 797-token factory prefill enforced by CI
- Delivery self-check gate
- Five-ring safety net
- 17 built-in skills plus custom Markdown skills
- MCP server support via JSON config
- Local voice input in Chinese and English
- Drag-and-drop and paste attachments
- Task time-box
- Searchable local session memory
- Bilingual English/Chinese UI
- Experimental local Web UI
- Optional OpenAI-compatible cloud model
- Public 288-cell benchmark data

## Integrations
Ollama, MCP, OpenAI-compatible endpoints

## Platforms
WINDOWS, MACOS, LINUX, WEB, API, CLI

## Pricing
Open Source

## Version
v1.9.2

## Links
- Website: https://github.com/Mingbird/Mingbird-agent
- Repository: https://github.com/Mingbird/Mingbird-agent
- EveryDev.ai: https://www.everydev.ai/tools/mingbird
