ai-rete-rag
ai·rete·rag pairs a Rete rule engine with retrieval-augmented generation to deliver deterministic, auditable decisions explained in plain language from your own documents.
At a Glance
Enough to build against and see the whole product, including the audit trail.
Engagement
Available On
Listed Oct 2026
About ai-rete-rag
ai·rete·rag is a hosted decision-intelligence platform that combines a Rete rule engine with retrieval-augmented generation (RAG). Rules handle the deterministic "what" — producing auditable, repeatable verdicts — while retrieval grounds the "why" in your own policy documents, contracts, and guidelines. The platform targets regulated industries such as financial services, healthcare, legal, insurance, and e-commerce where explainability and auditability are non-negotiable.
What It Is
ai·rete·rag is middleware that sits between your data and your users. It accepts structured facts via a REST API, evaluates them against YAML-authored domain rules using a pure-Python Rete network, retrieves the most relevant document chunks from a ChromaDB vector store, and synthesises a coherent verdict plus explanation using Claude. The result is a single API call that returns a decision, a confidence score, and a full audit trail — without requiring teams to write prompts or manage LLM logic directly.
Three-Layer Architecture
The platform is built around three independently configurable layers:
- Rete Rule Engine — A pure-Python Rete network with alpha nodes (fact filtering), beta nodes (fact joining), and terminal nodes (action firing). Rules are authored in YAML, support salience-based conflict resolution, and can be hot-reloaded without restarting the server. Every firing is recorded for replay.
- Retrieval-Augmented Generation — ChromaDB stores chunked embeddings (using the
all-MiniLM-L6-v2sentence-transformer model) of uploaded policy documents. At decision time, the top-k most relevant chunks are retrieved and passed to the language model, with relevance scores returned alongside each chunk. - Orchestrator — Inspects each request and selects the optimal mode automatically: rules-only for speed, RAG-only when no rules exist, or hybrid when both are available. Supports three response modes:
verdict,verdict_with_explanation, andfull_audit.
Three Wiring Patterns
The platform supports three composable patterns for combining rules and retrieval, all running live:
- Rules → Retrieval — Rules narrow retrieval scope before documents are fetched, so a cardiac case only pulls cardiology sources, reducing hallucination risk.
- Retrieval → Rules — Documents are parsed into facts (entities, dates, obligations) and asserted into the working memory session; rules then fire on what was read.
- Decision → Narrative — The engine fires first and produces a decision trace; retrieval is invoked only to generate a human-readable explanation grounded in source documents.
Audit Trail and Conflict Detection
Every decision links back to the exact rules that fired and explains why every other rule did not — down to the specific value that missed a threshold. A built-in conflict detection feature runs static analysis to flag when two rules with different verdicts could both match the same case, surfacing the issue before it reaches production. The platform also includes a browsable policy rule catalog per domain showing conditions, salience, and verdicts.
Target Domains and Setup Path
The platform ships with eight built-in demo domains — loan underwriting, fraud screening, clinical vitals, blockchain/AML, insurance, legal/compliance, operations, and e-commerce — each runnable live in the workspace without signup. Bringing a custom domain requires uploading documents and authoring YAML rules; the API is unified across all verticals. Access is via a hosted API (no local setup required); an API key is created in the account settings and used as a Bearer token on POST /api/v1/decide. A self-hosted deployment option is available on the Enterprise plan.
Community Discussions
Be the first to start a conversation about ai-rete-rag
Share your experience with ai-rete-rag, ask questions, or help others learn from your insights.
Pricing
Free
Enough to build against and see the whole product, including the audit trail.
- 1 domain
- 1,000 decisions / month
- 10 MB document storage
- All response modes incl. full_audit
- Public API access
Supporter
Ten times the free quota — keeps the lights on.
- 3 domains
- 10,000 decisions / month
- 50 MB document storage
- All response modes incl. full_audit
- Public API access
- Team members (shared quota)
- Community support
Builder
For a side project or an internal tool that has started getting real traffic.
- 5 domains
- 25,000 decisions / month
- 250 MB document storage
- All response modes incl. full_audit
- Public API access
- Team members (shared quota)
- Email support
Standard
For solo builders and small teams running real workloads.
- 10 domains
- 100,000 decisions / month
- 1 GB document storage
- All response modes incl. full_audit
- Public API access
- Team members (shared quota)
- Email support
Pro
For teams shipping decision-critical products.
- Unlimited domains
- 500,000 decisions / month
- 10 GB document storage
- All response modes incl. full_audit
- Rule change tracking on every decision
- Priority support (< 4 h SLA)
Enterprise
For regulated industries with advanced compliance needs.
- Unlimited everything
- Self-hosted deployment option
- Custom LLM endpoints
- Dedicated success engineer
- Custom SLA
- Compliance review & BAA on request
Capabilities
Key Features
- Rete rule engine with deterministic, auditable logic
- Retrieval-augmented generation grounded in uploaded documents
- YAML rule authoring — no code required
- Salience-based conflict resolution
- Hot-reload rules without server restart
- Full firing trace stored per decision
- ChromaDB vector store with sentence-transformer embeddings
- Per-domain vector collections
- Configurable chunk size and overlap
- Relevance scores returned with every chunk
- Automatic mode selection (rules-only, RAG-only, hybrid)
- Three response modes: verdict, verdict_with_explanation, full_audit
- Fact extraction from unstructured text
- Static conflict detection across rules
- Browsable policy rule catalog per domain
- Eight built-in demo domains
- Single unified REST API for all verticals
- Usage quota tracking via GET /api/v1/usage
- Self-hosted deployment option (Enterprise)
- Custom LLM endpoints (Enterprise)
