# Xtriever

> A hybrid retrieval engine for RAG that runs fully on-device — on iPhones, Android phones, and laptops — with no server or network required, written in Rust.

Xtriever is an open-source hybrid retrieval engine for retrieval-augmented generation (RAG), written in Rust and designed to run entirely on the device — no server, no network connection required. It supports iPhones, Android phones, and laptops, and exposes the same engine through Rust, Python, Swift, and Kotlin surfaces. The project is licensed under Apache 2.0 and is available on GitHub.

## What It Is

Xtriever combines lexical and dense search into a four-stage pipeline: BM25 lexical search (via tantivy), dense vector search (all-MiniLM-L6-v2 embeddings at 384 dimensions, quantized to int8), reciprocal rank fusion of the two result lists, and cross-encoder re-ranking (ms-marco-MiniLM-L-6-v2). An optional fifth stage — a learned-to-rank (LTR) ranker — is planned but not yet implemented; the `xtriever-ltr` crate is currently a placeholder. The two default models are eight-bit GGUF artifacts totaling 51.5 MB, running entirely on the CPU.

## On-Device Architecture

The engine is built around the constraint that everything must fit and run on a mobile device without a network call:

- Indexes and models can be memory-mapped read-only, and an index can open in place inside a read-only app bundle.
- On an iPhone 16e searching all of Simple English Wikipedia (~428,000 passages, 585 MB index), peak memory stays at 335 MB — under the project's self-imposed 600 MB ceiling.
- Median fused-list latency on that device is 0.20 s; re-ranking adds 1.21 s (dominated by the cross-encoder's per-candidate forward passes).
- The same index on a 2021 MacBook Pro (M1 Pro) via Python takes 0.14 s fused and 0.99 s re-ranked.
- Android results match the host bit-for-bit on everything except the cross-encoder, whose scores are within 7e-6 of the host's.

## Retrieval Quality

The README documents nDCG@10 on three BEIR datasets using the project's own evaluation harness:

| Configuration | SciFact | NFCorpus | FiQA | Mean |
|---|---|---|---|---|
| BM25 alone | 0.686 | 0.323 | 0.250 | 0.420 |
| Dense alone | 0.646 | 0.315 | 0.369 | 0.444 |
| BM25 + dense, fused | 0.715 | 0.354 | 0.370 | 0.480 |
| Fused + re-ranked (default) | 0.722 | 0.362 | 0.390 | 0.491 |
| With sparse expansion | 0.722 | 0.358 | 0.407 | 0.495 |

An optional sparse expansion stage using `opensearch-neural-sparse-encoding-doc-v3-distill` helps corpora where questions are worded differently from their answers, and is off by default.

## Multi-Language Surfaces and Demos

The FFI layer is generated by Mozilla's uniffi from a single Rust crate, so Python, Swift, and Kotlin bindings expose identical operations and records with no retrieval logic of their own:

- **Rust**: `HybridIndex` API in `crates/xtriever-pipeline`
- **Python**: wheel built with maturin
- **Swift**: XCFramework with an async `XtrieverIndex`, distributed as a Swift package
- **Kotlin**: Gradle library module for 64-bit ARM Android 8.0+

Four demo applications ship with the repository: a minimal Python demo (under 80 lines), a command-line Python demo over Simple English Wikipedia, a SwiftUI iOS app, and a Jetpack Compose Android app — all searching the same ~240,000-article Wikipedia corpus offline.

## Current Status: Version 0.1.0, Not Yet Published

The README explicitly states: "Status: 0.1.0, not published." The engine builds from the repository, but nothing is yet on crates.io, PyPI, Swift Package Index, or Maven. The repository was created in September 2026 and last pushed on September 24, 2026. CI runs on Linux, macOS, and Windows. The wasm32 build target is a known gap, tracked but not yet resolved. Each feature goes through a structured spec-plan-implement cycle, with 27 completed spec directories committed under `specs/`.

## Features
- Hybrid BM25 + dense vector search with reciprocal rank fusion
- Cross-encoder re-ranking (ms-marco-MiniLM-L-6-v2)
- Fully on-device: no server or network required
- Runs on iPhone, Android, and laptop
- Memory-mapped indexes and models for low footprint
- int8-quantized embeddings (all-MiniLM-L6-v2, 384 dimensions)
- Optional sparse lexical expansion via OpenSearch neural sparse encoder
- Rust, Python, Swift, and Kotlin bindings via uniffi
- Graceful degradation: returns previous stage results on error or timeout
- Deterministic results: same index, query, and config produce identical hits
- Per-hit score explanations and stage reports
- SHA-256 and size verification for all models
- BEIR evaluation harness included
- Four demo apps: minimal Python, CLI Wikipedia, iOS SwiftUI, Android Jetpack Compose
- Apache 2.0 open-source license

## Integrations
tantivy (BM25 lexical search), candle (Hugging Face, in-process model inference), uniffi (Mozilla, cross-language FFI), maturin (Python wheel build), all-MiniLM-L6-v2 (sentence-transformers), ms-marco-MiniLM-L-6-v2 (cross-encoder), opensearch-neural-sparse-encoding-doc-v3-distill, cargo-nextest, cargo-deny, XcodeGen (iOS), cargo-ndk (Android), Simple English Wikipedia

## Platforms
WINDOWS, MACOS, LINUX, ANDROID, IOS, API, DEVELOPER_SDK, CLI

## Pricing
Open Source

## Version
0.1.0

## Links
- Website: https://github.com/mirth/xtriever
- Documentation: https://github.com/mirth/xtriever
- Repository: https://github.com/mirth/xtriever
- EveryDev.ai: https://www.everydev.ai/tools/xtriever
