Xtriever
A hybrid retrieval engine for RAG that runs fully on-device — on iPhones, Android phones, and laptops — with no server or network required, written in Rust.
At a Glance
Fully free and open source under the Apache License 2.0. Build from source; not yet published to package registries.
Engagement
Available On
Alternatives
Listed Sep 2026
About Xtriever
Xtriever is an open-source hybrid retrieval engine for retrieval-augmented generation (RAG), written in Rust and designed to run entirely on the device — no server, no network connection required. It supports iPhones, Android phones, and laptops, and exposes the same engine through Rust, Python, Swift, and Kotlin surfaces. The project is licensed under Apache 2.0 and is available on GitHub.
What It Is
Xtriever combines lexical and dense search into a four-stage pipeline: BM25 lexical search (via tantivy), dense vector search (all-MiniLM-L6-v2 embeddings at 384 dimensions, quantized to int8), reciprocal rank fusion of the two result lists, and cross-encoder re-ranking (ms-marco-MiniLM-L-6-v2). An optional fifth stage — a learned-to-rank (LTR) ranker — is planned but not yet implemented; the xtriever-ltr crate is currently a placeholder. The two default models are eight-bit GGUF artifacts totaling 51.5 MB, running entirely on the CPU.
On-Device Architecture
The engine is built around the constraint that everything must fit and run on a mobile device without a network call:
- Indexes and models can be memory-mapped read-only, and an index can open in place inside a read-only app bundle.
- On an iPhone 16e searching all of Simple English Wikipedia (~428,000 passages, 585 MB index), peak memory stays at 335 MB — under the project's self-imposed 600 MB ceiling.
- Median fused-list latency on that device is 0.20 s; re-ranking adds 1.21 s (dominated by the cross-encoder's per-candidate forward passes).
- The same index on a 2021 MacBook Pro (M1 Pro) via Python takes 0.14 s fused and 0.99 s re-ranked.
- Android results match the host bit-for-bit on everything except the cross-encoder, whose scores are within 7e-6 of the host's.
Retrieval Quality
The README documents nDCG@10 on three BEIR datasets using the project's own evaluation harness:
| Configuration | SciFact | NFCorpus | FiQA | Mean |
|---|---|---|---|---|
| BM25 alone | 0.686 | 0.323 | 0.250 | 0.420 |
| Dense alone | 0.646 | 0.315 | 0.369 | 0.444 |
| BM25 + dense, fused | 0.715 | 0.354 | 0.370 | 0.480 |
| Fused + re-ranked (default) | 0.722 | 0.362 | 0.390 | 0.491 |
| With sparse expansion | 0.722 | 0.358 | 0.407 | 0.495 |
An optional sparse expansion stage using opensearch-neural-sparse-encoding-doc-v3-distill helps corpora where questions are worded differently from their answers, and is off by default.
Multi-Language Surfaces and Demos
The FFI layer is generated by Mozilla's uniffi from a single Rust crate, so Python, Swift, and Kotlin bindings expose identical operations and records with no retrieval logic of their own:
- Rust:
HybridIndexAPI incrates/xtriever-pipeline - Python: wheel built with maturin
- Swift: XCFramework with an async
XtrieverIndex, distributed as a Swift package - Kotlin: Gradle library module for 64-bit ARM Android 8.0+
Four demo applications ship with the repository: a minimal Python demo (under 80 lines), a command-line Python demo over Simple English Wikipedia, a SwiftUI iOS app, and a Jetpack Compose Android app — all searching the same ~240,000-article Wikipedia corpus offline.
Current Status: Version 0.1.0, Not Yet Published
The README explicitly states: "Status: 0.1.0, not published." The engine builds from the repository, but nothing is yet on crates.io, PyPI, Swift Package Index, or Maven. The repository was created in September 2026 and last pushed on September 24, 2026. CI runs on Linux, macOS, and Windows. The wasm32 build target is a known gap, tracked but not yet resolved. Each feature goes through a structured spec-plan-implement cycle, with 27 completed spec directories committed under specs/.
Community Discussions
Be the first to start a conversation about Xtriever
Share your experience with Xtriever, ask questions, or help others learn from your insights.
Pricing
Open Source
Fully free and open source under the Apache License 2.0. Build from source; not yet published to package registries.
- Full hybrid retrieval pipeline (BM25 + dense + fusion + re-ranking)
- Rust, Python, Swift, and Kotlin bindings
- On-device inference with no server required
- BEIR evaluation harness
- Four demo applications
Capabilities
Key Features
- Hybrid BM25 + dense vector search with reciprocal rank fusion
- Cross-encoder re-ranking (ms-marco-MiniLM-L-6-v2)
- Fully on-device: no server or network required
- Runs on iPhone, Android, and laptop
- Memory-mapped indexes and models for low footprint
- int8-quantized embeddings (all-MiniLM-L6-v2, 384 dimensions)
- Optional sparse lexical expansion via OpenSearch neural sparse encoder
- Rust, Python, Swift, and Kotlin bindings via uniffi
- Graceful degradation: returns previous stage results on error or timeout
- Deterministic results: same index, query, and config produce identical hits
- Per-hit score explanations and stage reports
- SHA-256 and size verification for all models
- BEIR evaluation harness included
- Four demo apps: minimal Python, CLI Wikipedia, iOS SwiftUI, Android Jetpack Compose
- Apache 2.0 open-source license
