nanosamur.ai
An open-source, self-hosted speech AI platform for organizations that need to capture, transcribe, and process sensitive conversations entirely within their own infrastructure.
At a Glance
Complete open-source speech AI platform, self-hosted via Docker Compose, licensed under Apache 2.0.
Engagement
Available On
Alternatives
Listed Sep 2026
About nanosamur.ai
nanosamur.ai is a complete, open-source speech AI platform built for organizations that cannot send sensitive audio to third-party cloud services. Licensed under Apache 2.0, the entire stack — browser UI, desktop application, APIs, orchestration, and speech services — is publicly available on GitHub and deployable via Docker Compose on infrastructure you control.
What It Is
nanosamur.ai is a self-hosted, model-agnostic speech processing platform that handles live transcription, speaker diarization, refinement, and final batch processing without relying on external speech APIs. It targets government, defence, and healthcare use cases where audio and transcripts must remain inside controlled infrastructure, including air-gapped or restricted networks. The platform ships as a Community Edition that can be evaluated locally and adapted to Kubernetes for production deployments.
Architecture and Components
The stack is distributed across several independently deployed services connected by Kafka for event streaming and PostgreSQL for persistence:
- SamuraiBFF — HTTP/WebSocket API, browser UI (ClojureScript), authentication, and orchestration
- Xamurai — Python speech services covering realtime transcription, refinement, finalization, and session audio recording
- SamuraiPersistor — Kafka-to-PostgreSQL transcript writer
- nanosamurai-sdk — Python SDK and CLI
Audio flows from browser or Electron client over WebSocket as PCM16LE mono 16 kHz, through Kafka topics, and into the appropriate speech-service pipeline. S3-compatible object storage (LocalStack locally; AWS S3, Ceph, or MinIO in deployments) holds session WAV recordings and speaker enrollment data.
Supported Models and Pipelines
The platform is model-agnostic and supports multiple model families across three processing stages:
- Whisper / Faster-Whisper / WhisperX — realtime, refined, and final; on by default
- Qwen3-ASR — realtime, refined, and final
- Nemotron — realtime
- Parakeet — refined and final (with embedded Sortformer diarization)
Multiple models can run in parallel on the same audio, producing separate labelled results for comparison. Speaker diarization is available via Pyannote or Sortformer, with optional speaker enrollment for identity-aware transcripts.
Deployment Model
Docker Compose is the supported public evaluation path. The quickstart requires Docker Engine or Docker Desktop, Docker Compose v2, an NVIDIA GPU with the NVIDIA container runtime, and a Hugging Face token for gated models. The full stack starts with three commands:
cp .env.example .env
docker compose pull
docker compose up -d
An optional observability compose file adds Grafana, Prometheus, Loki, and Tempo, with W3C trace context propagated across Kafka so asynchronous session work can be correlated end to end. The repository does not supply production Kubernetes manifests, but the containerized services can be adapted to Kubernetes.
Target Audience and Use Cases
The platform is explicitly designed for workflows where data sovereignty is non-negotiable:
- Government and defence — briefings, interviews, operational debriefs inside controlled or isolated networks
- Healthcare — clinical consultations and discussions governed by organisational data policies
The security model binds all host ports to 127.0.0.1 by default. Authenticated deployments must supply their own Keycloak instance; the evaluator stack uses fixed development credentials safe only for localhost use.
Integration and Extension Points
Beyond the browser UI and Electron desktop app, nanosamur.ai exposes REST, real-time WebSocket, Python SDK, Kafka events, and webhook contracts. The Community Edition does not ship workflow execution or webhook delivery services, but the agentic-workflow and webhook contracts are public, allowing teams to implement and plug in their own services.
Community Discussions
Be the first to start a conversation about nanosamur.ai
Share your experience with nanosamur.ai, ask questions, or help others learn from your insights.
Pricing
Community Edition
Complete open-source speech AI platform, self-hosted via Docker Compose, licensed under Apache 2.0.
- Full platform source code under Apache 2.0
- Browser UI and Windows Electron app
- Realtime, refined, and batch transcription
- Speaker diarization
- Multi-model support (Whisper, Qwen, Nemotron, Parakeet)
Capabilities
Key Features
- Live and batch speech transcription
- Speaker diarization (Pyannote and Sortformer)
- Speaker enrollment and identity-aware transcripts
- Multi-model parallel processing (Whisper, Qwen, Nemotron, Parakeet)
- Realtime WebSocket audio streaming
- Refinement and finalization pipelines
- Session WAV recording and storage
- PostgreSQL transcript persistence
- Browser UI and Windows Electron desktop app
- Python SDK and CLI
- REST and WebSocket APIs
- Kafka event streaming
- Agentic workflow and webhook contracts
- Optional Grafana, Prometheus, Loki, and Tempo observability stack
- W3C distributed trace context across async services
- S3-compatible object storage support
- Multitenancy support
- Self-hosted, air-gap capable deployment
- Docker Compose quickstart
- Apache 2.0 open-source license
