EveryDev.ai
Subscribe
Home
Tools

4,072+ AI tools

  • New
  • Trending
  • Featured
  • Compare
  • Arena
Categories
  • Agents2782
  • Coding1973
  • Infrastructure825
  • Projects603
  • Marketing598
  • Research520
  • Analytics468
  • Design462
  • MCP419
  • Testing346
  • Security323
  • Data305
  • Integration224
  • Prompts220
  • Communication210
  • Extensions196
  • Learning179
  • Voice175
  • Commerce160
  • DevOps135
  • Web95
  • Finance31
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Tools
    3. nanosamur.ai
    nanosamur.ai icon

    nanosamur.ai

    Speech Recognition

    An open-source, self-hosted speech AI platform for organizations that need to capture, transcribe, and process sensitive conversations entirely within their own infrastructure.

    Visit Website

    At a Glance

    Pricing
    Open Source

    Complete open-source speech AI platform, self-hosted via Docker Compose, licensed under Apache 2.0.

    Engagement

    Available On

    Windows
    Linux
    Web
    API
    SDK

    Resources

    WebsiteDocsGitHubllms.txt

    Topics

    Speech RecognitionAudioAI Infrastructure

    Alternatives

    UltravoxInworld AIVibeVoice
    Developer
    nanosamurainanosamurai builds nanosamur.ai, a complete open-source spee…

    Listed Sep 2026

    About nanosamur.ai

    nanosamur.ai is a complete, open-source speech AI platform built for organizations that cannot send sensitive audio to third-party cloud services. Licensed under Apache 2.0, the entire stack — browser UI, desktop application, APIs, orchestration, and speech services — is publicly available on GitHub and deployable via Docker Compose on infrastructure you control.

    What It Is

    nanosamur.ai is a self-hosted, model-agnostic speech processing platform that handles live transcription, speaker diarization, refinement, and final batch processing without relying on external speech APIs. It targets government, defence, and healthcare use cases where audio and transcripts must remain inside controlled infrastructure, including air-gapped or restricted networks. The platform ships as a Community Edition that can be evaluated locally and adapted to Kubernetes for production deployments.

    Architecture and Components

    The stack is distributed across several independently deployed services connected by Kafka for event streaming and PostgreSQL for persistence:

    • SamuraiBFF — HTTP/WebSocket API, browser UI (ClojureScript), authentication, and orchestration
    • Xamurai — Python speech services covering realtime transcription, refinement, finalization, and session audio recording
    • SamuraiPersistor — Kafka-to-PostgreSQL transcript writer
    • nanosamurai-sdk — Python SDK and CLI

    Audio flows from browser or Electron client over WebSocket as PCM16LE mono 16 kHz, through Kafka topics, and into the appropriate speech-service pipeline. S3-compatible object storage (LocalStack locally; AWS S3, Ceph, or MinIO in deployments) holds session WAV recordings and speaker enrollment data.

    Supported Models and Pipelines

    The platform is model-agnostic and supports multiple model families across three processing stages:

    • Whisper / Faster-Whisper / WhisperX — realtime, refined, and final; on by default
    • Qwen3-ASR — realtime, refined, and final
    • Nemotron — realtime
    • Parakeet — refined and final (with embedded Sortformer diarization)

    Multiple models can run in parallel on the same audio, producing separate labelled results for comparison. Speaker diarization is available via Pyannote or Sortformer, with optional speaker enrollment for identity-aware transcripts.

    Deployment Model

    Docker Compose is the supported public evaluation path. The quickstart requires Docker Engine or Docker Desktop, Docker Compose v2, an NVIDIA GPU with the NVIDIA container runtime, and a Hugging Face token for gated models. The full stack starts with three commands:

    cp .env.example .env
    docker compose pull
    docker compose up -d
    

    An optional observability compose file adds Grafana, Prometheus, Loki, and Tempo, with W3C trace context propagated across Kafka so asynchronous session work can be correlated end to end. The repository does not supply production Kubernetes manifests, but the containerized services can be adapted to Kubernetes.

    Target Audience and Use Cases

    The platform is explicitly designed for workflows where data sovereignty is non-negotiable:

    • Government and defence — briefings, interviews, operational debriefs inside controlled or isolated networks
    • Healthcare — clinical consultations and discussions governed by organisational data policies

    The security model binds all host ports to 127.0.0.1 by default. Authenticated deployments must supply their own Keycloak instance; the evaluator stack uses fixed development credentials safe only for localhost use.

    Integration and Extension Points

    Beyond the browser UI and Electron desktop app, nanosamur.ai exposes REST, real-time WebSocket, Python SDK, Kafka events, and webhook contracts. The Community Edition does not ship workflow execution or webhook delivery services, but the agentic-workflow and webhook contracts are public, allowing teams to implement and plug in their own services.

    nanosamur.ai - 1

    Community Discussions

    Be the first to start a conversation about nanosamur.ai

    Share your experience with nanosamur.ai, ask questions, or help others learn from your insights.

    Pricing

    OPEN SOURCE

    Community Edition

    Complete open-source speech AI platform, self-hosted via Docker Compose, licensed under Apache 2.0.

    • Full platform source code under Apache 2.0
    • Browser UI and Windows Electron app
    • Realtime, refined, and batch transcription
    • Speaker diarization
    • Multi-model support (Whisper, Qwen, Nemotron, Parakeet)

    Capabilities

    Key Features

    • Live and batch speech transcription
    • Speaker diarization (Pyannote and Sortformer)
    • Speaker enrollment and identity-aware transcripts
    • Multi-model parallel processing (Whisper, Qwen, Nemotron, Parakeet)
    • Realtime WebSocket audio streaming
    • Refinement and finalization pipelines
    • Session WAV recording and storage
    • PostgreSQL transcript persistence
    • Browser UI and Windows Electron desktop app
    • Python SDK and CLI
    • REST and WebSocket APIs
    • Kafka event streaming
    • Agentic workflow and webhook contracts
    • Optional Grafana, Prometheus, Loki, and Tempo observability stack
    • W3C distributed trace context across async services
    • S3-compatible object storage support
    • Multitenancy support
    • Self-hosted, air-gap capable deployment
    • Docker Compose quickstart
    • Apache 2.0 open-source license

    Integrations

    Docker Compose
    Kafka
    PostgreSQL
    LocalStack (S3-compatible)
    AWS S3
    Ceph RADOS Gateway
    MinIO
    Keycloak
    Grafana
    Prometheus
    Loki
    Tempo
    Alloy
    NVIDIA DCGM exporter
    Hugging Face (model downloads)
    Faster-Whisper
    WhisperX
    Qwen3-ASR
    Nemotron
    Parakeet
    Pyannote
    Sortformer
    vLLM
    API Available
    View Docs

    Ratings & Reviews

    No ratings yet

    Be the first to rate nanosamur.ai and help others make informed decisions.

    Developer

    nanosamurai

    nanosamurai builds nanosamur.ai, a complete open-source speech AI platform designed for organisations that cannot send sensitive conversations to third-party cloud services. The project delivers a production-grade, distributed architecture covering realtime transcription, speaker diarization, refinement, and agentic workflow integration — all running inside infrastructure the operator controls. The full stack, including browser UI, desktop app, APIs, and speech services, is published under the Apache 2.0 license on GitHub.

    Read more about nanosamurai
    WebsiteGitHub
    1 tool in directory

    Similar Tools

    Ultravox icon

    Ultravox

    Real-time voice AI platform with speech-native models for building and scaling conversational voice agents.

    Inworld AI icon

    Inworld AI

    Production-grade voice AI APIs offering top-ranked text-to-speech, speech-to-speech, speech-to-text, and LLM routing for developers building natural conversational applications.

    VibeVoice icon

    VibeVoice

    An open-source family of frontier voice AI models from Microsoft, including long-form TTS, multi-speaker speech synthesis, real-time streaming TTS, and long-form ASR with speaker diarization.

    Browse all tools

    Related Topics

    Speech Recognition

    AI tools that convert spoken language into text.

    52 tools

    Audio

    AI tools that generate or edit audio — music, sound effects, voice and speech, and podcast production.

    39 tools

    AI Infrastructure

    Infrastructure designed for deploying and running AI models.

    409 tools
    Browse all topics
    Back to all toolsSuggest an edit
    ratings
    discussions