nanosamurai
nanosamur.ai is a complete open-source speech AI platform for organizations that cannot send sensitive conversations to a third party. It captures, transcribes, refines, stores, and processes speech inside infrastructure controlled by the organization.
At a Glance
- Organizations that cannot send sensitive conversations to third parties
- Government and defence
- Healthcare
- Teams operating private clouds, Kubernetes clusters, restricted networks, or air-gapped environments
- +1 more
AI Tools by nanosamurai
(1)nanosamur.ai
Self Hosted Speech AI Platform
Discussions
No discussions yet
Be the first to start a discussion about nanosamurai
Latest News
We've added support for new ASR models: Whisper, Parakeet TDT, Nemotron 3.5 ASR and Qwen 3 ASR
Model Arena - Qwen3-ASR vs faster-whisper in real time
Qwen3-ASR joins nanosamur.ai realtime transcription
Why speech AI belongs inside your infrastructure
Products & Services
Apache-2.0 open-source, Docker Compose-based speech AI platform with browser UI, SamuraiBFF API and orchestration, a Windows-first Electron wrapper, recording storage, persistence, and selectable realtime, refined, and batch transcription pipelines.
Independent, labelled realtime tracks for Faster-Whisper and Qwen3-ASR; the same audio stream can be fanned out to multiple providers so operators can compare latency, revisions, strengths, and errors. Nemotron 3.5 ASR is also listed as a supported realtime family.
Semi-realtime refinement and post-session final processing using model-specific tracks. The published model matrix lists WhisperX, Qwen3-ASR, and Parakeet TDT for refined/final processing.
Python SDK and command-line source for consuming and integrating the platform.
Market Position
nanosamur.ai positions itself as a complete, self-hostable, model-agnostic speech platform rather than only a transcription model or inference API. Its differentiators are keeping the full audio-to-record data path under the operator's control, supporting multiple models and stages, and including UI, desktop capture, orchestration, persistence, integrations, and observability. In the creator's description it is more like Ollama plus an Ollama-Cloud-style service for voice models, while its model-arena capability is aimed at comparing providers such as Qwen3-ASR and Faster-Whisper.
Leadership
Founders
newcrobuzon
The Hacker News creator/poster said they had spent the previous couple of years consulting for organizations handling sensitive data in air-gapped environments, and built and open-sourced nanosamur.ai from those experiences.
Founding Story
The creator said the project grew out of consulting work for organizations with sensitive data that had to operate in air-gapped environments. The initial vision was to provide an Ollama-like, model-agnostic speech-to-text stack that could run locally, on premises, or in cloud/Kubernetes infrastructure without sending conversations to an external provider.
Business Model
Revenue Model
The available materials describe an Apache-2.0 open-source Community Edition that users run on their own hardware or infrastructure; no paid pricing or commercial revenue model is stated.
Target Markets
- Organizations that cannot send sensitive conversations to third parties
- Government and defence
- Healthcare
- Teams operating private clouds, Kubernetes clusters, restricted networks, or air-gapped environments
- Speech infrastructure, self-hosted AI, and MLOps teams
- Government and defence interviews, briefings, meetings, and operational debriefs
- Healthcare consultations and clinical discussions
- Sensitive business conversations requiring recordings and transcripts to remain inside organizational infrastructure
- Realtime captions, live search, prompts, and automations
- Post-session audit, playback, search, and durable records
- Model evaluation and benchmarking on identical audio streams