EveryDev.ai
Subscribe
Home
Tools

4,072+ AI tools

  • New
  • Trending
  • Featured
  • Compare
  • Arena
Categories
  • Agents2782
  • Coding1973
  • Infrastructure825
  • Projects603
  • Marketing598
  • Research520
  • Analytics468
  • Design462
  • MCP419
  • Testing346
  • Security323
  • Data305
  • Integration224
  • Prompts220
  • Communication210
  • Extensions196
  • Learning179
  • Voice175
  • Commerce160
  • DevOps135
  • Web95
  • Finance31
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Tools
    3. fusion-runtime
    fusion-runtime icon

    fusion-runtime

    Voice Assistant

    A self-hosted, open-source voice agent runtime that runs speech-to-text, an LLM, and text-to-speech together in one Python process on your own GPU, streaming into each other for low-latency replies.

    Visit Website

    At a Glance

    Pricing
    Open Source

    Fully free and open-source under Apache-2.0. Install via pip, self-host on your own hardware.

    Engagement

    Available On

    CLI
    API
    Linux
    macOS
    Windows

    Resources

    WebsiteDocsGitHubllms.txt

    Topics

    Voice AssistantAutonomous SystemsLocal Inference

    Alternatives

    SutandoHeardKuma Voice
    Developer
    SamarthUrs18Bengaluru

    Listed Sep 2026

    About fusion-runtime

    fusion-runtime is an Apache-2.0-licensed, self-hosted voice agent runtime that runs speech-to-text, a language model, and text-to-speech together in a single Python process on hardware you control. Released as v0.1.1 in late September 2026, it is in early development and installable today via pip install fusion-runtime.

    What It Is

    fusion-runtime eliminates the need to rent each stage of a voice pipeline from separate cloud APIs. Instead of calling Deepgram for STT, OpenAI or Groq for the LLM, and ElevenLabs or Cartesia for TTS, the runtime runs all three stages locally with open-weight models — Whisper for speech-to-text, llama.cpp/vLLM/SGLang for the language model, and Kokoro for text-to-speech — streaming each stage's output into the next so a reply starts playing while it is still being generated. Audio and transcripts stay on your machine by default; no third-party inference call is required in the default path.

    How the Pipeline Works

    The pipeline runs as a single process: microphone audio passes through Silero VAD (speech-only filtering), then faster-whisper (rolling-window transcription), then turn detection (silence plus resume logic), then the LLM, then Kokoro TTS, which speaks each sentence as it is written. A barge-in watcher runs alongside: if the caller talks over the reply, generation and playback stop immediately. Echo cancellation happens on the device — in the browser or terminal client — so the server hears clean audio and can distinguish a real interruption from the agent's own playback.

    • STT: faster-whisper (any CTranslate2-converted Whisper model, tiny through large-v3)
    • LLM: llama.cpp in-process, or any OpenAI-compatible endpoint (vLLM, SGLang, llama-server, Ollama, hosted APIs)
    • TTS: Kokoro (ONNX), with multiple voices
    • VAD: Silero VAD

    Update: v0.1.1 — Agents with Tool Calling

    Version 0.1.1, published 26 September 2026, adds tool-calling support. A tool is a plain Python function decorated with @tool; the model reads its name, docstring, and type hints, calls it when needed, and answers with the result. Key behaviors:

    • What the model says before a lookup ("Let me check.") is spoken while the tool runs — no dead air.
    • Talking over the wait cancels the lookup.
    • A tool that fails or times out tells the model what went wrong rather than ending the call.
    • Lookup results stay in the conversation history so the model doesn't re-invent answers.
    • Tools require a model server that supports function calling: vLLM, SGLang, llama-server --jinja, or a hosted API. The in-process llama.cpp runtime cannot yet call tools, and frun up says so at startup.

    Measured Performance and Concurrency

    The homepage publishes benchmark figures from the runtime's own per-turn telemetry, measured on an RTX 3090 with Qwen 7B q4, Whisper small, and Kokoro all on one card, through the browser client on 21 September 2026:

    MetricMedian
    Processing (turn end to first audio)~490 ms
    Stopwatch from last syllable991 ms
    Speech-to-text119 ms
    LLM first token27 ms
    TTS real-time factor0.09 (~11× faster than real time)
    LLM tokens/sec127

    For concurrent callers (measured 26 September 2026, same 3090, three turns each): with the model in-process via llama.cpp, four callers is the practical ceiling before queuing degrades response time significantly. With vLLM or SGLang batching the LLM, twelve callers answered in about one second; at sixteen callers, the speech stages (not the LLM) become the bottleneck, as Kokoro is not yet batched across callers.

    Deployment and Setup Path

    Installation requires Python 3.11–3.13. An agent is defined in a single Python file; frun models pull downloads exactly the models it names; frun up starts the server. The frun doctor command checks Python, libraries, GPU support, models, port, and audio, and reports how to fix any issues.

    For browser embedding, the runtime serves its own client JavaScript — two script tags and no build step. Browsers require https:// for microphone access, so production deployments need TLS and wss://. Authentication uses short-lived, single-use session tokens minted by the backend; API keys are never held in the browser.

    Deployment targets include any NVIDIA GPU machine, Docker containers, or systemd units on GPU VMs. The docs are written against RunPod as a reference host. A hosted cloud version is planned but not yet built.

    Who It Fits and Honest Tradeoffs

    The project's own documentation explicitly names the tradeoffs: open-weight models are behind the best closed APIs today. Whisper is not Nova-3; Qwen 0.5B/7B is not GPT-4o; Kokoro is not ElevenLabs. The runtime is described as a good fit for narrow-domain agents (reservations, order status, IVR replacement), privacy- or data-residency-sensitive teams, cost-sensitive high-volume deployments, and local-first developers (kiosks, robots, desktop apps). It is not yet suited for open-ended assistants requiring frontier-model reasoning, phone/telephony deployments, or workloads beyond about twelve simultaneous callers on a single GPU.

    fusion-runtime - 1

    Community Discussions

    Be the first to start a conversation about fusion-runtime

    Share your experience with fusion-runtime, ask questions, or help others learn from your insights.

    Pricing

    OPEN SOURCE

    Open Source

    Fully free and open-source under Apache-2.0. Install via pip, self-host on your own hardware.

    • Full voice agent runtime (STT + LLM + TTS in one process)
    • Tool calling support
    • Browser client included
    • CLI (frun up, talk, doctor, models, key, token)
    • vLLM and SGLang integration for concurrent callers

    Capabilities

    Key Features

    • Speech-to-text, LLM, and TTS in one process on one GPU
    • Streaming pipeline: reply starts playing while still being generated
    • Barge-in / interruption detection with echo cancellation
    • Tool calling via plain Python functions with @tool decorator
    • No dead air during tool lookups — pre-lookup speech is spoken immediately
    • Silero VAD for speech-only filtering
    • faster-whisper STT (any CTranslate2 Whisper model)
    • Kokoro TTS (ONNX) with multiple voices
    • llama.cpp in-process LLM or any OpenAI-compatible endpoint (vLLM, SGLang, Ollama)
    • Browser client served by the runtime itself — two script tags, no build step
    • Short-lived single-use session tokens for browser authentication
    • frun CLI: up, talk, doctor, models pull/list, key new, token
    • Per-turn telemetry: TTFA, tokens/sec, per-stage timings
    • Prometheus metrics endpoint
    • Structured JSON logging
    • Configurable turn detection (silence wait, interrupt threshold)
    • Multi-language STT via multilingual Whisper models
    • Docker and systemd deployment support
    • Library/SDK mode for use inside existing Python processes
    • Apache-2.0 license, commercial use permitted

    Integrations

    vLLM
    SGLang
    llama.cpp
    llama-server
    Ollama
    OpenAI API
    Groq API
    faster-whisper (CTranslate2)
    Kokoro TTS (ONNX)
    Silero VAD
    Hugging Face Hub
    Qwen2.5 models
    Whisper models
    PyTorch
    FastAPI
    WebSocket
    Prometheus
    Docker
    RunPod
    NVIDIA CUDA
    API Available
    View Docs

    Ratings & Reviews

    No ratings yet

    Be the first to rate fusion-runtime and help others make informed decisions.

    Developer

    SamarthUrs18

    fusion-runtime builds an open-source, self-hosted voice agent runtime that runs speech-to-text, an LLM, and text-to-speech together in one Python process on your own GPU. The project is developed by SamarthUrs18 and released under the Apache-2.0 license, enabling commercial embedding and rebranding. The runtime targets narrow-domain voice agents, privacy-sensitive deployments, and local-first developers who need full control over their voice stack without per-minute API costs.

    Bengaluru
    Read more about SamarthUrs18
    WebsiteGitHub
    1 tool in directory

    Similar Tools

    Sutando icon

    Sutando

    An open-source, self-hosted AI agent for macOS that uses voice, vision, and autonomous action to control your computer, join meetings, make phone calls, and build itself.

    Heard icon

    Heard

    Heard turns AI coding agent terminal activity into spoken voice updates, letting developers step away from the screen without losing track of what their agents are doing.

    Kuma Voice icon

    Kuma Voice

    A self-hosted, watch-only voice assistant prototype that streams microphone audio from Apple Watch to a FastAPI backend, supporting web search, Notion notes, and Google Calendar.

    Browse all tools

    Related Topics

    Voice Assistant

    AI voice assistants that perform tasks through voice commands.

    53 tools

    Autonomous Systems

    AI agents that can perform complex tasks with minimal human guidance.

    445 tools

    Local Inference

    Tools and platforms for running AI inference locally without cloud dependence.

    215 tools
    Browse all topics
    Back to all toolsSuggest an edit
    ratings
    discussions