EveryDev.ai
Subscribe
Home
Tools

4,376+ AI tools

  • New
  • Trending
  • Featured
  • Rate tools
  • Compare
  • Arena
Categories
  • Agents3274
  • Coding2275
  • Infrastructure1000
  • Projects696
  • Marketing636
  • Research587
  • MCP532
  • Design508
  • Analytics506
  • Testing394
  • Security376
  • Data327
  • Integration244
  • Prompts244
  • Communication235
  • Extensions217
  • Voice193
  • Learning190
  • Commerce170
  • DevOps153
  • Web103
  • Finance36
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • Users
  • Rate Tools
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Tools
    3. wllama
    wllama icon

    wllama

    Local Inference
    Featured

    WebAssembly binding for llama.cpp that runs GGUF LLM inference directly in the browser.

    Visit Website

    At a Glance

    Pricing
    Open Source

    wllama, a WebAssembly binding for llama.cpp, is released under the MIT License and can be used for free.

    Engagement

    Available On

    Linux
    Web
    API
    VS Code
    SDK

    Resources

    WebsiteDocsGitHubllms.txt

    Topics

    Local InferenceAI Development LibrariesAI Infrastructure

    Alternatives

    BaseRTDeepSpeedcuTile Rust
    Developer
    Xuan-Son NguyenFrance

    Listed Oct 2026

    About wllama

    wllama is an open-source WebAssembly binding for llama.cpp, created and maintained by Xuan-Son Nguyen. It lets web applications run GGUF language models directly in the browser, with no backend required, and is distributed as the npm package @wllama/wllama. A hosted demo app is available on Hugging Face Spaces.

    What It Is

    wllama is a TypeScript library that compiles llama.cpp to WebAssembly so LLM inference can run client-side. Developers load a GGUF model from the Hugging Face hub or from a URL, then call an OpenAI-compatible API to create chat completions or embeddings. Inference runs inside a worker so it does not block UI rendering.

    Capabilities

    The V3 release adds WebGPU, multimodal (image and audio file input) and tool calling support. With WebGPU, all layers are offloaded to the GPU by default, and the n_gpu_layers parameter can adjust or disable this. The library automatically switches between single-thread and multi-thread builds based on browser support, and has no runtime dependencies. Examples cover completions, embeddings with cosine distance, multimodal completion, tool calling and decision models.

    Model Handling and Limits

    Because of ArrayBuffer size restrictions, files are limited to 2GB, so larger models must be split with llama-gguf-split. Splitting into chunks of up to 512MB also allows parallel downloads. Quantized Q4, Q5 or Q6 models are recommended. The project notes that WebAssembly overhead can reduce performance by 25% to 50% compared with native llama.cpp, that smartphones may be buggy, and that Safari is not supported because it lacks Memory64 support. Multi-threading requires Cross-Origin-Embedder-Policy and Cross-Origin-Opener-Policy headers.

    Setup Path

    Install with npm, import the Wllama class, point it at the wasm file paths, load a model and request a chat completion. The wasm binaries can also be built from source using Docker.

    wllama - 1

    Community Discussions

    Be the first to start a conversation about wllama

    Share your experience with wllama, ask questions, or help others learn from your insights.

    Pricing

    OPEN SOURCE

    Open Source (MIT)

    wllama, a WebAssembly binding for llama.cpp, is released under the MIT License and can be used for free.

    • MIT License
    • Pre-built npm package @wllama/wllama
    • On-browser LLM inference via WebAssembly, no backend or GPU needed
    • WebGPU support
    • Multimodal support (image and audio file input)

    Capabilities

    Key Features

    • Run LLM inference in the browser via WebAssembly
    • OpenAI-compatible fully-typed API
    • WebGPU support
    • Multimodal support (image and audio input)
    • Tool calling
    • Embeddings support
    • GGUF model loading from Hugging Face or URL
    • Model splitting with parallel downloads
    • Auto switch between single-thread and multi-thread builds
    • Inference runs in a worker without blocking UI
    • No runtime dependencies
    • Custom logger support

    Integrations

    llama.cpp
    Hugging Face Hub
    npm
    React
    API Available
    View Docs

    Ratings & Reviews

    No ratings yet

    Be the first to rate wllama and help others make informed decisions.

    Rate other tools you’ve used

    Developer

    Xuan-Son Nguyen

    The sky is blue. The sun is yellow. Here we go. There and back again.

    France
    Read more about Xuan-Son Nguyen
    WebsiteGitHubLinkedInX / Twitter
    1 tool in directory

    Similar Tools

    BaseRT icon

    BaseRT

    BaseRT is the fastest LLM inference runtime for Apple Silicon, offering up to 6.4x faster prefill than llama.cpp and 3.9x faster than MLX, with an OpenAI-compatible server and multi-language bindings.

    DeepSpeed icon

    DeepSpeed

    An open-source deep learning optimization library by Microsoft that enables efficient training and inference of large-scale AI models through ZeRO, 3D-Parallelism, and other system innovations.

    cuTile Rust icon

    cuTile Rust

    A tile-based system for writing memory-safe, data-race-free GPU kernels in idiomatic Rust, extending Rust's ownership discipline across the GPU launch boundary.

    Browse all tools

    Related Topics

    Local Inference

    Tools and platforms for running AI inference locally without cloud dependence.

    239 tools

    AI Development Libraries

    Programming libraries and frameworks that provide machine learning capabilities, model integration, and AI functionality for developers.

    352 tools

    AI Infrastructure

    Infrastructure designed for deploying and running AI models.

    434 tools
    Browse all topics
    Back to all toolsSuggest an edit
    ratings
    discussions