EveryDev.ai
Subscribe
Home
Developers

3,067+ AI companies

  • Radar
  • Trending
AI Tools by Topic
  • AI Coding Assistants
  • Agent Frameworks
  • MCP Servers
  • AI Prompt Tools
  • Vibe Coding Tools
  • AI Design Tools
  • AI Database Tools
  • AI Website Builders
  • AI Testing Tools
  • LLM Evaluations
Follow Us
  • X / Twitter
  • LinkedIn
  • Reddit
  • Discord
  • Threads
  • Bluesky
  • Mastodon
  • YouTube
  • GitHub
  • Instagram
Get Started
  • About
  • Editorial Standards
  • Corrections & Disclosures
  • Community Guidelines
  • Advertise
  • Contact Us
  • Newsletter
  • Submit a Tool
  • Start a Discussion
  • Write A Blog
  • Share A Build
  • Terms of Service
  • Privacy Policy
Explore with AI
  • ChatGPT
  • Gemini
  • Claude
  • Grok
  • Perplexity
Agent Experience
  • llms.txt
Theme
With AI, Everyone is a Dev. EveryDev.ai © 2026
    1. Home
    2. Developers
    3. FlashML-org

    FlashML-org

    Bring frontier-scale AI models to edge devices through efficient, bandwidth-adaptive inference engines.

    Visit Website

    At a Glance

    1Tool Listed
    1Product
    4Capabilities
    Discussions
    Berkeley, CaliforniaHeadquarters
    2026Est.
    11Employees
    Focus Areas
    Local Inference
    LLM Orchestration
    AI Infrastructure
    Connect
    Latest News
    FreeToken: Efficient Edge-Native MoE Serving Paper ReleasedAug 24, 2026
    FreeToken Open-Sourced on GitHubAug 16, 2026
    Markets
    • AI Developers
    • Researchers
    • Privacy-conscious LLM users

    AI Tools by FlashML-org

    (1)
    View FreeToken
    FreeToken tool icon

    FreeToken

    Edge MoE Inference Engine

    Local InferenceLLM OrchestrationAI Infrastructure

    Discussions

    No discussions yet

    Be the first to start a discussion about FlashML-org

    Latest News

    08/24/2026

    FreeToken: Efficient Edge-Native MoE Serving Paper Released

    arxiv.org
    08/16/2026

    FreeToken Open-Sourced on GitHub

    github.com

    Products & Services

    1
    FreeToken
    2026-08-16

    An edge-native MoE serving engine supporting large models like DeepSeek-V4 and GLM.

    Market Position

    FlashML's FreeToken provides superior performance for MoE models compared to Ollama and llama.cpp by optimizing for bandwidth-limited consumer hardware.

    Leadership

    Founders

    SY

    Shuo Yang

    PhD student at UC Berkeley Sky Lab; lead developer of FreeToken.

    IS

    Ion Stoica

    Professor at UC Berkeley; Co-founder of Databricks and Anyscale.

    MZ

    Matei Zaharia

    Associate Professor at UC Berkeley; Co-founder of Databricks.

    SH

    Song Han

    Associate Professor at MIT; expert in AI hardware efficiency.

    KK

    Kurt Keutzer

    Professor at UC Berkeley; expert in deep learning systems.

    CX

    Chenfeng Xu

    Researcher at UC Berkeley; focused on efficient ML.

    Executive Team

    SY

    Shuo Yang

    Lead Researcher

    PhD candidate at Berkeley Sky Lab.

    IS

    Ion Stoica

    Faculty Advisor

    Professor at UC Berkeley; co-founder of Databricks.

    Founding Story

    Developed as a collaboration between researchers at UC Berkeley, MIT, and UT Austin to enable local inference of massive Mixture-of-Experts models on consumer GPUs.

    Business Model

    Revenue Model

    Open-source software; free for community use.

    Pricing Tiers

    Community
    Free

    Full access to the open-source code and application.

    Target Markets

    Industries & Segments
    • AI Developers
    • Researchers
    • Privacy-conscious LLM users
    Use Cases
    • Local LLM hosting
    • Private AI inference
    • Developer workstations

    Quick Facts

    Headquarters
    Berkeley, California
    Founded
    2026
    Entity Type
    Open Source Project / Research Organization
    Employees
    11
    Office Locations
    Berkeley
    Cambridge
    Austin

    History & Milestones

    2026-08-16

    FreeToken officially launched and open-sourced on GitHub.

    2026-08-24

    Released technical paper 'FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution'.

    Key Capabilities

    4
    Bandwidth-adaptive CPU-GPU execution
    Dynamic LRU expert cache
    3-4x faster decode than Ollama
    Native Windows and Linux support

    Integrations & Partnerships

    Platform Integrations

    • Ubuntu
    • Arch Linux
    • Windows
    • CLI

    Connect

    Website
    flashml.ai/
    GitHub
    FlashML-org
    Discord
    xzwSnMdsX
    Slack
    slack.com/oauth/flashml

    AI Topics

    3

    FlashML-org focuses on these topics:

    Local Inference(1)
    LLM Orchestration(1)
    AI Infrastructure(1)
    Back to all developersSuggest an edit