FlashML-org
Bring frontier-scale AI models to edge devices through efficient, bandwidth-adaptive inference engines.
At a Glance
- AI Developers
- Researchers
- Privacy-conscious LLM users
AI Tools by FlashML-org
(1)FreeToken
Edge MoE Inference Engine
Discussions
No discussions yet
Be the first to start a discussion about FlashML-org
Latest News
Products & Services
An edge-native MoE serving engine supporting large models like DeepSeek-V4 and GLM.
Market Position
FlashML's FreeToken provides superior performance for MoE models compared to Ollama and llama.cpp by optimizing for bandwidth-limited consumer hardware.
Leadership
Founders
Shuo Yang
PhD student at UC Berkeley Sky Lab; lead developer of FreeToken.
Ion Stoica
Professor at UC Berkeley; Co-founder of Databricks and Anyscale.
Matei Zaharia
Associate Professor at UC Berkeley; Co-founder of Databricks.
Song Han
Associate Professor at MIT; expert in AI hardware efficiency.
Kurt Keutzer
Professor at UC Berkeley; expert in deep learning systems.
Chenfeng Xu
Researcher at UC Berkeley; focused on efficient ML.
Executive Team
Shuo Yang
Lead Researcher
PhD candidate at Berkeley Sky Lab.
Ion Stoica
Faculty Advisor
Professor at UC Berkeley; co-founder of Databricks.
Founding Story
Developed as a collaboration between researchers at UC Berkeley, MIT, and UT Austin to enable local inference of massive Mixture-of-Experts models on consumer GPUs.
Business Model
Revenue Model
Open-source software; free for community use.
Pricing Tiers
Full access to the open-source code and application.
Target Markets
- AI Developers
- Researchers
- Privacy-conscious LLM users
- Local LLM hosting
- Private AI inference
- Developer workstations