JustVugg
To democratize access to frontier-class AI models by enabling them to run on consumer-grade hardware through optimized inference systems.
At a Glance
- AI Researchers
- Individual Developers
- Privacy-focused organizations
- Hardware enthusiasts
AI Tools by JustVugg
(1)colibri
Open Source MoE Inference Engine
Discussions
No discussions yet
Be the first to start a discussion about JustVugg
Latest News
Colibri v1.7.0 Released: Qwen3.6 support and performance optimizations
Critical Security Release: Six memory-safety issues patched in v1.6.2
DeepSeek V4 Flash support landed in v1.5.0
Kimi K3 (2.8T) and Inkling models now run on Colibri v1.3.0
Products & Services
A tiny inference engine for running large MoE models (744B to 2.8T) on consumer hardware by treating storage, RAM, and VRAM as a single hierarchy.
Deterministic, dependency-free memory for AI agents where memory lives in a Markdown file.
Open-source Durable Objects for live applications.
GPT-2-style LLM built from scratch in C/CUDA with hand-written backprop and FlashAttention.
Market Position
Positions itself as a leaner, MoE-specialized alternative to llama.cpp and vLLM, specifically targeting models that are significantly larger than the available RAM.
Leadership
Founders
Vincenzo Fornaro
AI Engineer and high-performance systems developer. Creator of the Colibri inference engine. Previously AI Engineer at SELEA.
Executive Team
Vincenzo Fornaro
Founder & Lead Developer
Expert in AI inference, computer vision, and systems programming in C, C++, and Go.
Founding Story
Started as a solo project by Vincenzo Fornaro on a 12-core laptop with 25GB of RAM to prove that massive models like GLM-5.2 could run locally without expensive datacenters.
Business Model
Revenue Model
Open source (Apache 2.0). Sponsorships and community donations.
Pricing Tiers
Full access to source code and prebuilt binaries on GitHub.
Target Markets
- AI Researchers
- Individual Developers
- Privacy-focused organizations
- Hardware enthusiasts
- Running frontier LLMs on consumer hardware
- Local and private AI inference
- Systems research for AI performance
- Accessibility of massive MoE models
- Community of contributors and testers across various hardware configurations.