LittleBit
Official implementation of LittleBit and LittleBit-2, methods for compressing large language models to sub-1-bit weights via latent factorization.
At a Glance
About LittleBit
LittleBit is the official research implementation from SamsungLabs of LittleBit (NeurIPS 2025) and its follow-up LittleBit-2 (ICML 2026). It compresses large language models into the sub-1-bit regime and provides training and evaluation scripts in Python.
What It Is
LittleBit is a model-compression codebase for LLM researchers and practitioners. It factorizes each dense weight matrix into low-rank latent factors, binarizes those factors, and restores magnitude information through lightweight learned scales. This enables compression down to 0.1 bits per weight while keeping the original model architecture at inference time.
Training and Evaluation Workflow
Models are compressed through Quantization-Aware Training, run on a single GPU or across multiple GPUs with DeepSpeed. Training uses SmoothSign quantization with the LittleBitLinear module and optional residual factorization. An evaluation script scores local checkpoints or Hugging Face Hub models on perplexity tasks (wikitext2, c4) and zero-shot tasks such as boolq, piqa, hellaswag, winogrande, arc_easy, arc_challenge and openbookqa. Older checkpoints without a littlebit_config.json can be evaluated by passing quantization arguments explicitly.
LittleBit-2 Initialization
LittleBit-2 addresses latent geometry misalignment at initialization. It applies Internal Latent Rotation with Joint Iterative Quantization (Joint-ITQ) to align SVD-derived latent factors with the binary hypercube before QAT. It is opt-in via the --use_itq flag and changes only initialization, so the deployed factorized layer is unchanged.
Supported Models and Setup
Supported model families are OPT, Llama (including Llama 2/3), Phi-4, Qwen2.5 and QwQ, Gemma 2 and 3, and Qwen3. The README recommends Python 3.12, CUDA 12.4 with PyTorch 2.8.0, and transformers 4.51.x to reproduce paper results. The code is released under the CC BY-NC 4.0 license, which restricts use to non-commercial purposes.
Community Discussions
Be the first to start a conversation about LittleBit
Share your experience with LittleBit, ask questions, or help others learn from your insights.
Pricing
Free
Capabilities
Key Features
- Sub-1-bit LLM compression from 1.0 down to 0.1 bits per weight
- Latent factorization of weight matrices with binarized factors and learned scales
- LittleBit-2 Joint-ITQ initialization via --use_itq
- Quantization-Aware Training with SmoothSign and optional residual factorization
- Single-GPU and multi-GPU DeepSpeed training
- Evaluation on perplexity and zero-shot tasks
- Evaluation of local or Hugging Face Hub checkpoints
- Support for legacy checkpoints via explicit quantization arguments
- No inference-time architecture change
