First release is live: littlebit-qwen3-4b on Hugging Face. Sub-1-bit builds in progress

[ Sub-1-bit LLM compression ]

Big models.
Little bits.

We are compressing Qwen3 to 0.55 bits per weight with LittleBit-2: same architecture, binarized latent factors, recovered by distillation from the original model. Our first release, Qwen3-4B for your laptop, is out now.

Target bits per weight
0.55
Linear layers, vs BF16
~29× smaller
Inference overhead from ITQ
0

// The method

Compression that keeps the model's shape

W
≈
sign(U)
h·g·ℓ
sign(V)ᵀ

Dense FP16 weight → low-rank factors of ±1, plus three thin scale vectors.

LittleBit-2 · ICML 2026

Latent factorization, binarized

Each weight matrix is split into low-rank factors with SVD. The factors are binarized to ±1, and lightweight learned scales restore the magnitude. LittleBit-2 adds Joint-ITQ: a rotation that lines the latent factors up with the binary hypercube before training starts, so less signal is lost at the sign step.

The rotation folds into the factors. At inference the layer is exactly the original LittleBit layer, with no extra compute.

// Pipeline

From Qwen3 to 0.55 bits

Five steps, one script. Every flag below is a default in train.sh.

  1. 01

    Latent factorization

    SVD splits every linear layer into rank-r factors sized to hit the target bit budget. Embeddings and lm_head stay BF16.

    --quant_mod LittleBitLinear --eff_bit 0.55
  2. 02

    Joint-ITQ rotation

    An internal rotation aligns the factors with the binary hypercube. This is the LittleBit-2 initialization.

    --use_itq True
  3. 03

    SmoothSign binarization

    Factors become ±1 with a smooth surrogate gradient, so the sign step stays trainable.

    --quant_func SmoothSign
  4. 04

    Residual compensation

    A second binarized path learns what the first one missed.

    --residual True
  5. 05

    Distillation from the original

    Quantization-aware training on C4 and WikiText-2 at 2048 tokens. The BF16 model is the teacher, kept on GPU for speed.

    --l2l_loss_scale 10.0 --teacher_offload False

// Releases

The models

Community releases on Hugging Face. Each model card says exactly how the model was built. Eval numbers are added when a run finishes.

AvailableMLX · GGUF

littlebit-qwen3-4b

Qwen3-4B at 4 bits for Apple Silicon (MLX) and llama.cpp / Ollama (GGUF). Standard 4-bit quantization, not sub-1-bit. 2.1 GB, ~74 tokens/s on an M1 Max.

Base
Qwen3-4B
Bits
4.5 bpw
MLX on Hugging Face → GGUF on Hugging Face →
Next1× H200

littlebit-qwen3-8b-0.55bpw

First sub-1-bit release, one epoch of distillation. Throughput measured: 2.84 s per step, about 16.5 hours per epoch.

Base
Qwen3-8B
Target
0.55 bpw
Plannedmulti-GPU

littlebit-qwen3-14b-0.55bpw

Five epochs, matching the paper's recipe.

Base
Qwen3-14B
Target
0.55 bpw

Results

WikiText-2 / C4 perplexity · zero-shot accuracy
MetricQwen3-8B (BF16)LittleBit 0.55 bpw
WikiText-2 PPL ↓pendingpending
C4 PPL ↓pendingpending
ARC-Easypendingpending
ARC-Challengependingpending
HellaSwagpendingpending
PIQApendingpending
WinoGrandependingpending

// How small

Pick a model. Pick a budget.

Base model
Bits per weight

Linear layers only. Embeddings and lm_head stay BF16 and add a fixed amount on top. Scale vectors are not counted.

BF16
–
LittleBit
–
– smaller linear layers, – weights

// Get started

Download it. Evaluate it. Reproduce it.

Download

Run the first release locally. MLX on a Mac, or GGUF with llama.cpp or Ollama.

mlx_lm.generate --model \
  LittleBitLLM/littlebit-qwen3-4b-mlx-4bit
Hugging Face →

Evaluate

Perplexity and zero-shot tasks against the BF16 baseline, from a local checkpoint or the Hub.

python eval.py --model_id <repo> \
  --ppl_task wikitext2,c4
Eval docs · coming soon

Reproduce

The full RunPod recipe: setup, train, eval, upload. Override anything with env vars.

MODEL_ID=Qwen/Qwen3-14B EPOCHS=5 \
  bash train.sh
Get the recipe →

// Built on

Standing on published work

Independent community release. Not an official Samsung Research model.

The little bits
under big models