Qwen3-14B at 0.55 bpw: training run in progress. Updates on X

[ Sub-1-bit LLM compression ]

Big models.
Little bits.

Qwen3 compressed to 0.55 bits per weight with LittleBit-2. Same architecture, binarized latent factors, recovered by distillation from the original model.

Bits per weight
0.55
Linear layers, vs BF16
~29× smaller
Inference overhead from ITQ
0

// The method

Compression that keeps the model's shape

W
≈
sign(U)
h·g·ℓ
sign(V)ᵀ

Dense FP16 weight → low-rank factors of ±1, plus three thin scale vectors.

LittleBit-2 · ICML 2026

Latent factorization, binarized

Each weight matrix is split into low-rank factors with SVD. The factors are binarized to ±1, and lightweight learned scales restore the magnitude. LittleBit-2 adds Joint-ITQ: a rotation that lines the latent factors up with the binary hypercube before training starts, so less signal is lost at the sign step.

The rotation folds into the factors. At inference the layer is exactly the original LittleBit layer, with no extra compute.

// Pipeline

From Qwen3 to 0.55 bits

Five steps, one script. Every flag below is a default in train.sh.

  1. 01

    Latent factorization

    SVD splits every linear layer into rank-r factors sized to hit the target bit budget. Embeddings and lm_head stay BF16.

    --quant_mod LittleBitLinear --eff_bit 0.55
  2. 02

    Joint-ITQ rotation

    An internal rotation aligns the factors with the binary hypercube. This is the LittleBit-2 initialization.

    --use_itq True
  3. 03

    SmoothSign binarization

    Factors become ±1 with a smooth surrogate gradient, so the sign step stays trainable.

    --quant_func SmoothSign
  4. 04

    Residual compensation

    A second binarized path learns what the first one missed.

    --residual True
  5. 05

    Distillation from the original

    Quantization-aware training on C4 and WikiText-2 at 2048 tokens. The BF16 model is the teacher, kept on GPU for speed.

    --l2l_loss_scale 10.0 --teacher_offload False

// Releases

The models

Community releases on Hugging Face. Eval numbers are published with each model card once the run finishes.

Pilot1× H100

littlebit-qwen3-0.6b

Pipeline check and throughput measurement. Short run, 0.05 epoch.

Base
Qwen3-0.6B
Target
0.55 bpw
Training8× H100 SXM

littlebit-qwen3-8b-0.55bpw

First full run at one epoch. Gate for the main release.

Base
Qwen3-8B
Target
0.55 bpw
Main release8× H100 SXM

littlebit-qwen3-14b-0.55bpw

Five epochs, matching the paper's recipe.

Base
Qwen3-14B
Target
0.55 bpw

Results

WikiText-2 / C4 perplexity · zero-shot accuracy
MetricQwen3-14B (BF16)LittleBit 0.55 bpw
WikiText-2 PPL ↓pendingpending
C4 PPL ↓pendingpending
ARC-Easypendingpending
ARC-Challengependingpending
HellaSwagpendingpending
PIQApendingpending
WinoGrandependingpending

// How small

Pick a model. Pick a budget.

Base model
Bits per weight

Linear layers only. Embeddings and lm_head stay BF16 and add a fixed amount on top. Scale vectors are not counted.

BF16
–
LittleBit
–
– smaller linear layers, – weights

// Get started

Download it. Evaluate it. Reproduce it.

Download

Weights and model card on Hugging Face. Load them with the LittleBit loader.

git clone <littlebit-repo>
Hugging Face · coming soon

Evaluate

Perplexity and zero-shot tasks against the BF16 baseline, from a local checkpoint or the Hub.

python eval.py --model_id <repo> \
  --ppl_task wikitext2,c4
Eval docs · coming soon

Reproduce

The full RunPod recipe: setup, train, eval, upload. Override anything with env vars.

MODEL_ID=Qwen/Qwen3-14B EPOCHS=5 \
  bash train.sh
See the pipeline →

// Built on

Standing on published work

Independent community release. Not an official Samsung Research model.

The little bits
under big models