Post-Training Ternarization of Qwen3 Language Models
Rotation, ternarization and error compensation on Qwen3-4B, end to end: capability, effective bits per weight, storage and inference, measured.
Technical report · August 2026 · Malik, Devan, MehraWhat we publish
Ternary models, learned quantization boundaries, and what it takes to keep reasoning intact at 1.58 bits.
Papers
Rotation, ternarization and error compensation on Qwen3-4B, end to end: capability, effective bits per weight, storage and inference, measured.
Technical report · August 2026 · Malik, Devan, MehraWhat survives when a pretrained 752M model is pushed to three states: not a uniformly weaker model, but a stratified one.
Technical paper · 2026 · Malik, Devan, MehraAttention stays at 16 bits, the MLPs go ternary, and a teacher restores what the cut removed.
Paper · 2026 · Malik, Vishnuprasad, Devan, MehraLearned quantization boundaries let a ternary network decide, per layer, where a weight becomes zero.
Paper · February 2026 · Malik, Mehra, Vishnuprasad, Pundir, Tyagi, ArutkeerthiWhy the next generation of models will run on the chips that already exist, and what that changes.
White paper · January 2026 · Malik, Mehra, Vishnuprasad, Pundir, TyagiIn one line
Every weight is −1, 0 or +1. The boundaries that decide which are learned, not fixed, so each layer chooses its own sparsity and the circuits that carry reasoning survive.
Full methodology, hardware and raw logs accompany every number we publish. A number without them does not ship.