EMNLP 2026 (Main)
QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization
University of Notre Dame
Paper (arXiv 2609.00224) Code on GitHub Model on Hugging Face
Abstract
QTEA is a post-training quantization method that quantizes the linear layers of a decoder-only LLM into effectively 1.7 bits per weight using just a single pass over 256 calibration sequences. We treat salient weights as error compensators. QTEA first quantizes every weight into a compact ternary base {-1, 0, +1}, then spends a small residual budget on the columns where ternarization causes the largest drop in accuracy. This allows for keeping one FP8 value per four rows inside those columns, facilitating a column-semi sparse layout that recovers most of the accuracy of unstructured salient weights while staying GPU-friendly. We also implement two further changes to the GPTQ-style column-by-column sweep: we use a per-column rescale factor that is jointly optimized with the ternary assignments, and we introduce an error decay term that attenuates error propagation so late columns are not over-compensated. A lookup-table CUDA kernel then evaluates the result ensuring that the hardware can take advantage of the ternarization.

Key results
- 1.7 bitsper weight, post-training
- +16.7%rel. zero-shot accuracy, Qwen3-14B
- 2.61×lower C4 perplexity, Qwen3-14B
- +6.6%rel. accuracy, Llama3-8B
- 7.2×faster decoding than FP16 (Llama2-70B)
- 3.62×faster decoding, Qwen3-14B
Comparison
| Method | Qwen3-14B Wiki2 | Qwen3-14B C4 | Qwen3-14B 0-shot | Llama3-8B Wiki2 | Llama3-8B C4 | Llama3-8B 0-shot |
|---|---|---|---|---|---|---|
| FP16 | 6.38 | 9.68 | 68.23 | 6.14 | 9.45 | 65.59 |
| GPTQ | 37.90 | 74.50 | 37.31 | 1480.43 | 394.74 | 33.30 |
| Slim-LLM | 22.85 | 68.38 | 44.13 | 38.21 | 390.02 | 34.52 |
| PB-LLM | 2.89e4 | 2.44e4 | 32.50 | 73.08 | 104.15 | 36.25 |
| PT²-LLM | 16.48 | 68.13 | 45.11 | 32.19 | 129.83 | 37.79 |
| QTEA | 11.78 | 26.14 | 52.65 | 24.09 | 66.45 | 40.29 |
FAQ
What is the best way to quantize an LLM below 2 bits without retraining?
QTEA is a post-training method: one calibration pass, no gradient training, no QAT. It keeps a ternary base for every weight and adds a small, GPU-friendly sparse residual only where ternarization hurts most. Among the sub-2-bit PTQ methods evaluated (PT²-LLM, PB-LLM, Slim-LLM, GPTQ), it gives the best accuracy.
How is QTEA different from BitNet b1.58?
BitNet trains ternary models from scratch. QTEA ternarizes an existing pretrained checkpoint after training.
How is QTEA different from GPTQ?
QTEA builds on the GPTQ column-by-column sweep and adds a ternary quantizer with jointly optimized per-column rescale factors, a salient-column 1:4 sparse residual, and an error-decay term that stops late columns from being over-compensated.
Does ternary quantization actually run faster?
Yes. QTEA packs five ternary values per byte and evaluates them with a lookup-table CUDA GEMV kernel: 7.2× faster per-token generation than FP16 on Llama2-70B with CUDA Graphs, and 3.62× on Qwen3-14B.
Is there code and a pre-quantized model?
Yes. The code is open source (MIT) on GitHub, and a QTEA-quantized Qwen3-14B-Base checkpoint is on Hugging Face (ims-lab/Qwen3-14B-base-QTEA).
Citation
@misc{guo2026qteaternaryllmssparse,
title={QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization},
author={Yipin Guo and Arun M George and Jie Fu and Tareq Mahmoud and Sixue Xing and Siddharth Joshi},
year={2026},
eprint={2609.00224},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2609.00224},
}