EMNLP 2026 (Main)

QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization

Yipin Guo, Arun M George, Jie Fu, Tareq Mahmoud, Sixue Xing, Siddharth Joshi

University of Notre Dame

Abstract

QTEA is a post-training quantization method that quantizes the linear layers of a decoder-only LLM into effectively 1.7 bits per weight using just a single pass over 256 calibration sequences. We treat salient weights as error compensators. QTEA first quantizes every weight into a compact ternary base {-1, 0, +1}, then spends a small residual budget on the columns where ternarization causes the largest drop in accuracy. This allows for keeping one FP8 value per four rows inside those columns, facilitating a column-semi sparse layout that recovers most of the accuracy of unstructured salient weights while staying GPU-friendly. We also implement two further changes to the GPTQ-style column-by-column sweep: we use a per-column rescale factor that is jointly optimized with the ternary assignments, and we introduce an error decay term that attenuates error propagation so late columns are not over-compensated. A lookup-table CUDA kernel then evaluates the result ensuring that the hardware can take advantage of the ternarization.

QTEA method overview

Key results

Comparison

Sub-2-bit PTQ comparison (paper Table 1). PPL ↓, zero-shot average accuracy ↑.
MethodQwen3-14B Wiki2Qwen3-14B C4Qwen3-14B 0-shotLlama3-8B Wiki2Llama3-8B C4Llama3-8B 0-shot
FP166.389.6868.236.149.4565.59
GPTQ37.9074.5037.311480.43394.7433.30
Slim-LLM22.8568.3844.1338.21390.0234.52
PB-LLM2.89e42.44e432.5073.08104.1536.25
PT²-LLM16.4868.1345.1132.19129.8337.79
QTEA11.7826.1452.6524.0966.4540.29

FAQ

What is the best way to quantize an LLM below 2 bits without retraining?

QTEA is a post-training method: one calibration pass, no gradient training, no QAT. It keeps a ternary base for every weight and adds a small, GPU-friendly sparse residual only where ternarization hurts most. Among the sub-2-bit PTQ methods evaluated (PT²-LLM, PB-LLM, Slim-LLM, GPTQ), it gives the best accuracy.

How is QTEA different from BitNet b1.58?

BitNet trains ternary models from scratch. QTEA ternarizes an existing pretrained checkpoint after training.

How is QTEA different from GPTQ?

QTEA builds on the GPTQ column-by-column sweep and adds a ternary quantizer with jointly optimized per-column rescale factors, a salient-column 1:4 sparse residual, and an error-decay term that stops late columns from being over-compensated.

Does ternary quantization actually run faster?

Yes. QTEA packs five ternary values per byte and evaluates them with a lookup-table CUDA GEMV kernel: 7.2× faster per-token generation than FP16 on Llama2-70B with CUDA Graphs, and 3.62× on Qwen3-14B.

Is there code and a pre-quantized model?

Yes. The code is open source (MIT) on GitHub, and a QTEA-quantized Qwen3-14B-Base checkpoint is on Hugging Face (ims-lab/Qwen3-14B-base-QTEA).

Citation

@misc{guo2026qteaternaryllmssparse,
  title={QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization},
  author={Yipin Guo and Arun M George and Jie Fu and Tareq Mahmoud and Sixue Xing and Siddharth Joshi},
  year={2026},
  eprint={2609.00224},
  archivePrefix={arXiv},
  primaryClass={cs.LG},
  url={https://arxiv.org/abs/2609.00224},
}