搜索:GPTQ

共命中 4 条(服务端检索)
GGUF vs GPTQ vs AWQ vs EXL2: LLM Model Formats Explained (2026)
GGUF, GPTQ, AWQ, EXL2, and EXL3 solve the same problem in different ways. This guide separates file containers from quan…
大模型 MarkTechPost · 9-19 阅读 41·访客 38
模型量化的精度账:从 FP16 到 INT4 该怎么选
从 FP16 压到 INT4,体积与显存降到约四分之一,量化是大模型部署的标配。本文讲清量化原理、PTQ 与 QAT 两条路线、RTN/GPTQ/AWQ 的思想差异,以及按显存逐级下降的选型决策表。
原创 大模型 精选 · 原创 · 2天前 阅读 1·访客 1
G^2PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large lan…
大模型 HuggingFace Daily Papers · 9-25 阅读 9·访客 9
Softmax Reparameterization for Output-Head Quantization
Large vocabularies make output heads a substantial inference cost in small language models. We propose softmax reparamet…
行业动态 HuggingFace Daily Papers · 9-25 阅读 3·访客 3