IT lexicon AI & ML Quantization

Quantization

AI & ML På svenska → Updated: 2026-05-23

Shrink an ML model by storing weights in fewer bits — 32-bit floats → 8-bit ints, or even 4-bit.

4x smaller memory, 4x faster inference, with small quality loss. Makes running 70B models on a gaming GPU possible. The backstory of GGUF, AWQ, GPTQ.

← Back to the lexicon