Quantization Explained: Fitting Giant LLMs onto Small GPUs

Abstract visualization of a large microchip dissolving into compact glowing data cubes that flow into a small GPU chip, representing neural network quantization

A 70-billion-parameter language model in 16-bit precision needs about 140 GB of GPU memory just to hold its weights — more than five RTX 4090s. Yet people routinely run 70B models on two 24 GB cards, and 7B models on a laptop. The trick is quantization: storing the model's numbers in fewer bits. This article explains what quantization actually does to a neural network, why naive rounding destroys quality, how modern methods like GPTQ and AWQ decide which precision each weight deserves, and what breaks when you push to 4 bits and below.

The memory wall

Inference for large language models is memory-bound, not compute-bound. Generating each token requires streaming every weight from GPU memory into the processor once. A 70B model at 16-bit precision (FP16 or BF16) is 70 × 109 parameters × 2 bytes ≈ 140 GB. An NVIDIA RTX 4090 has 24 GB. The arithmetic is unforgiving: the weights alone do not fit, and that is before activations and the KV cache.

Quantization attacks the “bytes per parameter” term directly. Drop from 16 bits to 8 and the same model needs ~70 GB. Drop to 4 bits and it needs ~35 GB — suddenly two consumer GPUs, or one high-end card, can hold it. Throughput improves too: moving half the bytes per token roughly doubles tokens per second, since memory bandwidth was the bottleneck. The question is entirely about what the rounding costs in quality, and the last few years of research have been about making that cost small.

What quantization actually does

At its core, quantization maps a continuous range of floating-point values onto a small grid of integers. The standard scheme is affine (uniform) quantization: each block of values gets a scale s and a zero-point z, and the mapping is:

x_quant = round(x / s) + z, with dequantization x_hat = s · (x_quant − z)

A concrete example: suppose a group of weights ranges from −1.0 to 1.0 and we quantize to 8-bit integers (256 levels). Then s = 2.0 / 255 ≈ 0.00784. A weight of 0.5 becomes round(0.5 / 0.00784) = 64, and dequantizes to 64 × 0.00784 ≈ 0.502. The error is at most half a step — tiny. So far, so harmless.

The catch is the word group. One scale for the whole model would be catastrophic, because different layers — and different channels within a layer — live on wildly different scales. Practical schemes quantize in small groups (per-channel, or blocks of 32–128 weights), each with its own scale. The scales themselves are stored in higher precision; their overhead is small because they are shared across many weights. This is why you see formats named like Q4_K_M in llama.cpp: they encode exactly how weights are grouped and how many bits each group gets.

Why naive rounding breaks: the outlier problem

If you round every weight to 8 bits with one scale per tensor, model quality falls off a cliff. The reason was characterized precisely by Dettmers and colleagues in LLM.int8() (2022): a tiny fraction of feature dimensions in large transformers develop outlier activations — values tens or hundreds of times larger than the rest.

Think about what one shared scale does to a group containing an outlier. The scale must stretch to cover the outlier's magnitude, so the 256 grid levels get spread over a huge range, and all the normal values collapse into a handful of bins near zero. Their fine structure — which is where most of the model's knowledge lives — is destroyed. It is like choosing the zoom level on a map to fit the whole country when you needed street detail.

LLM.int8()'s fix is conceptually clean: find the outlier feature dimensions (typically under 1% of them), keep those in 16-bit precision, and quantize everything else to 8 bits. The outliers are few enough that the memory cost is negligible, but protecting them preserves nearly all the quality. The deeper lesson stuck: not all weights and activations matter equally, and good quantization is about spending bits where they matter.

The precision ladder: 16 → 8 → 4

Each step down the ladder roughly halves the memory, with a different method suited to each rung:

  • FP16/BF16 → INT8 (weights and activations). With outlier-aware handling (LLM.int8(), or SmoothQuant, which migrates quantization difficulty from activations into weights where it is easier to absorb), 8-bit inference is nearly lossless. Memory halves; on supported hardware, throughput roughly doubles.
  • INT8 → INT4 (weights only). This is the sweet spot for fitting big models on small GPUs. GPTQ (Frantar et al., 2022) quantizes weights one layer at a time, using approximate second-order information — essentially the curvature of the loss — to decide the rounding order and to compensate: when one weight is rounded down, remaining weights are nudged to absorb the error. AWQ (Lin et al., 2023) takes a different route: it observes that a small set of salient weights (about 1%) matters disproportionately, and scales up their channels before quantizing so the rounding error lands where it hurts least.

The memory math for a 70B model: ~140 GB at 16-bit, ~70 GB at 8-bit, ~35–40 GB at 4-bit (the extra few GB are scales, zero-points, and overhead). Quality-wise, 4-bit weight quantization on a 70B model typically costs a small, measurable degradation — visible on perplexity benchmarks but often imperceptible in chat. The degradation grows as models get smaller: a 7B model has less redundancy to spare, so 4-bit hurts it more than it hurts a 70B model. Below 4 bits (3-bit, 2-bit, 1-bit research like BitNet), quality falls faster and the methods get exotic.

Weights are the easy part

Everything above quantizes weights, which are static — you can calibrate them once, offline, with a small dataset, and bake the result into the model file. Activations are harder: they depend on the input, change every forward pass, and contain the outliers. That is why the standard recipe for consumer GPUs is weight-only quantization: weights in 4-bit, activations computed in 16-bit. The weights are dequantized on the fly in small tiles as the matrix multiplication runs, so the big memory saving is kept while the arithmetic stays precise.

The third memory consumer is the KV cache, which grows with context length: every generated token's keys and values are stored for all future tokens. At 128K context, the KV cache of a 70B model can rival the weights in size. KV-cache quantization (methods like KIVI) compresses these cached values to 4 or even 2 bits, which is what makes very long contexts feasible on limited hardware. It is a separate axis from weight quantization, and the two compose.

What breaks, and how to tell

Quantization error is not uniform across capabilities. Perplexity — the standard language-modelling metric — usually degrades gracefully, but downstream behaviour can have sharper edges: arithmetic reasoning and precise factual recall tend to suffer before fluent chat does. A quantized model may write beautifully while quietly getting worse at multi-step math. This is why serious evaluations test task accuracy (MMLU, GSM8K, HumanEval), not just perplexity.

Calibration data matters more than people expect. Methods like GPTQ and AWQ need a small sample of text to measure which weights are salient and how activations behave. Calibrate on the wrong distribution — say, code when you will serve chat — and quality drops. The practical rule: calibrate on data resembling your actual workload.

Putting it together: 70B on two gaming GPUs

A worked example. Take Llama 3.1 70B, quantize weights to 4-bit with a group size of 128:

  • Weights: 70 × 109 × 0.5 bytes ≈ 35 GB, plus ~3–5 GB of scales and overhead.
  • Two RTX 4090s (24 GB each) hold this comfortably with room for activations and a modest KV cache.
  • Expected quality: within a few percent of the 16-bit model on standard benchmarks — often indistinguishable in interactive use.

The tooling has commoditized this. llama.cpp (GGUF formats) runs quantized models on CPU and GPU with no Python stack; Ollama wraps it in a one-command server; Hugging Face Transformers integrates bitsandbytes for 8-bit and 4-bit loading in a few lines; vLLM serves quantized models at high throughput with paged KV-cache management. The pipeline from “download weights” to “chat with 70B locally” is now an afternoon project.

Further reading

  • Dettmers, T., et al. (2022). LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale. arXiv:2208.07339
  • Frantar, E., et al. (2022). GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers. arXiv:2210.17323
  • Lin, J., et al. (2023). AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration. arXiv:2306.00978
  • Xiao, G., et al. (2023). SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models. arXiv:2211.10438

Similar Posts

Leave a Reply