Skip to content
AI & Automation

LLM Quantization: How to Make Language Models Smaller and Faster

How LLM quantization works: precision formats from FP16 to 4-bit, weight and activation quantization, methods such as GPTQ and AWQ, formats such as GGUF, KV cache quantization, quality trade-offs, hardware compatibility, evaluation and deployment.

Quick answer

Quantization stores model weights, and optionally activations and the KV cache, in fewer bits so models need less memory and bandwidth and often run faster. BF16 or FP16 is the usual baseline; FP8 and INT8 typically cost little quality on supported hardware; 4-bit formats such as GPTQ, AWQ, GGUF variants and newer 4-bit floating-point formats save more but need careful evaluation. Choose formats your hardware and serving engine accelerate, calibrate with representative data and compare quality, latency and cost against the unquantized model on your own tasks.

Where This Fits

Quantization is one lever in AI inference optimization. Hardware considerations are in GPU optimization for AI, local and device deployment in AI edge deployment and running open-weight models in LLM self-hosting.

Why Quantization Helps

A model with 8 billion parameters needs roughly 16 GB just for weights in 16-bit precision, about 8 GB in 8-bit and about 4 to 5 GB in 4-bit formats including overhead. Smaller weights fit on smaller or fewer GPUs, leave more memory for the KV cache (and therefore more concurrent users or longer contexts) and reduce the bytes read per generated token, which speeds up memory-bound decoding.

Precision Formats

FormatBits per weightTypical useNotes
BF16 / FP1616Baseline serving and trainingReference quality
FP88Weights and activations on newer GPUsOften small quality impact; needs hardware support for speed
INT88Weights, sometimes activationsWidely supported
INT4 (GPTQ, AWQ)4Weight-only for GPU servingLarger memory savings; evaluate quality
4-bit floating point (NVFP4, MXFP4)4Newest GPU architecturesHardware-specific
GGUF variants2 to 8llama.cpp, CPUs, laptops, edgeMany levels; lower levels lose more quality

Methods

Post-training quantization converts a trained model without retraining, using a small calibration dataset. GPTQ quantizes weights layer by layer to minimize output error; AWQ scales weights to protect those most important to activations. Quantization-aware training simulates low precision during training or fine-tuning for better quality at low bit widths, at higher cost. QLoRA fine-tunes adapters on top of a 4-bit base model, reducing fine-tuning memory.

Serving engines support many formats; vLLM, for example, lists FP8, MXFP8/MXFP4, NVFP4, INT8, INT4, GPTQ/AWQ, GGUF and others. Check that your engine accelerates, rather than merely loads, the format you choose.

Want to run larger models on smaller hardware?

ZSpace Labs evaluates quantization options against your quality and latency targets. See AI infrastructure services.

Start a Project

Quality Trade-Offs

Quality loss is uneven. Quantized models may perform nearly identically on common tasks but degrade on reasoning, long contexts, code, maths, less common languages or strict output formats. Smaller models tend to be more sensitive than larger ones. Calibration data matters: calibrate on text resembling your workload. Never rely on a single benchmark score.

Evaluation on your own tasks decides whether a quantized model is good enough.

KV Cache Quantization

For long contexts and high concurrency, the KV cache can use more memory than the weights. Some engines support storing it in FP8 or other reduced formats, increasing the number of tokens that fit. Evaluate long-context tasks specifically, since errors can accumulate over long sequences.

Hardware Compatibility

Speed gains require hardware and kernel support. FP8 compute is available on recent data-centre GPU generations; 4-bit floating-point formats need the newest architectures; integer formats have broad support through optimized kernels. On CPUs and Apple silicon, llama.cpp and GGUF are common. On phones and embedded devices, mobile runtimes have their own supported formats; see AI edge deployment.

Evaluating a Quantized Model

  • Run your application's evaluation set on both baseline and quantized models
  • Check segments: long inputs, languages, structured outputs, reasoning tasks
  • Measure memory, time to first token, tokens per second and throughput
  • Test at realistic concurrency, not single requests
  • Compare cost per successful task, not just per token
  • Keep the baseline available for rollback

Advantages and Limitations

Quantization often makes self-hosting and edge deployment practical and can reduce serving costs considerably. It can quietly reduce quality on specific tasks, gains depend on hardware support and quantized community models vary in reliability. Treat it like a model change: evaluate, roll out gradually and monitor.

How to Quantize Step by Step

  • 1. Define quality and performance targets
  • 2. Pick formats your hardware and engine accelerate
  • 3. Prefer official quantized releases or quantize with representative calibration data
  • 4. Evaluate on your tasks against the baseline
  • 5. Benchmark memory, latency and throughput at realistic load
  • 6. Roll out gradually with monitoring
  • 7. Re-evaluate when models, engines or hardware change

Quantization for Fine-Tuning

Quantization also changes fine-tuning economics. QLoRA, described in Dettmers et al., fine-tunes low-rank adapters on top of a 4-bit quantized base model, making it possible to adapt larger models on a single GPU. Quality of the resulting model should be evaluated against your tasks like any fine-tune. When deploying, you can serve the adapter on a quantized or full-precision base, and results can differ slightly, so evaluate the exact serving configuration.

Quantization on Edge Devices

On phones, laptops and embedded hardware, quantization is often required rather than optional, because memory and power are tight. Mobile and edge runtimes support specific integer and floating-point formats, and neural processing units may only accelerate certain operations and bit widths. Test the exact format on target devices for accuracy, latency, memory and battery or thermal behaviour under sustained use. See AI edge deployment.

Choosing a Quantization Level

SituationReasonable starting point
Production GPU serving, quality-sensitiveFP8 or INT8 where hardware supports it, evaluate
Model does not fit at 8-bit4-bit weight-only (GPTQ or AWQ), evaluate carefully
Laptops and CPUsGGUF at a mid-level quantization, test lower levels
Phones and embedded devicesFormats supported by the device runtime and accelerator
Long contexts at high concurrencyConsider KV cache quantization too

Worked Example

An illustrative scenario, not a client case: a team wants to self-host a mid-sized open-weight model on GPUs that cannot fit it in 16-bit at the required concurrency. An FP8 version meets quality targets on their evaluation set and fits with room for KV cache. A 4-bit version uses even less memory but drops on structured extraction tasks, so they deploy FP8 for production and use the 4-bit version only for internal experimentation.

Common Mistakes

  • Choosing the lowest bit width without evaluation
  • Formats the hardware cannot accelerate
  • Calibration data unlike the real workload
  • Testing single requests instead of realistic load
  • Downloading quantized models from unverified sources

Considering quantized models for production?

Talk to ZSpace Labs about a quantization evaluation on your tasks and hardware.

Start a Project

Conclusion

Quantization trades bits for efficiency. Choose formats your hardware accelerates, calibrate on realistic data and let evaluation on your own tasks decide how far to go.

FAQ

Common questions

Representing a model's weights, and sometimes activations or KV cache, with fewer bits than the original format, for example 8 or 4 bits instead of 16, to reduce memory and often increase speed.

Related services
Relevant industries
Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.

Keep exploring
AI & Automation
8 min read

LLM Self-Hosting: How to Run Open-Weight Models on Your Own Infrastructure

How to self-host open-weight language models: when it makes sense, open-weight vs open-source, licences, hardware selection, serving software, security, scaling, monitoring, maintenance and total cost of ownership compared with hosted APIs.

Read article
AI & Automation
7 min read

GPU Optimization for AI: How to Use Compute Resources Efficiently

How to use GPUs efficiently for AI workloads: understanding memory and utilization, batching, parallelism, precision, scheduling and sharing, profiling, workload placement and right-sizing, without relying on universal performance claims.

Read article
AI & Automation
7 min read

AI Edge Deployment: How to Run AI Models on Local and Edge Devices

How to deploy AI models on edge and local devices: edge hardware options, local inference runtimes, connectivity, privacy and latency benefits, model compression, fleet updates, monitoring and synchronization with cloud systems.

Read article