Quantization
Reducing model weights to lower-precision integers so frontier LLMs fit and run on modest local hardware without a cloud API.
grounded in: Trend theme 'Local LLM inference on modest hardware — squeezing frontier models onto slow/small machines' + the GLM 5.2 local-run walkthrough on the HN front page (2026-07-11).
Connected concepts
vLLM / PagedAttention, Local-first inference, 2-node NVIDIA GB10 (DGX Spark)
Explore it live in the knowledge graph →