Model offloading
Fitting a model too large for available accelerator memory by spilling layers or the KV cache to CPU RAM, disk, or a second node so it still runs on constrained hardware.
grounded in: Trend 'GLM 5.2 local run: getting a large model running on underpowered hardware' + 'Local LLM inference on modest hardware' theme (2026-07-11 HN feed)
Connected concepts
Quantization, Local-first inference, Distributed inference, vLLM / PagedAttention
Explore it live in the knowledge graph →