Model pruning
Removing redundant individual weights or whole structural blocks from a trained network to shrink its stored size and speed inference on constrained or GPU-less hardware, ideally with negligible accuracy loss.
grounded in: latest AI/tech trends theme 'Local & efficient LLM inference: running big models cheaply on old or GPU-less hardware and storing weights efficiently'
Connected concepts
Quantization, Knowledge distillation, 1-bit LLMs, Mixture-of-Experts, Memory-bandwidth-bound decode
Explore it live in the knowledge graph →