raghu@dark-factory :~/kb/model-pruning $ cat

Model pruning

Removing redundant individual weights or whole structural blocks from a trained network to shrink its stored size and speed inference on constrained or GPU-less hardware, ideally with negligible accuracy loss.

grounded in: latest AI/tech trends theme 'Local & efficient LLM inference: running big models cheaply on old or GPU-less hardware and storing weights efficiently'

Connected concepts

Quantization, Knowledge distillation, 1-bit LLMs, Mixture-of-Experts, Memory-bandwidth-bound decode

Explore it live in the knowledge graph →