raghu@dark-factory :~/kb/memory-bandwidth-bound $ cat

Memory-bandwidth-bound decode

LLM token generation is limited by how fast model weights stream from memory rather than by raw compute, so high-bandwidth unified memory is what lets large models run locally.

grounded in: Trend: 'Qwen3.5-122B on Mac Studio' and 'large models running on personal hardware' — the enabling fact is Apple Silicon's high memory bandwidth, an established inference property.

Connected concepts

Unified memory inference, CPU inference serving, Quantization, Model offloading, Inference economics

Explore it live in the knowledge graph →