Memory-mapped weights
Loading model weights via mmap so the OS page cache streams them from disk on demand, letting a machine run models larger than its physical RAM.
grounded in: latest AI/tech trends theme 'Local/low-end inference: Enthusiasm is high for running large models cheaply on old, GPU-less hardware'
Connected concepts
Model offloading, CPU inference serving, Memory-bandwidth-bound decode, Second-Life Hardware, Local-first inference
Explore it live in the knowledge graph →