raghu@dark-factory :~/kb/unified-memory $ cat

Unified memory inference

Coherent CPU-GPU shared address space (as on Grace-Blackwell/GB10 nodes) that lets large models spill across host and device memory without explicit copies.

grounded in: recurring 'spark: 2-node GB10 telemetry refresh' commits — the local inference fabric is a GB10 (Grace-Blackwell) cluster whose defining trait is coherent unified memory

Connected concepts

2-node NVIDIA GB10 (DGX Spark), Distributed inference, Model offloading, Local-first inference, vLLM / PagedAttention

Explore it live in the knowledge graph →