raghu@dark-factory :~/kb/distributed-inference $ cat

Distributed inference

Splitting a model's serving across multiple nodes (via tensor/pipeline parallelism) so a cluster can run models too large for one machine.

grounded in: The system's real 'spark' telemetry is a 2-node GB10 cluster (recent commits 'spark: 2-node GB10 telemetry refresh'), and cross-node model serving is the durable technique that makes such a cluster us

Connected concepts

vLLM / PagedAttention, 2-node NVIDIA GB10 (DGX Spark), Local-first inference

Explore it live in the knowledge graph →