Pipeline parallelism
A distributed-inference strategy that splits a model's layers into contiguous stages placed on different devices, so activations flow stage-to-stage like an assembly line rather than sharding each layer's tensors.
grounded in: HN NOTABLE 'Mesh LLM: Distributed AI inference layered' + the 'Minimal & distributed AI agents / spreading inference across machines' trend theme (2026-07-12)
Connected concepts
Tensor Parallelism, Distributed inference, Decentralized mesh inference, 2-node NVIDIA GB10 (DGX Spark)
Explore it live in the knowledge graph →