Open-weight LLMs
Models with publicly released weights (e.g. the GLM family) that can be self-hosted and run locally rather than called through a proprietary cloud API.
grounded in: Trend theme 'Local LLM inference on modest hardware' + NOTABLE 'GLM 5.2 local run', grounded in this system's local vLLM serving on the GB10 Spark cluster (spark.sh, spark_stats.py)
Connected concepts
vLLM / PagedAttention, Local-first inference, Quantization, Mixture-of-Experts, 2-node NVIDIA GB10 (DGX Spark)
Explore it live in the knowledge graph →