Small language model
A compact, few-billion-parameter model designed to run efficiently on-device or on modest local hardware rather than in the cloud, trading raw scale for latency, privacy, and cost.
grounded in: Trend theme 'On-device / local AI vs cloud' with Apple SpeechAnalyzer running a native on-device speech/ML model instead of a cloud API; this system already runs local vLLM inference on GB10 nodes.
Connected concepts
On-device speech recognition, Local-first inference, Quantization, CPU inference serving, Platform-native inference, Model offloading, vLLM / PagedAttention, Knowledge distillation
Explore it live in the knowledge graph →