Hacker News · item history
Apple Silicon and macOS VMs: 11–16× Faster LLM Inference with Llama.cpp
193
first seen points
305
peak points
305
latest points
34
observations
Posted on X
- 16x faster local inference on a Mac sounds huge until you check what it's 16x faster than: running the model inside a VM instead of on the host. The real number is "how fast is it against cloud API calls for the same task" — and that one's missing.