Ollama is now powered by MLX on Apple Silicon in preview
8.3 relevance
Score Breakdown
technical depth 8
novelty 8
actionability 9
community 9
strategic 7
personal 9
Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.
Ollama optimized for Apple Silicon, important for local LLM inference.
Summary
Ollama 0.19 preview on Apple Silicon uses MLX to achieve up to 1810 tokens/s prefill and 112 tokens/s decode with Qwen3.5-35B-A3B in NVFP4 format, doubling speed over 0.18. It leverages M5's GPU Neural Accelerators and unified memory, with enhanced caching for coding agents like Claude Code. Requires Macs with >32GB RAM for optimal performance.