Unified memory beats raw GPU compute for local AI
Apple's M3 Ultra delivers 819 GB/s bandwidth, proving unified memory outperforms discrete VRAM for running large local models without latency.
Apple's M3 Ultra delivers 819 GB/s bandwidth, proving unified memory outperforms discrete VRAM for running large local models without latency.
Testing five models on an Intel i5 reveals 36 tokens per second is the ceiling for small LLMs due to 20 GB/s memory bandwidth limits.