Set OLLAMA_NUM_PARALLEL to 4 or 8 on your Mac and still watching Ollama queue requests one by one? The setting may be fine. The real bottleneck is hiding somewhere else.
Read the full article on Popular AI