Llama.cpp can do 40 tok/s on M2 Max, 0% CPU usage, using all 38 GPU cores by from on 2023-06-04 17:24 (#6C1Y5) Comments