Back to feed
Reddit r/LocalLLaMA·

I ported EXL3 to run well on Apple Silicon - PonyExl3

Signal
75
Hype
25
In three linesEXL3 codec ported to Apple Silicon using Metal backend. M5 Max achieves ~600 tok/s prefill and ~38 tok/s generation (Qwen 27B), outperforming RTX 4090 on some benchmarks (68.5-80 tok/s decode). GitHub repo with reproducible results.
Read source
Your take?
Open sourceCode generationInfrastructureQwen

Summary generated by Claude — human-verified