sycl : port multi-column MMVQ from CUDA backend (~45% speculative decoding speedup on Intel Arc) by masonmilby · Pull Request #21845 · ggml-org/llama.cpp
Signal
75
Hype
15
In three linesSYCL port of multi-column MMVQ speculative decoding from CUDA backend to llama.cpp. ~45% speedup on Intel Arc cards. Update recommended from version b9519 onwards.Read source
Your take?
Summary generated by Claude — human-verified