Back to feed
Reddit r/LocalLLaMA·

sycl : port multi-column MMVQ from CUDA backend (~45% speculative decoding speedup on Intel Arc) by masonmilby · Pull Request #21845 · ggml-org/llama.cpp

Signal
75
Hype
15
In three linesSYCL port of multi-column MMVQ speculative decoding from CUDA backend to llama.cpp. ~45% speedup on Intel Arc cards. Update recommended from version b9519 onwards.
Read source
Your take?
Open sourceCode generationInfrastructure

Summary generated by Claude — human-verified