Back to feed
Reddit r/LocalLLaMA·

Here are some tips on hitting nearly 200 tok/s for DeepSeek v4 Flash on Hopper

Signal
65
Hype
25
In three linesDeepSeek v4 Flash optimization on Hopper GPU: achieving 193 tok/s using Canada-Quant quantization and vLLM MTP patching. Author documents performance gains to reduce local inference costs versus API pricing ($0.1966/M tokens).
Read source
Your take?
DeepSeekCode generationAI AgentsInfrastructureOpen source

Summary generated by Claude — human-verified