Running Qwen3.6-35B-A3B on a laptop RTX 4060 (8GB) — what worked, what didn't, and a surprising speculative-decoding result
Signal
72
Hype
15
In three linesUser optimizes Qwen3.6-35B-A3B (35B/3B active MoE) on RTX 4060 8GB. Final config: --no-mmap critical (11→43 tok/s), ≥1.5GB VRAM headroom mandatory, CPU bottleneck dominant. Speculative decoding +26% (contradicts community benchmarks). Hybrid architecture (10 attention + 40 GDN layers) explains counterintuitive results.Read source
Your take?
Summary generated by Claude — human-verified