Back to feed
Reddit r/LocalLLaMA·

Luce Spark: a 35B MoE on a 16 GB GPU, without the offload tax

Signal
78
Hype
25
In three linesLuce Spark runs 33-35B MoE models on 16 GB GPU without offload penalty. Qwen 35B-A3B: 13.3 GiB (vs 20.5), Laguna XS.2 33B-A3B: 14.6 GiB (vs 18.8). Only active experts (~8/256) stay in VRAM; rest in system RAM with intelligent swapping. Self-tuning via learned routing profile. Open-source Apache 2.0.
Read source
Your take?
Open sourceInfrastructureLlamaQwen

Summary generated by Claude — human-verified