Luce Spark: a 35B MoE on a 16 GB GPU, without the offload tax
Signal
78
Hype
25
In three linesLuce Spark runs 33-35B MoE models on 16 GB GPU without offload penalty. Qwen 35B-A3B: 13.3 GiB (vs 20.5), Laguna XS.2 33B-A3B: 14.6 GiB (vs 18.8). Only active experts (~8/256) stay in VRAM; rest in system RAM with intelligent swapping. Self-tuning via learned routing profile. Open-source Apache 2.0.Read source
Your take?
Summary generated by Claude — human-verified