POLARIS: Guiding Small Models to Write Long Stories
Signal
78
Hype
25
In three linesPOLARIS is a reinforcement optimization method (GRPO) to improve long-form text generation in small models. Applied to Qwen3.5-9B with 1.4K prompt-story pairs and 4 A100 GPUs, it uses a frontier LLM judge and human-reference injection. POLARIS-9B rivals models 3× larger and generalizes to stories 3× longer than training data.Read source
Your take?
Summary generated by Claude — human-verified