PragReST: Self-Reinforcing Counterfactual Reasoning for Pragmatic Language Understanding
Signal
78
Hype
25
In three linesPragReST is a self-supervised framework improving LLM pragmatic reasoning through counterfactual reasoning traces. Without human-labeled data, it combines supervised fine-tuning and reinforcement learning. On 4 benchmarks (PragMega, Ludwig, MetoQA, AltPrag), it gains +5.37% and +5.50% absolute for Qwen3-8B and Qwen3-14B.Read source
Your take?
Summary generated by Claude — human-verified