Back to feed
arXiv cs.CL·

Localizing Prompt Ambiguity in Large Language Models with Probe-Targeted Attribution

Signal
78
Hype
15
In three linesPRIG, a gradient attribution method, localizes ambiguity in LLM prompts by training a linear probe to distinguish clear from ambiguous prompts, then attributes the probe score to token representations. Evaluated on synthetic datasets (coding, math, writing) and a human-written gold benchmark, PRIG achieves 0.840 AUROC on combined synthetic benchmark and 0.891 AUROC on gold set.
Read source
Your take?
Prompt engineeringEvalsPapers

Summary generated by Claude — human-verified