Back to feed
arXiv cs.CL·

Prompt Perturbation for Reliable LLM Evaluation over Comparison Graphs

Signal
72
Hype
15
In three linesMethod to evaluate LLMs via pairwise comparisons by resolving intransitivity (cycles A≻B≻C≻A). Prompt perturbation framework generates prompt variants, identifies structural inconsistencies in comparison graphs, then applies filtered ranking methods to stabilize leaderboards.
Read source
Your take?
EvalsPrompt engineeringBenchmarks

Summary generated by Claude — human-verified