Speaking in Self-Assessing Tongues: On the Verbalized Confidence of LLMs in Machine Translation
Signal
72
Hype
15
In three linesStudy of LLM verbalized confidence reliability in machine translation. Five methods for extracting per-token confidence without internal signal access are compared against predicted probabilities. Results: similar performance for error detection and calibration, but little correlation between internal and verbalized methods.Read source
Your take?
Summary generated by Claude — human-verified