Back to feed
arXiv cs.CL·

Speaking in Self-Assessing Tongues: On the Verbalized Confidence of LLMs in Machine Translation

Signal
72
Hype
15
In three linesStudy of LLM verbalized confidence reliability in machine translation. Five methods for extracting per-token confidence without internal signal access are compared against predicted probabilities. Results: similar performance for error detection and calibration, but little correlation between internal and verbalized methods.
Read source
Your take?
EvalsReasoning

Summary generated by Claude — human-verified