Back to feed
arXiv cs.CL·

Possible or Definite? A Benchmark for Evaluating Diagnostic Uncertainty Preservation in Clinical Text

Signal
78
Hype
15
In three linesBenchmark of 1,200 clinical documents with 9,184 uncertainty annotations across five levels. LLMs poorly preserve uncertainty expressions (less than 50% of cases) and struggle with nuanced distinctions between adjacent levels. Reveals a failure mode missed by standard metrics.
Read source
Your take?
BenchmarksAI safetyEvals

Summary generated by Claude — human-verified