Back to feed
arXiv cs.AI·

Stress-testing medical large language models reveals latent safety pathology beyond benchmark accuracy

Signal
78
Hype
15
In three linesAI-MASLD, a stress-audit framework, evaluates 7 medical LLMs on 240 clinical cases with narrative perturbations. All perform well at baseline but diverge under realistic stress. Quantized models hide functional collapse; medical fine-tuning degrades logical stability and fairness. An open-weight model matches or exceeds proprietary alternatives on all safety dimensions.
Read source
Your take?
BenchmarksAI safetyEvalsAlignment

Summary generated by Claude — human-verified