Stress-testing medical large language models reveals latent safety pathology beyond benchmark accuracy
Signal
78
Hype
15
In three linesAI-MASLD, a stress-audit framework, evaluates 7 medical LLMs on 240 clinical cases with narrative perturbations. All perform well at baseline but diverge under realistic stress. Quantized models hide functional collapse; medical fine-tuning degrades logical stability and fairness. An open-weight model matches or exceeds proprietary alternatives on all safety dimensions.Read source
Your take?
Summary generated by Claude — human-verified