Back to feed
arXiv cs.LG·

Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models

Signal
78
Hype
15
In three linesEvaluation framework for LLM adversarial robustness based on computational cost (FLOPs) rather than query count. Study of 10 models with 3 attack strategies reveals: alignment has non-monotonic effects, scaling reduces gradient attacks but not template attacks, transfer possible across models, cost varies up to 5× across harm categories.
Read source
Your take?
AI safetyAlignmentEvalsBenchmarks

Summary generated by Claude — human-verified