Back to feed
arXiv cs.CL·

BenSyc: Benchmarking Conversational Sycophancy and Human Alignment in LLMs for Bengali Contexts

Signal
78
Hype
15
In three linesBenSyc is the first benchmark for evaluating conversational sycophancy in Bengali social contexts. Built from 170k Reddit comments, it tests 15+ LLMs on alignment classification and response generation. Best models achieve only 61.8% Macro-F1 on binary detection, revealing difficulty distinguishing empathetic support from excessive validation.
Read source
Your take?
BenchmarksAlignmentAI safetyEvals

Summary generated by Claude — human-verified