Back to feed
arXiv cs.CL·

What Do People Actually Want From AI? Mapping Preference Plurality

Signal
78
Hype
25
In three linesAnalysis of 1,500 open-ended responses from PRISM dataset (75 countries) on human preferences for AI systems. Finding: requested values vary widely across individuals (only truthfulness reaches 49%), with divergent definitions of the same concept. Current RLHF methods fail to capture this plurality by aggregating it into a single reward model.
Read source
Your take?
AlignmentReinforcement learningEvalsAI safety

Summary generated by Claude — human-verified