Back to feed
arXiv cs.CL·

Scenario-based Probing and Steering Cultural Values in Large Language Models--Extended Version

Signal
72
Hype
18
In three linesarXiv paper on probing cultural values in LLMs through scenario-based behavioral dilemmas mapped to World Values Survey axes. Authors apply activation steering across three open-source models to measure and shift implicit preferences along Inglehart-Welzel dimensions, revealing latent entanglement where interventions on one cultural axis induce shifts on another.
Read source
Your take?
PapersAlignmentEvalsAI safety

Summary generated by Claude — human-verified