Back to feed
arXiv cs.AI·

READER: Robust Evidence-based Authorship Decoding via Extracted Representations

Signal
72
Hype
18
In three linesREADER is a provenance framework to identify source LLM from black-box responses without predefined prompts. Using activation mapping on a frozen proxy LLM and Bayesian evidence accumulation, it achieves 31-42% top-1 accuracy on single response and 70-84% on 50 responses on Agent500 (50-target dataset).
Read source
Your take?
AI AgentsEvalsAI safety

Summary generated by Claude — human-verified