Back to feed
arXiv cs.AI·

DiagFlowBench: Evaluating How Language Models Handle Off-Procedure Inputs in Grounded Diagnostic Dialogue

Signal
75
Hype
15
In three linesDiagFlowBench evaluates how language models handle off-procedure inputs in industrial diagnostic dialogue. A dataset of 1,676 multi-turn conversations derived from 50 diagnostic flowcharts reveals models often select a real but contextually inadequate step rather than hallucinate, exposing a vulnerability: plausible but wrong advice grounded in documentation.
Read source
Your take?
BenchmarksEvalsReasoningAI safetyRAG

Summary generated by Claude — human-verified