Back to feed
arXiv cs.CL·

Fine-tuning LLMs for Passive Depression Severity Estimation from AI Mental Health Dialogue

Signal
72
Hype
25
In three linesFine-tuning Qwen3.5-27B to predict PHQ-9 depression scores directly from transcripts of conversations with an AI mental health application. 6,283 users (3,111 ground-truth labels + Claude Opus pseudolabels). Performance: MAE=2.6, RMSE=4.0, r=0.80, AUC=0.91 at PHQ-9≥10 clinical threshold.
Read source
Your take?
Fine-tuningReasoningQwenClaudeBenchmarks

Summary generated by Claude — human-verified