Fine-tuning LLMs for Passive Depression Severity Estimation from AI Mental Health Dialogue
Signal
72
Hype
25
In three linesFine-tuning Qwen3.5-27B to predict PHQ-9 depression scores directly from transcripts of conversations with an AI mental health application. 6,283 users (3,111 ground-truth labels + Claude Opus pseudolabels). Performance: MAE=2.6, RMSE=4.0, r=0.80, AUC=0.91 at PHQ-9≥10 clinical threshold.Read source
Your take?
Summary generated by Claude — human-verified