Predicting Dropout in Psychotherapy for Major Depressive Disorder: A Machine Learning Approach to Identifying At-Risk Patients
In brief
Machine-learning model predicts dropout from depression therapy with only modest accuracy
In a trial of 290 outpatients with major depression, the best algorithm-a random-forest model using SMOTE resampling-reached an F1-score of about 0.30, indicating limited ability to flag likely dropouts. Younger age, unemployment, multiple comorbidities, no partner and moderate severity were the strongest signals, and performance fell further in a separate validation cohort. More data and richer predictors will be needed before such tools can guide clinical practice.
- Journal
- Clinical psychology & psychotherapy (Q1)
- Published
- 1 January 2026
- Study design
- Randomized controlled trial
- Evidence level
- Level 1, High (CEBM 1b)
- Authors
- Susanne Bremer-Hoeve, Suzanne C van Bronswijk, Aartjan T F Beekman, Maartje Miggiels, Sanne J E Bruijniks, Maarten K van Dijk
- PMID
- 42590821
- DOI
- 10.1002/cpp.70320
Why clinicians should know about it
- Picked for Psychiatry and Mental Health (top studies of the week, 16 August 2026): Machine‑learning predicts psychotherapy dropout in MDD
Abstract
INTRODUCTION: Dropout from psychotherapy for major depressive disorder (MDD) undermines treatment effectiveness for both patients and clinicians. Early identification of individuals at risk of discontinuation is therefore essential. Machine learning methods, particularly those incorporating resampling techniques to address imbalanced clinical data, may offer improved predictive capacity. This study evaluates several machine-learning approaches to determine their potential in identifying patients at elevated risk of dropout. METHODS: Data were drawn from a randomised controlled trial of outpatients with MDD (n = 290) receiving short-term psychodynamic supportive psychotherapy (SPSP) or cognitive behavioural therapy (CBT). Candidate predictors were restricted to variables available in both the development and external validation datasets, resulting in a limited number of overlapping predictors. Multiple machine learning models and resampling strategies were applied to predict dropout, and the best-performing models were subsequently externally validated using an independent clinical dataset (n = 96). RESULTS: Models demonstrated low predictive performance overall. Resampling strategies influenced outcomes, with SMOTE consistently improving F1-scores compared with no resampling. The combination of random forest with SMOTE achieved the highest F1-score (0.303). Five key predictors of dropout emerged: younger age, unemployment, a higher number of comorbid Axis I disorders, absence of a partner, and moderate depression severity. All models showed reduced performance during external validation. DISCUSSION: The models showed very limited predictive performance, which may in part reflect the restricted predictor set resulting from the requirement that variables be assessed in both trials. While clinical application is premature at this stage, the findings provide a valuable proof of concept and highlight key methodological challenges inherent in cross-trial dropout prediction. Despite this, the identified predictors are routinely available in clinical practice, suggesting that future models, if developed with richer predictor sets and validated in larger samples, could ultimately support therapists in the early identification of patients at risk of dropout and in proactively addressing disengagement.
Abstract as published, via PubMed.
For healthcare professionals. The summary is generated by AI from the published abstract, and the evidence level is assigned automatically from the study design on the Oxford CEBM hierarchy. Neither is medical advice. Read the full paper before changing practice.