Skip to main content

Predicting Dropout in Psychotherapy for Major Depressive Disorder: A Machine Learning Approach to Identifying At-Risk Patients

In brief

Machine-learning model predicts dropout from depression therapy with only modest accuracy

In a trial of 290 outpatients with major depression, the best algorithm-a random-forest model using SMOTE resampling-reached an F1-score of about 0.30, indicating limited ability to flag likely dropouts. Younger age, unemployment, multiple comorbidities, no partner and moderate severity were the strongest signals, and performance fell further in a separate validation cohort. More data and richer predictors will be needed before such tools can guide clinical practice.

Journal
Clinical psychology & psychotherapy (Q1)
Published
1 January 2026
Study design
Randomized controlled trial
Evidence level
Level 1, High (CEBM 1b)
Authors
Susanne Bremer-Hoeve, Suzanne C van Bronswijk, Aartjan T F Beekman, Maartje Miggiels, Sanne J E Bruijniks, Maarten K van Dijk
PMID
42590821
DOI
10.1002/cpp.70320

Why clinicians should know about it

Abstract

INTRODUCTION: Dropout from psychotherapy for major depressive disorder (MDD) undermines treatment effectiveness for both patients and clinicians. Early identification of individuals at risk of discontinuation is therefore essential. Machine learning methods, particularly those incorporating resampling techniques to address imbalanced clinical data, may offer improved predictive capacity. This study evaluates several machine-learning approaches to determine their potential in identifying patients at elevated risk of dropout. METHODS: Data were drawn from a randomised controlled trial of outpatients with MDD (n = 290) receiving short-term psychodynamic supportive psychotherapy (SPSP) or cognitive behavioural therapy (CBT). Candidate predictors were restricted to variables available in both the development and external validation datasets, resulting in a limited number of overlapping predictors. Multiple machine learning models and resampling strategies were applied to predict dropout, and the best-performing models were subsequently externally validated using an independent clinical dataset (n = 96). RESULTS: Models demonstrated low predictive performance overall. Resampling strategies influenced outcomes, with SMOTE consistently improving F1-scores compared with no resampling. The combination of random forest with SMOTE achieved the highest F1-score (0.303). Five key predictors of dropout emerged: younger age, unemployment, a higher number of comorbid Axis I disorders, absence of a partner, and moderate depression severity. All models showed reduced performance during external validation. DISCUSSION: The models showed very limited predictive performance, which may in part reflect the restricted predictor set resulting from the requirement that variables be assessed in both trials. While clinical application is premature at this stage, the findings provide a valuable proof of concept and highlight key methodological challenges inherent in cross-trial dropout prediction. Despite this, the identified predictors are routinely available in clinical practice, suggesting that future models, if developed with richer predictor sets and validated in larger samples, could ultimately support therapists in the early identification of patients at risk of dropout and in proactively addressing disengagement.

Abstract as published, via PubMed.

View on PubMedFull text at the publisherOpen in the app

For healthcare professionals. The summary is generated by AI from the published abstract, and the evidence level is assigned automatically from the study design on the Oxford CEBM hierarchy. Neither is medical advice. Read the full paper before changing practice.