Characterization and Validation of EHR Computable Phenotypes for Long COVID Using Patient-Reported Symptoms: Insights from the Nationwide RECOVER Program
- Journal
- Journal of the American Medical Informatics Association : JAMIA (Q1)
- Published
- 24 July 2026
- Study design
- Unclassified
- Evidence level
- Level 5, Expert Opinion (CEBM 5)
- Authors
- Victor M Castro, Vivian Gainer, Nich Wattanasin, Andrew Cagan, Ana Holzbach, James Chan, et al.
- PMID
- 42496643
- DOI
- 10.1093/jamia/ocag115
Why clinicians should know about it
- Picked for Health Informatics (paper of the day, 25 July 2026).
Abstract
OBJECTIVE: Long COVID (LC) remains poorly understood, and there is a critical need for advanced computational tools to better identify and characterize patients. In this study, we use summarized symptom reports by RECOVER-Adult cohort participants linked to EHR data to characterize patients and train a computable phenotype algorithm of LC. MATERIALS AND METHODS: The study included adult participants with linked FHIR-sourced EHR data. We characterized EHR diagnoses, procedures, medications, lab tests, and vital sign features associated with LC. A computable phenotyping algorithm was trained and validated against patient-reported symptoms. MAIN OUTCOME AND MEASURES: We assessed model discrimination and calibration in a held-out test set. We describe important model features and evaluate model discrimination and calibration. RESULTS: The study included 1,501 RECOVER-Adult cohort participants with linked EHR data. 376 (25%) met criteria for highly symptomatic LC based on the RECOVER Long COVID Research Index (LCRI). EHR features associated with LC included clinician diagnosis of shortness of breath, malaise and fatigue, and cardiac dysrhythmias; documented treatment with albuterol, gabapentin, or duloxetine; or elevated heart rate. The algorithm identifying patients with highly symptomatic LC had an AUROC of 0.80 (95% confidence interval (CI) 0.74-0.85), and AUPRC of 0.58 (95% CI, 0.47-0.69). CONCLUSION AND RELEVANCE: These findings demonstrate that, using EHR data, a machine-learning model can accurately select patients with sets of self-reported LC symptoms. The model could help identify patients within a health system with the highest probability of the condition and facilitate screening, recruitment for clinical trials, and etiologic studies.
Abstract as published, via PubMed.
For healthcare professionals. The summary is generated by AI from the published abstract, and the evidence level is assigned automatically from the study design on the Oxford CEBM hierarchy. Neither is medical advice. Read the full paper before changing practice.