Real-world evaluation of a transformer-based natural language processing system for identifying social determinants of health from routine clinical documentation
In brief
NLP tool identifies only 5% of abuse and 0% of drug use from clinical notes
In a real-world test of the SODA transformer-based system on 414 patients, the algorithm captured merely 5% of abuse mentions and none of drug-use references, with modest sensitivity for alcohol (55%) and financial constraints (16%). The poor performance reflects both limited documentation of social risks and mismatches between note language and survey definitions, highlighting the need for better charting and locally tuned models before such tools can reliably guide care.
- Journal
- JAMIA open (Q1)
- Published
- 31 July 2026
- Study design
- Unclassified
- Evidence level
- Level 5, Expert Opinion (CEBM 5)
- Authors
- Xiangren Wang, Jingchuan Guo, Yi Guo, Yao An Lee, Mattia Prosperi, Jiang Bian, et al.
- PMID
- 42542644
- DOI
- 10.1093/jamiaopen/ooag147
Why clinicians should know about it
- Picked for Health Informatics (paper of the day, 3 August 2026).
Abstract
OBJECTIVES: To evaluate the real-world performance of a transformer-based natural language processing (NLP) system for extracting Social Determinants of Health (SDoH) from clinical notes, using survey-based SDoH measures a reference comparators. MATERIALS AND METHODS: This study was conducted at the University of Florida Health in adults with at least 2 clinical encounters in the prior year. A research survey was completed by 1001 participants; sampling targeted 50% Black patients to support subgroup analyses. Comparative analyses were restricted to the 414 participants who also had Epic SDoH survey data and clinical notes available for NLP extraction. We compared concepts extracted by the SOcial DeterminAnts (SODA) NLP pipeline against the independently administered research survey, which served as the primary reference standard, and against the structured Epic-embedded SDoH questionnaire. Nine domains were evaluated: abuse, alcohol use, drug use, education, financial constraints, housing, physical activity, social cohesion, and transportation. Sensitivity, specificity, positive predictive value, negative predictive value, and F1 scores were calculated by domain. RESULTS: The NLP pipeline more consistently aligned with negative survey responses than with patient-reported social needs, although performance was lower in some domains, especially alcohol use and financial constraints. Sensitivity was higher only for alcohol use (55%); the lowest values were for abuse (5%), drug use (0%), and financial constraints (16%). These results cannot be attributed to SODA extraction alone. They reflect some combination of social information not being recorded in clinical notes, content that was recorded but not extracted, and mismatches between extracted concepts and survey definitions, and the present analysis cannot separate these contributions. The 2 surveys agreed only modestly with each other, so no single instrument provides a definitive ground truth. CONCLUSION: The pipeline more consistently aligned with negative survey responses than with patient-reported social needs. Because the 2 surveys agreed only modestly, the reference itself is imperfect, and apparent NLP performance depends in part on which survey is used as the comparator. Apparent gaps in NLP performance reflect both how social risks are recorded in clinical notes and how patients disclose them across different survey settings, in addition to limits of the extraction pipeline. Improving documentation practices, integrating locally tuned large language models, and monitoring subgroup performance may all be needed to make SDoH identification tools reliably detect social needs across patient populations.
Abstract as published, via PubMed.
For healthcare professionals. The summary is generated by AI from the published abstract, and the evidence level is assigned automatically from the study design on the Oxford CEBM hierarchy. Neither is medical advice. Read the full paper before changing practice.