Mode
Text Size
Log in / Sign up

LLM-derived embeddings outperform tabular models for end-of-therapy outcome prediction in tuberculosis patientsNew Data Models Improve Prediction of Tuberculosis Treatment Success

AI-generated summary of the cited source, checked by automated accuracy review. How we work

Key Takeaway
Note that LLM-derived embeddings improve EOT outcome prediction, but routine data provides limited value for relapse prediction.

This observational analysis of Phase 3 clinical trial data included 2,918 participants with tuberculosis. The study compared time-resolved prediction models using tabular data and large language model-derived embeddings against baseline cavitation and sputum smear data to predict end-of-therapy (EOT) outcomes and post-treatment relapse.

For EOT outcomes, tabular models showed improved performance after month 3 with ROC-AUC up to 0.84. LLM-derived embedding models outperformed tabular models at months 3 and 4 with a Delta ROC-AUC of 0.12 and 0.14, respectively. For relapse prediction, tabular models showed only modest improvement through month 3 (ROC-AUC 0.58-0.63), but longitudinal modeling with sparse variable inclusion improved prediction at months 4-6 with Delta ROC-AUC of 0.09, 0.19, and 0.12.

Risk stratification using model-derived data showed improved 4-month relapse-free survival compared to baseline cavitation and sputum smear (94.8%/78.5% vs. 91.4%/82.9%). No safety or tolerability data were reported. A noted limitation is the reduced interpretability of models using longitudinal modeling and sparse variable inclusion. While routine clinical data can improve risk stratification, it offers limited predictive value for relapse specifically, suggesting a need for more specific biomarkers.

How this fits prior evidence

How this fits prior evidence: This finding addresses a gap in identifying predictive markers for tuberculosis outcomes. While previous coverage noted that WGS-based tools provide high specificity and sensitivity for rifampicin, isoniazid, and fluoroquinolone resistance, this study explores the utility of longitudinal clinical data and LLM-derived embeddings for predicting treatment success and relapse. It does not directly relate to the 13% prevalence of tuberculosis in children with severe acute malnutrition or the 0.81 AUC for drug-induced liver injury risk models.

Researchers analyzed data from 2,918 people in Phase 3 tuberculosis trials. They compared traditional methods, like looking at initial lung damage and sputum samples, against new computer models that use information collected throughout the course of treatment. The goal was to see if these models could better predict if a patient would finish therapy successfully or suffer a relapse.

The study found that models using information from the first few months of treatment performed better than standard methods at predicting successful completion. Specifically, models using large language model-derived data outperformed standard models. However, predicting a relapse was more difficult. While some advanced models showed slight improvements in identifying relapse risk, the results were less certain than the results for treatment completion.

Because this was an analysis of existing trial data and not a new clinical trial, these findings are not yet ready to change daily medical practice. The researchers noted that while routine data helps identify who might finish treatment, it still offers limited value for predicting a relapse. More specific markers are needed to accurately predict when a patient might fall ill again after treatment ends.

What this means for you:
New data models can better predict treatment success for tuberculosis, but predicting relapse remains difficult.

Common questions

How does this help patients with tuberculosis?

These models use data collected during treatment to help doctors predict if a patient will successfully finish their therapy. While the current models are better at predicting successful completion than standard methods, they are still limited when it comes to predicting if a patient will have a relapse after treatment ends.

Is this a new treatment for tuberculosis?

No, this is not a new medication or treatment. It is an analysis of data from existing Phase 3 clinical trials. The study looked at how different computer models can interpret patient data to better predict health outcomes.

Can these models predict if a patient will relapse?

The models showed some improvement in predicting relapse compared to standard methods, but the results were only modest. The study suggests that current routine clinical data provides limited predictive value for relapse, and more specific markers are needed for that purpose.

Study Details

Study typePhase3
Sample sizen = 2,918
EvidenceLevel 2
PublishedSep 2026
View Original Abstract ↓
Relapse after apparently successful tuberculosis (TB) therapy remains difficult to predict, and how relapse risk evolves throughout treatment remains unclear. Using harmonised clinical data from two Phase 3 trials (2,918 participants), we performed time-resolved prediction of end-of-therapy (EOT) outcomes and post-treatment relapse by training models using tabular data at monthly intervals from baseline to therapy end. Prediction of EOT outcomes improved after month 3 (ROC-AUC up to 0.84), driven by sputum-smear and solid culture. In contrast, relapse prediction among participants with favourable EOT outcomes and completed follow-up improved only modestly through month 3 (ROC-AUC 0.58-0.63) before declining, with age, sex, clinical symptoms and bacterial burden contributing most strongly to prediction. Models trained on large language model-derived embeddings, a more flexible representation of the same variables, matched tabular relapse models in performance throughout therapy, and outperformed tabular EOT outcome models at months 3 and 4 ({Delta}ROC-AUC: 0.12 and 0.14), with longitudinal modelling and sparse variable inclusion only improving relapse prediction at months 4-6 ({Delta}ROC-AUC: 0.09, 0.19 and 0.12), however with reduced interpretability. Models incorporating data after baseline provided incremental improvements in post-treatment relapse risk stratification compared with baseline cavitation and sputum smear alone (4-month relapse-free survival: 94.8%/78.5% for model-derived low/high risk groups, vs. 91.4%/82.9% for baseline easy-/hard-to-treat groups). Overall, these findings suggest that while routine clinical data collected after baseline can improve post-treatment risk stratification, it offers limited predictive value for relapse, underscoring the need for relapse-specific biomarkers.
Free Newsletter

Clinical research that matters. Delivered to your inbox.

Join thousands of clinicians and researchers. No spam, unsubscribe anytime.