metricas

Gastroenterología y Hepatología

Sugerencias
Gastroenterología y Hepatología Development and validation of interpretable machine learning models to detect un...
Información de la revista
Cita
Cita
Compartir
Descargar PDF
Más opciones de artículo
Visitas
360
Original Article
Acceso a texto completo
Disponible online el 1 de junio de 2026

Development and validation of interpretable machine learning models to detect unconfirmed hepatitis C using electronic health records (LiverTAI)

Desarrollo y validación de modelos de aprendizaje automático interpretables para detectar hepatitis C no confirmada utilizando historias clínicas electrónicas (LiverTAI)
Visitas
360
Gloria Sánchez-Antolína, Gema de la Pozab, Lorena Hidalgoc, María Victoria Aguilerad,e,f,g, Berta Cuyàsf,h,i, Francisco Ledesmaj, Eva Sanzk, Víctor Fanjull, Clara L. Oestel, José Luis Calleja Panerom,
Autor para correspondencia
joseluis.calleja@uam.es

Corresponding author.
a Hospital Universitario Río Hortega, IBioVall, Universidad de Valladolid, Valladolid, Spain
b Hospital Universitario de Fuenlabrada, Madrid, Spain
c Hospital Universitario Infanta Sofía, Madrid, Spain
d Hospital Universitario y Politécnico La Fe, Valencia, Spain
e Instituto de Investigación Sanitaria La Fe (IIS La Fe), Valencia, Spain
f CIBEREHD, Instituto de Salud Carlos III (ISCIII), Madrid, Spain
g Faculty of Medicine, Valencia University, Valencia, Spain
h Hospital de la Santa Creu i Sant Pau, Barcelona, Spain
i Universitat Autònoma de Barcelona, Barcelona, Spain
j Former AbbVie Spain Employee, Madrid, Spain
k AbbVie Spain, Madrid, Spain
l Savana Research, Madrid, Spain
m Hospital Universitario Puerta de Hierro, Majadahonda, Madrid, Spain
Ver más
Este artículo ha recibido
Información del artículo
Resumen
Texto completo
Bibliografía
Descargar PDF
Estadísticas
Figuras (4)
fig0005
fig0010
fig0015
fig0020
Tablas (3)
Table 1. Demographic and clinical characteristics of all patient groups according to HCV serology status.
Tablas
Table 2. Validation metrics of the full and reduced models on the validation set.
Tablas
Table 3. Variable importance and coefficients for the reduced Extreme Gradient Boosting (XGB) and logistic regression (LR) predictive models.
Tablas
Material adicional (1)
Abstract
Objective

Hepatitis C virus (HCV) infection poses a global health threat with many undiagnosed cases despite advances in diagnosis and treatment. The LiverTAI project in Spain utilized electronic health records (EHRs) to study factors associated with HCV infection, aiming to develop a predictive model for identifying unconfirmed HCV cases in the hospital setting.

Patients and methods

Clinical data from EHRs of six hospitals in Spain were analyzed using machine learning and natural language processing through EHRead®. Patients were categorized as HCV positive, negative, or unknown. A semi-supervised learning framework allowed to incorporate labeled and unlabeled patient data extracted from clinical narratives. Propensity score matching was applied to reduce bias. Seven classification algorithms were used to predict HCV status based on 117 selected features, including demographics, risk factors, comorbidities, and clinical events. Model performance was confirmed through independent geographic validation.

Results

Among 2,440,358 screened patients, 44,235 were included in the training set, and 11,286 in the validation set. The Extreme Gradient Boosting model showed the best performance (AUC–ROC 0.794), followed by the logistic regression model (AUC–ROC 0.779). Key predictors included HCV risk factors (age, male sex, HIV, drug use), liver-related issues (cirrhosis, hepatocellular carcinoma), and extrahepatic conditions (neuropsychiatric, cardiovascular, immune-related disorders, cancer, and inflammatory processes).

Conclusions

LiverTAI identified new patients with potential HCV infection in routine hospital EHRs, providing a proof of concept for risk-stratified opportunistic screening. The model supports more efficient in-hospital testing strategies, though further prospective validation is required to confirm generalizability and clinical utility.

Keywords:
Hepatitis C virus (HCV)
Electronic health records (EHRs)
Natural language processing
Machine learning
Predictive model
Unconfirmed HCV cases
Resumen
Objetivo

La infección por virus de la hepatitis C (VHC) constituye una amenaza global, con numerosos casos no diagnosticados a pesar de los avances en diagnóstico y tratamiento. El proyecto LiverTAI en España utilizó historias clínicas electrónicas (HCE) para estudiar factores asociados con VHC y desarrollar un modelo predictivo capaz de identificar casos no confirmados en entornos hospitalarios.

Pacientes y métodos

Se analizaron HCE de seis hospitales en España mediante aprendizaje automático y procesamiento de lenguaje natural utilizando EHRead®. Los pacientes se clasificaron como positivos, negativos o desconocidos para VHC. Se empleó un enfoque de aprendizaje semi-supervisado que incorporó datos etiquetados y no etiquetados extraídos de narrativas clínicas, aplicando pareamiento por puntaje de propensión para reducir sesgos. Siete algoritmos de clasificación predijeron el estado de VHC a partir de 117 características, incluyendo demografía, factores de riesgo, comorbilidades y eventos clínicos. El desempeño se validó de forma geográfica independiente.

Resultados

De 2.440.358 pacientes evaluados, 44.235 se incluyeron en entrenamiento y 11.286 en validación. Extreme Gradient Boosting mostró el mejor desempeño (AUC-ROC 0,794), seguido de regresión logística (AUC-ROC 0,779). Los principales predictores fueron factores de riesgo de VHC (edad, sexo masculino, VIH, consumo de drogas), complicaciones hepáticas (cirrosis, carcinoma hepatocelular) y condiciones extrahepáticas (trastornos neuropsiquiátricos, cardiovasculares, inmunológicos, cáncer e inflamación).

Conclusiones

LiverTAI identificó pacientes con posible infección por VHC en HCE hospitalarias, demostrando la viabilidad del cribado oportunista estratificado por riesgo. El modelo apoya estrategias más eficientes de testeo hospitalario, aunque requiere validación prospectiva adicional para confirmar su aplicabilidad clínica.

Palabras clave:
Virus de la hepatitis C (VHC)
Historias clínicas electrónicas (HCE)
Procesamiento de lenguaje natural (PLN)
Aprendizaje automático
Modelo predictivo
Casos de VHC no confirmados
Resumen gráfico
Texto completo
Introduction

Hepatitis C virus (HCV) remains a worldwide health concern, with an estimated global prevalence of 1.8% (1.4–2.3%),1 and an active infection prevalence of 0.22% in Spain.2 The World Health Organization's Global Health Sector Strategy (GHSS) aims to reduce new HCV infections by 90% and HCV-related mortality by 65% by 2030 compared with 2015 levels.3 Despite the availability of curative treatments like direct-acting antiviral (DAA),4 HCV infection is silent and more prevalent among populations with limited access to healthcare, leading to underdiagnosis and increased risk of serious chronic HCV-associated disorders such as hepatocellular carcinoma and cirrhosis.5 Underdiagnosis remains a hurdle to eliminating HCV,4 and robust risk-based testing procedures need to be established since universal screening is still unviable in most healthcare regions.4 In Spain, it was estimated that 29.4% of active HCV infections remained undiagnosed in 2018, based on data from the national population-based seroprevalence survey conducted in 2017–2018.2 More recent modelling studies indicate that, despite a substantial reduction in overall prevalence following the widespread implementation of DAAs, a clinically relevant proportion of HCV infections in the general population continues to remain undiagnosed.6

HCV transmission is linked to percutaneous drug use, iatrogenic infections, high-risk sexual activities, vertical transmission, intranasal drug use, piercings, and tattoos.4 Testing has therefore geared towards high-risk populations such as injecting drug users (IDU) or incarcerated individuals,4 commonly using surveys based on known risk factors; however, this approach has limitations such as non-response bias or high dropout rates.7 Furthermore, it does not allow to identify previously unknown risk factors. Evidence from Nigeria8 and the United States9 suggests additional transmission factors remain undiscovered, underscoring the need for broader patient populations (i.e., not selected only based on previously known risk factors) and comprehensive multivariable analyses.

Machine learning (ML) methods to detect previously unknown disease patterns have enabled predictive models dealing with multivariable correlation in large and diverse patient datasets,10 yet many of these models have focused on HCV-related liver disease progression, such as fibrosis, cirrhosis, and hepatocellular carcinoma (HCC),11,12 rather than identifying undiagnosed HCV infection. Most models rely exclusively on structured data sources such as laboratory values, pharmacy claims or administrative codes, often derived from private or single-center datasets, limiting generalizability and interpretability,13–16 often without external validation. Consequently, representative real-world evidence (RWE) for community-level HCV identification remains scarce.

This sub-study is part of the LiverTAI project,17 which aims to identify factors associated with HCV infection and to characterize demographics, clinical features, patient journey, and linkage to care among HCV-screened patients in Spain. This analysis sought to develop a predictive model for potentially unconfirmed HCV infection in clinical settings using cutting-edge ML tools, including natural language processing (NLP) and a semi-supervised learning framework that incorporates both labeled and unlabeled EHR data. Propensity score matching (PSM) enhanced fairness and reduced bias associated with data completeness and healthcare utilization. We also provide a comprehensive list of clinical features associated with HCV infection. This proof of concept demonstrates the potential of real-world data from a universal healthcare system to support opportunistic HCV case finding. With prospective validation, the model could be integrated into hospital systems to flag patients for confirmatory testing during routine care.

Patients and methodsStudy design and population

This was a multicenter and observational study using real world data for a retrospective analysis of patients who were attended within the Spanish National Healthcare Network between January 1, 2014, and December 31, 2018. The target study population consisted of all adult patients with at least one record in which any of the HCV-related terms (regarding general, HCV test, HCV genotype, or HCV treatment) was mentioned during the study period (Table S1). Of note, the specific context in which these HCV-related terms appeared within the EHRs, could be both, as needed to discard HCV as an etiologic factor for another disorder or per protocol in the presence of specific diseases. The earliest mention of an HCV-related term within the study period was defined as the index date (date of inclusion in the study) for each patient. Descriptive analyses were performed at the index date and a predictive model for undiagnosed HCV patient detection was developed.

This manuscript followed the transparent reporting of a multivariable prediction model for individual prognosis or diagnosis – artificial intelligence (TRIPOD-AI) guidelines. Further details on the study protocol are available in the Supplementary Methods.

Data source and extraction

The study was based on the secondary use of clinical data in the patients’ EHRs of six hospitals (Table S2). The unstructured information in the EHRs of included patients was collected from outpatient clinic reports, discharge reports, and emergency reports including all available departments. Structured pharmacy reports were only available for two hospitals. All information was extracted and analyzed using the EHRead® technology following previously described methods.18 Further technical details on the NLP pipeline and variable construction process, including the full list of study variables (Table S3), are provided in the Supplementary Methods.

Predictive model training and validation

To develop the predictive models, data from the six study hospitals were first grouped into two independent datasets. The training set included hospitals from central Spain (Madrid and Valencia), and the validation set covered the northern half of Spain (Valladolid and Barcelona). Positive and negative classes for predictive modeling were defined based on documented HCV infection status (outcome). Within the study population, patients with HCV seropositivity (as captured in the free-text narratives of physicians) were considered the positive class, labeled as HCVab+. HCV seropositivity was determined when any of the following were detected in the EHRs: a positive HCV serology/PCR, a reported HCV genotype, or an HCV-specific treatment. In contrast, patients with a confirmed negative HCV serology result were considered the negative class and labeled as HCVab−. For some patients, conclusive information regarding their HCV infection status could not be found. These patients for which HCV infection status could not be ascertained were considered unlabeled patients (HCVabx) and remained unclassified. To ensure temporal consistency and minimize the risk of reverse causation, variables reflecting post-diagnostic information were excluded from the predictor set. Specifically, features likely to represent downstream consequences of an established HCV diagnosis –such as HCV-specific treatments, genotyping results, or diagnostic procedures (e.g., FibroScan®) explicitly performed in an HCV context– were removed. Accordingly, predictor variables were restricted to information documented prior to or at the index date and to features plausibly representing antecedent risk factors or concurrent conditions, rather than consequences of a confirmed HCV diagnosis.

Potential information bias in the training set between HCVab+ and HCVab− patients was mitigated using propensity score matching (PSM). Propensity scores were estimated using logistic regression based on three covariates: (i) length of the lookback period, (ii) number of EHR records during that period, and (iii) attendance at HCV-related core departments (Gastroenterology, Internal Medicine, and Infectious Diseases). These variables were selected to account for differences in healthcare utilization and clinical follow-up between patients with and without documented HCV serology. Matching was then performed using 1:1 nearest neighbor matching without replacement and without a caliper. This approach aimed to balance non-biological factors related to data availability and healthcare contact while preserving clinically relevant variability for model development.

While matching resulted in a balanced sample (50% prevalence) which facilitates model learning, we assessed real-world generalizability using an independent validation set where no matching was applied and the natural prevalence of HCV infection was preserved. Seven full semi-supervised classification algorithms [K-neighbors (KN), Decision Tree (DT), Random Forest (RF), Extreme Gradient Boosting (XGB), Support Vector Machine (SVM), logistic regression (LR), and Gaussian Naïve Bayes (NB)] were applied to predict HCV status based on selected features, including demographics, risk factors, comorbidities, and clinical events. We used a semi-supervised self-training strategy to incorporate unlabeled patients with unknown HCV status. A supervised classifier iteratively labeled high-confidence cases and retrained until convergence, improving generalizability across all model families. We then built parsimonious models with fewer predictors and evaluated them in an independent validation set. Robustness was assessed through sensitivity analyses comparing models with vs. without PSM and supervised vs. semi-supervised learning, confirming that bias-mitigation steps did not compromise performance.

For a detailed explanation of the predictive model training and validation, please refer to the Supplementary Methods.

Data analysis and description

Categorical variables were presented as frequencies and numerical variables as median and inter-quartile range (Q1, Q3). Data was analyzed and represented using “R” software (version 4.0.2) and Python (version 3.7.12).

ResultsDescriptive analysis

We analyzed a total of 49,704,746 de-identified EHRs corresponding to 2,440,358 patients during the study period. Among them, 55,521 were included in the study. As shown in Fig. 1, in the training set, 32,093 patients out of 44,235 had a documented HCV serology status (labeled patients): 10,937 (34.1%) were HCV seropositive (HCVab+) and 21,156 (65.9%) were HCV seronegative (HCVab−). In the validation set, 8708 out of 11,286 patients were labeled, where 3283 (37.7%) were HCVab+ and 5,425 (62.3%) were HCVab−. We also considered a further 12,142 unlabeled patients for the training set and 2578 patients in the validation set, for whom HCV serology status was unknown (HCVabx).

Figure 1.

Data source and population. Data from HCV patients’ EHRs from six tertiary hospitals from the Spanish National Healthcare Network were extracted and analyzed using EHRead® technology based on NLP. The source population was 2,440,358 patients throughout the study period (January 2014 to December 2018). Four hospitals from the central region of Spain were included in the training set (blue) and two hospitals from the northern half of Spain comprised the validation set (red). The number of labeled vs. unlabeled patients is shown for each set, where the labeled patients were either HCV seropositive (HCVab+) or seronegative (HCVab−) and unlabeled (HCVabx) patients were those for which no information regarding HCV serology status was detected. EHR: electronic health record; HCV: hepatitis C virus; NLP: natural language processing.

The demographic and clinical characteristics of each patient cohort according to HCV serology status (including those HCVab− patients after the PSM) are described in Table 1 and Table S4. Patients’ distribution and metrics before and after PSM are shown in Fig. S1 and Table S5.

Table 1.

Demographic and clinical characteristics of all patient groups according to HCV serology status.

  Training setValidation set
  HCVab+  HCVab−  HCVab−(matched)  HCVabx  HCVab+  HCVab−  HCVabx 
Total  10937  21156  10937  12142  3283  5425  2578 
Age (years)
Median (Q1, Q3)  54.7 (46.3, 65.9)  41.7 (33.5, 60)  42.9 (33.7, 60)  51.5 (38.9, 65.2)  57.5 (49.1, 72.7)  52.9 (37.5, 68.9)  56.7 (45.7, 72.5) 
Sex
Male, n (%)  6181 (56.5)  7974 (37.7)  4537 (41.5)  5933 (48.9)  2026 (61.7)  2613 (48.2)  1470 (57) 
Female, n (%)  4756 (43.5)  13182 (62.3)  6400 (58.5)  6209 (51.1)  1257 (38.3)  2812 (51.8)  1108 (43) 
HCV risk factors
Drugs (IDU), n (%)  356 (3.3)  110 (0.5)  59 (0.5)  181 (1.5)  280 (8.5)  107 (2)  186 (7.2) 
Blood transfusions, n (%)  1423 (13)  2715 (12.8)  1222 (11.2)  1275 (10.5)  487 (14.8)  729 (13.4)  272 (10.6) 
Piercings, n (%)  741 (6.8)  1252 (5.9)  548 (5)  604 (5)  246 (7.5)  462 (8.5)  180 (7) 
Tattoos, n (%)  116 (1.1)  136 (0.6)  89 (0.8)  84 (0.7)  55 (1.7)  27 (0.5)  11 (0.4) 

HCVab+: HCV seropositive; HCVab−: HCV seronegative; HCVabx: unknown HCV serology status; HCVab− (matched): HCV seronegative after applying the PSM; IDU: injecting drug users.

Full predictive models

The full list of selected study predictors contained 117 variables. Table S6 shows cross-validation metrics of all models after the PSM on the training set. The best-performing model was XGB (AUC–ROC mean, SD: 0.809, 0.027) and LR showed similar performance (AUC–ROC mean, SD: 0.791, 0.025). Additional sensitivity analyses comparing models trained with and without PSM and supervised vs. semi-supervised approaches are presented in Table S7. To assess model generalizability, we validated both models using the validation set (Table 2). We obtained robust AUC–ROCs (mean, 95% CI) for the full XGB (0.794, 0.784–0.803) and full LR (0.779, 0.769–0.788). Validation curves and calibration plots are shown in Fig. 2A and Fig. S2, respectively.

Table 2.

Validation metrics of the full and reduced models on the validation set.

Model  Accuracy  Precision  Recall  F1-score  F2-score  AUC–ROC 
Full XGB  0.713 (0.703–0.723)  0.594 (0.583–0.604)  0.757 (0.741–0.772)  0.665 (0.654–0.676)  0.717 (0.704–0.73)  0.794 (0.784–0.803) 
Full LR  0.701 (0.691–0.71)  0.584 (0.573–0.595)  0.718 (0.702–0.734)  0.644 (0.633–0.655)  0.686 (0.673–0.7)  0.779 (0.769–0.788) 
Reduced XGB  0.676 (0.666–0.685)  0.562 (0.55–0.573)  0.643 (0.627–0.66)  0.6 (0.588–0.611)  0.625 (0.611–0.639)  0.752 (0.741–0.762) 
Reduced LR  0.63 (0.619–0.64)  0.507 (0.496–0.518)  0.62 (0.602–0.636)  0.558 (0.545–0.57)  0.594 (0.579–0.608)  0.724 (0.713–0.734) 

Mean (95% CI) are presented for each metric. AUC–ROC: area under the curve–receiver operating characteristics; LR: logistic regression; XGB: Extreme Gradient Boosting.

Figure 2.

Performance curves for predictive models. Performance curves for the gradient boosting trees and logistic regression models are displayed for both full (A) and reduced (B) models. For each, the receiver operator characteristic curves (ROC) are shown on the left and the precision-recall curves appear on the right. The ROC curve illustrates the diagnostic ability of a binary classifier system as its discrimination threshold is varied and the precision-recall curve shows the tradeoff between precision and recall for different thresholds. The higher area under the curve in both cases, the better performance the model has. Dashed lines represent random classifiers.

Fig. 3 depicts the clinical features with the highest OR in the selected models, grouped by clinical domain. XGB feature importance (FI) and LR coefficients and OR obtained by the models for each predictor are shown in Table S8, which is explained further in the Supplementary Results in the “Results” section.

Figure 3.

Schematic representation of key clinical predictors of positive HCV serology grouped by disease cluster. Visual summary of the clinical features with the highest odds ratios (OR) in the full logistic regression model, organized by disease cluster. Within each cluster, features are ordered from highest to lowest OR. The schematic is intended to illustrate the main quantitative findings presented in Table 2 and Supplementary Table S7. In all features included in the figure, LR OR >1. *Indicates categorical features with LR OR ≥5 and numerical features with LR OR ≥1. COPD: chronic obstructive pulmonary disorder; HBV: hepatitis B virus; HDV: hepatitis D virus; HIV: human immunodeficiency virus; HSV: herpes simplex virus; LR: logistic regression; OR: odds ratio.

Reduced predictive models

We identified the most relevant features related to the outcome and selected the optimal feature number as described in the “Methods”, obtaining reduced models that included 10 predictors (Table 3). CV of the retrained, reduced yielded an AUC–ROC mean (SD) of 0.784 (0.023) for XGB and an AUC–ROC (SD) of 0.763 (0.022) for LR (Table S9). Validation metrics (AUC [95% CI]) for the reduced models were 0.752 (0.741–0.762) for XGB and 0.724 (0.713–0.734) for LR (Table 2, Fig. 2B, and Fig. S2). As a proof of concept, the reduced XGB model was implemented to identify potential HCVab+ patients among the HCVabx patients from the validation set. Its output reflected 1164 (45% of the validation HCVabx group) patients that were potentially HCVab+ patients, with the characteristics shown in Table S10. Briefly, 69.2% were male, with a mean age of 56 years, 15.2% presented drug addiction, 13.8% had cirrhosis, and 11.8% had portal hypertension, among others.

Table 3.

Variable importance and coefficients for the reduced Extreme Gradient Boosting (XGB) and logistic regression (LR) predictive models.

  XGB feature importance  LR coefficient  LR odds ratio 
Cirrhosis  0.438  3.28  26.59 
Drug addiction (IDU)  0.115  2.88  17.88 
HIV  0.148  2.73  15.28 
Hepatocellular carcinoma  0.039  2.01  7.46 
Schizophrenia  0.035  1.72  5.56 
HBV  0.022  1.26  3.52 
Male sex  0.015  0.45  1.56 
Portal hypertension  0.009  0.3  1.36 
Age (years)  0.137  0.04  1.04 
Ulcerative colitis  0.042  −1.43  0.24 
Intercept  NA  −2.19  0.11 

The higher the feature importance, the greater the weight of a variable in the XGB model. A positive coefficient (or OR >1) in the LR indicates the presence of a characteristic is associated with a higher risk of HCV infection, whereas a negative coefficient (or OR <1) denotes the characteristic is associated with a lower risk of HCV infection.

Discussion

We developed and validated semi-supervised ML-based predictive models to detect patients with unconfirmed HCV in a hospital setting using available clinical information in free text from EHRs. To our knowledge, this is the first model that has been developed and validated using patients’ EHR data from a public healthcare system using NLP and ML. Our approach applies a semi-supervised learning framework to leverage unlabeled data and incorporates explicit bias-mitigation strategies and multicenter validation. It was entirely data-driven and based on a large sample size and a rich feature set, including both clinical and non-clinical variables extracted through artificial intelligence techniques. Several factors showed positive association with HCV infection (cirrhosis, IDU, HIV coinfection, hepatocellular carcinoma, schizophrenia, HBV, portal hypertension, age, male sex) and one negative association (ulcerative colitis). Models performance was robust, reflecting high classifying power and including less evident variables that correlate with positive HCV serology that could support targeted testing.

Recent ML models for undiagnosed HCV detection mainly rely primarily on structured data (claims, laboratory, or registry datasets) and conventional supervised learning methods, particularly tree-based ensembles such as Random Forest or XGBoost, reporting AUCs of 0.77–0.95.13–16 Most were developed in U.S. or private healthcare systems, where access patterns and testing practices may introduce inclusion bias and limit generalizability to public or universal systems. They also generally lack external validation, bias-mitigation strategies, and integration of unstructured clinical text. In contrast, LiverTAI expands predictive modeling to a universal healthcare context and integrates richer real-world data sources by leveraging NLP to extract both structured and unstructured EHR data across multiple hospitals in Spain's universal healthcare system, applying semi-supervised learning to incorporate unlabeled patient information, and using PSM to reduce non-biological data bias. This combined framework represents a methodological advance that broadens the applicability and interpretability of existing models, while achieving performance (AUC0.79) comparable to models based solely on structured datasets.

We used PSM to account for HCV testing variability while retaining relevant clinical differences between HCVab+ and HCVab− patients during model training, and semi-supervised training to incorporate and learn from the features of HCVabx patients. Together, these approaches minimized bias and the risk of overfitting that could arise from relying solely on advanced cases (i.e., patients with chronic HCV infection) vs. healthy controls for model development. Sensitivity analyses showed that both strategies maintained comparable discrimination metrics, indicating that they did not compromise model performance while supporting fairness and internal validity. The full XGB model (AUC–ROC=0.809) validation on a geographically independent dataset showed the highest performance (AUC–ROC=0.794), while LR models showed similar metrics and were used for interpretability. However, given that models were trained on a matched set and applied to a validation set with a lower HCVab+ prevalence, calibration plots show that our models tend to systematically overestimate risk. Therefore, patient stratification should rely on predicted classes rather than interpreting raw predicted scores as absolute measures of risk. Although not intended for diagnosis, we also classified unlabeled patients from the validation set using the reduced XGB model to identify possible candidates for HCV testing, with potentially positive HCV serology tests. Recall is prioritized over precision at the expense of false positives (as shown by F2-scores), which would nevertheless show up during HCV testing. Further validation could be performed prospectively through active screening campaigns of patients detected as potential positive HCV candidates by the model presented here.

Using our methodology, we detected multiple potential predictors of chronic HCV infection, including conditions previously associated with HCV-related extrahepatic manifestations affecting neuropsychiatric, cardiovascular, and immune-mediated disorders, as well as cancer and inflammatory processes.5,19 Some of them have been found in the literature, either as potential HCV risk factors or in which HCV infection plays a role. Severe mental illness is a well-recognized risk factor for HCV infection.12,20 In line with this, we observed a strong association between schizophrenia (LR OR 4.57) and HCV seropositivity (LR OR 4.57), consistent with previous epidemiological evidence. This association likely reflects social vulnerability, increased exposure to behavioral risk factors, and barriers to healthcare access, rather than disease-specific biological mechanisms. Other neuropsychiatric conditions, including multiple sclerosis (LR OR 5.66), and Parkinson's disease (LR OR 2.43), were also identified as potential predictors of HCV infection; however, these associations may reflect different underlying mechanisms and should be interpreted cautiously. Cardiovascular associations included peripheral arterial disease (LR OR 1.53) and myocarditis (LR OR 1.51), which is in line with other studies.21 A strong association with essential mixed cryoglobulinemia (LR OR 14.93) was also found, as expected, since HCV is considered the major cause of this disorder.22 Systematic reviews have shown that chronic HCV infection is significantly associated with increased risk (approximately 2-fold) of developing lung cancer,23 which is the extrahepatic cancer with the highest LR OR (1.46) in our study. Gastric cancer has also been associated with HCV,24 and had the fourth highest LR OR (1.28) in our study. Although researchers have reported a link between cutaneous porphyria and HCV infection,25 this feature with an LR OR of 3.21 was excluded from the list of predictors (FI=0) in XGB. However, these associations should be interpreted with caution, as our cross-sectional design precludes causal inference and residual confounding cannot be excluded. Some negative associations, such as the one observed with ulcerative colitis, may reflect adjustment effects within multivariable models rather than true inverse relationships. An alternative and clinically plausible explanation is that patients with ulcerative colitis often undergo structured infectious disease screening as part of routine clinical assessment prior to initiating immunosuppressive or biologic therapies, including HCV testing as recommended by clinical guidelines.26 This systematic screening may increase the likelihood of documented negative HCV serology in this population, potentially contributing to the observed negative association in real-world data.

Key strengths include use of routine physician-generated data, ensuring independence from data collection, and the study's setting within the Spanish national healthcare system, granting universal healthcare access to Spanish citizens and avoiding the bias towards insured patients inherent in other ML-based predictive studies to diagnose HCV.13,14 Therefore, our predictive model underscores a commitment to equity and justice by mitigating healthcare access selection biases, thus prioritizing inclusivity and fair representation for vulnerable populations. By including fairness as a central consideration in model design, deployment, and evaluation, we ensure benefit for all patients. Moreover, this large sample size from Southern Europe, a region with relatively high HCV prevalence, guarantees enough variability to capture potentially unknown features alongside classical risk factors of HCV infection. Additionally, using NLP to extract data from the free-text narratives in EHRs enabled us to access an untapped source of information that enriched our predictive models with data from routine clinical practice.27 Extracting unstructured clinical data from EHRs beyond using ICD codes or other structured data alone13,27 improved patient diversity than controlled studies or clinical trials.28 Using NLP to recover data can also reduce bias and maximize generalizability of EHR research.29 In this sense, compared with the classical methodology for extracting clinical information, this tool allows us to quickly and efficiently search through large amounts of medical reports to support clinical research. Furthermore, our study's LR models are more interpretable than previous predictive models and were validated using an independent dataset to detect undiagnosed HCV patients.13 Finally, although the present study was designed as a methodological proof of concept rather than an implementation or clinical decision-making study, its findings support development of decision algorithms for HCV testing, with potential application to other diseases. Future work should include prospective validation and assessment of clinical utility.

Several limitations apply. First, data were derived from retrospective, cross-sectional hospital records, which may be prone to reporting and information bias, and clinically supervised term extraction may limit the features from which ML-based models can learn. Future work should explore data-driven feature discovery and ablation techniques to identify latent variables and assess their contribution on performance and generalizability. Second, the study population was intentionally enriched with patients whose EHRs contained HCV-related mentions or testing information. Therefore, the model was not designed to identify completely asymptomatic or undocumented infections, but rather to detect unconfirmed cases within hospital settings, an opportunistic, hospital-based screening scenario. Consequently, disease prevalence in our dataset was higher than in the general population. Moreover, it is important to note that the study period overlaps with the large-scale expansion of free access to DAAs in Spain from 2015 onwards, which substantially increased treatment uptake and linkage to care. This shift may have influenced documentation practices in EHRs, particularly regarding treatment-related variables and the recording of HCV status. As a result, some predictors may partly reflect evolving care pathways and data capture mechanisms rather than stable biological associations. Third, although Spain's National Health System ensures relatively homogeneous access to care, variability among hospitals is expected. To account for potential institutional differences, we pooled data from multiple hospitals for model training and validated performance in independent centers located in different regions. This spatial validation strategy strengthens robustness within the national context; however, broader external validation in other countries and healthcare systems is needed to confirm generalizability. Fourth, structured pharmacy data were available in only two of the six hospitals. Nevertheless, relevant pharmacological information was consistently retrieved from free-text narratives written by healthcare professionals in discharge summaries, progress notes, and outpatient or day-hospital records across all departments, providing comprehensive coverage and mitigating the potential impact of missing structured pharmacy fields. Fifth, not fully standardized EHR formats highlight the value of NLP for harmonizing heterogeneous, unstructured information. Some observed associations, such as those involving schizophrenia or ulcerative colitis, should be interpreted cautiously, as the cross-sectional and retrospective design precludes causal inference and residual confounding cannot be excluded. Finally, although LiverTAI demonstrated robust performance and interpretability across hospitals, further work is needed to evaluate its clinical impact, integration into workflows, cost-effectiveness, and to incorporate a formal fairness assessment to ensure equitable performance across patient subgroups. Future efforts should include prospective validation through pragmatic or impact studies, supported by decision-curve analysis and cost-effectiveness modelling. Broader implementation will also require NLP linguistic adaptation and a minimal EHR infrastructure, potentially limiting applicability in low-resource settings. This study should therefore be viewed as a proof of concept for methodological feasibility and bias mitigation rather than clinical deployment.

In conclusion, this study demonstrates that NLP-extracted EHR data combined with semi-supervised ML can identify hospitalized patients at risk of unconfirmed HCV infection. By leveraging free-text clinical narratives and bias-mitigation strategies, LiverTAI enables risk-stratified case finding within hospital workflows, potentially supporting testing efficiency. Addressing community underdiagnosis will require complementary population-based strategies. Future work should prioritize prospective validation with standardized data capture, quality auditing, and integration of structured variables to enhance reproducibility and clinical reliability, alongside international collaborations to confirm generalizability and global applicability.

Authors’ contributions

Conception and design: All authors. Collection and assembly of data: VF, CLO. Data analysis: VF, CLO. Interpretation: All authors. Manuscript writing: VF, CLO, FL, ES. Manuscript revision and approval: All authors. JLC is responsible for the overall content [as guarantor].

Ethical considerations

This study was classified as a “non-post-authorization study” by the Spanish Agency of Medicines and Health Products (AEMPS) and was approved by the Independent Ethics Committee of each participating hospital. The cohort study was conducted in compliance with legal and regulatory requirements and followed generally accepted research practices described in the Helsinki Declaration in its latest edition, Good Pharmacoepidemiology Practices, and applicable local regulations. Patient consent was waived, since data were retrospectively analyzed from patients’ EHRs, anonymized, and aggregated in an irreversibly dissociated manner. Because this was a retrospective study based on real-world data, data collection and assessment for both potential predictors and outcomes were performed in a blind manner. Both the outcome and predictors were automatically extracted from EHRs, reflecting routine clinical practice, and therefore outcome assessment was based on clinical decisions made by each physician, with no influence from the predictors. No patient or public involvement was undertaken in the design, conduct, reporting, interpretation, or dissemination of this study.

Funding

This study was funded by AbbVie.

Conflicts of interest

JL Calleja Panero has been a consultant and speaker for AbbVie and Gilead Sciences. E Sanz is an employee at AbbVie and may have AbbVie stocks. V Fanjul is an employee at Savana Research. F Ledesma worked as an employee for AbbVie SLU during the development of the present research and manuscript writing but is no longer a member of the company. CL Oeste worked as an employee for Savana Research during the development of the present research and manuscript writing but is no longer a member of the company. G Sánchez-Antolín, G de la Poza, V Aguilera, L Hidalgo, and B Cuyàs have nothing to disclose.

Data availability

Data are available on reasonable request to the authors. Thereafter, the committee of the project, together with the Ethics Committee of the hospitals involved, will assess the proposal and potentially proceed to the data sharing. The data used in this study consist of de-identified electronic health records (EHRs) processed with EHRead® technology. Due to privacy regulations and proprietary restrictions, raw EHR data and analytical code cannot be made publicly available. However, de-identified aggregated data and model specifications may be shared upon reasonable request, subject to institutional and ethical approval. This retrospective, non-interventional study was not registered in any public registry, as it involved the secondary use of routinely collected EHR data, with no prospective assignment of interventions or impact on patient care. In accordance with TRIPOD-AI guidance, registration is not mandatory for this type of predictive modeling study.

Acknowledgements

Support to conduct the study was provided by Savana who took part with AbbVie in the study design, research, analysis, data collection, writing, and data interpretation, all funded by AbbVie (award/grant number: N/A). No honoraria or payments were made for authorship. AbbVie was responsible for reviewing and for the approval of the publication. The authors thank Sara Paris from AbbVie for taking part in the study design, Regina Santos de LaMadrid from AbbVie for writing and data interpretation, and Miren Taberna MD PhD, David Casadevall MD PhD, Hugo Casero, Natalia Polo from Savana for study design, research, analysis, data collection, and writing within the AbbVie funding.

Appendix B
Supplementary data

The following are the supplementary data to this article:

Icono mmc1.doc

References
[1]
N. Salari, M. Kazeminia, N. Hemati, M. Ammari-Allahyari, M. Mohammadi, S. Shohaimi.
Global prevalence of hepatitis C in general population: a systematic review and meta-analysis.
Travel Med Infect Dis, 46 (2022),
[2]
A. Estirado Gómez, S. Justo Gil, A. Limia, A. Avellón, A. Arce Arnáez, R. González-Rubio, et al.
Prevalence and undiagnosed fraction of hepatitis C infection in 2018 in Spain: results from a national population-based survey.
Eur J Public Health, 31 (2021), pp. 1117-1122
[3]
Global Health Sector Strategy on Viral Hepatitis 2016–2021.
Towards ending viral hepatitis [Internet].
World Health Organization, (2016),
[4]
C.W. Spearman, G.M. Dusheiko, M. Hellard, M. Sonderup.
Hepatitis C.
Lancet, 394 (2019), pp. 1451-1466
[5]
L. Kuna, J. Jakab, R. Smolic, G.Y. Wu, M. Smolic.
HCV extrahepatic manifestations.
J Clin Transl Hepatol, 7 (2019), pp. 1-11
[6]
C. Thomadakis, I. Gountas, K. Gountas, N. Nuño, B. Brime, R. Sendino, et al.
Prevalence of active HCV infection in Spain in 2022 using multiparameter evidence synthesis.
J Viral Hepat, 33 (2026),
[7]
K.L. Cheung, P.M. Ten Klooster, C. Smit, H. De Vries, M.E. Pieterse.
The impact of non-response bias due to sampling in public health studies: a comparison of voluntary versus mandatory recruitment in a Dutch national survey on adolescent health.
BMC Public Health, 17 (2017), pp. 276
[8]
O. Obienu, S. Nwokediuko, A. Malu, O.A. Lesi.
Risk factors for hepatitis C virus transmission obscure in Nigerian patients.
Gastroenterol Res Pract, 2011 (2011), pp. 1-4
[9]
E.Y. Ho, N.B. Ha, A. Ahmed, W. Ayoub, T. Daugherty, G. Garcia, et al.
Prospective study of risk factors for hepatitis C virus acquisition by Caucasian, Hispanic, and Asian American patients.
[10]
J.C. Ahn, A. Connell, D.A. Simonetto, C. Hughes, V.H. Shah.
Application of artificial intelligence for the diagnosis and treatment of liver diseases.
Hepatology, 73 (2021), pp. 2546-2563
[11]
E. Audureau, F. Carrat, R. Layese, C. Cagnot, T. Asselah, D. Guyader, et al.
Personalized surveillance for hepatocellular carcinoma in cirrhosis – using machine learning adapted to HCV status.
J Hepatol, 73 (2020), pp. 1434-1445
[12]
H.I. Shousha, A.H. Awad, D.A. Omran, M.M. Elnegouly, M. Mabrouk.
Data mining and machine learning algorithms using IL28B genotype and biochemical markers best predicted advanced liver fibrosis in chronic hepatitis C.
Jpn J Infect Dis, 71 (2018), pp. 51-57
[13]
O.M. Doyle, N. Leavitt, J.A. Rigg.
Finding undiagnosed patients with hepatitis C infection: an application of artificial intelligence to patient claims data.
Sci Rep, 10 (2020),
[14]
J. Rigg, O. Doyle, N. McDonogh, N. Leavitt, R. Ali, A. Son, et al.
Finding undiagnosed patients with hepatitis C virus: an application of machine learning to US ambulatory electronic medical records.
BMJ Health Care Inform, 30 (2023),
[15]
M.A. Hezari, M. Baes, A.A. Hezari, M. Hassanbabaei.
Advanced predictive modeling for hepatitis C diagnosis using machine learning.
Clin Mol Epidemiol, 1 (2024), pp. 12
[16]
N. Dagan, O. Magen, M. Leshchinsky, M. Makov-Assif, M. Lipsitch, B.Y. Reis, et al.
Prospective evaluation of machine learning for public health screening: identifying unknown hepatitis C carriers.
[17]
J.L. Calleja Panero, G. De La Poza, L. Hidalgo, M.V. Aguilera Sancho-Tello, X. Torras, R. Santos de Lamadrid, et al.
Patient journey of individuals tested for HCV in Spain: LiverTAI, a retrospective analysis of EHRs through natural language processing.
Gastroenterol Hepatol, 46 (2023), pp. 491-503
[18]
L. Canales, S. Menke, S. Marchesseau, A. D’Agostino, C. del Rio-Bermudez, M. Taberna, et al.
Assessing the performance of clinical natural language processing systems: development of an evaluation methodology.
JMIR Med Inform, 9 (2021),
[19]
P. Cacoub, C. Comarmond.
Considering hepatitis C virus infection as a systemic disease.
Semin Dial, 32 (2019), pp. 99-107
[20]
R. Wei, J. Wang, X. Wang, G. Xie, Y. Wang, H. Zhang, et al.
Clinical prediction of HBV and HCV related hepatic fibrosis using machine learning.
EBioMedicine, 35 (2018), pp. 124-132
[21]
Y.-H. Hsu, C.-H. Muo, C.-Y. Liu, W.-C. Tsai, C.-C. Hsu, F.-C. Sung, et al.
Hepatitis C virus infection increases the risk of developing peripheral arterial disease: a 9-year population-based cohort study.
J Hepatol, 62 (2015), pp. 519-525
[22]
G. Lauletta, S. Russi, V. Conteduca, L. Sansonno.
Hepatitis C virus infection and mixed cryoglobulinemia.
Clin Dev Immunol, 2012 (2012), pp. 1-11
[23]
B. Ponvilawan, N. Charoenngam, P. Rujirachun, P. Wattanachayakul, S. Tornsatitkul, T. Rittiphairoj.
Chronic hepatitis C virus infection is associated with an increased risk of lung cancer: a systematic review and meta-analysis.
[24]
Y. Yang, Z. Jiang, W. Wu, L. Ruan, C. Yu, Y. Xi, et al.
Chronic hepatitis virus infection are associated with high risk of gastric cancer: a systematic review and cumulative analysis.
Front Oncol, 11 (2021),
[25]
J.P. Gisbert, L. García-Buey, J. María Pajares, R. Moreno-Otero.
Prevalence of hepatitis C virus infection in porphyria cutanea tarda: systematic review and meta-analysis.
J Hepatol, 39 (2003), pp. 620-627
[26]
J.M.F. Chebli, P.D. Gaburri, L.A. Chebli, T.C. da Rocha Ribeiro, A.L. Pinto, O. Ambrogini Júnior, et al.
A guide to prepare patients with inflammatory bowel diseases for anti-TNF-α therapy.
Med Sci Monit, 20 (2014), pp. 487-498
[27]
J.R. Ayala Solares, F.E. Diletta Raimondi, Y. Zhu, F. Rahimian, D. Canoy, J. Tran, et al.
Deep learning for electronic health records: a comparative review of multiple deep neural architectures.
J Biomed Inform, 101 (2020),
[28]
M.R. Cowie, J.I. Blomster, L.H. Curtis, S. Duclaux, I. Ford, F. Fritz, et al.
Electronic health records to facilitate clinical research.
Clin Res Cardiol, 106 (2017), pp. 1-9
[29]
S. Khurshid, C. Reeder, L.X. Harrington, P. Singh, G. Sarma, S.F. Friedman, et al.
Cohort design and natural language processing to reduce bias in electronic health records research.
NPJ Digit Med, 5 (2022), pp. 47
asdasdasd
Opciones de artículo
Herramientas
Material suplementario