metricas

Open Respiratory Archives

Suggestions
Open Respiratory Archives Synthetic Lung-cancer Cohorts Generated by a Large Language Model: Epidemiologic...
Journal Information
Vol. 8. Issue 1.
(January - March 2026)
Cite
Cite
Share
Download PDF
More article options
Visits
764
Vol. 8. Issue 1.
(January - March 2026)
Short Communication
Full text access

Synthetic Lung-cancer Cohorts Generated by a Large Language Model: Epidemiological Validity Assessment

Cohortes sintéticas de cáncer de pulmón generadas por inteligencia artificial: evaluación de la validez epidemiológica
Visits
764
Álvaro Fuentes-Martína,
Corresponding author
alvarofuentesmartin@gmail.com

Corresponding author.
, Julio Mayolb, Bárbara Segura Méndezc, Ángel Cilleruelo-Ramosd
a Servicio de Cirugía Torácica, Hospital Clínico Universitario de Valladolid, Universidad de Valladolid, Spain
b Hospital Clínico San Carlos, IdISSC, Universidad Complutense de Madrid, Spain
c Servicio de Cirugía Cardiaca, Hospital Universitario de Salamanca, Spain
d Servicio de Cirugía Torácica, Hospital Clínico Universitario de Valladolid, Universidad de Valladolid, Spain
This item has received
Article information
Abstract
Full Text
Bibliography
Download PDF
Statistics
Figures (2)
fig0005
fig0010
Additional material (1)
Abstract

Large language models (LLMs) are increasingly used in medicine for clinical reasoning and educational simulation. This study assessed the epidemiological plausibility of a synthetic lung-cancer cohort generated by ChatGPT-4.0. A total of 102 virtual cases were created in Spanish using structured prompts including demographic, histologic, and molecular variables. When descriptively compared with international datasets (GLOBOCAN 2020, SEER, and biomarker meta-analyses), the cohort reproduced general disease patterns but showed statistically significant deviations (p<0.05): early-stage disease and EGFR-positive tumors were overrepresented, while advanced stages, ALK rearrangements, and extreme PD-L1 values were underrepresented. These discrepancies likely reflect biases in model training data and the probabilistic nature of generative language models. Despite this quantified generative bias, the utility of these cohorts for non-epidemiological tasks like educational simulation is discussed, provided methodological transparency is maintained.

Keywords:
Synthetic cohorts
Large language models
Thoracic oncology
Resumen

Los modelos de lenguaje de gran escala (LLM) se utilizan cada vez más en medicina para el razonamiento y la simulación clínica. Este estudio evaluó la plausibilidad epidemiológica de una cohorte sintética de cáncer de pulmón generada por ChatGPT-4.0. Se crearon un total de 102 casos sintéticos mediante prompts estructurados que incluían variables demográficas, histológicas y moleculares. Al compararla con bases de datos epidemiológicas, la cohorte reprodujo patrones generales de la enfermedad, aunque mostró desviaciones estadísticamente significativas (p<0,05): sobrerrepresentación de estadios iniciales y de EGFR frente a la infrarepresentación de estadios avanzados, reordenamientos ALK y valores extremos de PD-L1. Estas discrepancias reflejan sesgos en el entrenamiento y la naturaleza probabilística de los modelos generativos. A pesar de este sesgo generativo cuantificado, se discute la utilidad de estas cohortes para tareas no epidemiológicas como la educación médica, siempre que se mantenga la transparencia metodológica.

Palabras clave:
Cohortes sintéticas
Modelos de lenguaje de gran escala
Oncología torácica
Graphical abstract
Full Text

Large language models (LLMs) are increasingly applied in medicine, supporting clinical reasoning and educational simulation.1,2 In oncology, they have been explored as decision-support tools3,4 and for automated clinical case generation. One of their most promising uses is the creation of synthetic clinical cohorts – realistic yet fictitious datasets that emulate real patient populations while preserving privacy.5 Despite their potential, the epidemiological representativeness of LLM-generated cohorts remains unvalidated. This study aimed to evaluate the epidemiological plausibility of a synthetic lung-cancer cohort generated by ChatGPT-4.0, comparing key demographic, histologic, and molecular variables against international data.

We conducted a descriptive, exploratory study to assess the internal consistency and external plausibility of data generated by ChatGPT-4.0 (OpenAI). A convenience sample of 102 virtual patients was generated, a size deemed sufficient for an initial exploratory descriptive assessment using Spanish-language prompts structured in a Role–Task–Format framework. Generation occurred in batches of five patients per iteration, reflecting the model's operational text limit. No corrections or parameter adjustments were introduced between batches to preserve methodological consistency.

Prompts instructed ChatGPT to create clinically coherent profiles including demographic, oncologic, and molecular variables: age, sex, histologic subtype, TNM stage, and biomarkers (Epidermal Growth Factor Receptor [EGFR], Anaplastic Lymphoma Kinase [ALK], and Programed Death-Ligand 1 [PD-L1]) (Supplementary Annex 1). These variables were extracted manually and compared descriptively with global reference datasets (GLOBOCAN 2020, SEER) and biomarker meta-analyses6–11 (Fig. 1). Continuous variables were summarized as mean±SD and categorical variables as frequencies. Observed cohort frequencies (e.g., stage, biomarkers) were compared against expected population benchmarks (derived from Refs. 6–11) using Chi-squared (χ2) (p<0.05 was considered statistically significant).

Fig. 1.

Percentage comparison between the synthetic cohort and real-world data.

The synthetic cohort included 102 virtual patients, with a mean age 66.5±6.0 years (range 51–79). In comparison, global oncology registries7,8 indicate a median diagnostic age of approximately 70 years, suggesting a slightly younger synthetic population. Sex distribution was balanced (51 men, 51 women), a distribution significantly deviating from the expected male predominance (≈65–70%)7 (χ2=14.24, p<0.01).

Histologic distribution comprised adenocarcinoma 52%, squamous-cell carcinoma 41%, and small-cell carcinoma 7%. No large-cell carcinoma was generated. Although the adenocarcinoma proportion was similar to that observed in population-based studies6 (≈45–55%), small-cell carcinoma was underrepresented (10–15% expected), and large-cell carcinoma (≈5–7%) was absent, a distribution significantly deviating from population data (χ2=11.82, p<0.01).

Staging analysis showed a predominance of early-stage disease: 65% (Stages I–II), 17% (Stage III), and 18% (Stage IV). Real-world data indicate that ≈40% of non-small-cell lung-cancer (NSCLC) are diagnosed at Stage IV. In our synthetic NSCLC cohort (N=95), only 18% were Stage IV, confirming a significant over-representation of localized stages (χ2=19.34, p<0.01).

Among the 95 simulated NSCLC cases, 45% harbored activating EGFR mutations, a rate significantly higher than Western prevalence (≈15%9; χ2=68.24, p<0.01). No ALK rearrangements were identified (0%), a significant deviation from the expected 3–5% prevalence10 (χ2=3.96, p<0.05). PD-L1 expression was uniformly intermediate (1–49%), with no negative or highly positive cases (≥50%) – a distribution significantly deviating from clinical cohorts11 (χ2=142.5, p<0.01).

Taken together, the synthetic cohort reproduced general disease patterns but showed notable deviations in key variables, particularly stage distribution and molecular profile, indicating statistically significant generative bias relative to real-world epidemiological distributions.

This exploratory study evaluated the epidemiological consistency of a synthetic lung-cancer cohort generated by ChatGPT-4.0. Although the model produced clinically plausible, coherent profiles, it exhibited systematic biases – most notably a predominance of early-stage disease and EGFR-positive tumors. These deviations likely reflect the probabilistic nature of generative models and the uneven representation of clinical scenarios in training data.

The over-representation of early-stage cases suggests a narrative bias favoring curative, well-structured clinical stories over terminal presentations. The under-representation of small-cell and absence of ALK-positive cases may relate to their lower frequency and visibility in the scientific literature. Likewise, the homogeneous PD-L1 pattern indicates that ChatGPT tends to assign intermediate values under uncertainty, reflecting a limitation in quantitative realism.

Such discrepancies align with prior observations of LLM “hallucinations,” i.e., systematic deviations from expected patterns due to the statistical weighting of learned text.12 These findings underscore the need to impose explicit population-level constraints when generating synthetic datasets. Without controlled proportions or post-generation validation, representativeness cannot be assumed.

From an educational standpoint, however, generating structured, realistic clinical cases retains considerable value. Synthetic cohorts can support virtual tumor-board exercises, simulation-based teaching, and preliminary algorithm testing. Their utility lies more in training and hypothesis generation than in precise epidemiological representation.

Reproducibility represents an additional methodological challenge. Generative models evolve continuously, and outputs obtained with one version may not be replicable with subsequent iterations, even when using identical prompts. Clear documentation of the model version, generation date, and prompting framework is therefore essential to maintain methodological transparency and facilitate reproducibility.

Another limitation concerns the batch-generation process. Because the cohort was created in groups of five patients – an operational restriction of ChatGPT-4.0 – sequential generation may have introduced distributional bias. Although no systematic drift was visually observed, future research should quantify inter-batch variability to confirm data stability.

Despite these limitations, LLM-generated cohorts demonstrate the feasibility of producing complex, internally coherent clinical datasets without compromising confidentiality. With appropriate methodological safeguards and validation procedures, this approach could enhance educational realism and foster innovation in thoracic-oncology training.

The synthetic lung-cancer cohort generated with ChatGPT-4.0 reproduced general disease patterns but showed statistically significant deviations in stage distribution and biomarker prevalence. This quantified generative bias limits its epidemiological representativeness, the findings illustrate both the promise and the boundaries of LLM-based data generation. Responsible implementation requires methodological transparency, explicit acknowledgment of biases, and quantitative validation. With continued refinement, large language models could become valuable complementary tools for simulation, medical education, and early-phase research in thoracic oncology.

Ethical statement

This study did not involve real patients or human subjects. Instead, it was based entirely on synthetic data generated through an artificial intelligence model (ChatGPT-4.0), and no identifiable or confidential patient information was used. Therefore, the requirement for informed consent was waived. Nevertheless, the study protocol was reviewed and approved by the Clinical Research Ethics Committee of our institution (Reference: PI-25-146-C), and the project was conducted in accordance with the ethical principles of the Declaration of Helsinki (2013 revision). The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.

Declaration of generative AI and AI-assisted technologies in the writing process

During the preparation of this work, the authors used the Generative Pre-trained Transformer 4 (ChatGPT-4) not only for grammar review and translation, but also for the structured generation of a synthetic cohort of 102 virtual patients with lung cancer, based on predefined clinical, molecular, and psychosocial parameters. This simulated dataset was used for research purposes within the framework of this study. After using this tool, the authors reviewed, validated, and edited the output as necessary, and take full responsibility for the content of the publication.

Funding

This research received no external funding.

Authors’ contributions

All authors contributed substantially to the design of the study, data analysis, manuscript drafting, and critical revision of its content. All authors have read and approved the final version of the manuscript.

Conflicts of interest

The authors declare no conflicts of interest.

Data availability

All data used in this study were synthetically generated using the ChatGPT-4.0 model (OpenAI) and do not correspond to real individuals. The full methodology used for data generation is described in the Methods section. Examples of the synthetic cases are provided in the Supplementary Material. No real patient data were accessed or used in this study.

Appendix B
Supplementary data

The following are the supplementary data to this article:

Icono mmc1.doc

References
[1]
M.M. Zitu, T.D. Le, T. Duong, S. Haddadan, M. Garcia, R. Amorrortu, et al.
Large language models in cancer: potentials, risks, and safeguards.
BJR Artif Intell, 2 (2024),
[2]
Á. Fuentes-Martín, Á. Cilleruelo-Ramos, B. Segura-Méndez, J. Mayol.
Can an artificial intelligence model pass an examination for medical specialists?.
Arch Bronconeumol, 59 (2023), pp. 534-536
[3]
J. Zabaleta, B. Aguinagalde, I. Lopez, A. Fernandez-Monge, J.A. Lizarbe, M. Mainer, et al.
Utility of artificial intelligence for decision making in thoracic multidisciplinary tumor boards.
J Clin Med, 14 (2025), pp. 399
[4]
V. Nardone, F. Marmorino, M.M. Germani, N. Cichowska-Cwalińska, V.S. Menditti, P. Gallo, et al.
The role of artificial intelligence on tumor boards: perspectives from surgeons, medical oncologists and radiation oncologists.
Curr Oncol, 31 (2024), pp. 4984-5007
[5]
Q. Jin, Z. Wang, C.S. Floudas, F. Chen, C. Gong, D. Bracken-Clarke, et al.
Matching patients to clinical trials with large language models.
Nat Commun, 15 (2024), pp. 9074
[6]
Y. Zhang, S. Vaccarella, E. Morgan, M. Li, J. Etxeberria, E. Chokunonga, et al.
Global variations in lung cancer incidence by histological subtype in 2020: a population-based study.
Lancet Oncol, 24 (2023), pp. 1206-1218
[7]
SEER Cancer Stat Facts: Lung and Bronchus Cancer. National Cancer Institute. Available from: https://seer.cancer.gov/statfacts/html/lungb.html. [Accessed 2 April 2025].
[8]
J. Subramanian, D. Morgensztern, B. Goodgame, M.Q. Baggstrom, F. Gao, J. Piccirillo, et al.
Distinctive characteristics of non-small cell lung cancer (NSCLC) in the young: a surveillance, epidemiology, and end results (SEER) analysis.
J Thorac Oncol, 5 (2010), pp. 23-28
[9]
R. Rosell, T. Moran, C. Queralt, et al.
Screening for epidermal growth factor receptor mutations in lung cancer.
N Engl J Med, 361 (2009), pp. 958-967
[10]
V. Cognigni, F. Pecci, A. Lupi, G. Pinterpe, C. De Filippis, C. Felicetti, et al.
The landscape of ALK-rearranged non-small cell lung cancer: a comprehensive review of clinicopathologic, genomic characteristics, and therapeutic perspectives.
Cancers (Basel), 14 (2022), pp. 4765
[11]
D.M. Hwang, T. Albaqer, R.C. Santiago, J. Weiss, J. Tanguay, M. Cabanero, et al.
Prevalence and heterogeneity of PD-L1 expression by 22C3 assay in routine population-based and reflexive clinical testing in lung cancer.
J Thorac Oncol, 16 (2021), pp. 1490-1500
[12]
N. McKenna, T. Li, L. Cheng, M.J. Hosseini, M. Johnson, M. Steedman.
Sources of hallucination by Large Language Models on inference tasks.
In the 17th conference of the European chapter of the association for computational linguistics: findings of EACL 2023, Association for Computational Linguistics (ACL), (2023), pp. 2758-2774 http://dx.doi.org/10.18653/v1/2023.findings-emnlp.182
Copyright © 2025. Sociedad Española de Neumología y Cirugía Torácica (SEPAR)
Download PDF
asdasdasd
Article options
Tools
Supplemental materials