Large language models (LLMs) are increasingly used in medicine for clinical reasoning and educational simulation. This study assessed the epidemiological plausibility of a synthetic lung-cancer cohort generated by ChatGPT-4.0. A total of 102 virtual cases were created in Spanish using structured prompts including demographic, histologic, and molecular variables. When descriptively compared with international datasets (GLOBOCAN 2020, SEER, and biomarker meta-analyses), the cohort reproduced general disease patterns but showed statistically significant deviations (p<0.05): early-stage disease and EGFR-positive tumors were overrepresented, while advanced stages, ALK rearrangements, and extreme PD-L1 values were underrepresented. These discrepancies likely reflect biases in model training data and the probabilistic nature of generative language models. Despite this quantified generative bias, the utility of these cohorts for non-epidemiological tasks like educational simulation is discussed, provided methodological transparency is maintained.
Los modelos de lenguaje de gran escala (LLM) se utilizan cada vez más en medicina para el razonamiento y la simulación clínica. Este estudio evaluó la plausibilidad epidemiológica de una cohorte sintética de cáncer de pulmón generada por ChatGPT-4.0. Se crearon un total de 102 casos sintéticos mediante prompts estructurados que incluían variables demográficas, histológicas y moleculares. Al compararla con bases de datos epidemiológicas, la cohorte reprodujo patrones generales de la enfermedad, aunque mostró desviaciones estadísticamente significativas (p<0,05): sobrerrepresentación de estadios iniciales y de EGFR frente a la infrarepresentación de estadios avanzados, reordenamientos ALK y valores extremos de PD-L1. Estas discrepancias reflejan sesgos en el entrenamiento y la naturaleza probabilística de los modelos generativos. A pesar de este sesgo generativo cuantificado, se discute la utilidad de estas cohortes para tareas no epidemiológicas como la educación médica, siempre que se mantenga la transparencia metodológica.
Large language models (LLMs) are increasingly applied in medicine, supporting clinical reasoning and educational simulation.1,2 In oncology, they have been explored as decision-support tools3,4 and for automated clinical case generation. One of their most promising uses is the creation of synthetic clinical cohorts – realistic yet fictitious datasets that emulate real patient populations while preserving privacy.5 Despite their potential, the epidemiological representativeness of LLM-generated cohorts remains unvalidated. This study aimed to evaluate the epidemiological plausibility of a synthetic lung-cancer cohort generated by ChatGPT-4.0, comparing key demographic, histologic, and molecular variables against international data.
We conducted a descriptive, exploratory study to assess the internal consistency and external plausibility of data generated by ChatGPT-4.0 (OpenAI). A convenience sample of 102 virtual patients was generated, a size deemed sufficient for an initial exploratory descriptive assessment using Spanish-language prompts structured in a Role–Task–Format framework. Generation occurred in batches of five patients per iteration, reflecting the model's operational text limit. No corrections or parameter adjustments were introduced between batches to preserve methodological consistency.
Prompts instructed ChatGPT to create clinically coherent profiles including demographic, oncologic, and molecular variables: age, sex, histologic subtype, TNM stage, and biomarkers (Epidermal Growth Factor Receptor [EGFR], Anaplastic Lymphoma Kinase [ALK], and Programed Death-Ligand 1 [PD-L1]) (Supplementary Annex 1). These variables were extracted manually and compared descriptively with global reference datasets (GLOBOCAN 2020, SEER) and biomarker meta-analyses6–11 (Fig. 1). Continuous variables were summarized as mean±SD and categorical variables as frequencies. Observed cohort frequencies (e.g., stage, biomarkers) were compared against expected population benchmarks (derived from Refs. 6–11) using Chi-squared (χ2) (p<0.05 was considered statistically significant).
The synthetic cohort included 102 virtual patients, with a mean age 66.5±6.0 years (range 51–79). In comparison, global oncology registries7,8 indicate a median diagnostic age of approximately 70 years, suggesting a slightly younger synthetic population. Sex distribution was balanced (51 men, 51 women), a distribution significantly deviating from the expected male predominance (≈65–70%)7 (χ2=14.24, p<0.01).
Histologic distribution comprised adenocarcinoma 52%, squamous-cell carcinoma 41%, and small-cell carcinoma 7%. No large-cell carcinoma was generated. Although the adenocarcinoma proportion was similar to that observed in population-based studies6 (≈45–55%), small-cell carcinoma was underrepresented (10–15% expected), and large-cell carcinoma (≈5–7%) was absent, a distribution significantly deviating from population data (χ2=11.82, p<0.01).
Staging analysis showed a predominance of early-stage disease: 65% (Stages I–II), 17% (Stage III), and 18% (Stage IV). Real-world data indicate that ≈40% of non-small-cell lung-cancer (NSCLC) are diagnosed at Stage IV. In our synthetic NSCLC cohort (N=95), only 18% were Stage IV, confirming a significant over-representation of localized stages (χ2=19.34, p<0.01).
Among the 95 simulated NSCLC cases, 45% harbored activating EGFR mutations, a rate significantly higher than Western prevalence (≈15%9; χ2=68.24, p<0.01). No ALK rearrangements were identified (0%), a significant deviation from the expected 3–5% prevalence10 (χ2=3.96, p<0.05). PD-L1 expression was uniformly intermediate (1–49%), with no negative or highly positive cases (≥50%) – a distribution significantly deviating from clinical cohorts11 (χ2=142.5, p<0.01).
Taken together, the synthetic cohort reproduced general disease patterns but showed notable deviations in key variables, particularly stage distribution and molecular profile, indicating statistically significant generative bias relative to real-world epidemiological distributions.
This exploratory study evaluated the epidemiological consistency of a synthetic lung-cancer cohort generated by ChatGPT-4.0. Although the model produced clinically plausible, coherent profiles, it exhibited systematic biases – most notably a predominance of early-stage disease and EGFR-positive tumors. These deviations likely reflect the probabilistic nature of generative models and the uneven representation of clinical scenarios in training data.
The over-representation of early-stage cases suggests a narrative bias favoring curative, well-structured clinical stories over terminal presentations. The under-representation of small-cell and absence of ALK-positive cases may relate to their lower frequency and visibility in the scientific literature. Likewise, the homogeneous PD-L1 pattern indicates that ChatGPT tends to assign intermediate values under uncertainty, reflecting a limitation in quantitative realism.
Such discrepancies align with prior observations of LLM “hallucinations,” i.e., systematic deviations from expected patterns due to the statistical weighting of learned text.12 These findings underscore the need to impose explicit population-level constraints when generating synthetic datasets. Without controlled proportions or post-generation validation, representativeness cannot be assumed.
From an educational standpoint, however, generating structured, realistic clinical cases retains considerable value. Synthetic cohorts can support virtual tumor-board exercises, simulation-based teaching, and preliminary algorithm testing. Their utility lies more in training and hypothesis generation than in precise epidemiological representation.
Reproducibility represents an additional methodological challenge. Generative models evolve continuously, and outputs obtained with one version may not be replicable with subsequent iterations, even when using identical prompts. Clear documentation of the model version, generation date, and prompting framework is therefore essential to maintain methodological transparency and facilitate reproducibility.
Another limitation concerns the batch-generation process. Because the cohort was created in groups of five patients – an operational restriction of ChatGPT-4.0 – sequential generation may have introduced distributional bias. Although no systematic drift was visually observed, future research should quantify inter-batch variability to confirm data stability.
Despite these limitations, LLM-generated cohorts demonstrate the feasibility of producing complex, internally coherent clinical datasets without compromising confidentiality. With appropriate methodological safeguards and validation procedures, this approach could enhance educational realism and foster innovation in thoracic-oncology training.
The synthetic lung-cancer cohort generated with ChatGPT-4.0 reproduced general disease patterns but showed statistically significant deviations in stage distribution and biomarker prevalence. This quantified generative bias limits its epidemiological representativeness, the findings illustrate both the promise and the boundaries of LLM-based data generation. Responsible implementation requires methodological transparency, explicit acknowledgment of biases, and quantitative validation. With continued refinement, large language models could become valuable complementary tools for simulation, medical education, and early-phase research in thoracic oncology.
Ethical statementThis study did not involve real patients or human subjects. Instead, it was based entirely on synthetic data generated through an artificial intelligence model (ChatGPT-4.0), and no identifiable or confidential patient information was used. Therefore, the requirement for informed consent was waived. Nevertheless, the study protocol was reviewed and approved by the Clinical Research Ethics Committee of our institution (Reference: PI-25-146-C), and the project was conducted in accordance with the ethical principles of the Declaration of Helsinki (2013 revision). The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.
Declaration of generative AI and AI-assisted technologies in the writing processDuring the preparation of this work, the authors used the Generative Pre-trained Transformer 4 (ChatGPT-4) not only for grammar review and translation, but also for the structured generation of a synthetic cohort of 102 virtual patients with lung cancer, based on predefined clinical, molecular, and psychosocial parameters. This simulated dataset was used for research purposes within the framework of this study. After using this tool, the authors reviewed, validated, and edited the output as necessary, and take full responsibility for the content of the publication.
FundingThis research received no external funding.
Authors’ contributionsAll authors contributed substantially to the design of the study, data analysis, manuscript drafting, and critical revision of its content. All authors have read and approved the final version of the manuscript.
Conflicts of interestThe authors declare no conflicts of interest.
Data availabilityAll data used in this study were synthetically generated using the ChatGPT-4.0 model (OpenAI) and do not correspond to real individuals. The full methodology used for data generation is described in the Methods section. Examples of the synthetic cases are provided in the Supplementary Material. No real patient data were accessed or used in this study.




